Shaduf.
Aging Clocks Catalog/Evidence and methods

Aging Clocks Catalog · How claims are assessed

How the evidence is judged

Reliable measurement, prediction, intervention response and better health decisions are different achievements. No single statistic establishes them all.

Foundation reviewed ; reliability and calibration evidence reviewed . A bounded documentary review, not a diagnostic or treatment service.

The important question is not whether a clock is “validated” in the abstract. It is validated for which claim, in which setting, with which implementation? Consistent measurement, external prognosis, treatment response, clinical utility and surrogate validity are distinct. A research biomarker can be useful without meeting the last two standards. [M03; M05; M26]

A claim needs the right kind of evidence

Separate standards—not a ladder of marketing badges
ClaimEvidence that addresses itEvidence that is not enough
“The result is reproducible”Matched collection-to-report replicates, absolute error and relevant batch/platform conditions.Calendar-age correlation, another lab’s ICC or triplicate measurement of just one assay stage.
“The result predicts future outcomes”Prospective external evaluation, suitable covariates, calibration and performance uncertainty.Development fit, cross-sectional disease association or selecting a model on the outcomes used to evaluate it.
“The intervention changed the result”A prespecified controlled contrast, reliable longitudinal measurement, missingness strategy and multiplicity handling.A before/after change, favorable subgroup selection or separate significance tests in each arm.
“Changing the result improves health”Causal evidence connecting intervention, score change and relevant health benefit, considering alternative pathways.Baseline prognosis or pathway enrichment from the same assay used to calculate the score.
“Using the test is clinically useful”A validated action/threshold and benefits, harms and costs relative to a realistic decision without the test.Statistical significance, a high area-under-curve statistic or bundled advice that could be given without testing.
“The result is a surrogate endpoint”Context-specific evidence that treatment effects on the biomarker predict treatment effects on a clinical endpoint.Observational association, response to one intervention, expert enthusiasm or “surrogate” used simply to mean proxy.

Assessment standards informed by the inspected methods, controlled studies and FDA–NIH BEST definitions; not a claim that every paper was designed to satisfy every row. [M03; M05; M11; M20; M25; M26; SC1]

Prediction questionDoes the baseline score help forecast an outcome?

A conditional association or discrimination result can address prognosis. It does not show that lowering the score prevents that outcome.

Treatment questionDoes a treatment’s marker effect predict its clinical effect?

Surrogate validity concerns effects in a defined context. The two questions require different evidence; one cannot be substituted for the other.

Prognosis and surrogacy are not synonyms. Observational association alone is generally insufficient for validated-surrogate status. [M26]

Three separate questions before interpreting a change

  1. Does it exceed the relevant error?

    First match the output and complete pipeline. Identical numerical inputs test software reproducibility; a split specimen tests only steps after the split; separate collections include collection and intervening physiology. Paired version bridges estimate platform/pipeline disagreement. State bias and absolute difference variance for the relevant experiment rather than substitute calendar-age error or an ICC. [R03; R05]

  2. Does it reflect the intended biological construct?

    A real inflammatory or cell-mixture change is not automatically sustained aging. Examine standardized repeated visits, untreated trajectories and independent functional or clinical measures. The inspected timing study’s dense series was one person and its additional paired data came from a stress experiment; it supports attention to context, not a universal daily age fluctuation. [R08; R09; R11]

  3. Does using it improve a decision?

    Identify the action triggered by the score and compare benefits, harms and costs with a realistic no-test decision. Even a resolved, biologically relevant change does not establish clinical utility. Surrogacy separately requires treatment effects on the marker to predict clinical effects in a defined context. [M26]

No single reliability coefficient answers all three. Under a stable, approximately normal, unbiased error process, 1.96 × SD(error difference) is a statistical repeatability bound—not a clinically important difference. In general, Var(e₂ − e₁) = Var(e₂) + Var(e₁) − 2 Cov(e₁, e₂); the equal-independent-error shortcut is a special case. This is a measurement model, not an estimated threshold for any product. The matched-platform example estimates a mean version offset, not same-platform intervisit variation.

A group contrast is not an individual response classifier

For a hypothetical balanced trial with independent participants, n per arm and individual-change measurement-error SD sD, the measurement-error component of the difference-in-mean-changes has standard error sD × √(2/n). This excludes biological heterogeneity, clustering, missingness and systematic batch bias. More participants can reduce random error in a group contrast; they do not remove an arm-confounded shift. A statistically detectable group effect can therefore be smaller than an individual repeatability bound without contradiction.

The CALERIE control back-conversion illustrates why output definition also matters: its PCGrimAge summary is a group age-gap change standardized by baseline dispersion, not individual measurement error or raw-age change. [R09]

The missing uncertainty for clinical and proteomic scores

Clinical-laboratory composites: for a fixed linear score, input-error contribution is w′Σw: weights combined with the covariance matrix of input errors in the correct transformed units. The inspected clinical Phenotypic Age formula is linear in those transformed inputs, including log CRP, when calendar age is treated as known. A serial-use interval additionally requires collection/physiological and cross-visit covariance for the proposed timing. Those matrices were not established for the inspected services. Adding or averaging analyte CVs cannot recover them. [M03; R15]

Proteomic clocks: a linear score likewise needs weighted covariance; for a nonlinear clock, full-pipeline replicated predictions under the actual specimen, panel, preprocessing and model version are the more direct empirical check. Average protein ICCs or a count of reproducible proteins do not supply final-score error. A small preanalytical experiment found altered inflammatory-protein measurements despite analytical controls, but calculated no aging clock; it cannot establish error in ProtAge, PAC or rentosertib’s six outputs. [R10; R11]

Clinical-laboratory and proteomic individual-change thresholds therefore remain unresolved for the specific services and pipelines inspected. Neither methylation-platform variance nor external risk association fills the missing score-level evidence.

Numbers only mean something with their comparator

Mean absolute error for calendar age does not measure longitudinal reliability or outcome prediction. A hazard ratio describes conditional association under a fitted model. A C-statistic measures ranking; it does not establish that a predicted 10% risk occurs 10% of the time. That last question concerns calibration. Clinical utility additionally asks what happens when a decision changes because of the result.

The NMR example reports five-year C-statistics of 0.837 versus 0.772, and ten-year values of 0.830 versus 0.790, in FINRISK. It has a real conventional-risk comparison, but cohort-based scaling limits immediate individual classification and the models were not simply the same base with one extra predictor. The worked application preserves that boundary. [M11]

What a favorable summary must not hide

Negative findings, contradictions and limits retained in this catalog
EvidenceWhat remains visibleConsequence
CALERIEModest DunedinPACE response; principal-component PhenoAge and GrimAge did not show the same response. [M25]Do not present all clocks as agreeing or declare the responsive measure uniquely true.
Retinal age gapAll-cause mortality association, but no significant cardiovascular- or cancer-mortality association in the inspected abstract. [M14]Outcome-specific nulls narrow the claim; no diagnosis or screening benefit is established.
2025 clock benchmarkAbstract: 39 biomarkers, more than 20,000 people, little relationship between age accuracy and mortality-prediction capacity. Full methods unavailable. [M06]No detailed ranking or universal winner is adopted.
Six-clock IPF analysisRegistry, 21/22-site count, architecture and comparison-family discrepancies; selected 42-person cohort, numerical supplements unavailable. [SC1; SC2; SC4]Do not claim prespecification, numerical replication, intention-to-treat clock analysis, six independent replications or generalized rejuvenation.
Public implementationCurrent ProteoClock restrictions apply to galkin_2025 weights; other models have distinct access paths. [SC3]Neither “all proteomics is closed” nor “all six are freely reproducible” follows.
Commercial evidenceModel publications and external cohorts are not independent validation of every current service version. [M21; M22; M23; M24]Keep product-to-paper equivalence, individual error and benefit of the purchased package as separate unresolved claims.

What this review did

This is a purposive, claim-driven documentary review with foundation evidence/access observations dated 21 September 2026 and a reliability/calibration follow-up dated 22 September 2026. It follows selected original methods and validations; separates parent and ancillary trials; inspects public product/legal documents; and reads code where it changes feasibility. It is not a systematic review, market census, pooled meta-analysis or product certification. Search visibility was not treated as popularity or market share.

“Inspected” means the stated relevant text was read—not every citation, figure or supplement. Some emerging-family entries and the 2025 benchmark rely only on abstracts/metadata. The foundation used publisher-provided ProtAge and CALERIE text delivered on ResearchGate; unrelated citing-paper snippets were excluded. The follow-up also inspected CALERIE’s university-hosted publisher supplement and Tay’s Supplementary Table S4 visually. Source access depth remains specific to each entry, not an assertion that all supplements were read.

All study numbers are author-reported unless explicitly identified as arithmetic. No participant-level dataset was obtained, clock package run, restricted model acquired, product purchased or seller contacted. The follow-up reconstructed agreement limits for all 18 selected specimen/model rows in Tay’s rounded S4 aggregates, keeping age and pace units separate, and checked consistency with reported rounding. It added conditional normal-theory intervals, back-converted rounded CALERIE group summaries and checked fixed synthetic clinical/calibration examples. It did not reproduce raw-pair nonparametric tests, methylation processing, imputation or clock fitting. [R05; R09; R15]

How to read the local source labels

M identifies scientific methods/validation or the official terminology framework; SC identifies the six-clock case, parent trial and model/access sources; C identifies seller/contract documents; T identifies implementation documentation; added R labels identify the reliability/calibration evidence review. They are source identifiers, not quality scores. A local source note tells you the inspected depth and relevant conflicts. A source’s availability does not give another source’s claims its authority.

Original papers support scientific claims; sellers describe their own offerings and terms; current code documentation describes access, not what an older trial necessarily ran. A “latest” documentation address is mutable. The date attached to the review is not a new publication date, a purchase guarantee or a claim that every underlying experiment was recent.

The paired-platform reconstruction assumes independent participants and approximately normal, homogeneous differences within each specimen/model. Rounded aggregates cannot test those assumptions, outliers, proportional bias or subgroup behavior. The intervals are pointwise rather than simultaneous over all rows. S4 and main-text SDs disagree in some rows, and a normalization name/reference inconsistency remains; S4 was used without inventing unrounded data or silently resolving preprocessing. [R05]

What would materially strengthen the conclusions?

For an individual retest: independent, matched collection-to-report repeatability, untreated longitudinal variation, exact report-version correspondence and a demonstrated decision advantage. An unrelated age-error statistic would not resolve the gap.

For the IPF case: reconciled registry histories, dated protocol/analysis plan, ancillary arm denominators, numerical S2–S5 tables and exact study-specific model artifacts. Reproducing tests from published predictions would be useful, but would still not equal reproducing all predictions from raw proteins.

For clinical risk use: external calibration and threshold/net-benefit evaluation against a realistic current comparator, followed by evidence on the decisions and outcomes produced by using the score. One more adjusted association would not by itself settle utility.

Reader safeguard: uncertainty is part of the conclusion, not a footnote to remove. The catalog remains useful by narrowing a claim to what was actually tested, not by assuming the unavailable evidence would be favorable.

Sources and reading limits

Source labels distinguish primary research, seller documents and implementation notes. Inherited M/SC/T source observations are from 21 September 2026; added R source observations are from 22 September 2026. Calculation checks do not constitute model or clinical replication.

M03. Levine et al. (2018), An epigenetic biomarker of aging for lifespan and healthspan Clinical selection, units/coefficient table, methylation stage and validation passages inspected; no raw-data reanalysis.

M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.

M06. Ying et al. (2025), A unified framework for systematic curation and evaluation of aging biomarkers Author-institution abstract/metadata only; detailed methods and rankings not adopted.

M11. Deelen et al. (2019), A metabolic profile of all-cause mortality risk Primary results, conventional comparator, FINRISK evaluation and scaling limitation inspected; reported models not rerun.

M14. Zhu et al. (2022), Retinal age gap as a predictive biomarker for mortality risk Complete primary abstract, including outcome-specific nulls; full methods and deployment validation not inspected.

M20. Oh et al. (2023), Organ aging signatures in the plasma proteome Primary construction, tissue-enrichment, platform and cognitive-progression comparator passages; no decision-impact trial identified in that material.

M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.

M22. Sehgal et al. (2025), Systems Age Version of record: 15 September 2025. Abstract/metadata and primary update linkage inspected; detailed rankings and current product equivalence withheld.

M23. Shokhirev et al. (2024), CheekAge: a next-generation buccal epigenetic aging clock Primary abstract, development/replicate description and conflicts; Tally-funded, with company-employee authors. No exact consumer-version match.

M24. Shokhirev et al. (2024), CheekAge is predictive of mortality in human blood Primary incomplete-feature blood adaptation and mortality-model sections; company-affiliated external-cohort study, not a prospective consumer cheek-test trial.

M25. Waziry et al. (2023), CALERIE DNA methylation analysis Relevant primary methods/results, analysis population and null outcomes inspected through publisher-supplied full text delivered on ResearchGate; not a longevity-outcome trial.

M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.

SC1. Zhavoronkov et al. (2026), Proteomic clocks in a phase 2a trial Published 7 September 2026. Primary results, methods, availability, conflicts and supplement descriptions inspected; numerical supplements not retrieved. Developer-led; Insilico Medicine interests disclosed.

SC2. Xu et al. (2025), A generative AI-discovered TNIK inhibitor for IPF: randomized phase 2a trial Primary registration, disposition, endpoints, safety and protocol-access text inspected; protocol, statistical analysis plan and registry history not obtained. Sponsor/developer interests apply.

SC3. ProteoClock README First-party access description observed 21 September 2026; inspected blob f41e85c59b1ffeec6b8da66fb7029dd4cd77f30a. Exact release-to-trial pin and full licence tree not audited.

SC4. Argentieri et al. (2024), Proteomic aging clock predicts mortality and disease risk Primary training, external validation and covariate passages inspected through publisher-provided text delivered on ResearchGate. No model execution.

R03. McEwen et al. (2018), EPIC/450K clock assessment Relevant methods/results/discussion inspected, including normalization and the post-bisulfite-conversion technical split. No raw-array reproduction; its age-error-based heuristic is not adopted as an individual detection threshold.

R05. Tay et al. (2025), DNAm age differences across EPIC versions and specimens Relevant methods/results/conflicts and Supplementary Table S4 inspected. Same extracted DNA, 16 paired samples per specimen, ages 40–60, 14 Chinese participants; EPICv2 imputation was part of the pipeline. Rounded aggregate reconstruction, not raw-data reproduction. Main-text/S4 and normalization-reference discrepancies remain unresolved. Disclosures include clock-foundation and advisory roles.

R08. Koncevičius et al. (2024), circadian variation in epigenetic age Relevant design/results and limits inspected, including the one-person dense series and stress-experiment origin of the separate paired dataset. No universal daily amplitude or personal threshold inferred.

R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods, sample handling and university-hosted publisher article and supplement, Tables S3/S4 inspected on 22 September 2026. Rounded group summaries were back-converted, not individual-level models rerun. Baseline dispersion is not technical error; trial/model-author evidence is not independent consumer-service validation.

R10. Haslam et al. (2022), plasma-proteomics reproducibility Primary abstract and disclosures inspected; full methods not recovered. Assay-profile findings, not an executed clock-level covariance or repeatability analysis.

R11. Huang et al. (2021), preanalytical variability of inflammatory-protein measurement Relevant delay, donor, control and results sections inspected. Small EDTA-plasma, panel-specific study with Olink affiliation; no aging-clock predictions calculated and no finding about a particular IPF trial’s processing.

R15. BioAge calculation sources: phenoage_calc.R, kdm_calc.R and hd_calc.R. Complete files reread on 22 September 2026; the inspected blobs are recorded on the tools page. Fixed synthetic arithmetic was checked separately; the R package was not run and no independent author-supplied reference vector or service-specific covariance was obtained.

Next: Inspect the three worked evidence chains or separate public code from reproducible access.

Search published pools, pages, reports, and evidence.