Aging Clocks Catalog · Evidence in use
Three decisions, worked through
Buying a test, choosing an intervention endpoint and adding risk information require different evidence—even when every output is called an aging clock.
1. Should I buy a test—or interpret a retest as progress?
The decision: compare finger-prick blood methylation, cheek methylation, IgG glycans and a routine-blood score. Curiosity, exploratory self-tracking and a clinical decision do not require the same evidence. Begin with what the result could change, not which product promises the youngest age.
The evidence chain: identify the measured assay, named algorithm/version, target, report, advice service and applicability evidence. The four-offering audit found model-specific publications, but did not establish that purchasing any of these testing-and-advice packages improves long-term outcomes compared with a realistic no-test alternative. A product’s independent service validation is not established by a related model paper. [M05; M12; M21; M23; M24]
Worked example—hypothetical: a blood estimate of 54, cheek estimate of 48 and pace of 0.90 are not three votes about one true age. They cannot be averaged. When chronological age goes from 50 to 51 and the same named age score goes from 52 to 50, score change is −2 and raw-gap change is −3. Neither demonstrates a treatment effect or life-years gained.
Decision that follows: demand a matched pipeline and applicable within-person uncertainty before interpreting change. A seller’s retest interval is not that evidence. No extra test may be needed for the intended decision; a clinical concern should not be replaced by a clock score. The interpretation guide shows the arithmetic and uncertainty assumptions.
2. Can clocks help assess an intervention?
The real setting: rentosertib—also described using INS018_055/ISM001-055 development nomenclature—is an AI-assisted TNIK-inhibitor drug-development programme followed by an actual human trial in idiopathic pulmonary fibrosis (IPF). Six computational clocks were applied to measured human serum. The clock analysis was computational; the trial was not wholly simulated. [SC1; SC2]
The parent trial lasted 12 weeks, comparing placebo with 30 mg once daily, 30 mg twice daily and 60 mg once daily. The primary endpoint was treatment-emergent adverse events; lung-function outcomes were secondary. Treatment-related events/discontinuations included liver-related events and diarrhea. A favorable clock pattern is not a prescribing or stand-alone safety argument. [SC2]
| Arm | Randomized | Completed parent trial | Ancillary clock count |
|---|---|---|---|
| Placebo | 17 | 15 | Not independently recovered |
| 30 mg once daily | 18 | 16 | Not independently recovered |
| 30 mg twice daily | 18 | 12 | Not independently recovered |
| 60 mg once daily | 18 | 12 | Not independently recovered |
| Total | 71 | 55 | 42 |
What was measured and reported?
Serum at baseline and weeks 2, 4 and 12 was analyzed with Olink Explore 3072; 2,841 proteins remained after quality control. The report describes normalization without bridging/anchoring controls. Serum–plasma equivalence, platform calibration and longitudinal measurement reliability cannot be assumed. [SC1]
The ancillary paper, published 7 September 2026, reports treatment-associated decreases in several predictions, the strongest aggregate pattern at week 4, and 21 comparisons below q = 0.10. These are author-reported findings—not independently recomputed counts. One-sided paired Wilcoxon and active-versus-placebo change Mann–Whitney tests were used with Benjamini–Hochberg adjustment. [SC1]
Six outputs are not six independent replications. The clocks share participants and disease-responsive proteins, and their targets differ. The paper cannot distinguish aging-related from disease-related proteomic effects. Selection after completion and consent also means this is not automatically an intention-to-treat analysis.
Arithmetic correction: 6 clocks × 3 post-baseline visits × 3 active regimens = 54 active-versus-placebo comparisons total, or 18 per regimen—not 54 per arm. The actual adjustment family still requires the numerical tables.
The six models—and what remains unpinned
| Trial label | Target / original evidence | Unresolved implementation boundary |
|---|---|---|
| ProtAge | 204-protein chronological-age LightGBM model; external CKB and FinnGen evaluation. [SC4] | Ancillary prose calls it deep learning. Prose error or different artifact remains unresolved; current access is author-request, not unrestricted bundled weights. |
| OrganAge_chrono | Chronological-age member of Goeminne’s proteomic organ family, distinct from Oh’s models. [SC7] | Exact all-organ/whole-body configuration, features, scaling and missing-protein policy are not pinned. |
| OrganAge_mortality | Mortality-oriented Goeminne family member; age-like output does not make its target chronological age. [SC7] | Related construction and shared trial participants are not independent clinical replication. |
| PAC | Age + 128 proteins; mortality-oriented Gompertz model and reference-age mapping. Reported 70:30 test split is within UK Biobank. [SC6] | Risk-equivalent age is not remaining lifespan; exact ancillary scaling/coefficient artifacts unverified. |
| ipfP3GPT | Related Galkin work includes a pathway-aware age predictor and a separate expression-generation transformer; external COVID-19 data had incomplete panel coverage. [SC5] | Alias-to-artifact mapping needs care. Current galkin_2025 / best_clock.pt weights are restricted to authorized UK Biobank RAP use. |
| PAOPAC | Ancillary paper cites a Xu et al. preprint, DOI 10.64898/2026.04.24.720503. | Original preprint not retrieved. Architecture, feature list, independent validation and relation to the package’s Han model are unverified. Embedded date is not verified first posting; do not call it peer reviewed. |
Four are presented as age-prediction implementations and two as mortality-oriented. None becomes a longitudinal pace-trained estimator merely by subtracting two predictions. [SC1]
Documentary conflicts and what has not been reproduced
Registration: the parent publication identifies NCT05938920. The ancillary body links that identifier, but reference 34 cites NCT05975983. Registry pages rendered only interface shells in the review; detailed histories were not inspected. Keep the conflict rather than borrowing eligibility or prespecification from the second identifier. [SC1; SC2]
Sites and chronology: ancillary body/methods say 21 versus 22 sites. “Multicentre Chinese trial” is supported without choosing a count. Trial completion predates some cited 2025/2026 model versions: prospectively collected samples do not prove prespecified final analyses. The dated protocol and statistical analysis plan require academic/research request through a secure environment; they were not obtained. [SC1; SC2]
Tables and artifacts: deposit OMIX008341 was not retrieved. Supplement descriptions identify S2 predictions, S3 paired tests, S4 treatment-versus-placebo tests and S5 summary; numerical rows were unavailable. No p/q-value reanalysis, raw-protein-to-score reproduction, power or equivalence calculation was performed. Exact trial models and preprocessing remain unpinned. [SC1]
Access: the current README restricts galkin_2025 weights to registered UK Biobank researchers inside its Research Analysis Platform and prohibits local download/copy/storage. ProtAge and the Han model have author-request access. Other models are not necessarily closed; unrestricted end-to-end six-model reproduction is simply not established. [SC3]
The defensible conclusion: proteomic clocks merit further study as exploratory IPF intervention readouts. The broadest clock pattern need not match the strongest clinical result. Shared-proteome pathway enrichment is not independent evidence of rejuvenation; a nonsignificant week-4/week-12 contrast does not demonstrate a plateau. Healthy-person benefit, generalized geroprotection and surrogate validity remain unestablished. [SC1; M26]
What would make an intervention change interpretable?
For the IPF clock analysis, the missing measurement record includes plate/batch allocation by arm and visit; collection, processing and storage history; feature-level quality control and missingness; fixed imputation, normalization, scalers and model versions; technical or bridge replicate outputs; and placebo longitudinal covariance. These are requirements for estimating error in the trial’s actual scores—not evidence that a processing artifact occurred. Protein-level assay controls alone cannot supply a clock-level change threshold. [SC1; R11]
Prespecify raw-score, age-gap, residual or pace change and any calendar-age input. Keep calibration fixed across the comparison; a visit-specific recentering can erase a common shift. Balance time and treatment across assay batches and evaluate any version bridge independently. If time and platform are perfectly confounded without bridge specimens, adjustment alone cannot separate their effects. This measurement review does not reproduce the six-clock trial, close its registry or supplement gaps, recover ancillary arm counts, or provide a trial-specific minimum detectable change. [R04; SC1; SC2]
A controlled counterpoint: clocks need not agree
The 2023 CALERIE methylation analysis was post hoc: 197 of 220 randomized adults had the methylation data analyzed. It reported a modest group-level reduction in DunedinPACE, but not the same response in principal-component PhenoAge and GrimAge. Achieved calorie restriction was smaller than prescribed. These nulls must remain beside the responsive finding. [M25]
The CALERIE PC-clock outcomes were standardized age gaps, not raw ages. Its 24-month control example shows a declining PCGrimAge gap alongside an increasing raw clock age at a nominal two-year interval. Baseline scale SDs are not technical error SDs. The laboratory was blinded, treatment groups were distributed across chips, and visits from a participant were kept together where possible—conditions unlike an unbridged baseline/follow-up platform switch. [R09]
A randomized group contrast can be resolved even when an individual two-visit response is poorly resolved: independent random error in a group mean averages down. Arm-confounded batch bias does not. Neither CALERIE’s group result nor a platform offset from an unrelated experiment establishes one commercial customer’s response threshold. The group-versus-individual distinction keeps those claims separate. [R09]
Different targets, sensitivity, duration, measurement or true lack of change could contribute. Disagreement does not prove every clock invalid—or the responsive one uniquely true. No longer-life outcome was directly demonstrated, and a mortality benefit estimated from separate observational associations is not a randomized outcome.
Study-team decision: prespecify the biomarker question, controlled contrast and multiplicity plan; keep clinical/functional endpoints; report all planned clocks. Establishing a surrogate requires treatment effects on the marker to predict clinical effects in the intended context, not just baseline prognosis or responsiveness. [M26]
3. Does another assay add useful risk information?
The decision: with age and routine predictors already available, does a new score improve relevant prediction enough to justify the added assay? ProtAge’s external-cohort associations make it a credible research candidate, but do not establish an individual action threshold, absolute risk or benefit from lowering the score. No unverified incremental C-index is supplied here. [SC4]
Oh et al.’s organ models use circulating proteins selected partly through tissue-enrichment information. A cognition-optimized brain score remained associated with progression after baseline clinical status, age, pTau181 and an Alzheimer’s polygenic score were included. It is not a direct brain-age measurement, a generic brain-score substitute or proof of improved care. The historical comparator is not a claim that pTau181 is the best current clinical comparator. [M20]
| Horizon | 14-biomarker score C-statistic | Conventional-risk model | Difference: arithmetic only |
|---|---|---|---|
| Five years | 0.837 | 0.772 | 0.065 |
| Ten years | 0.830 | 0.790 | 0.040 |
Deelen et al. developed a mortality score in European cohorts and evaluated it in FINRISK. Conventional factors included major clinical/behavioral predictors; chronological age was the survival time scale. These were not simply identical base models with and without one added predictor. The C-statistic reflects ranking, not calibration or a demonstrated decision benefit. [M11]
A useful positive result with a firm boundary: cohort-based scaling prevented immediate individual risk classification. This is a mortality-risk score—not MetaboAge or GlycanAge—and not a portable personal calculator. A deployment decision still needs current comparators, calibration, threshold consequences and evidence that using the result improves outcomes. [M11]
Sources and reading limits
Source labels distinguish primary research, seller documents and implementation notes. Inherited M/SC/T source observations are from 21 September 2026; added R source observations are from 22 September 2026. Calculation checks do not constitute model or clinical replication.
M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.
M12. Krištić et al. (2014; online 2013), Glycans are a novel biomarker of chronological and biological ages Primary development, external-population and small longitudinal-subset results; not validation of a pinned 2026 product.
M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.
M23. Shokhirev et al. (2024), CheekAge: a next-generation buccal epigenetic aging clock Primary abstract, development/replicate description and conflicts; Tally-funded, with company-employee authors. No exact consumer-version match.
M24. Shokhirev et al. (2024), CheekAge is predictive of mortality in human blood Primary incomplete-feature blood adaptation and mortality-model sections; company-affiliated external-cohort study, not a prospective consumer cheek-test trial.
SC1. Zhavoronkov et al. (2026), Proteomic clocks in a phase 2a trial Published 7 September 2026. Primary results, methods, availability, conflicts and supplement descriptions inspected; numerical supplements not retrieved. Developer-led; Insilico Medicine interests disclosed.
SC2. Xu et al. (2025), A generative AI-discovered TNIK inhibitor for IPF: randomized phase 2a trial Primary registration, disposition, endpoints, safety and protocol-access text inspected; protocol, statistical analysis plan and registry history not obtained. Sponsor/developer interests apply.
SC3. ProteoClock README First-party access description observed 21 September 2026; inspected blob f41e85c59b1ffeec6b8da66fb7029dd4cd77f30a. Exact release-to-trial pin and full licence tree not audited.
SC4. Argentieri et al. (2024), Proteomic aging clock predicts mortality and disease risk Primary training, external validation and covariate passages inspected through publisher-provided text delivered on ResearchGate. No model execution.
SC5. Galkin et al. (2025), AI-driven toolset for IPF and aging research Primary clock, expression-generation and external-disease-data passages; developer-affiliated evidence. No weights obtained or run.
SC6. Kuo et al. (2024), Proteomic aging clock (PAC) Primary cohort, selection, Gompertz target and train/test passages; full coefficients and ancillary-trial artifact not verified.
SC7. Goeminne et al. (online 2024; issue 2025), Plasma protein-based organ-specific aging and mortality models Primary abstract and publisher results excerpts, not full coefficient tables or frozen trial implementations.
M25. Waziry et al. (2023), CALERIE DNA methylation analysis Relevant primary methods/results, analysis population and null outcomes inspected through publisher-supplied full text delivered on ResearchGate; not a longevity-outcome trial.
M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.
M20. Oh et al. (2023), Organ aging signatures in the plasma proteome Primary construction, tissue-enrichment, platform and cognitive-progression comparator passages; no decision-impact trial identified in that material.
M11. Deelen et al. (2019), A metabolic profile of all-cause mortality risk Primary results, conventional comparator, FINRISK evaluation and scaling limitation inspected; reported models not rerun.
R04. Zhuang et al. (2025), EPIC-version differences in methylation tools Relevant methods/results and analysis code inspected, including missing probes and separate/pooled reference adjustment. Code read, not executed; a clock name does not establish complete input coverage or absolute agreement.
R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods, sample handling and university-hosted publisher article and supplement, Tables S3/S4 inspected on 22 September 2026. Rounded group summaries were back-converted, not individual-level models rerun. Baseline dispersion is not technical error; trial/model-author evidence is not independent consumer-service validation.
R11. Huang et al. (2021), preanalytical variability of inflammatory-protein measurement Relevant delay, donor, control and results sections inspected. Small EDTA-plasma, panel-specific study with Olink affiliation; no aging-clock predictions calculated and no finding about a particular IPF trial’s processing.
Next: Inspect the claim-by-claim evidence standards or build a defensible study shortlist.