Aging Clocks Catalog · Evidence in use
Three decisions, worked through
Buying a test, choosing an intervention endpoint and adding risk information require different evidence—even when every output is called an aging clock.
1. Should I buy a test—or interpret a retest as progress?
The decision: compare finger-prick blood methylation, cheek methylation, IgG glycans and a routine-blood score. Curiosity, exploratory self-tracking and a clinical decision do not require the same evidence. Begin with what the result could change, not which product promises the youngest age.
The evidence chain: identify the measured assay, named algorithm/version, target, report, advice service and applicability evidence. The four-offering audit found model-specific publications, but did not establish that purchasing any of these testing-and-advice packages improves long-term outcomes compared with a realistic no-test alternative. A product’s independent service validation is not established by a related model paper. [M05; M12; M21; M23; M24]
Worked example—hypothetical: a blood estimate of 54, cheek estimate of 48 and pace of 0.90 are not three votes about one true age. They cannot be averaged. When chronological age goes from 50 to 51 and the same named age score goes from 52 to 50, score change is −2 and raw-gap change is −3. Neither demonstrates a treatment effect or life-years gained.
Decision that follows: demand a matched pipeline and applicable within-person uncertainty before interpreting change. A seller’s retest interval is not that evidence. No extra test may be needed for the intended decision; a clinical concern should not be replaced by a clock score. The interpretation guide shows the arithmetic and uncertainty assumptions.
2. Can clocks help assess an intervention?
The real setting: rentosertib—also described using INS018_055/ISM001-055 development nomenclature—is an AI-assisted TNIK-inhibitor drug-development programme followed by an actual human trial in idiopathic pulmonary fibrosis (IPF). Six computational clocks were applied to measured human serum. The clock analysis was computational; the trial was not wholly simulated. [SC1; SC2]
The parent trial lasted 12 weeks, comparing placebo with 30 mg once daily, 30 mg twice daily and 60 mg once daily. The primary endpoint was treatment-emergent adverse events; lung-function outcomes were secondary. Treatment-related events/discontinuations included liver-related events and diarrhea. A favorable clock pattern is not a prescribing or stand-alone safety argument. [SC2]
Who contributed to the clock analysis?
| Arm | Randomized | Completed | Proteomic consent | Four-visit clock cohort | Observed week-12 FVC* |
|---|---|---|---|---|---|
| Placebo | 17 | 15 | 11 | 11 | 14 |
| 30 mg once daily | 18 | 16 | 11 | 11 | 13 |
| 30 mg twice daily | 18 | 12 | 11 | 11 | 12 |
| 60 mg once daily | 18 | 12 | 10 | 9 | 10 |
| Total | 71 | 55 | 43 | 42 | 49 |
*FVC (forced vital capacity) counts are arithmetic from the parent report’s missing week-12 counts: 3, 5, 6 and 8. They are neither a further step in the proteomic flow nor the nominal intention-to-treat analysis population. The reporting summary identifies the excluded week-12 proteomic case as a 60 mg recipient. Clock-arm counts are 11/11/11/9 at each analyzed visit; no participant-level join was performed. [SC1; SC2; S03]
Selection still matters: recovering the denominators does not restore randomization after conditioning on completion, post-trial consent and complete samples. The direction and magnitude of selection bias remain unknown. The later review response clarifies that six participants with grade ≥3 adverse events, including two placebo recipients, were retained; the article also reports a sensitivity excluding them. That sensitivity does not recover the noncompleters or nonconsenters. A sole focus on the last missing sample would overlook the earlier selection. [SC1; S05]
What was measured and reported?
Serum at baseline and weeks 2, 4 and 12 was analyzed with Olink Explore 3072; 2,841 proteins remained after quality control. The report describes normalization without bridging/anchoring controls. Serum–plasma equivalence, platform calibration and longitudinal measurement reliability cannot be assumed. [SC1]
The ancillary paper, published 7 September 2026, reports treatment-associated decreases in several predictions, the strongest aggregate pattern at week 4, and 21 of 54 comparisons below q = 0.10. These are author-reported findings—not independently recomputed counts. One-sided paired Wilcoxon and active-versus-placebo change Mann–Whitney tests were used with Benjamini–Hochberg adjustment. [SC1]
Six outputs are not six independent replications. The clocks use the same selected people, visits and serum measurements; related model signals are correlated, and their targets differ. Exact feature overlap and change-score covariance have not been recomputed. The paper cannot distinguish aging-related from disease-related proteomic effects. Selection after completion and consent also means this is not automatically an intention-to-treat analysis.
Arithmetic correction: 6 clocks × 3 post-baseline visits × 3 active regimens = 54 active-versus-placebo comparisons total, or 18 per regimen—not 54 per arm. The actual adjustment family still requires the numerical tables.
A positive robustness check: the final article reports 100,000 patient-label randomizations and p < 0.0001 for its significance-count analysis. This is stronger evidence against an exchangeable-label null than treating six outputs as independent votes. Patient-level relabeling can retain dependence among a person’s repeated outputs. It remains an author-reported analysis, not an independent rerun. [SC1; S05]
The label constraints, recomputation of tests and false-discovery-rate families, processing choices and numerical null distribution remain unverified. The analysis concerns treatment-label association within selected data; it neither restores excluded participants nor distinguishes disease response from generalized aging.
The six models—and what remains unpinned
| Trial label | Target / original evidence | Unresolved implementation boundary |
|---|---|---|
| ProtAge | 204-protein chronological-age LightGBM model; external CKB and FinnGen evaluation. [SC4] | Ancillary prose calls it deep learning. Prose error or different artifact remains unresolved; current access is author-request, not unrestricted bundled weights. |
| OrganAge_chrono | Chronological-age member of Goeminne’s proteomic organ family, distinct from Oh’s models. [SC7] | Exact all-organ/whole-body configuration, features, scaling and missing-protein policy are not pinned. |
| OrganAge_mortality | Mortality-oriented Goeminne family member; age-like output does not make its target chronological age. [SC7] | Related construction and shared trial participants are not independent clinical replication. |
| PAC | Age + 128 proteins; mortality-oriented Gompertz model and reference-age mapping. Reported 70:30 test split is within UK Biobank. [SC6] | Risk-equivalent age is not remaining lifespan; exact ancillary scaling/coefficient artifacts unverified. |
| ipfP3GPT | Related Galkin work includes a pathway-aware age predictor and a separate expression-generation transformer; external COVID-19 data had incomplete panel coverage. [SC5] | Alias-to-artifact mapping needs care. Current galkin_2025 / best_clock.pt weights are restricted to authorized UK Biobank RAP use. |
| PAOPAC | Ancillary paper cites a Xu et al. preprint, DOI 10.64898/2026.04.24.720503; full preprint methods remain unverified. | The author README is now inspectable: Olink NPX plus Age metadata, a Windows/CPython-3.9 compiled module and a separate binary. Neither was run. Architecture, features, exact trial match and public-deployment permission remain unverified; metadata alone does not identify the effect of calendar age. [S14] |
Four are presented as age-prediction implementations and two as mortality-oriented. None becomes a longitudinal pace-trained estimator merely by subtracting two predictions. [SC1]
Registrations, public protocol and remaining reproducibility gaps
Two registrations, not two aliases. NCT05938920 identifies INS018-055-003, the published 71-person parent trial. NCT05975983 identifies a different protocol, INS018-055-004; its retrieved registry snapshot lists 40 participants. The ancillary body links the first identifier but reference 34 cites the second. The citation mismatch remains; the second trial cannot supply this cohort’s denominators or prespecification. [SC1; SC2; S04; S06; S07]
The WHO registry mirrors were refreshed on 5 January 2026 (003) and 24 November 2025 (004); they are snapshots, not complete histories or current recruitment checks. The 003 mirror lists first enrolment on 19 June 2023 and registration on 28 June, whereas the publications describe conduct starting 19 July. The mirror’s geography also differs from the Chinese conduct description. Different milestones might explain this, but that was not verified: neither an unqualified prospective-registration verdict nor US recruitment for the published cohort follows. Ancillary body/methods still say 21 versus 22 sites; the parent says 21. [SC1; SC2; S04; S06; S07]
A public protocol is now available. The registry-hosted parent protocol and its embedded statistical analysis section were inspected. It documents exploratory proteomics, but not advance specification of the final six clocks, one-sided tests, adjustment families or significance-count randomization. This replaces the earlier request-only access account. [S03; S04; S05]
| Public document | Why it matters |
|---|---|
| Sampling at baseline and weeks 2, 4, 8 and 12 | The final clock analysis uses four visits and omits week 8. The review response describes available four-visit data as prespecified; it does not reconcile the broader protocol schedule. |
| Baseline plus at least one follow-up for biomarker inclusion | Not the same population rule as four-visit complete cases. The document describes exploratory/descriptive biomarker analyses without formal tests, unlike the final clock testing. |
| Mixed revision dates and labels | Revision history lists version 4.0 on 2 February 2024, while an amendment table uses version 3.0 for that date; later headers say version 4.0 on 2 August 2024. Choosing the earlier date would not settle the chronology. |
| Embedded clinical statistical section | Two-sided 0.05 tests, no multiplicity adjustment and repeated-measures/likelihood handling differ from final proteomic testing and the parent report’s ANCOVA/multiple-imputation description. A reconciled final plan is still needed. |
These are documentary differences, not proof of erroneous clinical results. Both displayed revision dates follow the reported trial start; August follows its completion. The July-2026 reporting form says downstream biomarker analysts were not blinded to group, distinct from parent-trial blinding. The final clock analysis remains exploratory. [S03; S04; S05]
Numerical tables and artifacts: descriptions identify S2 predictions, S3 paired tests, S4 active-versus-placebo comparisons and S5 summary. Actual spreadsheet rows were not retrieved; OMIX008341 remains an access dependency. No rentosertib p/q-value, patient-label randomization or raw-protein-to-score computation was independently rerun. The stated ProteoClock 1.0.0 version is not an immutable specification of all six trial artifacts. [SC1]
Access: the ProteoClock notice still restricts galkin_2025 weights to registered UK Biobank researchers inside RAP, prohibiting local download/copy/storage. ProtAge and the Han model have author-request paths. The newly inspected PAOPAC README does not make the full workflow freely reproducible. [SC3; S14]
The defensible conclusion: proteomic clocks merit further study as exploratory IPF intervention readouts. Significance counts do not establish a superior dose or regimen, and the broadest clock pattern need not match the strongest clinical result. Shared-proteome pathway enrichment is not independent evidence of rejuvenation; a nonsignificant week-4/week-12 contrast establishes neither a plateau nor equivalence. Healthy-person benefit, generalized geroprotection and surrogate validity remain unestablished. [SC1; M26]
The reported weak clock–FVC association does not isolate an aging-specific effect: noise, timing, small samples and incomplete disease measurement remain explanations. It is not a mediated fraction. The authors removed an earlier protein-overlap figure after a circularity concern; it must not be revived as final independent validation. [SC1; S05]
What would make an intervention change interpretable?
For the IPF clock analysis, the missing measurement record includes plate/batch allocation by arm and visit; collection, processing and storage history; feature-level quality control and missingness; fixed imputation, normalization, scalers and model versions; technical or bridge replicate outputs; and placebo longitudinal covariance. These are requirements for estimating error in the trial’s actual scores—not evidence that a processing artifact occurred. Protein-level assay controls alone cannot supply a clock-level change threshold. [SC1; R11]
Prespecify raw-score, age-gap, residual or pace change and any calendar-age input. Keep calibration fixed across the comparison; a visit-specific recentering can erase a common shift. Balance time and treatment across assay batches and evaluate any version bridge independently. If time and platform are perfectly confounded without bridge specimens, adjustment alone cannot separate their effects. The follow-up recovered the ancillary arm counts and public protocol, but did not reproduce the six-clock trial, reconcile the complete historical record, obtain numerical S2–S5 rows, or provide a trial-specific minimum detectable change. [R04; SC1; SC2]
A controlled counterpoint: response beside the nulls
CALERIE randomized 220 healthy, nonobese adults for two years; 218 started treatment. The post hoc methylation analysis used 197 people with baseline and at least one follow-up: 128 assigned calorie restriction and 69 control. At 12 months, the counts were 125/66; at 24 months, 117/68. The reported intention-to-treat comparisons preserve assignment among those with required molecular data—not complete outcomes for all 220. Average achieved restriction was about 12%, versus 25% prescribed. These are study descriptions, not treatment advice. [S08; S09]
| Model and quantity | 12 months | 24 months |
|---|---|---|
| PCPhenoAge age-gap change | −0.03 [−0.19, 0.12] | 0.05 [−0.11, 0.20] |
| PCGrimAge age-gap change | −0.04 [−0.16, 0.07] | 0.05 [−0.07, 0.17] |
| DunedinPACE pace change | −0.29 [−0.45, −0.13] | −0.25 [−0.41, −0.09] |
Table S4, author-reported estimates. PC-clock outputs are changes in clock-minus-calendar-age gaps; pace has its own scale. Baseline standard deviations are not technical-error SDs, and these pointwise intervals were not multiplicity-adjusted by the catalog. Negative values indicate a lower score/pace change relative to control. [S08; S09]
The DunedinPACE response and PC-clock nulls belong together. Neither null establishes equivalence or makes the responsive measure uniquely true. Multiplying the pace estimates by the published baseline SD of 0.09 gives approximately −0.0261 [−0.0405, −0.0117] pace units at 12 months and −0.0225 [−0.0369, −0.0081] at 24 months. This is arithmetic on rounded group summaries, not a personal response threshold. [S08]
The retained 24-month control example shows that a declining PCGrimAge age gap can coexist with increasing raw clock age. Laboratory blinding, treatment allocation across chips and processing repeat samples together where possible distinguish CALERIE from an unbridged platform switch. Random error in group means can average down; arm-confounded batch bias does not. [S08; S09]
Multiplicity and printed precision: the catalog tested six primary-model intention-to-treat comparisons and an expanded set of all 22 S4 intention-to-treat comparisons using Bonferroni and Holm at familywise 0.05. Both DunedinPACE visits remain below threshold in the six-test sensitivity. The 12-month result remains below it in the 22-test sensitivity too; at 24 months, the central printed value passes but its nearest-rounding envelope crosses 0.05 under both corrections. The unrounded p-value is needed to settle that broader-family classification. These are bounded sensitivity calculations, not the authors’ original adjustment procedure or a raw-model rerun. Inspect the values and rounding assumption. [S08]
What the adherence and cell-adjustment sensitivities change
Treatment-on-treated (TOT) estimates use instrumental variables and are displayed for 20% restriction, not the average achieved exposure. Assignment was randomized; adherence was not. A causal dose interpretation needs additional assumptions, including instrument relevance and exclusion. Adjusting for estimated cell changes also changes the target and may remove real biological response. [S08; S09]
For 24-month DunedinPACE TOT, the cell-adjusted estimate is −0.30 [−0.58, −0.02], p = 0.033, versus −0.40 [−0.67, −0.12], p = 0.004 without that adjustment, in baseline-SD units. The former does not meet the paper’s p < 0.005 convention despite its pointwise interval excluding zero. Cell-adjusted intention-to-treat remains negative, but its printed p = 0.005 cannot settle a strict less-than-0.005 decision. This weakens a claim of unchanged robustness under every sensitivity; it does not overturn the main intention-to-treat result. [S08]
Mediation was analyzed—not validated surrogacy. CALERIE modeled whether DunedinPACE changes mediated clinical/risk-marker changes. Figure S4 describes small mediated fractions: many proportion-mediated intervals excluded zero, while intervals for the average causal mediation effect (ACME) included zero except for log CRP and clinical Phenotypic Age (the blood-chemistry composite). Those summaries are not interchangeable. The mediator was not randomized, and the models did not establish that pace change preceded clinical change. The evidence guide explains the causal assumptions and current-code reconciliation limits. No mediation model was independently rerun. [S08; S09]
No mortality or longer-life benefit was directly demonstrated in this molecular report. A mortality benefit extrapolated from external observational pace associations is not a randomized outcome. Different targets, duration, measurement and biological responses can explain clock disagreement without a universal verdict on clocks. [S08]
Study-team decision: define the response claim, lock the estimator and controlled contrast, report nulls and sensitivity limits, and retain independent clinical/functional endpoints. Trial responsiveness, statistical mediation and prediction of treatment effects on clinical outcomes are different achievements. A useful disease-response biomarker need not be described as generalized rejuvenation. [SC1; S08; M26]
3. Does another assay add useful risk information?
The decision: with age and routine predictors already available, does a new score improve relevant prediction enough to justify the added assay? ProtAge’s external-cohort associations make it a credible research candidate, but do not establish an individual action threshold, absolute risk or benefit from lowering the score. No unverified incremental C-index is supplied here. [SC4]
Oh et al.’s organ models use circulating proteins selected partly through tissue-enrichment information. A cognition-optimized brain score remained associated with progression after baseline clinical status, age, pTau181 and an Alzheimer’s polygenic score were included. It is not a direct brain-age measurement, a generic brain-score substitute or proof of improved care. The historical comparator is not a claim that pTau181 is the best current clinical comparator. [M20]
| Horizon | 14-biomarker score C-statistic | Conventional-risk model | Difference: arithmetic only |
|---|---|---|---|
| Five years | 0.837 | 0.772 | 0.065 |
| Ten years | 0.830 | 0.790 | 0.040 |
Deelen et al. developed a mortality score in European cohorts and evaluated it in FINRISK. Conventional factors included major clinical/behavioral predictors; chronological age was the survival time scale. These were not simply identical base models with and without one added predictor. The C-statistic reflects ranking, not calibration or a demonstrated decision benefit. [M11]
A useful positive result with a firm boundary: cohort-based scaling prevented immediate individual risk classification. This is a mortality-risk score—not MetaboAge or GlycanAge—and not a portable personal calculator. A deployment decision still needs current comparators, calibration, threshold consequences and evidence that using the result improves outcomes. [M11]
Sources and reading limits
Source labels distinguish primary research, seller documents and implementation notes. Foundation observations are from 21 September 2026; R reliability/calibration observations are from 22 September. The S intervention sources and the updated SC1/SC2 inspections are dated 26 September 2026. Author reports, documentary verification and the catalog’s aggregate calculations are distinguished; none is raw-data or model replication.
M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.
M12. Krištić et al. (2014; online 2013), Glycans are a novel biomarker of chronological and biological ages Primary development, external-population and small longitudinal-subset results; not validation of a pinned 2026 product.
M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.
M23. Shokhirev et al. (2024), CheekAge: a next-generation buccal epigenetic aging clock Primary abstract, development/replicate description and conflicts; Tally-funded, with company-employee authors. No exact consumer-version match.
M24. Shokhirev et al. (2024), CheekAge is predictive of mortality in human blood Primary incomplete-feature blood adaptation and mortality-model sections; company-affiliated external-cohort study, not a prospective consumer cheek-test trial.
SC1. Zhavoronkov et al. (2026), Proteomic clocks in a phase 2a trial Published 7 September 2026. Follow-up inspection on 26 September covered relevant full results/methods, Figure 3, extended-data captions, availability and supplement descriptions, including the 100,000-randomization caption. Actual S2–S5 spreadsheet rows and permutation were not independently reproduced. Developer-led; Insilico and model-author interests disclosed.
SC2. Xu et al. (2025), A generative AI-discovered TNIK inhibitor for IPF: randomized phase 2a trial Published 3 June 2025. Follow-up inspection on 26 September covered disposition, FVC/safety, statistical methods and registration. Public parent-protocol evidence is now available (S04); the final analysis-plan chronology and raw clinical analysis remain unreproduced. Sponsor/developer interests apply.
SC3. ProteoClock README First-party access description observed 21 September 2026; inspected blob f41e85c59b1ffeec6b8da66fb7029dd4cd77f30a. Exact release-to-trial pin and full licence tree not audited.
SC4. Argentieri et al. (2024), Proteomic aging clock predicts mortality and disease risk Primary training, external validation and covariate passages inspected through publisher-provided text delivered on ResearchGate. No model execution.
SC5. Galkin et al. (2025), AI-driven toolset for IPF and aging research Primary clock, expression-generation and external-disease-data passages; developer-affiliated evidence. No weights obtained or run.
SC6. Kuo et al. (2024), Proteomic aging clock (PAC) Primary cohort, selection, Gompertz target and train/test passages; full coefficients and ancillary-trial artifact not verified.
SC7. Goeminne et al. (online 2024; issue 2025), Plasma protein-based organ-specific aging and mortality models Primary abstract and publisher results excerpts, not full coefficient tables or frozen trial implementations.
M25. Waziry et al. (2023), CALERIE DNA methylation analysis Relevant primary methods/results, analysis population and null outcomes inspected through publisher-supplied full text delivered on ResearchGate; not a longevity-outcome trial.
M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.
M20. Oh et al. (2023), Organ aging signatures in the plasma proteome Primary construction, tissue-enrichment, platform and cognitive-progression comparator passages; no decision-impact trial identified in that material.
M11. Deelen et al. (2019), A metabolic profile of all-cause mortality risk Primary results, conventional comparator, FINRISK evaluation and scaling limitation inspected; reported models not rerun.
R04. Zhuang et al. (2025), EPIC-version differences in methylation tools Relevant methods/results and analysis code inspected, including missing probes and separate/pooled reference adjustment. Code read, not executed; a clock name does not establish complete input coverage or absolute agreement.
R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods, sample handling and university-hosted publisher article and supplement, Tables S3/S4 inspected on 22 September 2026. Rounded group summaries were back-converted, not individual-level models rerun. Baseline dispersion is not technical error; trial/model-author evidence is not independent consumer-service validation.
R11. Huang et al. (2021), preanalytical variability of inflammatory-protein measurement Relevant delay, donor, control and results sections inspected. Small EDTA-plasma, panel-specific study with Olink affiliation; no aging-clock predictions calculated and no finding about a particular IPF trial’s processing.
S03. Ancillary Reporting Summary Four-page publisher form, dated July 2026; sample exclusion, downstream analysis blinding, software versions and public protocol link inspected. A retrospective reporting form, not an advance analysis plan.
S04. Public parent protocol INS018-055-003 Relevant revision history, sampling and embedded statistical sections inspected in the 110-page registry-hosted file. Mixed February/August 2024 version headers and the planned week-8 sample remain unresolved. Not a complete historical chain or an independently authenticated final statistical analysis plan.
S05. Ancillary transparent peer-review exchange Relevant author/reviewer discussions of selection, adverse events, permutation, removed overlap analysis, sampling and disease-versus-aging interpretation inspected. Author explanations are not independent replication; removed analyses are not final evidence.
S06. WHO ICTRP representation of NCT05938920 Displayed record inspected; refreshed 5 January 2026, protocol INS018-055-003. Selected registry fields, not a complete change history or verification of current recruitment status.
S07. WHO ICTRP representation of NCT05975983 Displayed record inspected; refreshed 24 November 2025, protocol INS018-055-004. A different trial, not an alternative registration for the 71-person parent cohort.
S08. Waziry et al. (2023), CALERIE article and supplement University-hosted publisher text; relevant methods, Tables S4/S6 and Figure S4 inspected on 26 September 2026. Forty rounded comparison summaries support the catalog’s bounded arithmetic, not participant-level model or mediation replication. Trial/model-author interests do not establish independent consumer-service validation.
S09. CALERIE author code Inspected tree db56e58a7546934f3d9f2c5764cc8ddaa5894a9c. Data and mediation files read completely; relevant primary-analysis sections read. No Stata model executed. Current-code concerns do not establish which revision produced the published figure or prove an error in its numerical results.
S14. PAOPAC author README Complete text inspected on 26 September 2026, blob ba5d4a8ee7ce2107d8fe34f1dbe24b40dc2da5f0. Input/platform requirements were read; compiled module and separate binary model were not obtained or run. Full preprint methods, exact trial-artifact correspondence and permission for public deployment remain unverified.
Next: Inspect the claim-by-claim evidence standards or build a defensible study shortlist.