Aging Clocks Catalog · For study teams
Choose a clock for a study
Choose the question first. Then rule out incompatible specimens, targets, populations and implementations.
A defensible shortlist begins by excluding mismatches. A strong model can be wrong for your study because it predicts the wrong target, requires another specimen or depends on a reference fit you cannot reproduce. The recommendations below are research starting points, not a universal ranking. [M03; M05; SC4; SC6; T01]
Five gates before comparing performance
- Define the inference and comparator.
Age resemblance, future outcome prediction and intervention response require different tests. Specify what a clock must add to chronological age, ordinary predictors or direct clinical/functional outcomes.
- Match the measurement.
Record specimen, assay generation, feature coverage, the imputer, normalization and reference calibration—not just the platform family. Protein-name overlap is not Olink–SomaScan equivalence; cheek methylation is not automatically whole-blood methylation.
- Match the population.
Check age range, sex, ancestry, geography, disease and medication context. External validation in one cohort does not validate every subgroup or use.
- Match the uncertainty to the task.
For change, seek absolute within-person error and untreated longitudinal variation—not just calendar-age accuracy or a high reliability coefficient.
- Freeze an accessible implementation.
Pin version, coefficients/weights, preprocessing, reference constants, code and permissions before analysis. A model name alone is not a reproducible specification.
These are design safeguards drawn from the methods and feasibility audit, not laboratory practices independently audited on site. [M03; M05; M20; T01; SC3]
Write the pipeline beside the clock name
| Record explicitly | What must be fixed or documented |
|---|---|
| Model version and output | Original or principal-component estimator; coefficient/weight version; raw score, age gap, named-reference residual or pace; units and any calendar-age input. |
| Specimen and assay generation | Tissue/cell preparation, array or protein-panel generation, manifest, feature IDs and duplicate-probe rule. “EPIC” or a protein name alone is incomplete. |
| Imputation | Missing features, permitted missingness patterns, imputer version and training reference. A successfully computed result does not prove supported coverage. |
| Normalization | Before-clock and within-clock transformations, frozen scalers and processing version. Do not choose the pipeline after seeing which yields a favorable change. |
| Reference calibration | Reference cohort, coefficients and final scale; whether fitted externally, pooled or separately by visit/platform. Keep bridge-fitting data separate from its evaluation. |
These are proposed reporting requirements grounded in the inspected platform and calibration workflows. For example, missing final predictor probes and a PC model’s larger input set are different coverage problems; “only a few missing probes” is not an error bound. [R04] A separate age calibration can remove a shared offset without establishing absolute agreement.
A task-by-model shortlist
| Study situation | Defensible comparison | Do not substitute |
|---|---|---|
| Blood DNAm; future morbidity or mortality | Compare GrimAge2 or DNAmPhenoAge with age, routine predictors and an age-trained reference. Evaluate held-out discrimination, calibration and incremental prediction. [M03; M04] | The smallest calendar-age mean absolute error is not the mortality winner. Keep whole-blood and cheek pipelines separate. |
| Randomized intervention; pace or response | DunedinPACE is one longitudinally grounded candidate. Prespecify a between-group longitudinal contrast, confidence intervals, missingness and multiplicity; retain independent clinical outcomes. [M05; M25] | Do not select the clock that declines most after seeing the results, or silently replace original models with principal-component variants. |
| Existing routine laboratory values | Compare fixed clinical Phenotypic Age or a fixed KDM specification with the original values and a relevant clinical model. [M03; T01] | Do not invent an absent biomarker, infer a KDM reference fit, or treat a proprietary routine-blood product as published Phenotypic Age. |
| Proteomics; organ-related risk | Compare global and organ-related models on matched samples, targets and protein panels; seek external prediction beyond ordinary risk measures. [M20; SC4; SC7] | “OrganAge” is not one implementation. A blood protein proxy is not an organ biopsy, and serum/plasma interchangeability must be tested. |
| Cells, organisms or postmortem tissue | Use tissue-, cell- or species-matched methods with independent biological replicates and appropriate functional endpoints. [M09; M18] | Thousands of cells from a few donors are not thousands of independent people. A cell-state shift is not human clinical rejuvenation. |
Reliability is not a single number
DunedinPACE reports strong same-platform technical-replicate intraclass correlations (ICCs), with lower cross-platform reliability. OMICmAge reports an ICC of 0.998 from 30 replicates. These are useful findings in studied settings—not a commercial service’s universal minimum detectable change. The DunedinPACE reliability datasets differ in design and population, so their ICC differences do not isolate a causal platform penalty. [M05; M21]
Cross-sectional age error is not within-person repeatability. Mean absolute error against calendar age includes calibration, age range and the relation of the training target to birthdays. ICC depends on between-person variance and whether consistency or absolute agreement is being assessed. Neither supplies an individual difference SD by itself. Processing one stored sample twice does not capture collection, shipping or day-to-day physiology. [M05; R03]
PCPhenoAge and PCGrimAge deserve consideration when longitudinal precision matters, but they are different estimators—not silent upgrades to original-clock baselines. Better technical precision does not establish better biological relevance or zero platform bias. [R01]
A concrete version discontinuity: in Tay et al.’s 16 paired same-DNA samples, buffy-coat PCPhenoAge averaged −1.96 score-years for EPICv2 minus EPICv1. Participants were aged 40–60, 14 were Chinese, and the pipeline included EPICv2 imputation. This is combined platform/pipeline evidence, not a universal correction or intervisit error estimate. The worked comparison gives the conditional mean-offset interval. [R05]
For multi-visit studies: balance or randomize assay batches across time and treatment, retain blinded controls and paired bridge specimens, and report all planned clocks—including nulls. Validate a platform bridge on held-out specimens spanning the intended range. A fully confounded platform/time change cannot be separated by statistical adjustment alone. These are proposed safeguards; the catalog has not audited a laboratory’s implementation.
Protect the validation from leakage
Keep model development and the claimed validation separate. Feature selection, normalization or residualization using validation outcomes can contaminate the result; any adaptation needs a separately labeled evaluation. Split repeated or family-related data at the participant/family level, single-cell data at least at donor level, and assess device/site effects for imaging. A large test set does not cure a wrong split.
In the brain single-nucleus example, the study includes 73,941 nuclei from 31 postmortem donors. The donor—not each nucleus—is the independent person. Its cell-type resolution does not establish transfer to living-person diagnosis. [M09]
A compact record to carry into an analysis plan
Carry the implementation record into the plan alongside validation split, outcome and horizon, ordinary comparator, uncertainty and subgroup checks, rights and runnable artifacts. For an intervention, specify the raw-score/gap/residual/pace contrast, calendar-age handling, assay balancing, missing-data strategy and multiplicity family. For clinical or proteomic retests, missing score-level covariance or repeated predictions remain a reason to withhold an individual threshold—not to borrow methylation-platform variance.
These details matter more than a brand name or “best clock” badge. The benchmark abstract covers 39 biomarkers in more than 20,000 participants and reports little relationship between age accuracy and mortality-prediction capacity. Its full methods were not obtained, so detailed rankings are not adopted. [M06]
Sources and reading limits
Source labels distinguish primary research, seller documents and implementation notes. Inherited M/SC/T source observations are from 21 September 2026; added R source observations are from 22 September 2026. Calculation checks do not constitute model or clinical replication.
M03. Levine et al. (2018), An epigenetic biomarker of aging for lifespan and healthspan Clinical selection, units/coefficient table, methylation stage and validation passages inspected; no raw-data reanalysis.
M04. Lu et al. (2022), DNA methylation GrimAge version 2 Primary version, training and multi-cohort validation passages inspected; no commercial-version equivalence inferred.
M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.
M06. Ying et al. (2025), A unified framework for systematic curation and evaluation of aging biomarkers Author-institution abstract/metadata only; detailed methods and rankings not adopted.
M09. Muralidharan et al. (2025), Human Brain Cell-Type-Specific Aging Clocks Primary abstract/opening results and availability text; complete donor-split and per-cell methods not audited.
M18. Lu et al. (2023), Universal DNA methylation age across mammalian tissues Primary abstract, target/scope and limitation passages; no treatment-equivalence claim across species.
M20. Oh et al. (2023), Organ aging signatures in the plasma proteome Primary construction, tissue-enrichment, platform and cognitive-progression comparator passages; no decision-impact trial identified in that material.
M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.
M25. Waziry et al. (2023), CALERIE DNA methylation analysis Relevant primary methods/results, analysis population and null outcomes inspected through publisher-supplied full text delivered on ResearchGate; not a longevity-outcome trial.
T01. BioAge R package README and relevant source files inspected, not executed. DESCRIPTION 0.1.0 declares GPL-3; package paper was identified, not independently read.
SC3. ProteoClock README First-party access description observed 21 September 2026; inspected blob f41e85c59b1ffeec6b8da66fb7029dd4cd77f30a. Exact release-to-trial pin and full licence tree not audited.
SC4. Argentieri et al. (2024), Proteomic aging clock predicts mortality and disease risk Primary training, external validation and covariate passages inspected through publisher-provided text delivered on ResearchGate. No model execution.
SC6. Kuo et al. (2024), Proteomic aging clock (PAC) Primary cohort, selection, Gompertz target and train/test passages; full coefficients and ancillary-trial artifact not verified.
SC7. Goeminne et al. (online 2024; issue 2025), Plasma protein-based organ-specific aging and mortality models Primary abstract and publisher results excerpts, not full coefficient tables or frozen trial implementations.
R01. Higgins-Chen et al. (2022), principal-component clock reliability Abstract and extended-data/replicate descriptions inspected, not full methods. Improved agreement in specified experiments is model-development evidence, not a portable consumer threshold.
R03. McEwen et al. (2018), EPIC/450K clock assessment Relevant methods/results/discussion inspected, including normalization and the post-bisulfite-conversion technical split. No raw arrays rerun; the age-error-based heuristic is not adopted as an individual detection threshold.
R04. Zhuang et al. (2025), EPIC-version differences in methylation tools Relevant methods/results and analysis code inspected, including missing probes and separate/pooled reference adjustment. Code read, not executed; a clock name does not establish complete input coverage or absolute agreement.
R05. Tay et al. (2025), DNAm age differences across EPIC versions and specimens Relevant methods/results/conflicts and Supplementary Table S4 inspected. Same extracted DNA, 16 paired samples per specimen, ages 40–60, 14 Chinese participants; EPICv2 imputation was part of the pipeline. Rounded aggregate reconstruction, not raw-data reproduction. Main-text/S4 and normalization-reference discrepancies remain unresolved. Disclosures include clock-foundation and advisory roles.
Next: Work through intervention endpoints or inspect implementation and access gates.