Aging Clocks Catalog · For reports and retests
Age, pace, uncertainty and conflicting results
A number can change because biology changed, the measurement changed, or a different quantity was reported. Start by separating those possibilities.
Study evidence and teaching examples are labeled separately. The matched-platform comparison and CALERIE control trajectory are published study findings with aggregate calculations. The other worked examples use invented inputs. None supplies a commercial product’s individual-change threshold or a retest interval.
First ask whether the numbers measure the same thing
In a hypothetical comparison, suppose a blood age report is 54, a cheek age report is 48, and a pace report is 0.90. Averaging them would combine potentially different targets, tissues and reference populations with a non-age output. Selecting 48 as the “true” result validates nothing. A pace score of 0.8 is not a statement that someone is 20% younger. [M03; M05]
| Quantity | Arithmetic | Interpretation |
|---|---|---|
| Raw clock score | Predicted age = 55 score-years (hypothetical) | The estimator’s output before subtracting chronological age or a fitted reference prediction. |
| Raw age gap | Predicted 55 − chronological 60 = −5 score-years | Difference from calendar age, without a fitted reference adjustment. |
| Reference residual | Predicted 55 − reference prediction 58 = −3 score-years | Departure from the specified fitted relationship at chronological age 60—not the same quantity as the raw gap. |
| Pace output | A longitudinally trained score, such as DunedinPACE, has its own reference calibration. | It is not automatically obtained by subtracting two age predictions. |
Record whether a report gives the raw score, score minus chronological age, a residual from a named age regression, or a pace. Residuals depend on the reference population and fitted relationship; a model name alone does not identify the quantity. [M04; M05; R04]
Same DNA, different platform: a score can fall
Matched-platform evidence, not a before/after intervention. Tay et al. used the same extracted DNA for EPICv1 and EPICv2 measurements in 16 paired samples per specimen. The screening participants were aged 40–60; 14 were Chinese. The pipeline included EPICv2 missing-value imputation. This comparison measures the combined platform/pipeline difference, not the isolated effect of the array hardware. [R05]
Buffy-coat PCPhenoAge
Same extracted DNA, two EPIC generations. Buffy coat is a white-cell-rich blood fraction. This is the principal-component methylation estimator, not the clinical-laboratory Phenotypic Age formula.
Mean: −1.96 score-years
Reconstructed 95% interval for the mean offset: approximately −2.40 to −1.52 score-years.
A platform/pipeline change can therefore create an apparent reduction without intervening biology. Do not add 1.96 years to other EPICv2 reports. The small cohort, specimen, imputation and laboratory conditions do not support a universal correction or a “two-year error” label. A useful bridge needs paired specimens through both complete pipelines and a held-out check of both bias and remaining disagreement. [R05; R04]
Separate calibration can hide a common offset
Synthetic example: take ages [40, 50, 60, 70], first scores [43, 49, 58, 70], and second scores two units lower for everyone: [41, 47, 56, 68]. Fit a separate score-versus-age regression, with an intercept, to each set.
Raw mean difference: −2
Every second score is two score-years lower. The common offset remains visible in the original units.
Residual mean difference: 0
The second regression absorbs the offset into its intercept. The two sets of residuals are identical by construction.
The identity is simple: if Y₂ = Y₁ + c, separate age regressions absorb c, whether it came from technical bias or a genuine common biological shift. Agreement of the residuals therefore does not validate absolute longitudinal agreement. This is an algebraic illustration, not participant data. Separate and pooled calibration choices are also explicit in the inspected EPIC-comparison analysis code. [R04]
Separate calibration can serve within-cohort association questions; it cannot estimate a cohort-wide shift that it removes. Keep an external or prespecified reference for a longitudinal comparison, preserve the original units, and test any bridge independently. If platform and visit are perfectly confounded without bridge samples, statistical adjustment alone cannot identify their separate effects. [R04; R05]
A score change and a gap change are not the same
| Quantity | First visit | Second visit | Change |
|---|---|---|---|
| Chronological age | 50 | 51 | +1 year |
| Predicted age | 52 | 50 | −2 score-years |
| Raw age gap | +2 | −1 | −3 score-years |
The score fell by 2, while the gap fell by 3 because calendar age also advanced. Neither is automatically a validated pace, a causal treatment effect or years of life gained.
An actual control trajectory: the gap fell while raw age rose
CALERIE study evidence. The methylation analysis included 197 participants, with 69 ad-libitum controls at baseline and visit-specific missingness. Participants were predominantly White, relatively young and healthy. At 24 months, the reported PCGrimAge control age-gap change was −0.26 baseline SD (95% CI −0.37 to −0.14); the baseline scale SD was 2.82 years. The outcome was clock age minus chronological age—not raw clock age. [R09]
| Quantity | Calculation | Meaning |
|---|---|---|
| Control age-gap change | −0.26 × 2.82 = −0.7332 gap-years 95% interval: −1.0434 to −0.3948 | The fitted group gap declined. The normalizer is between-person baseline dispersion, not technical error. |
| Elapsed calendar time | +2 years, assuming a nominal exactly two-year interval | An explicit condition for the arithmetic below, not verification of every participant’s visit date. |
| Implied raw clock-age change | 2 − 0.7332 = +1.2668 score-years | At that nominal interval, raw clock age increased even though the age gap fell. |
The extra decimals show the calculation from rounded inputs, not extra measurement precision. This is neither age reversal nor a universal expected untreated slope. Physiology, calibration and measurement may contribute to the trajectory. The group interval cannot be repurposed as one customer’s error distribution. [R09]
A visible decrease may still be compatible with no change
Assume, only for this hypothetical illustration, independent, unbiased normally distributed error with a standard deviation of 1.5 score-years at each visit, and unchanged conditions. The standard error of the difference is √(1.5² + 1.5²) = 2.121. A rough 95% interval around the observed −2 change is −6.158 to +2.158.
A high intraclass correlation (ICC) does not supply this individual error model. ICC depends on between-person variation; calendar-age mean absolute error describes yet another task. DunedinPACE’s technical and cross-platform results and OMICmAge’s 0.998 ICC from 30 replicates do not establish a universal collection-to-report threshold for a consumer service. [M05; M21]
More generally, Var(e₂ − e₁) = Var(e₂) + Var(e₁) − 2 Cov(e₁, e₂). The familiar √2 multiplier needs equal, independent errors; shared batch error or a version offset changes the problem. The paired-platform SD above must not replace the invented 1.5 in this separate teaching example. Which experiment measures which uncertainty?
Before describing a retest as biological change
- Identify the same quantity.
Match raw score/gap/residual/pace, model version, units, specimen, assay generation, imputation, normalization and reference calibration. A familiar brand name is not enough.
- Check the measurement conditions.
Compare collection timing and relevant preparation, acute illness or medication changes, processing, shipping, batch and laboratory. Check method changes and paired bridging evidence. A duplicate aliquot does not estimate all variation between visits.
- Find uncertainty for this use.
Seek absolute collection-to-report difference variance, bias and interval-specific untreated trajectories in the reported units. A protein ICC, cross-sectional age error or seller interval is not that evidence.
- Separate change from its cause.
A before/after difference can reflect untreated change, concurrent changes or regression to the mean. Selecting an unusually “old” first result can produce a less extreme second result without improvement.
These proposed checks follow the platform, replicate and control designs, not a validated threshold for every test. [R04; R05; R09]
No numerical regression-to-the-mean correction is justified here without the relevant selection and reliability data. A three- or six-month marketing interval does not demonstrate that individual change can be resolved at that interval. The commercial audit marks what is and is not publicly established.
A unit mistake can look like biology
In the fixed clinical Phenotypic Age example, entering CRP in mg/L as if it were mg/dL makes the input ten times too large and shifts the result by 2.43545 score-years, without changing the specimen. Conversion must happen before the natural logarithm; zero/nonpositive CRP cannot be silently logged. The static calculation gives the exact inputs and formula. [M03; T01]
Risk and treatment claims need another step
An age-scaled risk score is not a personal lifespan forecast. Even a hazard ratio needs baseline risk, horizon and model assumptions. With invented baseline risk 5% and HR 1.5, a proportional-hazards illustration gives 1 − 0.951.5 = 7.405%—not a simple addition of percentage points. This is not a validated clock-specific risk calculator.
Likewise, in a hypothetical trial with mean change −2 in treatment and −1 in placebo, the treatment-versus-placebo change contrast is −1. A significant within-treatment change and a nonsignificant control change do not establish a significant difference between groups. Clinical utility and surrogate validity require evidence beyond that contrast. [M26]
Sources and reading limits
Source labels distinguish primary research, seller documents and implementation notes. Inherited M/T source observations are from 21 September 2026; added R source observations are from 22 September 2026. The platform interval and CALERIE back-conversion are aggregate arithmetic, not raw-data or model replication.
M03. Levine et al. (2018), An epigenetic biomarker of aging for lifespan and healthspan Clinical selection, units/coefficient table, methylation stage and validation passages inspected; no raw-data reanalysis.
M04. Lu et al. (2022), DNA methylation GrimAge version 2 Primary version, training and multi-cohort validation passages inspected; no commercial-version equivalence inferred.
M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.
M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.
M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.
T01. BioAge R package README and relevant source files inspected, not executed. DESCRIPTION 0.1.0 declares GPL-3; package paper was identified, not independently read.
R04. Zhuang et al. (2025), EPIC-version differences in methylation tools Relevant methods/results and the analysis code inspected, including separate and pooled calibration; code read, not executed. The synthetic constant-offset example is catalog arithmetic, not this study’s participant result.
R05. Tay et al. (2025), DNAm age differences across EPIC versions and specimens Published 23 April 2025. Relevant methods/results/conflicts and Supplementary Table S4 (PDF page 3) inspected. Reconstruction uses rounded S4 aggregates, not individual predictions or raw methylation; the added mean interval is conditional on the assumptions shown beside it. Main-text/S4 SD discrepancies in other rows and a normalization-reference inconsistency remain unresolved; no source preprocessing was independently reproduced. Disclosures include clock-foundation and advisory roles.
R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods and university-hosted publisher article and supplement, Tables S3/S4 inspected. The control example uses rounded Table S3 group summaries; baseline SD is not technical error. Trial/model-author evidence, not independent consumer-service validation.
Next: See these distinctions in actual studies or inspect the evidence needed for stronger claims.