Shaduf.
Aging Clocks Catalog/Reliability, portability and meaningful change

Aging Clocks Catalog · For reports and retests

Age, pace, uncertainty and conflicting results

A number can change because biology changed, the measurement changed, or a different quantity was reported. Start by separating those possibilities.

Foundation reviewed ; reliability and calibration evidence reviewed . A bounded documentary review, not a diagnostic or treatment service.

Study evidence and teaching examples are labeled separately. The matched-platform comparison and CALERIE control trajectory are published study findings with aggregate calculations. The other worked examples use invented inputs. None supplies a commercial product’s individual-change threshold or a retest interval.

First ask whether the numbers measure the same thing

In a hypothetical comparison, suppose a blood age report is 54, a cheek age report is 48, and a pace report is 0.90. Averaging them would combine potentially different targets, tissues and reference populations with a non-age output. Selecting 48 as the “true” result validates nothing. A pace score of 0.8 is not a statement that someone is 20% younger. [M03; M05]

Raw score, age gap, reference residual and pace
QuantityArithmeticInterpretation
Raw clock scorePredicted age = 55 score-years (hypothetical)The estimator’s output before subtracting chronological age or a fitted reference prediction.
Raw age gapPredicted 55 − chronological 60 = −5 score-yearsDifference from calendar age, without a fitted reference adjustment.
Reference residualPredicted 55 − reference prediction 58 = −3 score-yearsDeparture from the specified fitted relationship at chronological age 60—not the same quantity as the raw gap.
Pace outputA longitudinally trained score, such as DunedinPACE, has its own reference calibration.It is not automatically obtained by subtracting two age predictions.

Record whether a report gives the raw score, score minus chronological age, a residual from a named age regression, or a pace. Residuals depend on the reference population and fitted relationship; a model name alone does not identify the quantity. [M04; M05; R04]

Same DNA, different platform: a score can fall

Matched-platform evidence, not a before/after intervention. Tay et al. used the same extracted DNA for EPICv1 and EPICv2 measurements in 16 paired samples per specimen. The screening participants were aged 40–60; 14 were Chinese. The pipeline included EPICv2 missing-value imputation. This comparison measures the combined platform/pipeline difference, not the isolated effect of the array hardware. [R05]

What was compared

Buffy-coat PCPhenoAge

Same extracted DNA, two EPIC generations. Buffy coat is a white-cell-rich blood fraction. This is the principal-component methylation estimator, not the clinical-laboratory Phenotypic Age formula.

EPICv2 minus EPICv1

Mean: −1.96 score-years

Reconstructed 95% interval for the mean offset: approximately −2.40 to −1.52 score-years.

Tay et al., Supplementary Table S4: 16 pairs; rounded mean −1.96 and paired-difference SD 0.82 score-years. The catalog’s interval is −1.96 ± t(0.975, 15) × 0.82 / √16. It assumes independent participants and approximately normal, homogeneous paired differences; those assumptions cannot be tested from rounded aggregates. This is a pointwise mean-offset interval, not an individual uncertainty band or same-platform intervisit error estimate. [R05]

A platform/pipeline change can therefore create an apparent reduction without intervening biology. Do not add 1.96 years to other EPICv2 reports. The small cohort, specimen, imputation and laboratory conditions do not support a universal correction or a “two-year error” label. A useful bridge needs paired specimens through both complete pipelines and a held-out check of both bias and remaining disagreement. [R05; R04]

Separate calibration can hide a common offset

Synthetic example: take ages [40, 50, 60, 70], first scores [43, 49, 58, 70], and second scores two units lower for everyone: [41, 47, 56, 68]. Fit a separate score-versus-age regression, with an intercept, to each set.

Before separate calibration

Raw mean difference: −2

Every second score is two score-years lower. The common offset remains visible in the original units.

After separate calibration

Residual mean difference: 0

The second regression absorbs the offset into its intercept. The two sets of residuals are identical by construction.

The identity is simple: if Y₂ = Y₁ + c, separate age regressions absorb c, whether it came from technical bias or a genuine common biological shift. Agreement of the residuals therefore does not validate absolute longitudinal agreement. This is an algebraic illustration, not participant data. Separate and pooled calibration choices are also explicit in the inspected EPIC-comparison analysis code. [R04]

Separate calibration can serve within-cohort association questions; it cannot estimate a cohort-wide shift that it removes. Keep an external or prespecified reference for a longitudinal comparison, preserve the original units, and test any bridge independently. If platform and visit are perfectly confounded without bridge samples, statistical adjustment alone cannot identify their separate effects. [R04; R05]

A score change and a gap change are not the same

One hypothetical person, two visits
QuantityFirst visitSecond visitChange
Chronological age5051+1 year
Predicted age5250−2 score-years
Raw age gap+2−1−3 score-years

The score fell by 2, while the gap fell by 3 because calendar age also advanced. Neither is automatically a validated pace, a causal treatment effect or years of life gained.

An actual control trajectory: the gap fell while raw age rose

CALERIE study evidence. The methylation analysis included 197 participants, with 69 ad-libitum controls at baseline and visit-specific missingness. Participants were predominantly White, relatively young and healthy. At 24 months, the reported PCGrimAge control age-gap change was −0.26 baseline SD (95% CI −0.37 to −0.14); the baseline scale SD was 2.82 years. The outcome was clock age minus chronological age—not raw clock age. [R09]

Back-conversion of rounded group summaries, not an individual error model
QuantityCalculationMeaning
Control age-gap change−0.26 × 2.82 = −0.7332 gap-years
95% interval: −1.0434 to −0.3948
The fitted group gap declined. The normalizer is between-person baseline dispersion, not technical error.
Elapsed calendar time+2 years, assuming a nominal exactly two-year intervalAn explicit condition for the arithmetic below, not verification of every participant’s visit date.
Implied raw clock-age change2 − 0.7332 = +1.2668 score-yearsAt that nominal interval, raw clock age increased even though the age gap fell.

The extra decimals show the calculation from rounded inputs, not extra measurement precision. This is neither age reversal nor a universal expected untreated slope. Physiology, calibration and measurement may contribute to the trajectory. The group interval cannot be repurposed as one customer’s error distribution. [R09]

A visible decrease may still be compatible with no change

Assume, only for this hypothetical illustration, independent, unbiased normally distributed error with a standard deviation of 1.5 score-years at each visit, and unchanged conditions. The standard error of the difference is √(1.5² + 1.5²) = 2.121. A rough 95% interval around the observed −2 change is −6.158 to +2.158.

A two-score-year decrease with uncertainty crossing zero The observed change is minus two score-years. Under hypothetical independent normal errors with standard deviation 1.5 at each visit, the approximate 95 percent interval is minus 6.158 to plus 2.158. It includes zero and is not a threshold for a real test. Observed −2 −6.1580: no change+2.158 Lower scoreHigher score
Hypothetical arithmetic, not product evidence. The interval crosses zero. Correlated errors, batch shifts, model changes and biological variation need a different uncertainty model.

A high intraclass correlation (ICC) does not supply this individual error model. ICC depends on between-person variation; calendar-age mean absolute error describes yet another task. DunedinPACE’s technical and cross-platform results and OMICmAge’s 0.998 ICC from 30 replicates do not establish a universal collection-to-report threshold for a consumer service. [M05; M21]

More generally, Var(e₂ − e₁) = Var(e₂) + Var(e₁) − 2 Cov(e₁, e₂). The familiar √2 multiplier needs equal, independent errors; shared batch error or a version offset changes the problem. The paired-platform SD above must not replace the invented 1.5 in this separate teaching example. Which experiment measures which uncertainty?

Before describing a retest as biological change

  1. Identify the same quantity.

    Match raw score/gap/residual/pace, model version, units, specimen, assay generation, imputation, normalization and reference calibration. A familiar brand name is not enough.

  2. Check the measurement conditions.

    Compare collection timing and relevant preparation, acute illness or medication changes, processing, shipping, batch and laboratory. Check method changes and paired bridging evidence. A duplicate aliquot does not estimate all variation between visits.

  3. Find uncertainty for this use.

    Seek absolute collection-to-report difference variance, bias and interval-specific untreated trajectories in the reported units. A protein ICC, cross-sectional age error or seller interval is not that evidence.

  4. Separate change from its cause.

    A before/after difference can reflect untreated change, concurrent changes or regression to the mean. Selecting an unusually “old” first result can produce a less extreme second result without improvement.

These proposed checks follow the platform, replicate and control designs, not a validated threshold for every test. [R04; R05; R09]

No numerical regression-to-the-mean correction is justified here without the relevant selection and reliability data. A three- or six-month marketing interval does not demonstrate that individual change can be resolved at that interval. The commercial audit marks what is and is not publicly established.

A unit mistake can look like biology

In the fixed clinical Phenotypic Age example, entering CRP in mg/L as if it were mg/dL makes the input ten times too large and shifts the result by 2.43545 score-years, without changing the specimen. Conversion must happen before the natural logarithm; zero/nonpositive CRP cannot be silently logged. The static calculation gives the exact inputs and formula. [M03; T01]

Risk and treatment claims need another step

An age-scaled risk score is not a personal lifespan forecast. Even a hazard ratio needs baseline risk, horizon and model assumptions. With invented baseline risk 5% and HR 1.5, a proportional-hazards illustration gives 1 − 0.951.5 = 7.405%—not a simple addition of percentage points. This is not a validated clock-specific risk calculator.

Likewise, in a hypothetical trial with mean change −2 in treatment and −1 in placebo, the treatment-versus-placebo change contrast is −1. A significant within-treatment change and a nonsignificant control change do not establish a significant difference between groups. Clinical utility and surrogate validity require evidence beyond that contrast. [M26]

Sources and reading limits

Source labels distinguish primary research, seller documents and implementation notes. Inherited M/T source observations are from 21 September 2026; added R source observations are from 22 September 2026. The platform interval and CALERIE back-conversion are aggregate arithmetic, not raw-data or model replication.

M03. Levine et al. (2018), An epigenetic biomarker of aging for lifespan and healthspan Clinical selection, units/coefficient table, methylation stage and validation passages inspected; no raw-data reanalysis.

M04. Lu et al. (2022), DNA methylation GrimAge version 2 Primary version, training and multi-cohort validation passages inspected; no commercial-version equivalence inferred.

M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.

M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.

M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.

T01. BioAge R package README and relevant source files inspected, not executed. DESCRIPTION 0.1.0 declares GPL-3; package paper was identified, not independently read.

R04. Zhuang et al. (2025), EPIC-version differences in methylation tools Relevant methods/results and the analysis code inspected, including separate and pooled calibration; code read, not executed. The synthetic constant-offset example is catalog arithmetic, not this study’s participant result.

R05. Tay et al. (2025), DNAm age differences across EPIC versions and specimens Published 23 April 2025. Relevant methods/results/conflicts and Supplementary Table S4 (PDF page 3) inspected. Reconstruction uses rounded S4 aggregates, not individual predictions or raw methylation; the added mean interval is conditional on the assumptions shown beside it. Main-text/S4 SD discrepancies in other rows and a normalization-reference inconsistency remain unresolved; no source preprocessing was independently reproduced. Disclosures include clock-foundation and advisory roles.

R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods and university-hosted publisher article and supplement, Tables S3/S4 inspected. The control example uses rounded Table S3 group summaries; baseline SD is not technical error. Trial/model-author evidence, not independent consumer-service validation.

Next: See these distinctions in actual studies or inspect the evidence needed for stronger claims.

Search published pools, pages, reports, and evidence.