Shaduf.Research preview
Aging Clocks Catalog/Computational identity, expected outputs and frozen references

Aging Clocks Catalog · How claims are assessed

How the evidence is judged

Reliable measurement, prediction, intervention response and better health decisions are different achievements. No single statistic establishes them all.

Foundation reviewed ; reliability and calibration evidence reviewed ; intervention evidence reviewed . A bounded documentary review, not a diagnostic or treatment service.

The important question is not whether a clock is “validated” in the abstract. It is validated for which claim, in which setting, with which implementation? Consistent measurement, external prognosis, treatment response, clinical utility and surrogate validity are distinct. A research biomarker can be useful without meeting the last two standards. [M03; M05; M26]

A claim needs the right kind of evidence

Separate standards—not a ladder of marketing badges
ClaimEvidence that addresses itEvidence that is not enough
“The result is reproducible”Matched collection-to-report replicates, absolute error and relevant batch/platform conditions.Calendar-age correlation, another lab’s ICC or triplicate measurement of just one assay stage.
“The result predicts future outcomes”Prospective external evaluation, suitable covariates, calibration and performance uncertainty.Development fit, cross-sectional disease association or selecting a model on the outcomes used to evaluate it.
“The intervention changed the result”A prespecified controlled contrast, reliable longitudinal measurement, missingness strategy and multiplicity handling.A before/after change, favorable subgroup selection or separate significance tests in each arm.
“Changing the result improves health”Causal evidence connecting intervention, score change and relevant health benefit, considering alternative pathways.Baseline prognosis or pathway enrichment from the same assay used to calculate the score.
“Using the test is clinically useful”A validated action/threshold and benefits, harms and costs relative to a realistic decision without the test.Statistical significance, a high area-under-curve statistic or bundled advice that could be given without testing.
“The result is a surrogate endpoint”Context-specific evidence that treatment effects on the biomarker predict treatment effects on a clinical endpoint.Baseline or individual-level association, response to one intervention, or a single-trial mediation result; none alone validates prediction of clinical treatment effects.

Assessment standards informed by the inspected methods, controlled studies and FDA–NIH BEST definitions; not a claim that every paper was designed to satisfy every row. [M03; M05; M11; M20; M25; M26; SC1]

Prediction questionDoes the baseline score help forecast an outcome?

A conditional association or discrimination result can address prognosis. It does not show that lowering the score prevents that outcome.

Treatment questionDoes a treatment’s marker effect predict its clinical effect?

Surrogate validity concerns effects in a defined context. The two questions require different evidence; one cannot be substituted for the other.

Prognosis and surrogacy are not synonyms. Observational association alone is generally insufficient for validated-surrogate status. [M26]

Five different meanings of “the calculation checks out”

Computational-identity evidence reviewed strengthens the source and software checks behind the clinical example. It does not supply an independently established primary-method expected output or clinical validation. Keep these five claims separate when accepting a tool or interpreting a report.

What was checked—and what that check does not establish
Evidence layerConcrete evidence in this reviewBoundary
Source-method identityThe 2018 equation and restored BioAge source support 0.090165 for the named example; the later correction prints a conflicting denominator. [D01; D02; D03]Source history supports a declared mapping, not the exact execution artifact behind every historical study.
Internal regression checksInherited synthetic expectations, stable/nested mathematics and a separately written high-precision arithmetic path agree for the declared cases.Expectations derived from the same assumptions can catch a code regression without independently validating those assumptions.
Implementation parityThe verified 1,405-byte Biolearn function reproduces the three earlier software-variant outputs when executed in isolation. [D05]Agreement with this library function confirms its behavior—not that it is the authority for a disputed original method. No complete package or empirical workflow was run.
Independent expected outputNot recovered in a qualifying form. Inspected primary equations, tests, helper examples and export routes did not yield a complete applicable independent clinical input/output pair. [D01; D08]A second arithmetic path, a function’s own output or a DNAmPhenoAge test cannot be relabeled an independent reference for the clinical-laboratory calculation.
Clinical validationPopulation, assay, outcome, intervention and decision evidence must be evaluated under the claim-specific standards above.Deterministic agreement does not measure collection/intervisit error, establish patient benefit or validate a surrogate. A correctly calculated research score can still answer the wrong clinical question. [M26]

The exact unmet gate: a complete permitted input with its units/transforms, source and reference revision, expected output and precision, plus a provenance chain independent of the implementation being accepted. Original-author certification would be strong evidence, but a separately established primary-method reference computation can also qualify. A lone score, molecular coefficient list, shared spreadsheet derivation or self-generated test expectation is insufficient. This bounded search did not show that such evidence is absent everywhere. [D01; D08]

The named formula comparisons and existing KDM/HD export designs are useful progress. Neither an export recipe nor source-level parity closes the independent-output gate. Static explanations remain supported; a preset feature still needs separate reference, source-choice and rights review.

A software discrepancy is not automatically a published-study error. Establishing that stronger claim would require the study’s actual revision, preprocessing, reference and execution chain. The conflicting 2019 equation was read from publisher-extracted PDF text but could not be visually verified after rendering failed; the 2018 equations were visually checked. The maintainer’s attributed author confirmation is not independently authenticated correspondence. Those limits narrow the claim without erasing the documented code history. [D01; D02; D03]

Three separate questions before interpreting a change

  1. Does it exceed the relevant error?

    First match the output and complete pipeline. Identical numerical inputs test software reproducibility; a split specimen tests only steps after the split; separate collections include collection and intervening physiology. Paired version bridges estimate platform/pipeline disagreement. State bias and absolute difference variance for the relevant experiment rather than substitute calendar-age error or an ICC. [R03; R05]

  2. Does it reflect the intended biological construct?

    A real inflammatory or cell-mixture change is not automatically sustained aging. Examine standardized repeated visits, untreated trajectories and independent functional or clinical measures. The inspected timing study’s dense series was one person and its additional paired data came from a stress experiment; it supports attention to context, not a universal daily age fluctuation. [R08; R09; R11]

  3. Does using it improve a decision?

    Identify the action triggered by the score and compare benefits, harms and costs with a realistic no-test decision. Even a resolved, biologically relevant change does not establish clinical utility. Surrogacy separately requires treatment effects on the marker to predict clinical effects in a defined context. [M26]

No single reliability coefficient answers all three. Under a stable, approximately normal, unbiased error process, 1.96 × SD(error difference) is a statistical repeatability bound—not a clinically important difference. In general, Var(e₂ − e₁) = Var(e₂) + Var(e₁) − 2 Cov(e₁, e₂); the equal-independent-error shortcut is a special case. This is a measurement model, not an estimated threshold for any product. The matched-platform example estimates a mean version offset, not same-platform intervisit variation.

A group contrast is not an individual response classifier

For a hypothetical balanced trial with independent participants, n per arm and individual-change measurement-error SD sD, the measurement-error component of the difference-in-mean-changes has standard error sD × √(2/n). This excludes biological heterogeneity, clustering, missingness and systematic batch bias. More participants can reduce random error in a group contrast; they do not remove an arm-confounded shift. A statistically detectable group effect can therefore be smaller than an individual repeatability bound without contradiction.

The CALERIE control back-conversion illustrates why output definition also matters: its PCGrimAge summary is a group age-gap change standardized by baseline dispersion, not individual measurement error or raw-age change. [R09]

Noise does not always make an effect a lower bound. Under additive, mean-zero outcome error independent of treatment, an unstandardized randomized contrast has unchanged expectation and greater variance. Standardized effects may shrink when noise inflates the scale SD; systematic or treatment-dependent error can behave differently. This is reasoning under a specified measurement model, not an estimate of CALERIE’s error process or evidence of “at least this much” aging improvement. [S08]

The missing uncertainty for clinical and proteomic scores

Repeating a deterministic formula on identical values cannot estimate the covariance below. The source and parity checks establish computational properties; collection, assay and intervisit error require measurements under the relevant protocol.

Clinical-laboratory composites: for a fixed linear score, input-error contribution is w′Σw: weights combined with the covariance matrix of input errors in the correct transformed units. The inspected clinical Phenotypic Age formula is linear in those transformed inputs, including log CRP, when calendar age is treated as known. A serial-use interval additionally requires collection/physiological and cross-visit covariance for the proposed timing. Those matrices were not established for the inspected services. Adding or averaging analyte CVs cannot recover them. [M03; R15]

Proteomic clocks: a linear score likewise needs weighted covariance; for a nonlinear clock, full-pipeline replicated predictions under the actual specimen, panel, preprocessing and model version are the more direct empirical check. Average protein ICCs or a count of reproducible proteins do not supply final-score error. A small preanalytical experiment found altered inflammatory-protein measurements despite analytical controls, but calculated no aging clock; it cannot establish error in ProtAge, PAC or rentosertib’s six outputs. [R10; R11]

Clinical-laboratory and proteomic individual-change thresholds therefore remain unresolved for the specific services and pipelines inspected. Neither methylation-platform variance nor external risk association fills the missing score-level evidence.

Numbers only mean something with their comparator

Mean absolute error for calendar age does not measure longitudinal reliability or outcome prediction. A hazard ratio describes conditional association under a fitted model. A C-statistic measures ranking; it does not establish that a predicted 10% risk occurs 10% of the time. That last question concerns calibration. Clinical utility additionally asks what happens when a decision changes because of the result.

The NMR example reports five-year C-statistics of 0.837 versus 0.772, and ten-year values of 0.830 versus 0.790, in FINRISK. It has a real conventional-risk comparison, but cohort-based scaling limits immediate individual classification and the models were not simply the same base with one extra predictor. The worked application preserves that boundary. [M11]

CALERIE: what a rounded p-value can and cannot settle

The favorable DunedinPACE group response remains beside the PCPhenoAge and PCGrimAge null contrasts. The catalog’s sensitivity audit transcribed 40 published summaries: 22 intention-to-treat (ITT) tests from Table S4, six primary-model treatment-on-treated tests, and 12 cell-adjusted ITT/TOT tests from S6. These are comparison rows, not people or independent replications. [S08]

Two declared sensitivity families: six ITT comparisons (three primary models × two visits), and all 22 S4 ITT comparisons (11 implementations × two visits). Bonferroni and Holm were applied at familywise alpha 0.05. These are the catalog’s sensitivity choices—not the paper’s original p < 0.005 convention, proof of a prespecified family, or coverage of every exploratory, subgroup and mediation test. Both corrections can address dependent tests; neither repairs invalid input p-values, selection or model misspecification.

DunedinPACE adjusted p-values: bounded arithmetic, not refitted trial models
ITT family / correction12 months: printed-value result24 months: printed-value result24 months: rounding envelope*
6 primary tests / Bonferroni0.0028980.0120.0090–0.0150
6 primary tests / Holm0.0028980.0100.0075–0.0125
22 S4 ITT tests / Bonferroni0.0106260.0440.0330–0.0550
22 S4 ITT tests / Holm0.0106260.0420.0315–0.0525

*Assuming nearest rounding to the last displayed digit, printed 24-month p = 0.002 has the conservative envelope [0.0015, 0.0025]. The 12-month p = 4.83E-04 retains its displayed precision, half-width 0.0000005. The envelopes propagate printed precision through the corrections; they are not confidence intervals. A different printing rule would require a different envelope. Published pointwise effect intervals were not adjusted. [S08]

What survives these sensitivities

The earlier response and the six-test result

The 12-month DunedinPACE contrast stays below 0.05 in both families, including its rounding envelope. The 24-month contrast also stays below 0.05 throughout the six-primary-test sensitivity.

What needs more precision

The expanded-family result at 24 months

Central printed values pass, but both 22-test envelopes cross 0.05. The unrounded p-value is needed to decide this classification; the record does not support a definite pass or definite failure.

These results do not negate the group response or identify the best clock. The other S4 ITT models do not become significant through these corrections. Participant-level mixed models, instrumental-variable estimates, mediation and methylation predictions were not refitted. Correlated outcomes and two visits are not independent clinical replications. The application also retains the less favorable cell-adjusted TOT result, rather than claiming unchanged robustness under every sensitivity. [S08]

How the adjustment was checked

For m tests, Bonferroni uses min(1, m × p). Holm orders p-values, multiplies by the number of tests remaining, takes the cumulative maximum, caps at one and restores the original order. Lower and upper printed-precision bounds were propagated using exact decimal arithmetic. The bounded audit checked input identities, source scales, arithmetic examples, row-order invariance and deterministic reruns; it did not reconstruct unavailable unrounded observations. [S08]

Mediation is not the same as clinical-surrogate validation

CALERIE did analyze mediation. Its models connected DunedinPACE change with changes in clinical and blood-chemistry outcomes. Figure S4 describes small mediated fractions. Many intervals for the proportion mediated excluded zero; intervals for the average causal mediation effect (ACME) included zero except for log CRP and clinical Phenotypic Age. The latter is the clinical-laboratory composite, not PC DNAmPhenoAge. No precise indirect effects were digitized or re-estimated by the catalog. [S08]

Within one trial

An indirect-effect model

ACME concerns the modeled indirect effect; proportion mediated expresses a share of the total effect. Their estimates and uncertainty are different. Randomizing treatment does not randomize the mediator.

Across appropriate trials

Prediction of clinical treatment effects

A surrogate must support prediction of treatment effects on a defined clinical endpoint in a specified use. Baseline prognosis, individual association or a single mediation result cannot substitute for that evidence.

Causal mediation additionally needs assumptions about mediator–outcome confounding, treatment-induced confounders, model specification and measurement. The inspected code uses pace change and outcome change over the same baseline-to-24-month interval; it does not establish temporal precedence. A risk-marker mediation model in one trial does not establish fewer deaths, longer life or trial-level clinical surrogacy. [S08; S09; S11]

A marker need not itself cause benefit to be predictively useful. Conversely, evidence for one causal pathway does not establish the net clinical effects of every treatment that moves the marker. Statistical mediation is neither a universal prerequisite nor sufficient evidence for every surrogate use. A relevant validation should include independent interventions and discordant outcomes, not only favorable examples. [S10; S11]

Current-code questions are not proven publication errors

The inspected mediation script contains a MAP expression using a diastolic/systolic pressure ratio and overwrites a previously constructed b_ variable using the current outcome inside its loop. These are source-code reconciliation questions if run literally. Without the final data, execution log and evidence that this exact revision generated Figure S4, they are not verified corrections to published numerical results and do not invalidate the separately fitted primary clock contrasts. A source-pinned reproduction is needed; no Stata model was run here. [S09]

The rentosertib distinction: a low clock–FVC R-squared is not mediation and does not isolate an independent aging mechanism. The reported 100,000 patient-label randomizations are meaningful evidence against the specified exchangeable-label null, but do not randomize mechanisms, undo selection or validate a clinical surrogate. Benjamini–Hochberg is not automatically invalid because clocks correlate; its guarantees depend on the dependence conditions and valid input p-values. Actual numerical rows and the complete permutation implementation remain unverified. [SC1; S05; S10]

What a favorable summary must not hide

Negative findings, contradictions and limits retained in this catalog
EvidenceWhat remains visibleConsequence
CALERIEDunedinPACE group response beside PC-clock null contrasts, a rounding-dependent expanded-family 24-month boundary, and mediation without surrogate validation. [S08; S09]Do not claim uniform response, an individual threshold, a robust pass/failure from a rounded boundary, or randomized mortality benefit.
Retinal age gapAll-cause mortality association, but no significant cardiovascular- or cancer-mortality association in the inspected abstract. [M14]Outcome-specific nulls narrow the claim; no diagnosis or screening benefit is established.
2025 clock benchmarkAbstract: 39 biomarkers, more than 20,000 people, little relationship between age accuracy and mortality-prediction capacity. Full methods unavailable. [M06]No detailed ranking or universal winner is adopted.
Six-clock IPF analysisFlow and arms recovered (71 → 55 → 43 → 42; final 11/11/11/9); protocols 003 and 004 distinguished; public parent protocol inspected. Author-reported permutation strengthens the positive record. [SC1; SC2; S03; S04; S06; S07]Selection, citation/history and final-plan chronology, 21/22-site and architecture discrepancies, actual S2–S5 rows and trial artifacts remain unresolved. Do not call the final analysis prospectively locked or independently reproduced.
Public implementationProteoClock restrictions on galkin_2025 persist; the PAOPAC author README is now inspectable. [SC3; S14]Documentation or a binary pointer does not establish a verified trial-matched implementation or permission for public use.
Commercial evidenceModel publications and external cohorts are not independent validation of every current service version. [M21; M22; M23; M24]Keep product-to-paper equivalence, individual error and benefit of the purchased package as separate unresolved claims.

What this review did

This is a purposive, claim-driven documentary review with foundation evidence/access observations dated 21 September 2026, a reliability/calibration follow-up dated 22 September 2026, and intervention evidence dated 26 September 2026. It follows selected original methods and validations; separates parent and ancillary trials; inspects public product/legal documents; and reads code where it changes feasibility. It is not a systematic review, market census, pooled meta-analysis or product certification. Search visibility was not treated as popularity or market share.

“Inspected” means the stated relevant text was read—not every citation, figure or supplement. Some emerging-family entries and the 2025 benchmark rely only on abstracts/metadata. The foundation used publisher-provided ProtAge and CALERIE text delivered on ResearchGate; unrelated citing-paper snippets were excluded. The follow-up also inspected CALERIE’s university-hosted publisher supplement and Tay’s Supplementary Table S4 visually. Source access depth remains specific to each entry, not an assertion that all supplements were read.

All study numbers are author-reported unless explicitly identified as arithmetic. No participant-level dataset was obtained, clock package run, restricted model acquired, product purchased or seller contacted. The follow-up reconstructed agreement limits for all 18 selected specimen/model rows in Tay’s rounded S4 aggregates, keeping age and pace units separate, and checked consistency with reported rounding. It added conditional normal-theory intervals, back-converted rounded CALERIE group summaries and checked fixed synthetic clinical/calibration examples. It did not reproduce raw-pair nonparametric tests, methylation processing, imputation or clock fitting. [R05; R09; R15]

The intervention follow-up inspected official registry mirrors, the public parent protocol and embedded statistical section, a reporting form, peer-review exchanges, CALERIE numerical supplements and author code. It recovered cohort/arm flow and executed multiplicity/rounding sensitivities on 40 published summaries. It did not rerun rentosertib S2–S5 tests or patient-label randomizations, or CALERIE mixed, instrumental-variable or mediation models. Documentary verification, aggregate arithmetic and recomputing molecular predictions remain three different levels of work. [S03; S04; S05; S06; S07; S08; S09]

The computational-identity follow-up, observed on , inspected primary equations, revision history, helper/output paths, tests and existing export designs. It executed one byte-pinned Biolearn function in isolation and checked declared mathematics on three inherited synthetic cases. No whole clock package, reference fit, export notebook or empirical workflow was run. The independent expected-output gate remains open. [D01; D03; D05; D08]

How to read the local source labels

M identifies scientific methods/validation or the official terminology framework; SC identifies the six-clock case, parent trial and model/access sources; C identifies seller/contract documents; T identifies implementation documentation; R labels identify the reliability/calibration evidence review; S labels identify the intervention follow-up. Existing SC1/SC2 notes record their deeper follow-up inspection. They are source identifiers, not quality scores. A local source note tells you the inspected depth and relevant conflicts. A source’s availability does not give another source’s claims its authority.

Original papers support scientific claims; sellers describe their own offerings and terms; current code documentation describes access, not what an older trial necessarily ran. A “latest” documentation address is mutable. The date attached to the review is not a new publication date, a purchase guarantee or a claim that every underlying experiment was recent.

The paired-platform reconstruction assumes independent participants and approximately normal, homogeneous differences within each specimen/model. Rounded aggregates cannot test those assumptions, outliers, proportional bias or subgroup behavior. The intervals are pointwise rather than simultaneous over all rows. S4 and main-text SDs disagree in some rows, and a normalization name/reference inconsistency remains; S4 was used without inventing unrounded data or silently resolving preprocessing. [R05]

What would materially strengthen the conclusions?

For an individual retest: independent, matched collection-to-report repeatability, untreated longitudinal variation, exact report-version correspondence and a demonstrated decision advantage. An unrelated age-error statistic would not resolve the gap.

For the IPF case: a reconciled registry and protocol-revision history, the final approved analysis plan and laboratory/plate records, actual numerical S2–S5 rows, trial-specific artifacts and change-score covariance. The public protocol and ancillary arm counts are no longer missing. Reproducing statistical tests from predictions would still not equal recomputing those predictions from raw proteins.

For CALERIE: unrounded p-values for the expanded-family 24-month boundary and a source-pinned reproduction of the relevant models, including mediation. Neither its group intervals nor statistical mediation supplies individual-change uncertainty or clinical-surrogate validation.

For clinical risk use: external calibration and threshold/net-benefit evaluation against a realistic current comparator, followed by evidence on the decisions and outcomes produced by using the score. One more adjusted association would not by itself settle utility.

Reader safeguard: uncertainty is part of the conclusion, not a footnote to remove. The catalog remains useful by narrowing a claim to what was actually tested, not by assuming the unavailable evidence would be favorable.

Sources and reading limits

Source labels distinguish primary research, seller documents and implementation notes. Foundation observations are from 21 September 2026; R reliability/calibration observations are from 22 September. The S intervention sources and the updated SC1/SC2 inspections are dated 26 September 2026. Author reports, documentary verification and the catalog’s aggregate calculations are distinguished; none is raw-data or model replication.

M03. Levine et al. (2018), An epigenetic biomarker of aging for lifespan and healthspan Clinical selection, units/coefficient table, methylation stage and validation passages inspected; no raw-data reanalysis.

M05. Belsky et al. (2022), DunedinPACE Primary longitudinal target, normalization, technical/cross-platform reliability and relevant validation text inspected; model not executed.

M06. Ying et al. (2025), A unified framework for systematic curation and evaluation of aging biomarkers Author-institution abstract/metadata only; detailed methods and rankings not adopted.

M11. Deelen et al. (2019), A metabolic profile of all-cause mortality risk Primary results, conventional comparator, FINRISK evaluation and scaling limitation inspected; reported models not rerun.

M14. Zhu et al. (2022), Retinal age gap as a predictive biomarker for mortality risk Complete primary abstract, including outcome-specific nulls; full methods and deployment validation not inspected.

M20. Oh et al. (2023), Organ aging signatures in the plasma proteome Primary construction, tissue-enrichment, platform and cognitive-progression comparator passages; no decision-impact trial identified in that material.

M21. Chen et al. (2026), OMICmAge Published 25 February 2026. Primary development, external validation, replicate, limitation, access and conflict sections; sponsored/company-affiliated work with patent interests.

M22. Sehgal et al. (2025), Systems Age Version of record: 15 September 2025. Abstract/metadata and primary update linkage inspected; detailed rankings and current product equivalence withheld.

M23. Shokhirev et al. (2024), CheekAge: a next-generation buccal epigenetic aging clock Primary abstract, development/replicate description and conflicts; Tally-funded, with company-employee authors. No exact consumer-version match.

M24. Shokhirev et al. (2024), CheekAge is predictive of mortality in human blood Primary incomplete-feature blood adaptation and mortality-model sections; company-affiliated external-cohort study, not a prospective consumer cheek-test trial.

M25. Waziry et al. (2023), CALERIE DNA methylation analysis Relevant primary methods/results, analysis population and null outcomes inspected through publisher-supplied full text delivered on ResearchGate; not a longevity-outcome trial.

M26. FDA–NIH BEST: Validated Surrogate Endpoint Official definitions and evidentiary discussion; used as a framework, not a regulatory-status verdict for any aging test.

SC1. Zhavoronkov et al. (2026), Proteomic clocks in a phase 2a trial Published 7 September 2026. Follow-up inspection on 26 September covered relevant full results/methods, Figure 3, extended-data captions, availability and supplement descriptions, including the 100,000-randomization caption. Actual S2–S5 spreadsheet rows and permutation were not independently reproduced. Developer-led; Insilico and model-author interests disclosed.

SC2. Xu et al. (2025), A generative AI-discovered TNIK inhibitor for IPF: randomized phase 2a trial Published 3 June 2025. Follow-up inspection on 26 September covered disposition, FVC/safety, statistical methods and registration. Public parent-protocol evidence is now available (S04); the final analysis-plan chronology and raw clinical analysis remain unreproduced. Sponsor/developer interests apply.

SC3. ProteoClock README First-party access description observed 21 September 2026; inspected blob f41e85c59b1ffeec6b8da66fb7029dd4cd77f30a. Exact release-to-trial pin and full licence tree not audited.

SC4. Argentieri et al. (2024), Proteomic aging clock predicts mortality and disease risk Primary training, external validation and covariate passages inspected through publisher-provided text delivered on ResearchGate. No model execution.

R03. McEwen et al. (2018), EPIC/450K clock assessment Relevant methods/results/discussion inspected, including normalization and the post-bisulfite-conversion technical split. No raw-array reproduction; its age-error-based heuristic is not adopted as an individual detection threshold.

R05. Tay et al. (2025), DNAm age differences across EPIC versions and specimens Relevant methods/results/conflicts and Supplementary Table S4 inspected. Same extracted DNA, 16 paired samples per specimen, ages 40–60, 14 Chinese participants; EPICv2 imputation was part of the pipeline. Rounded aggregate reconstruction, not raw-data reproduction. Main-text/S4 and normalization-reference discrepancies remain unresolved. Disclosures include clock-foundation and advisory roles.

R08. Koncevičius et al. (2024), circadian variation in epigenetic age Relevant design/results and limits inspected, including the one-person dense series and stress-experiment origin of the separate paired dataset. No universal daily amplitude or personal threshold inferred.

R09. Waziry et al. (2023), CALERIE methylation-clock analysis Relevant methods, sample handling and university-hosted publisher article and supplement, Tables S3/S4 inspected on 22 September 2026. Rounded group summaries were back-converted, not individual-level models rerun. Baseline dispersion is not technical error; trial/model-author evidence is not independent consumer-service validation.

R10. Haslam et al. (2022), plasma-proteomics reproducibility Primary abstract and disclosures inspected; full methods not recovered. Assay-profile findings, not an executed clock-level covariance or repeatability analysis.

R11. Huang et al. (2021), preanalytical variability of inflammatory-protein measurement Relevant delay, donor, control and results sections inspected. Small EDTA-plasma, panel-specific study with Olink affiliation; no aging-clock predictions calculated and no finding about a particular IPF trial’s processing.

R15. BioAge calculation sources: phenoage_calc.R, kdm_calc.R and hd_calc.R. Complete files reread on 22 September 2026; the inspected blobs are recorded on the tools page. Fixed synthetic arithmetic was checked separately; the R package was not run and no independent author-supplied reference vector or service-specific covariance was obtained.

S03. Ancillary Reporting Summary Four-page publisher form, dated July 2026; sample exclusion, downstream analysis blinding, software versions and public protocol link inspected. A retrospective reporting form, not an advance analysis plan.

S04. Public parent protocol INS018-055-003 Relevant revision history, sampling and embedded statistical sections inspected in the 110-page registry-hosted file. Mixed February/August 2024 version headers and the planned week-8 sample remain unresolved. Not a complete historical chain or an independently authenticated final statistical analysis plan.

S05. Ancillary transparent peer-review exchange Relevant author/reviewer discussions of selection, adverse events, permutation, removed overlap analysis, sampling and disease-versus-aging interpretation inspected. Author explanations are not independent replication; removed analyses are not final evidence.

S06. WHO ICTRP representation of NCT05938920 Displayed record inspected; refreshed 5 January 2026, protocol INS018-055-003. Selected registry fields, not a complete change history or verification of current recruitment status.

S07. WHO ICTRP representation of NCT05975983 Displayed record inspected; refreshed 24 November 2025, protocol INS018-055-004. A different trial, not an alternative registration for the 71-person parent cohort.

S08. Waziry et al. (2023), CALERIE article and supplement University-hosted publisher text; relevant methods, Tables S4/S6 and Figure S4 inspected on 26 September 2026. Forty rounded comparison summaries support the catalog’s bounded arithmetic, not participant-level model or mediation replication. Trial/model-author interests do not establish independent consumer-service validation.

S09. CALERIE author code Inspected tree db56e58a7546934f3d9f2c5764cc8ddaa5894a9c. Data and mediation files read completely; relevant primary-analysis sections read. No Stata model executed. Current-code concerns do not establish which revision produced the published figure or prove an error in its numerical results.

S10. FDA surrogate endpoint resources Official context-of-use framework inspected on 26 September 2026. Evidence requirements, not a comprehensive regulatory-status determination for named aging clocks.

S11. Imai, Keele and Tingley (2010), A General Approach to Causal Mediation Analysis Primary methodological definitions, identification assumptions and sensitivity discussion inspected. Not additional clinical evidence about either intervention.

S14. PAOPAC author README Complete text inspected on 26 September 2026, blob ba5d4a8ee7ce2107d8fe34f1dbe24b40dc2da5f0. Input/platform requirements were read; compiled module and separate binary model were not obtained or run. Full preprint methods, exact trial-artifact correspondence and permission for public deployment remain unverified.

Computational-identity evidence observed . D labels distinguish source history, code inspection and bounded execution; they do not refresh unrelated evidence.

D01. Levine et al. (2018) and Supplementary Data 1. Clinical equations, units, Table S1 and distinction from DNAmPhenoAge inspected; relevant 2018 equations/tables visually checked. No full-precision original execution artifact or independently supplied clinical input/output fixture recovered.

D02. Liu et al., correction published 25 February 2019. The conflicting equation was read from publisher-extracted PDF text; image rendering failed, so visual verification is not claimed. No journal reconciliation or historical execution artifact obtained.

D03. BioAge history: 11 October 2023 change, 2 April 2026 restoration, subsequent revert/reapplication and resulting pinned expression. Diffs establish code chronology. The maintainer’s attribution to Levine is not authenticated correspondence or proof of every historical study’s implementation.

D05. Pinned Biolearn hematology function, blob d3df517dce9a2ebe392d5b976c6ba4208b4c72fd. The 1,405-byte source was verified and executed only in isolation on three inherited synthetic cases with NumPy/pandas. The complete package, remote models and empirical workflows were not run. This checks implementation parity, not independent primary-method truth.

D08. Biolearn clinical-layer tests and generic-model tests were read, not executed. Layout/unit/round-trip tests and molecular-model expectations did not provide a qualifying independent clinical Phenotypic Age fixture. This is a bounded search result, not a claim that testing or such an artifact is absent everywhere.

Next: Inspect the three worked evidence chains or separate public code from reproducible access.

Search published pools, pages, reports, and evidence.