What these tests actually do
The most common form measures DNA methylation: small chemical marks attached to DNA at specific positions, which change in patterned ways across the lifespan. A test measures the methylation state at a set of positions and feeds those values into a statistical model, which returns a single number described as biological age.
The key point is that the model is a prediction machine, and it was trained on something. It learned which combination of methylation values best predicts a target variable in a training dataset. Everything the output means derives from what that target was.
Other tests exist alongside methylation clocks. Some use blood biochemistry panels combined into a composite score. Some measure telomere length. Some use proteins or metabolites. All share the same structure: measure something, put it through a model, return an age. All share the same interpretive question.
What the model was trained to predict
Two broad generations exist and the distinction is more important than any difference between brands.
First generation clocks were trained to predict chronological age. The training target was the date on the birth certificate. Such a model is judged by how closely it recovers known ages, and a good one does so quite closely. The consequence is uncomfortable: the better the model gets at its training task, the less room is left for anything else, because the residual is precisely the part not explained by calendar age.
Second generation clocks were trained on health outcomes instead, typically mortality or a composite of clinical measures alongside age. These carry more information about health, because that is what they were built to detect, and they generally outperform the first generation at predicting outcomes in cohort studies.
A commercially sold test may use either, or a proprietary model that is not described. If the test does not say what its model was trained on, no interpretation of the result is possible, and that alone is grounds to disregard the output.
Why two tests disagree, and why the same sample can too
Send a sample to two providers and you may get results that differ by years. This is expected rather than scandalous. Different clocks use different positions, different models and different training targets, so they are not measuring the same construct. Agreement would be the surprise.
More troubling is within-test variability. Test-retest reliability, the extent to which the same sample or the same person measured again produces the same answer, is a recognised weakness of methylation clocks. Technical noise in the measurement, differences in the mix of cell types in a blood sample, sample handling and processing batch all shift the result.
Cell composition deserves particular attention. A blood sample contains a mixture of cell types, that mixture shifts with recent infection, stress, exercise and time of day, and methylation differs between cell types. Some clocks adjust for this and some do not. A result that moved because you had a cold last week is not a result about ageing.
The practical implication is direct. If the measurement error is of similar size to the changes people are trying to detect, then tracking your own score over time is largely tracking noise, and any intervention will appear to work roughly half the time.
What a result can and cannot support
These tests are genuinely valuable in research. At population scale, where individual noise averages out, second generation clocks predict outcomes and have become a useful tool for studying ageing biology. That is a real contribution.
What has not been established is that they are valid surrogate endpoints. Showing that a clock predicts mortality in a cohort is not the same as showing that changing the clock changes mortality. That second demonstration would require an intervention trial with clinical outcomes, and it has not been done for any clock. Our explainer on surrogate endpoints sets out exactly why the first does not imply the second.
For an individual, the position is weaker still. A single result carries measurement error large enough to matter, no clinical guideline attaches an action to it, and there is no evidence-based response to an unfavourable score other than the general advice that would apply anyway. A test whose result cannot change what you do is not a useful test, whatever it costs.
There is also a commercial structure worth naming. Where the same organisation sells the test and the intervention intended to improve the score, the incentive is to sell a measurement that responds to the product. That is not an allegation against any particular company. It is a structural conflict that a careful reader should account for. Our review of NAD precursors discusses a field where this pattern is common.
If a test result would worry you and there is no action attached to it, that is a reasonable ground for not taking it. The measurements with established clinical value in this space, blood pressure, glycaemic measures, lipids, and where indicated cardiorespiratory fitness, are available through ordinary clinical care and come attached to guidance about what to do with them.[1]