Longevity Magazine A review journal of healthspan, preventative medicine and ageing
Methodology

How we grade evidence

Four bands, two claim types, and a set of rules that cap a grade no matter how much has been published.

MethodologyLast checked 31 July 2026

In short

Every review in this journal grades one stated claim, not a substance. Grades run from A to D. A means the claim is supported by adequately powered randomised trials with clinical outcomes, replicated independently. B means one such trial exists, or the equivalent for a prognostic claim. C means human evidence exists but is short, small or confined to surrogate endpoints. D means the evidence is preclinical, or human data are too sparse or conflicted to support the claim. Surrogate-only evidence is capped at C however much of it there is.

A grade attaches to a claim, never to a substance

The most common failure in health writing is grading a thing rather than a statement about a thing. Exercise is not graded. The claim that measured cardiorespiratory fitness predicts mortality is graded, and it receives a different grade from the claim that raising your own fitness lowers your own risk.

So every review states the claim being graded at the top, in a single sentence, before any evidence is discussed. If a review makes several claims they are graded separately or the additional ones are discussed in the body with their own grade stated in words.

This has a consequence readers should expect. A compound can appear with a low grade while the underlying science is excellent, because the human question has not been asked. Grade D is not an accusation of bad science. It is a statement that the claim being made in public has not been tested in people.

The four evidence grade bands, A at the top to D at the bottomAReplicated, adequately powered, clinical outcomesBOne adequately powered trial, or consistent large cohortsCShort, small or surrogate-outcome human data onlyDPreclinical, or human data too sparse to support the claim
FigureThe grading ladder used throughout the journal. A grade attaches to a stated claim, never to a substance in general.

Two kinds of claim, graded on different criteria

Not every question can be randomised. Nobody will randomise people to decades of low fitness, and it would be unethical to try. Grading prognostic claims on interventional criteria would mean permanently understating evidence that is as good as evidence about that question can be.

So we distinguish two claim types and state which applies on every review.

Intervention claims assert that doing something changes an outcome. These are graded principally on randomised evidence, because randomisation is the only design that balances the factors nobody measured.

Prognostic claims assert that a measurement predicts an outcome. These are graded on the size, number and independence of the cohorts, the consistency of the association, whether a dose-response relationship is present and orderly, whether the association survives adjustment, and how well the field has addressed reverse causation.

A prognostic grade never licenses an interventional conclusion. Where a review carries a high prognostic grade, it states the interventional grade separately and explicitly, because the gap between the two is where most public confusion lives.

The four bands, in full

GradeIntervention claimsPrognostic claims
AAdequately powered randomised trials with clinical outcomes, reporting on a pre-registered primary endpoint, replicated by an independent group in a different population, with consistent direction and rough magnitude.Large cohorts across multiple populations and eras reporting the same association, with an orderly dose-response, surviving adjustment, with reverse causation adequately addressed.
BAt least one adequately powered randomised trial with a clinical or robust functional primary endpoint, reporting benefit on the pre-registered analysis, not yet independently replicated.Multiple large cohorts agreeing in direction, but with a material weakness such as self-reported exposure, inconsistent dose-response, or limited geographic range.
CHuman randomised evidence exists but is short, small, or confined to surrogate or intermediate outcomes. Or a large observational literature exists with no randomised test of the claim.Consistent associations from observational data where reverse causation or confounding cannot be adequately excluded, or where measurement of the exposure is poor.
DEvidence is preclinical, or human data are early phase, uncontrolled, or too conflicted to support the claim. Includes fields where target engagement cannot be demonstrated.Associations reported but inconsistent, or derived from small or unrepresentative samples, or the measurement itself is not validated.

Rules that cap a grade regardless of volume

Some features of an evidence base limit the grade no matter how many studies exist. Volume of publication is not a proxy for certainty, and a field can accumulate a hundred papers without answering its central question.

  • Surrogate outcomes only. Capped at C. A biomarker moving is evidence about the biomarker. Our explainer on surrogate endpoints sets out why volume does not fix this.
  • No demonstrable target engagement. Capped at D. If nobody can show the intervention reached and acted on its target, a null result is uninterpretable and a positive result cannot be attributed.
  • Animal evidence only. Capped at D, however strong and however well replicated. Multi-site replication in mammals raises confidence that the animal finding is real, not that it transfers.
  • Uncontrolled human reports. Contributes nothing to a grade. Self-selected users reporting their own outcomes carry no information about causation.
  • Sponsor-controlled evidence base. Capped at B where all substantive trials are funded or conducted by parties selling the intervention, until independent replication exists.
  • Trial duration far shorter than the claim. Capped at C where a claim about years or decades rests on trials of weeks.

What we weigh within a band

Within a band, several things move our confidence without changing the letter. We say so in the body of the review rather than inventing intermediate grades.

Pre-registration and whether the reported primary outcome matches the registered one. Allocation concealment and blinding, especially for subjective outcomes. Attrition and whether analysis kept participants in their assigned groups. Whether the comparator was a fair one. Whether the population resembles anyone the claim is aimed at. Whether the effect, if real, would be large enough to matter to a person. And whether the field has produced null results, because a literature with no null results is a literature with a publication problem.

Funding is recorded and weighed but never used alone to dismiss a study. Disclosed industry funding is a functioning system doing its job. The concern is a whole evidence base controlled by interested parties, which is why that appears as a cap rather than a judgement on any single trial.

What we do not do

We do not score substances out of ten, produce league tables, or rank interventions against one another. Different claims are supported by different kinds of evidence, and a single ranking would imply a comparability that does not exist.

We do not give doses, protocols or personal recommendations. Several interventions covered here are prescription only medicines, and prescribing is a clinical decision made by a doctor who knows the patient. Nothing on this site is medical advice, and this is stated on every page rather than buried in a disclaimer.

We do not name individual studies, authors, journals or numerical results in our reviews. This is a deliberate editorial constraint. Describing a literature qualitatively, by trial phase, population, duration, species, funding type and replication status, is more useful to a reader deciding how much confidence to place in a field, and it removes the temptation to lend false precision to a summary. Readers who want the primary sources are pointed to the institutional indexes where they can be found and read directly.

We do not accept payment, product, hospitality or advance sight of coverage from any party with an interest in a grade. The editorial policy sets out the full position.

Review cycle, and how grades change

Every graded review states, in its own words, what would raise the grade and what would lower it. That statement is written before the evidence arrives, which is the same discipline we ask of trialists. It makes the grade a testable position rather than an opinion.

Reviews are re-examined at least annually and sooner when a substantial trial reports. When a grade changes, the page records that it changed and why, rather than being quietly rewritten.

As of this revision, no claim reviewed in this journal holds grade A on an interventional basis. One prognostic claim does. That distribution is not editorial pessimism. It is what the field looks like when claims are graded against the evidence for the specific thing being asserted, and it is the most useful single fact this journal can offer a reader.

Frequently asked

What does a grade actually apply to?

One stated claim, written out in a single sentence at the top of the review. It never applies to a substance in general, so the same compound can carry different grades for different claims made about it, and a low grade is not a judgement on the quality of the underlying science.

Why can surrogate evidence never exceed grade C?

Because a marker moving is evidence about the marker. Validating a surrogate requires showing that changes in it reliably predict changes in the outcome across treatments and populations, which has been done for very few markers. Without that, any volume of surrogate evidence remains evidence about the surrogate.

Why is strong animal evidence capped at D?

Because the translation record from animal outcome to human clinical benefit is poor across medicine as a whole. Well replicated animal work raises confidence that the animal finding is real. It does not raise confidence that it transfers, and the cap makes that distinction visible rather than letting it be smoothed over.

Do you ever grade between bands?

No. Intermediate grades imply a precision the underlying evidence cannot support, and they encourage arguing about half-steps rather than about the evidence. Where our confidence sits near a boundary we say so in the body of the review.

How often are grades reviewed?

At least annually, and sooner when a substantial trial reports. Every review states in advance what would raise and lower its grade, which makes the position falsifiable. When a grade changes, the page records the change and the reason rather than being quietly rewritten.