Start with the question, not the conclusion
Every trial answers a narrower question than its headline suggests. The narrowing happens in four places at once: the population recruited, the intervention as delivered, the comparison it was tested against, and the outcome measured. Researchers call this PICO, and it is the single most useful discipline a lay reader can adopt.
A trial of a drug in adults aged fifty to seventy with a specific diagnosis, given for twelve weeks, compared with placebo, measuring a blood marker, has told you about exactly that. It has not told you about a healthy forty year old, about two years of use, about the drug compared with exercise, or about whether anyone lived longer.
Write those four things down before reading anything else. Half of the overclaiming in health journalism dissolves at that step, because the mismatch between the question asked and the conclusion reported becomes visible immediately.
Randomisation and blinding, and what they protect against
Randomisation exists to make the groups comparable in everything except the intervention, including in the ways nobody thought to measure. That is its unique power, and it is why an observational study, however large, cannot substitute for it. Adjustment can only correct for variables that were recorded; randomisation balances the ones that were not.
Look for how randomisation was done and whether allocation was concealed from whoever enrolled participants. Concealment is not the same as blinding. If the person recruiting knows which arm the next participant will land in, they can, without any intent to deceive, steer sicker patients away from the intervention arm.
Blinding then protects the outcome. Participants who know they received the active treatment report symptoms differently. Assessors who know report differently too. Blinding matters most for subjective outcomes and least for outcomes such as death, which is one reason hard endpoints are prized.
Where an intervention cannot be blinded, as with exercise or surgery in many cases, look for whether the people measuring the outcome were kept unaware of allocation. That is often possible even when blinding participants is not.
The primary endpoint, and whether it was declared in advance
A trial should nominate one primary outcome, in advance, in a public registry, before recruitment. That single declaration is what makes the statistical result interpretable, because it fixes the question before the answer is known.
Measure enough outcomes and some will move by chance alone. If a trial declares a primary endpoint, misses it, and the published paper leads on a secondary outcome or a subgroup, you are reading a hypothesis that happens to be dressed as a result. Outcome switching of this kind is common enough that checking the registry entry against the paper is worth the two minutes it takes.
Subgroup findings deserve particular suspicion. A result that holds only in women over sixty, or only in the participants with the highest baseline values, is a starting point for the next trial rather than a finding. If it was not pre-specified, treat it as a coincidence until someone reproduces it deliberately.
Watch also for composite endpoints, which bundle several outcomes into one count. A composite of death, heart attack and hospital admission can be driven almost entirely by its least serious component. Look at whether the components moved together.
Who was in it, and who was kept out
The exclusion criteria tell you who the result does not apply to, and they are usually more informative than the inclusion criteria. Trials routinely exclude people with multiple conditions, those over a certain age, those on interacting medicines, and those judged unlikely to complete follow-up.
The consequence is that the population in which a treatment is tested is often healthier and simpler than the population in which it will be used. This is a particular problem in ageing research, where the people most affected by an intervention are the least likely to be recruited into a trial of it.
Then check attrition. How many people were randomised and how many were analysed? If a substantial number dropped out, and especially if more dropped out of one arm, the groups are no longer the groups that were randomised. An analysis that keeps everyone in the arm they were assigned to, whatever they went on to do, is more conservative and generally more trustworthy than one that analyses only those who completed as directed.
Reading the numbers you are given
Two presentations of the same result can produce completely different impressions. A relative change describes the proportion by which risk shifted. An absolute change describes how many more or fewer people experienced the outcome. Relative figures always sound larger, and they are what press releases use.
Always ask what the underlying rate was. A large relative change applied to a rare event is a small absolute change, and it is the absolute change that determines whether anything meaningful happened to anybody.
Look for the interval around the estimate rather than the estimate alone. A confidence interval describes the range of results compatible with the data; a wide one signals that the trial was too small to say much, whatever the central figure suggests. A result described as statistically significant with an interval that nearly touches no effect is a weak result being reported strongly.
Statistical significance is not clinical significance. Given enough participants, trivial differences become statistically detectable. The question that matters is whether the size of the difference would change anything for a person.
Funding, registration, and the trials you never see
Find the funding statement and the conflict of interest declarations. They are usually at the end, in small type. Industry funding does not invalidate a trial, and disclosed funding is far better than the alternative, but it does shift the prior, because funders influence design, comparator choice, dose, duration and the decision to publish.
Publication bias is the deeper problem, and it is invisible in any single paper. Trials with positive results are more likely to be published, published faster, and published in more visible places. The literature you can read is therefore a biased sample of the research that was done. Trial registries exist partly so that unpublished studies leave a trace.
For readers wanting a single practical habit: look up the trial registration, compare the declared primary outcome with the one the paper leads on, and read the funding statement before the abstract. Those three steps catch most of what goes wrong. Systematic reviews, which pool trials and formally assess this kind of bias, are more reliable than any single study, and the Cochrane Library is where to find them.
This method underpins every grade in this journal. Our grading methodology sets out how these judgements are converted into a letter, and our review of metformin is the clearest worked example of an observational literature that cannot answer the question being asked of it.