Study designs, ranked

At the top sit well-conducted systematic reviews and meta-analyses, which gather every study meeting stated criteria and synthesise them — valuable precisely because they show the whole picture rather than the most interesting slice of it. Below them, randomised controlled trials: participants are allocated by chance, which distributes both known and unknown differences evenly and is the only design that reliably supports a causal claim. Blinding, where feasible, prevents expectation from colouring the result.

Below those are observational designs. Prospective cohorts follow groups over time and are strong for questions no trial can ethically address, but they can only adjust for confounders that were measured. Case-control studies work backwards from an outcome and are efficient for rare events, at the cost of recall and selection bias. Cross-sectional studies capture a single moment and cannot establish sequence. At the bottom, case reports and case series describe what happened to one or a few people — useful for raising a hypothesis or flagging a rare harm, useless for estimating how often anything occurs. Animal and laboratory work sits outside the hierarchy: essential to science, not evidence about people.

Absolute risk, relative risk and number needed to treat

Relative risk expresses change as a proportion of the starting risk. Absolute risk expresses it in actual events. If a treatment halves risk, that is meaningless until you know what it halved. Going from four events per thousand to two is a fifty per cent relative reduction and a two-per-thousand absolute reduction; going from four hundred per thousand to two hundred is also fifty per cent and is an entirely different proposition. Headlines almost always report relative figures, because they are larger.

Number needed to treat converts this into something intuitive: how many people must receive the treatment for one to benefit. A number needed to treat of twenty means nineteen people take the drug, and its side effects, without benefiting — which is not an argument against treatment, but is the honest shape of most preventive medicine. The mirror image, number needed to harm, applies the same logic to adverse effects. Ask for both, and the decision usually becomes clearer than any single statistic makes it.

P-values and confidence intervals in plain language

A p-value answers one narrow question: if there were genuinely no effect, how surprising would data like this be. A p-value below 0.05 has become a conventional threshold for calling a result statistically significant, but the threshold is arbitrary, and the p-value says nothing about how large an effect is or how important it might be. With a large enough sample, a trivial difference becomes statistically significant; with a small sample, an important one can fail to reach it.

A confidence interval is more informative because it shows a range of values compatible with the data. An intervention that reduces risk by 20% with an interval from 18% to 22% is a different finding from one reporting 20% with an interval from 1% to 39%, even though both may be reported as significant. Look at where the interval sits and how wide it is. Be wary too of multiple comparisons: a study testing twenty outcomes will usually find one that looks significant by chance alone, which is why pre-registered primary outcomes matter and why subgroup findings should be treated as hypotheses.

Associated with is not causes

This is the single most common error in health reporting. Observational studies show that things occur together; they cannot on their own show that one produced the other. Confounding is the usual culprit — some third factor drives both. People who take vitamins also tend to exercise, eat differently and attend screenings, and disentangling the vitamin from the lifestyle is genuinely hard. Reverse causation is the other classic: early illness changes behaviour before it is diagnosed, so the behaviour appears to cause the illness it was actually a response to.

Careful observational researchers know this and use design and analysis to address it, and some causal conclusions are firmly established from observational data — smoking and lung cancer, where the effect was enormous, consistent across many designs and populations, dose-dependent and biologically coherent. That is the bar. A single cohort study reporting a modest association between one food and one outcome does not clear it. When you read "linked to", "associated with" or "tied to", read it literally: these words are chosen because the stronger word is not supported.

Preprints, peer review and replication

A preprint is a manuscript posted publicly before peer review. Preprints accelerated science significantly during the pandemic and they remain valuable, but nobody independent has yet tried to find the flaws. Some preprints never pass review, and some change substantially before they do. Treat a preprint as a conversation in progress rather than a conclusion, and check whether it was subsequently published.

Peer review is a filter, not a guarantee. Reviewers rarely see raw data, cannot detect most fabrication, and journals vary enormously in rigour — with a substantial predatory publishing industry charging fees for essentially no review at all. What genuinely increases confidence is independent replication: a different team, a different population, ideally a different funder, reaching the same result. Replication problems have been documented across biomedical and psychological research, and a finding that has never been repeated is provisional however prestigious the journal. Conference abstracts sit lower still — brief, often not peer reviewed, and a meaningful share never reach full publication.

Funding, disclosures and registration

Industry funding does not make a study wrong, and most drug trials are industry funded because nobody else pays for phase 3. It does justify closer attention to design choices: which comparator was used and at what dose, which endpoint was chosen, how long follow-up ran, and whether the analysis population was defined in advance. Systematic evaluations have found industry-funded studies more likely to report favourable conclusions, which is generally attributed to these design and publication choices rather than fraud.

Two documents help. Author disclosure statements list financial relationships — worth reading, particularly for review articles and editorials, where opinion carries more weight than data. And trial registration, on a public registry before enrolment, records what was planned. Comparing the registered primary outcome with the published one reveals outcome switching, one of the more consequential and least visible problems in the literature. If a trial is not registered and the paper does not explain why, that is genuinely informative.

A two-minute checklist

Who was studied, and are they like you — age, sex, baseline risk, existing conditions. What was the design, and does it support the claim being made. How many participants, and for how long. Compared with what: placebo, standard care, or nothing. Absolute numbers as well as relative. Was the outcome one that matters to a person, or a marker standing in for it. Has it been replicated. Who funded it, and was it registered in advance.

Then one last question, the one that catches most bad claims: what would this look like if it were not true. A claim that cannot fail any test is not a scientific claim. If you want to go further, Cochrane reviews are the best single starting point for a synthesised answer, guideline bodies show how a body of evidence has been weighed, and abstracts on public indexes give you the design and the numbers in a couple of paragraphs. Our health news standards describe how we apply the same filters before writing anything.