Start with the phase

The single most informative fact about any medical finding is where it sits in the development pipeline. Preclinical work — cells in a dish, or mice — is where the great majority of exciting results live and die. Something being cured in a mouse model tells you the idea is worth testing; it does not predict a human benefit, and the attrition from this stage to an approved medicine is brutal. Phase 1 trials test safety and dosing in a small number of people and are not designed to show that anything works.

Phase 2 looks for a signal of efficacy in a few dozen to a few hundred participants, and is where enthusiasm typically peaks and evidence typically does not. Phase 3 is the expensive part: hundreds to thousands of participants, randomised, usually blinded, powered to detect an effect on an outcome that matters. A meaningful share of treatments that looked convincing in phase 2 fail in phase 3. When a headline does not tell you the phase, that omission is usually doing work, and you can check the trial's registered entry yourself on a public registry.

Absolute versus relative risk

This is the most reliable way that technically accurate reporting misleads. If a condition affects four people in a thousand and a treatment reduces that to two, the relative risk reduction is fifty per cent and the absolute reduction is two per thousand. Both statements are true. Only one of them helps you decide whether to take the drug, tolerate its side effects and pay for it.

A useful companion figure is the number needed to treat: how many people have to take the treatment for one of them to benefit. For interventions in high-risk groups that number can be pleasingly small; for the same intervention in low-risk people it is often in the hundreds. This is why the same medicine can be clearly worthwhile for one person and marginal for another, and why blanket claims that a drug does or does not work are usually the wrong shape. If a report gives you only relative numbers, treat that as a signal to look for the absolute ones before forming a view.

Surrogate endpoints and what they hide

A surrogate endpoint is a measurement used as a stand-in for the thing anyone actually cares about: a laboratory value instead of a heart attack, tumour shrinkage instead of survival, bone density instead of a fracture. Surrogates make trials faster and cheaper, and sometimes they track the real outcome well. Sometimes they do not, and medicine has learned this the hard way more than once — treatments that improved the marker while leaving patients no better off, or worse.

The practical rule is to notice which one you are being told about. When a report says a treatment improved a score, a level or an image, ask what happened to symptoms, function, hospital admissions or survival, and whether those were measured at all. Regulators sometimes approve on surrogates with a requirement for confirmatory outcome data later; that is a reasonable trade in serious disease with no alternatives, but it means the question is open, not settled. Cardiovascular outcome trials became standard in diabetes drug development precisely because glucose-lowering alone had proved to be an unreliable guide to benefit.

Size, replication and conflicts

Small studies produce dramatic results more often than large ones, in both directions, simply because noise is larger relative to signal. A striking effect from a few dozen participants is a hypothesis. When a large well-conducted trial disagrees with a small one, the large one is usually closer to right. Duration matters equally: a twelve-week trial cannot tell you about a treatment someone will take for thirty years, and effects on weight, blood pressure and mood are notorious for shrinking as follow-up lengthens.

Replication is what turns a finding into knowledge, and independent replication — different investigators, different population, ideally different funders — carries more weight than the same group repeating itself. Funding does not invalidate a study, and most drug trials are funded by the companies that make the drug because nobody else will pay for them. What it does mean is that design choices worth scrutinising: the comparator chosen, the dose of that comparator, the endpoint selected, and whether the protocol was registered before the data were collected. Registration matters because it makes it visible when the reported outcome is not the one originally planned.

What a genuine advance has looked like

mRNA vaccine platforms are the clearest recent example: two decades of unglamorous work on nucleoside modification and lipid nanoparticle delivery, then large randomised trials, then deployment at scale, then years of safety surveillance. The platform matters more than any single product, because the same approach can be redirected at other targets. GLP-1 receptor agonists are another: a class developed for type 2 diabetes that produced substantial weight loss, and then — the part that made it a genuine advance rather than a commercial one — cardiovascular outcome trials showing benefit on events, not just on numbers.

CAR-T cell therapy, in which a patient's own T cells are engineered to attack their cancer, produced durable remissions in some blood cancers that had exhausted other options, at the cost of serious toxicities and enormous complexity. Gene therapies for sickle cell disease, approved in the US and elsewhere, address the underlying genetics of a disease that had seen little fundamental progress for decades — with real questions still open about long-term durability, access and cost. Immune checkpoint inhibitors changed the outlook in several advanced cancers, though they work for a minority of patients and carry autoimmune toxicities. Every one of these took years and produced hard outcome data, which is the pattern worth looking for.

Applying this to the next headline

Run the checklist. What species. What phase. How many people, followed for how long. Compared with what — placebo, standard care, or nothing at all. Absolute or relative. Real outcome or surrogate. Published and peer reviewed, or a conference abstract or preprint. Who paid, and was the protocol registered in advance. That takes about two minutes with the abstract in front of you and disposes of most overstated coverage.

Then ask the question that matters to you: does this apply to me. Trials enrol defined populations, and a benefit demonstrated in people with established heart disease may not transfer to someone without it. If something looks genuinely relevant to your situation, the productive move is to bring it to your clinician with the specifics rather than to act on a headline. Our guide to reading studies goes deeper into designs, confidence intervals and the difference between association and cause.