SocraticHow to read the evidence
A fifteen-minute tutorial · no prior knowledge assumed

How to read the charts on this site.

Every page here shows the same few shapes: a dot with a bar, a dashed line, a row of squares, a probability. This page teaches each one with something you can move. By the end you can read any paper page or cell of the map without help.

1A study reports a number, and a range

Every result you will see is a dot with a bar through it. Drag the two sliders and read the sentence.

← favours the treatmentfavours the comparison →
best guess0.800.56 to 1.14can't rule out
0.10.250.524101 · no effect
0.05 Measles vaccine0.7 SGLT2 inhibitors, heart failure admissions0.75 Statins, major vascular events per 1 mmol/L LDL lowered0.9 Aspirin for people with no prior heart disease2 Obesity and type 2 diabetes15 Smoking and lung cancer
Could not tellthe best guess leans towards the treatment, but the bar crosses the no-effect line, so it is a hint, not a finding
How big
dramaticdramatic
dramaticlargetypicaltiny
modest, and typical · like statins
How sure
precisedecentvague
precisedecentvague
top ÷ bottom = 2.0 · a decent trial
Why

Most effective drugs on hard outcomes land here, 10 to 30% fewer events; a statin is about 0.75. The faint ticks on the line are effects you already know, from the measles vaccine near 0.05 to smoking and lung cancer near 15; hover one.

Precision of a ratio is relative, so the bar is judged by its top divided by its bottom, not by how wide it looks: 0.23 to 0.26 is exactly as precise as 3.7 to 4.3. Under about 1.4 comes from a large trial; over 2.5 could not tell you much whichever side the dot is on. Width shrinks with the square root of the evidence: four times the people, or four times the events, halves the bar. For deaths and heart attacks it is the number of events that counts, which is why those trials need five to twenty thousand people while a blood-sugar trial needs hundreds. The bar measures repeatability, not truth: a huge, precise trial with a leaky design can still be wrong.

2What a ratio means

Most of the numbers on the map are ratios: hazard ratios, risk ratios, odds ratios. They compare how often something happened in two groups.

Risk ratiohow many had itOdds ratiohad it, against did notHazard ratiohow fast it arrivedThree ways to divide the same two groups. Set the groups, then watch each one get built.
100 people given the treatment
100 people given the comparison

All three read the same way. 1 is no difference. Below 1, less in the treated group; above 1, more. Here the outcome is rare enough, 12 in 100, that the three agree.

Risk ratiothe share who had it, one group divided by the other
treated
8 of 100
comparison
12 of 100
8/100÷12/1000.67
33% fewer had it.
Odds ratiohad it against did not, one group divided by the other
treated
8 : 92
comparison
12 : 88
8:92÷12:880.64
Practically the same as the risk ratio here.
Hazard ratiohow fast it arrived over the follow-up, one group divided by the other
treated rate÷comparison rate0.65
Practically the same as the risk ratio here.

Which one gets reported? The study decides, not the author. Pick the shape of the study you are reading.

Risk ratio when everyone can be counted at the same finish line. Hazard ratio when the finish line moves. Odds ratio when the study worked backwards from the outcome.

How big
dramaticdramatic
dramaticlargetypicaltiny
large · like sglt2 inhibitors
4fewer per 100 peoplethe absolute difference
25treated for one to benefitnormal for prevention
Why

Why the map spaces ratios on a log scale. Halving and doubling are the same size of effect, so 0.5 and 2 sit the same distance from 1. The hazard ratio here is drawn as if the outcome arrived at a steady pace; real trials estimate it from when each event happened.

30 to 50% fewer events, or up to double; uncommon for a drug on a hard outcome, common for strong risk factors. The same zones as section 1: most effective drugs on hard outcomes sit between 0.7 and 0.9.

Relative and absolute are different questions: a 25% reduction of something rare is a small number of people, which is why trial reports give both. The number needed to treat: under 10 is exceptional, 20 to 100 is normal for prevention, hundreds is common and still worthwhile for a cheap, safe drug.

3Which side is good

Some results are ratios, some are differences, and for some outcomes lower is better while for others higher is. The graph sorts all of that out so the treatment's good side is always the left.

Ratiofor things that happen or not: a death, a heart attack, an infection. No effect is 1.Differencefor things you measure: HbA1c, weight, blood pressure. Treated minus comparison; no effect is 0.The outcome decides which one a study reports. Both are drawn the same way: the treatment's good side is on the left, whichever direction is good.
← favours the treatmentfavours the comparison →
the study foundHR 0.80
0.521 · no effect
Treatment betterdeaths: lower is better · 20% fewer deaths in the treated group
Does it matter
below the thresholdcounts clinically
10 to 30% fewer is what a good drug usually achieves on deaths
Why is the treatment's good side on the left?

It is the convention in medical forest plots, made standard by the Cochrane reviews: “favours treatment” on the left, “favours control” on the right. It came from ratios, where most outcomes are bad events, a ratio below 1 means fewer of them, and smaller numbers sit to the left. For outcomes where higher is better, published plots either flip the axis or leave you to work it out. We chose to keep the treatment's good side on the left every time, so there is one rule to learn.

4Many studies on one line

Put every study on the same axis and the shape of the evidence appears. This is called a forest plot. Click a study.

← favours the treatmentfavours the comparison →

Study C, 12,000 people. Its whole bar is left of the line: it favours the treatment. A big study, so a short bar; this is the size an outcome trial needs, five to twenty thousand people, because deaths and heart attacks are rare and the effects are modest.

Leans towards the treatmentthree clear of the line on the left, three that could not tell, none clearly against

5Which studies get to vote

Only randomised, registered trials cast votes on the map. Here is why a study where people chose their own treatment cannot be trusted to, however large.

Randomised triala coin decides who gets the treatment. It votes.Observational studypeople and their doctors chose. It can hint, it cannot vote.Registeredthe plan was public before the results. Required for a vote.Forty people, sixteen of them sicker. The drug does nothing. Watch what happens to the result when who gets it is not random.
Given the drug · 20 people, 12 sicker
40% had the outcome
Not given the drug · 20 people, 4 sicker
25% had the outcome
healthiersickerhad the outcome
The drug looks harmfulrisk ratio 1.60 · and it did nothing. The sicker people chose it, and sicker people have the outcome more. This is confounding, and no amount of arithmetic afterwards fully removes it.

That is why only randomised trials vote. A registry entry, filed before the trial ran, says what it planned to measure, so nobody can pick the flattering result afterwards. Observational studies still appear on the map as hints, in the “outside the trials” group, and they feed the risk-factor rows.

Versus placebodoes it work at all? A dummy pill or usual care as the comparison.Head-to-headis it better than the drug we already use? The map keeps this apart from placebo, because the same drug can win one and lose the other.
What if a trial did not set out to win?

Most head-to-head trials are superiority trials: the new drug must beat the old one. Some are non-inferiority trials: the aim is to show the new drug is not meaningfully worse, usually because it is cheaper or safer. There, “no difference” is success. The registry does not record which design a trial used, and the map does not yet tell them apart, so a non-inferiority trial that succeeded is currently read as a vote against. About one head-to-head trial in a hundred carries the word in its title. This is a known gap and is on the list.

6Every study gets one vote, and the votes become belief

Studies are not averaged. Each trial casts one vote in each cell it speaks to, and the votes become a probability. The paper page shows that probability with and without the paper you are reading.

One trial, one voteA trial reports many results. Only the one it promised to measure before it started, the main result, decides its vote.
Many papers, still one voteThe first report, a follow-up and a subgroup analysis are the same trial and the same patients, so together they get one square.
Many votes, one beliefEach trial's square goes into the cell. The squares become the number below.
The numberthe chance the treatment works, from 0 to 1. 0.5 is a coin toss.The barhow unsure the graph still is. More deciding studies, shorter bar.The statewhat the pattern of votes says: one study, replicated, or contested.Only randomised, registered trials vote, as in section 5. Inconclusive votes do not count. Click a square to change its vote; hover one to see the number without it.
supports contradicts inconclusive
Belief
0.673 support, 1 contradict · 1 inconclusive, not counted
Without the hovered onehover a square
no deciding studynothing counts yetone studya single trial can be wrongreplicatedindependent studies agreecontestedat least a quarter disagree

What moves it? Add a study and watch which of the three changes.

Agreeing studies push the number up and shorten the bar. One disagreeing study in four is enough to make the claim contested; only enough agreeing studies to push the disagreement under a quarter lift it. Inconclusive studies change nothing.

What if a trial's results disagree with each other?

The main result is the one the trial promised to measure before it started. Secondary results are everything else it measured along the way; they can add support, never contradict on their own. Every vote answers one question: does the treatment help with this outcome?

Supports: the treatment helpedbecause its main result favoured the treatment
Why

A trial's main, pre-declared result decides its vote. Secondary results can add support but never contradict on their own, because a trial tests many things and some will look bad by chance. A main result that found nothing counts against, because a trial built to find an effect and finding none is evidence.

What about “significant”?

A result is called significant when its p value is under 0.05, meaning a result that surprising would turn up by chance less than one time in twenty if there were no effect. That is a convention, not a law of nature: a lone trial at p = 0.04 is fragile, and the same finding twice is worth far more than one finding at p = 0.001. The map counts trials, not p values, for exactly that reason.

Why

This is a toy version of the rule; the real one, with its error models, is under Methodology. The second row is the whole point of the paper page: remove one study and see whether anything changes. On the map, a study is one evidence family, so a trial reported in five papers is still one square.

7Markers versus outcomes

Most diabetes trials measure blood sugar, not heart attacks. A drug can move the number and not the thing that matters. This is the most important idea on the page.

Markersomething measured in blood or on a scale: HbA1c, LDL, weight, blood pressure. Cheap, fast, moves in months. Also called a surrogate.Outcomesomething that happens to a person: a heart attack, a hospital stay, death. Slow and rare, so a trial needs thousands of people and years.Markers are used because the outcome trial is expensive. The bet is that moving the marker moves the outcome. Sometimes it does.
the drugRosiglitazone
markerHbA1c the good direction
outcomeheart attacks more, the outcome trial overturned it
Marker and outcome disagreedIt lowered blood sugar as well as any diabetes drug, and was prescribed to millions. Then the outcome trials showed more heart attacks. The marker moved the right way; the outcome moved the wrong way.

How the map uses this. A cell on a marker column is cheaper to fill but proves less. When the map proposes what to test next it says which kind of trial it is asking for, and when the outcome trial is huge it names the cheaper step: the marker trial that should come first, and why that is not the end of the story.

Expected gainhow much one more independent trial would move the belief. High when a cell has one trial or a real disagreement; near zero when it is settled.Settledthe belief is already at its ceiling. Another trial of the same kind changes nothing, so the map does not ask for one.Cheaper stepthe marker trial the map points at first when the outcome trial would take thousands of people and years.

8Reading a cell of the map

Everything above meets in one square. Here is what its colour, its border and its corner number say, and why a cell with many trials can rest on very few.

A rowone treatment, or one risk factor in the lower rows.A columnone outcome. Most ask “does the treatment improve this?” The safety columns ask the opposite: “does it cause more of this?”A cell is one treatment against one outcome. The map has four ways to colour it, and one that shows all four. Click a view.
fillthe lean of the resultsbrightnesshow many trials
#3
cornerrank of what to test nextborderhow sure the belief is

One glyph carries all four. Fill colour is the lean of the results, brightness is how many trials, the border is how sure the belief is, and the corner number is the rank of what to test next.

replicated: independent trials agreeone trial onlycontestedreplicated harm, safety columnsno registered trial: a gap

Why a cell with 25 trials can rest on two. Registering a trial is required; posting its results is not. Of the trials that post, many report per-arm numbers with no comparison, or an outcome the map cannot read as for or against. Open any cell and the funnel shows the loss at each step.

25registered to measure it
9posted a result16 never posted
3readable as for or against6 unreadable
2decided1 could not tell
Badges you will meet on a paper page
Retracted
the journal withdrew it. It no longer votes.
Expression of concern
the journal has doubts and is looking. It still votes, flagged.
Erratum
a correction was published; usually small.
Registry-linked trial
the paper reports a registered trial, so it can vote.
Outside this release
the trial exists in the registry but is not in the current map build.
Full text read
the extractor read the whole paper, not only the abstract.

9Words you will meet

Estimate
The single best number a study found. The dot.
Confidence interval
The range the study cannot rule out, usually the middle 95%. The bar. For a ratio, judge it by top divided by bottom: under about 1.4 is a large trial, over 2.5 could not tell you much.
Hazard ratio (HR)
How fast the outcome arrives in the treated group compared with the comparison group, over time. 0.8 means 20% slower. 1 means no difference.
Risk ratio (RR)
How often the outcome happened in the treated group compared with the comparison group. Same reading as a hazard ratio.
Odds ratio (OR)
A cousin of the risk ratio used in some designs. Reads the same way; a bit more extreme for common outcomes.
Mean difference
For measured quantities such as HbA1c or weight: treated minus comparison. 0 is no effect.
p value
How surprising the result would be if there were truly no effect. Small is surprising. It says nothing about how big the effect is; the interval does.
Primary outcome
The one result a trial promised to measure before it started. It decides the trial's vote.
Comparator
What the treatment was compared with: placebo or usual care, or another drug, which the map calls head-to-head.
Registered trial
A trial listed on a public registry such as ClinicalTrials.gov before it ran, with its planned outcomes. Registering is required; posting results is not, which is why many cells have trials with no readable result.
Evidence family
One independent source of evidence: one trial, however many papers report it. The belief counts families, so a trial reported five times is still one vote.
Belief
The graph's probability that a claim is true, from its evidence families. Contested means a real minority disagrees. Replicated means independent families agree.
Number needed to treat
How many people must take the treatment for one to benefit. Under 10 is exceptional, 20 to 100 is normal for prevention, hundreds is common and still worthwhile for cheap, safe drugs.
Standardised effect
When a measure has no natural unit, the effect is scaled by its spread: 0.2 is small, 0.5 medium, 0.8 large. Most drug effects on symptoms and scores are 0.2 to 0.5.
Blood pressure and LDL
Two anchors worth knowing: every 5 mmHg of systolic pressure lowered is worth about 10% fewer major cardiovascular events; every 1 mmol/L (39 mg/dL) of LDL lowered is worth about 22% fewer major vascular events.
Cell
One square of the map: one treatment against one outcome.
Randomised trial
A coin decides who gets the treatment, so the groups differ only in the treatment. The only kind of study that votes.
Observational study
People and their doctors chose. Sicker or richer or more careful people end up in one group, and that, not the treatment, can explain the result. Called confounding.
Placebo
A dummy treatment, so that neither the patients nor the doctors know who got the real one. Removes the effect of expecting to get better.
Surrogate, or marker
A measurement that stands in for an outcome: HbA1c for diabetes complications, LDL for heart attacks. Cheap to move; whether the outcome follows must be shown, not assumed.
Non-inferiority trial
A head-to-head trial whose aim is to show the new drug is not meaningfully worse than the old. “No difference” is success there, which the map does not yet tell apart from a failed superiority trial.
Follow-up
How long people were watched. A hazard ratio over one year and one over ten are not the same claim; longer follow-up finds slower harms and benefits.
Subgroup
A result for part of the trial, such as people over 65. With ten subgroups, one will look striking by chance, so a subgroup result is a hint until a trial is run in that group.
Why not one average
Meta-analyses pool studies into one diamond. The map does not, because pooling hides disagreement and lets one large trial outvote everything; each independent trial keeps its own vote instead.
Expected gain
How much one more independent trial would move a cell's belief. Drives what the map proposes to test next.
Settled
A claim whose belief is at its ceiling; another trial of the same kind would not move it.
Cheaper step
The marker trial the map points at first when the outcome trial would need thousands of people and years.

The votes and belief demos are simplified illustrations of the rules; the exact rules, with their error models, are under Methodology. Try it on a real paper: metformin, aspirin and cancer mortality.