How to read the charts on this site.
Every page here shows the same few shapes: a dot with a bar, a dashed line, a row of squares, a probability. This page teaches each one with something you can move. By the end you can read any paper page or cell of the map without help.
1A study reports a number, and a range
Every result you will see is a dot with a bar through it. Drag the two sliders and read the sentence.
Why
Most effective drugs on hard outcomes land here, 10 to 30% fewer events; a statin is about 0.75. The faint ticks on the line are effects you already know, from the measles vaccine near 0.05 to smoking and lung cancer near 15; hover one.
Precision of a ratio is relative, so the bar is judged by its top divided by its bottom, not by how wide it looks: 0.23 to 0.26 is exactly as precise as 3.7 to 4.3. Under about 1.4 comes from a large trial; over 2.5 could not tell you much whichever side the dot is on. Width shrinks with the square root of the evidence: four times the people, or four times the events, halves the bar. For deaths and heart attacks it is the number of events that counts, which is why those trials need five to twenty thousand people while a blood-sugar trial needs hundreds. The bar measures repeatability, not truth: a huge, precise trial with a leaky design can still be wrong.
2What a ratio means
Most of the numbers on the map are ratios: hazard ratios, risk ratios, odds ratios. They compare how often something happened in two groups.
All three read the same way. 1 is no difference. Below 1, less in the treated group; above 1, more. Here the outcome is rare enough, 12 in 100, that the three agree.
Which one gets reported? The study decides, not the author. Pick the shape of the study you are reading.
Risk ratio when everyone can be counted at the same finish line. Hazard ratio when the finish line moves. Odds ratio when the study worked backwards from the outcome.
Why
Why the map spaces ratios on a log scale. Halving and doubling are the same size of effect, so 0.5 and 2 sit the same distance from 1. The hazard ratio here is drawn as if the outcome arrived at a steady pace; real trials estimate it from when each event happened.
30 to 50% fewer events, or up to double; uncommon for a drug on a hard outcome, common for strong risk factors. The same zones as section 1: most effective drugs on hard outcomes sit between 0.7 and 0.9.
Relative and absolute are different questions: a 25% reduction of something rare is a small number of people, which is why trial reports give both. The number needed to treat: under 10 is exceptional, 20 to 100 is normal for prevention, hundreds is common and still worthwhile for a cheap, safe drug.
3Which side is good
Some results are ratios, some are differences, and for some outcomes lower is better while for others higher is. The graph sorts all of that out so the treatment's good side is always the left.
Why is the treatment's good side on the left?
It is the convention in medical forest plots, made standard by the Cochrane reviews: “favours treatment” on the left, “favours control” on the right. It came from ratios, where most outcomes are bad events, a ratio below 1 means fewer of them, and smaller numbers sit to the left. For outcomes where higher is better, published plots either flip the axis or leave you to work it out. We chose to keep the treatment's good side on the left every time, so there is one rule to learn.
4Many studies on one line
Put every study on the same axis and the shape of the evidence appears. This is called a forest plot. Click a study.
Study C, 12,000 people. Its whole bar is left of the line: it favours the treatment. A big study, so a short bar; this is the size an outcome trial needs, five to twenty thousand people, because deaths and heart attacks are rare and the effects are modest.
5Which studies get to vote
Only randomised, registered trials cast votes on the map. Here is why a study where people chose their own treatment cannot be trusted to, however large.
That is why only randomised trials vote. A registry entry, filed before the trial ran, says what it planned to measure, so nobody can pick the flattering result afterwards. Observational studies still appear on the map as hints, in the “outside the trials” group, and they feed the risk-factor rows.
What if a trial did not set out to win?
Most head-to-head trials are superiority trials: the new drug must beat the old one. Some are non-inferiority trials: the aim is to show the new drug is not meaningfully worse, usually because it is cheaper or safer. There, “no difference” is success. The registry does not record which design a trial used, and the map does not yet tell them apart, so a non-inferiority trial that succeeded is currently read as a vote against. About one head-to-head trial in a hundred carries the word in its title. This is a known gap and is on the list.
6Every study gets one vote, and the votes become belief
Studies are not averaged. Each trial casts one vote in each cell it speaks to, and the votes become a probability. The paper page shows that probability with and without the paper you are reading.
What moves it? Add a study and watch which of the three changes.
Agreeing studies push the number up and shorten the bar. One disagreeing study in four is enough to make the claim contested; only enough agreeing studies to push the disagreement under a quarter lift it. Inconclusive studies change nothing.
What if a trial's results disagree with each other?
The main result is the one the trial promised to measure before it started. Secondary results are everything else it measured along the way; they can add support, never contradict on their own. Every vote answers one question: does the treatment help with this outcome?
Why
A trial's main, pre-declared result decides its vote. Secondary results can add support but never contradict on their own, because a trial tests many things and some will look bad by chance. A main result that found nothing counts against, because a trial built to find an effect and finding none is evidence.
What about “significant”?
A result is called significant when its p value is under 0.05, meaning a result that surprising would turn up by chance less than one time in twenty if there were no effect. That is a convention, not a law of nature: a lone trial at p = 0.04 is fragile, and the same finding twice is worth far more than one finding at p = 0.001. The map counts trials, not p values, for exactly that reason.
Why
This is a toy version of the rule; the real one, with its error models, is under Methodology. The second row is the whole point of the paper page: remove one study and see whether anything changes. On the map, a study is one evidence family, so a trial reported in five papers is still one square.
7Markers versus outcomes
Most diabetes trials measure blood sugar, not heart attacks. A drug can move the number and not the thing that matters. This is the most important idea on the page.
How the map uses this. A cell on a marker column is cheaper to fill but proves less. When the map proposes what to test next it says which kind of trial it is asking for, and when the outcome trial is huge it names the cheaper step: the marker trial that should come first, and why that is not the end of the story.
8Reading a cell of the map
Everything above meets in one square. Here is what its colour, its border and its corner number say, and why a cell with many trials can rest on very few.
One glyph carries all four. Fill colour is the lean of the results, brightness is how many trials, the border is how sure the belief is, and the corner number is the rank of what to test next.
Why a cell with 25 trials can rest on two. Registering a trial is required; posting its results is not. Of the trials that post, many report per-arm numbers with no comparison, or an outcome the map cannot read as for or against. Open any cell and the funnel shows the loss at each step.
- Retracted
- the journal withdrew it. It no longer votes.
- Expression of concern
- the journal has doubts and is looking. It still votes, flagged.
- Erratum
- a correction was published; usually small.
- Registry-linked trial
- the paper reports a registered trial, so it can vote.
- Outside this release
- the trial exists in the registry but is not in the current map build.
- Full text read
- the extractor read the whole paper, not only the abstract.
9Words you will meet
- Estimate
- The single best number a study found. The dot.
- Confidence interval
- The range the study cannot rule out, usually the middle 95%. The bar. For a ratio, judge it by top divided by bottom: under about 1.4 is a large trial, over 2.5 could not tell you much.
- Hazard ratio (HR)
- How fast the outcome arrives in the treated group compared with the comparison group, over time. 0.8 means 20% slower. 1 means no difference.
- Risk ratio (RR)
- How often the outcome happened in the treated group compared with the comparison group. Same reading as a hazard ratio.
- Odds ratio (OR)
- A cousin of the risk ratio used in some designs. Reads the same way; a bit more extreme for common outcomes.
- Mean difference
- For measured quantities such as HbA1c or weight: treated minus comparison. 0 is no effect.
- p value
- How surprising the result would be if there were truly no effect. Small is surprising. It says nothing about how big the effect is; the interval does.
- Primary outcome
- The one result a trial promised to measure before it started. It decides the trial's vote.
- Comparator
- What the treatment was compared with: placebo or usual care, or another drug, which the map calls head-to-head.
- Registered trial
- A trial listed on a public registry such as ClinicalTrials.gov before it ran, with its planned outcomes. Registering is required; posting results is not, which is why many cells have trials with no readable result.
- Evidence family
- One independent source of evidence: one trial, however many papers report it. The belief counts families, so a trial reported five times is still one vote.
- Belief
- The graph's probability that a claim is true, from its evidence families. Contested means a real minority disagrees. Replicated means independent families agree.
- Number needed to treat
- How many people must take the treatment for one to benefit. Under 10 is exceptional, 20 to 100 is normal for prevention, hundreds is common and still worthwhile for cheap, safe drugs.
- Standardised effect
- When a measure has no natural unit, the effect is scaled by its spread: 0.2 is small, 0.5 medium, 0.8 large. Most drug effects on symptoms and scores are 0.2 to 0.5.
- Blood pressure and LDL
- Two anchors worth knowing: every 5 mmHg of systolic pressure lowered is worth about 10% fewer major cardiovascular events; every 1 mmol/L (39 mg/dL) of LDL lowered is worth about 22% fewer major vascular events.
- Cell
- One square of the map: one treatment against one outcome.
- Randomised trial
- A coin decides who gets the treatment, so the groups differ only in the treatment. The only kind of study that votes.
- Observational study
- People and their doctors chose. Sicker or richer or more careful people end up in one group, and that, not the treatment, can explain the result. Called confounding.
- Placebo
- A dummy treatment, so that neither the patients nor the doctors know who got the real one. Removes the effect of expecting to get better.
- Surrogate, or marker
- A measurement that stands in for an outcome: HbA1c for diabetes complications, LDL for heart attacks. Cheap to move; whether the outcome follows must be shown, not assumed.
- Non-inferiority trial
- A head-to-head trial whose aim is to show the new drug is not meaningfully worse than the old. “No difference” is success there, which the map does not yet tell apart from a failed superiority trial.
- Follow-up
- How long people were watched. A hazard ratio over one year and one over ten are not the same claim; longer follow-up finds slower harms and benefits.
- Subgroup
- A result for part of the trial, such as people over 65. With ten subgroups, one will look striking by chance, so a subgroup result is a hint until a trial is run in that group.
- Why not one average
- Meta-analyses pool studies into one diamond. The map does not, because pooling hides disagreement and lets one large trial outvote everything; each independent trial keeps its own vote instead.
- Expected gain
- How much one more independent trial would move a cell's belief. Drives what the map proposes to test next.
- Settled
- A claim whose belief is at its ceiling; another trial of the same kind would not move it.
- Cheaper step
- The marker trial the map points at first when the outcome trial would need thousands of people and years.
The votes and belief demos are simplified illustrations of the rules; the exact rules, with their error models, are under Methodology. Try it on a real paper: metformin, aspirin and cancer mortality.