Causal Inference: From Correlation to Cause
Move beyond 'correlation is not causation': how randomized trials, natural experiments, and confounder control establish real causes.
Statistics for Everyday Life · Lesson 4
Move beyond 'correlation is not causation': how randomized trials, natural experiments, and confounder control establish real causes.
"Correlation is not causation" is a warning, not a method. It tells you what not to believe, but it leaves the real question unanswered: how do we ever know that one thing causes another? Doctors, governments, and businesses cannot wait for certainty. They must act on causes. Will this drug help? Will this policy raise wages? Will this change to a website sell more books?
Over the past century, statisticians built tools to answer causal questions from data. This lesson is about three of them: the randomized experiment, the natural experiment, and the careful control of confounders. Together they turn a cautious slogan into a working discipline, one that has reshaped medicine, economics, and public policy.
A confounder is a third factor that influences both the supposed cause and the effect, creating a link that is real but misleading. Ice-cream sales and drowning both rise together, but neither causes the other; hot weather drives both. The whole difficulty of causal inference is ruling out confounders you can name and, worse, the ones you cannot.
The cleanest solution is to assign the treatment at random. If a coin flip decides who gets the new drug and who gets a placebo, then the two groups are, on average, alike in every other respect, measured or not. Any difference in outcome can then be attributed to the treatment. R. A. Fisher developed this logic for agriculture at Rothamsted in the 1920s, randomly assigning treatments to plots so that differences in soil would average out. Randomization is powerful precisely because it balances the confounders you never thought to record.
When deliberate randomization is impossible or unethical, sometimes nature, policy, or history assigns people to groups in a way that is close to random. A researcher who spots such a moment can compare the groups as if an experiment had been run. This is the natural experiment, and much of modern economics is built on finding good ones.
A company believes a new tutoring program raises test scores. Students who enrolled did score higher, but they may also have been more motivated, a confounder. To settle it, the company randomizes: among students who apply, a lottery decides who gets a place. Because the lottery is blind to motivation, the enrolled and non-enrolled groups start out comparable. If the enrolled group later scores higher, the program, not the motivation, is the credible cause.
Controlling for confounders in observational data is useful but limited. Suppose you study whether coffee causes heart trouble and you carefully adjust for age, weight, and exercise. Your estimate can still be wrong if smokers happen to drink more coffee and you failed to measure smoking. Statistical adjustment can only remove the confounders you actually recorded. Randomization removes them all at once, which is why it remains the gold standard when it is feasible.
In the 1920s at Rothamsted Experimental Station, R. A. Fisher made randomization the cornerstone of experimental design, and his 1935 book, The Design of Experiments, spread the method well beyond farming.
Decades earlier, in 1854, the London physician John Snow investigated a cholera outbreak. He mapped deaths around the Broad Street water pump and compared households supplied by two water companies, one drawing sewage-tainted water and one cleaner water. Because households were served by one company or the other for reasons unrelated to their health, the comparison worked like an early natural experiment, pointing to contaminated water rather than bad air.
In 2021 the Sveriges Riksbank Prize in Economic Sciences in Memory of Alfred Nobel went to David Card, Joshua Angrist, and Guido Imbens. Card was honored for empirical contributions to labour economics, including minimum-wage studies with Alan Krueger, and Angrist and Imbens for methodological contributions to analyzing causal relationships, formalizing what natural experiments can and cannot reveal.
For each claim, name a plausible confounder and then a design that would remove it. "People who take vitamins live longer." "Cities with more police have more crime." "Students who use the library get better grades." For each, ask: what randomization or natural experiment would let you see the true effect?
Think Like a Maester: The question is never whether two things move together, but whether anything other than the cause could have made them move together.
Moving from correlation to cause takes more than a warning; it takes a method. Randomized controlled trials, rooted in Fisher's agricultural work, break the link between treatment and confounders by assigning treatment at random. Natural experiments, from John Snow's cholera map to the work honored by the 2021 Nobel Prize, exploit near-random variation when true randomization is impossible. Controlling for confounders helps but only for the factors you measured. Choosing the right design is how data earns the right to speak of causes.
Mark this lesson complete to track your progress.