P-Hacking and the Replication Crisis
How good-faith research produces unreliable results: p-values, p-hacking, publication bias, and the replication crisis.
Statistics for Everyday Life · Lesson 5
How good-faith research produces unreliable results: p-values, p-hacking, publication bias, and the replication crisis.
Most people meet statistics through published research: a study 'finds' that a diet, a drug, or a teaching method works. We tend to trust that if a result cleared the bar of statistical significance and survived peer review, it must be solid. Increasingly, that trust is being tested from inside science itself.
Since the mid-2000s, statisticians and researchers have shown that a large share of published findings do not hold up when the study is repeated. The causes are rarely fraud. Far more often, honest researchers misunderstand their tools, make many small choices that quietly inflate the chance of a false positive, and publish in a system that rewards exciting results over careful ones.
A p-value is the probability of seeing data at least as extreme as yours if the null hypothesis (no real effect) were true. It is a statement about data assuming no effect. It is not the probability that the hypothesis is true, and not the probability your result happened 'by chance'. A small p-value, say below 0.05, means the data would be surprising under the null, no more. The threshold 0.05 is a convention, not a law of nature.
Analysing real data involves dozens of choices: which outliers to drop, which variables to control for, when to stop collecting data, which of several outcomes to report. Each choice can be defensible, but trying many combinations and keeping the one that crosses p < 0.05 is called p-hacking. It manufactures significance out of noise. Simmons, Nelson, and Simonsohn called this flexibility 'researcher degrees of freedom' in their 2011 paper 'False-Positive Psychology'.
Journals prefer positive, novel results. Studies that find 'no effect' often go unpublished, the so-called file-drawer problem. So the published literature is a biased sample of all the research done, over-representing flukes and exaggerated effects.
Suppose you test 20 genuinely useless supplements for an effect on mood, each at the 0.05 threshold. Even though none works, the expected number of 'significant' results is 20 x 0.05 = 1. Run enough independent tests and a false positive is nearly guaranteed. If only that one lucky test gets written up and published, readers see a 'breakthrough' with no idea that 19 silent failures sit behind it.
Not every failed replication means the original was wrong, and not every small p-value is p-hacking. A well-designed study with a pre-registered hypothesis, an adequate sample size, and a single planned analysis gives a p-value you can take at face value. The problem is not significance testing itself; it is undisclosed flexibility and selective reporting. Transparency, not abandoning statistics, is the fix.
In 2005 the epidemiologist John Ioannidis published 'Why Most Published Research Findings Are False' in PLoS Medicine, arguing with a simple model that in many fields the majority of claimed discoveries are likely false positives, given small studies, small effects, and flexible analysis. It became one of the most-cited papers in its field.
In 2015 the Open Science Collaboration repeated 100 psychology studies and reported the results in Science. Where 97 of the originals had reported significant results, only about 36 of the replications did, and effect sizes shrank by roughly half.
In 2016 the American Statistical Association issued a formal statement on p-values, its first such statement on a specific method. It warned that a p-value does not measure the probability that a hypothesis is true, that 'significance' should not by itself drive conclusions, and that p < 0.05 should never be treated as a bright line for truth.
Take a claimed finding and list the analytic choices the researchers could have made differently. Ask: was the hypothesis stated before the data were seen? How many outcomes were measured? Would the result survive a slightly different but equally reasonable analysis?
Think Like a Maester: A single significant result is a question, not an answer, so ask how many analyses it took to find.
Statistical significance is a useful signal, not a certificate of truth. Good-faith research can still produce unreliable findings when p-values are misunderstood, when many analyses are tried until one 'works', and when only positive results get published. Ioannidis's 2005 paper, the 2015 replication of 100 psychology studies, and the ASA's 2016 statement all point the same way: treat single significant results with humility, and place more trust in findings that are pre-registered, adequately powered, transparently reported, and independently replicated.
Mark this lesson complete to track your progress.