MegaMaester

Statistics for Everyday Life · Lesson 5

P-Hacking and the Replication Crisis

beginner16 min · 13 cards
Start here

P-Hacking and the Replication Crisis

How good-faith research produces unreliable results: p-values, p-hacking, publication bias, and the replication crisis.

Concept 1 of 10

Why this matters

Most people meet statistics through published research: a study 'finds' that a diet, a drug, or a teaching method works. We tend to trust that if a result cleared the bar of statistical significance and survived peer review, it must be solid. Increasingly, that trust is being tested from inside science itself.

Since the mid-2000s, statisticians and researchers have shown that a large share of published findings do not hold up when the study is repeated. The causes are rarely fraud. Far more often, honest researchers misunderstand their tools, make many small choices that quietly inflate the chance of a false positive, and publish in a system that rewards exciting results over careful ones.

Concept 2 of 10

Core concepts

What a p-value actually means

A p-value is the probability of seeing data at least as extreme as yours if the null hypothesis (no real effect) were true. It is a statement about data assuming no effect. It is not the probability that the hypothesis is true, and not the probability your result happened 'by chance'. A small p-value, say below 0.05, means the data would be surprising under the null, no more. The threshold 0.05 is a convention, not a law of nature.

P-hacking and researcher degrees of freedom

Analysing real data involves dozens of choices: which outliers to drop, which variables to control for, when to stop collecting data, which of several outcomes to report. Each choice can be defensible, but trying many combinations and keeping the one that crosses p < 0.05 is called p-hacking. It manufactures significance out of noise. Simmons, Nelson, and Simonsohn called this flexibility 'researcher degrees of freedom' in their 2011 paper 'False-Positive Psychology'.

Publication bias

Journals prefer positive, novel results. Studies that find 'no effect' often go unpublished, the so-called file-drawer problem. So the published literature is a biased sample of all the research done, over-representing flukes and exaggerated effects.

Concept 3 of 10

Worked example

Suppose you test 20 genuinely useless supplements for an effect on mood, each at the 0.05 threshold. Even though none works, the expected number of 'significant' results is 20 x 0.05 = 1. Run enough independent tests and a false positive is nearly guaranteed. If only that one lucky test gets written up and published, readers see a 'breakthrough' with no idea that 19 silent failures sit behind it.

Concept 4 of 10

Counterexample

Not every failed replication means the original was wrong, and not every small p-value is p-hacking. A well-designed study with a pre-registered hypothesis, an adequate sample size, and a single planned analysis gives a p-value you can take at face value. The problem is not significance testing itself; it is undisclosed flexibility and selective reporting. Transparency, not abandoning statistics, is the fix.

Concept 5 of 10

Case study: Ioannidis, the ASA, and the replication projects

In 2005 the epidemiologist John Ioannidis published 'Why Most Published Research Findings Are False' in PLoS Medicine, arguing with a simple model that in many fields the majority of claimed discoveries are likely false positives, given small studies, small effects, and flexible analysis. It became one of the most-cited papers in its field.

In 2015 the Open Science Collaboration repeated 100 psychology studies and reported the results in Science. Where 97 of the originals had reported significant results, only about 36 of the replications did, and effect sizes shrank by roughly half.

In 2016 the American Statistical Association issued a formal statement on p-values, its first such statement on a specific method. It warned that a p-value does not measure the probability that a hypothesis is true, that 'significance' should not by itself drive conclusions, and that p < 0.05 should never be treated as a bright line for truth.

Concept 6 of 10

Common misconceptions

  • 'p < 0.05 means there is a 95% chance the effect is real.' No. The p-value assumes the null is true and says nothing directly about that probability.
  • 'A failed replication proves the first study was fraud.' Usually it reflects noise, small samples, or hidden flexibility, not misconduct.
  • 'Peer review guarantees a result is reliable.' Review catches some errors but cannot see the analyses that were tried and quietly dropped.
  • 'Non-significant means no effect.' It often means the study was too small to detect one.
Concept 7 of 10

Interactive challenge — Spot the p-hack

Take a claimed finding and list the analytic choices the researchers could have made differently. Ask: was the hypothesis stated before the data were seen? How many outcomes were measured? Would the result survive a slightly different but equally reasonable analysis?

Think Like a Maester: A single significant result is a question, not an answer, so ask how many analyses it took to find.

Concept 8 of 10

Knowledge check

  1. In plain terms, what does a p-value of 0.03 tell you?
  2. Why does testing many hypotheses raise the chance of a false positive?
  3. What is publication bias, and how does it distort the literature?
  4. What did the 2015 Open Science Collaboration find when repeating 100 psychology studies?
  5. What was the American Statistical Association's 2016 statement warning against?
Concept 9 of 10

Lesson summary

Statistical significance is a useful signal, not a certificate of truth. Good-faith research can still produce unreliable findings when p-values are misunderstood, when many analyses are tried until one 'works', and when only positive results get published. Ioannidis's 2005 paper, the 2015 replication of 100 psychology studies, and the ASA's 2016 statement all point the same way: treat single significant results with humility, and place more trust in findings that are pre-registered, adequately powered, transparently reported, and independently replicated.

Quick check

What two historical activities are usually described as the early roots of statistics?