Beyond p-values: Effect Sizes and Intervals
Why p-values alone mislead: effect sizes tell you how big, and confidence intervals how precise. A plain-language guide beyond significance.
Statistics for Everyday Life · Lesson 5
Why p-values alone mislead: effect sizes tell you how big, and confidence intervals how precise. A plain-language guide beyond significance.
The previous lessons kept returning to a caution: a p-value tells you whether an effect is surprising under chance, but not how big or important it is. This lesson gives you the tools that fill that gap — effect sizes and confidence intervals — and they transform how you read any statistical result.
A p-value answers one narrow question and is easily over-read. It depends heavily on sample size: with enough data, trivial effects become 'significant'; with too little, real effects can be missed. On its own, 'p < 0.05' tells you almost nothing about whether a result matters.
An effect size measures the magnitude of a difference or relationship — how large the gap between groups is, in meaningful terms. A weight-loss pill might show a 'statistically significant' effect that amounts to half a pound; the effect size reveals it is practically useless. Always ask not just 'is there an effect?' but 'how big?'
A confidence interval gives a range of plausible values for the true effect, rather than a single point. A narrow interval signals a precise estimate; a wide one signals great uncertainty. An interval that ranges from 'tiny' to 'huge' is honest about how little we really know, in a way a lone p-value hides.
Two studies both report a 'significant' benefit. Study A: effect size large, confidence interval narrow and well away from zero — convincing. Study B: effect size tiny, interval barely excluding zero and stretching to trivially small — technically significant, practically weak. The p-values might look similar; the effect sizes and intervals tell you which result to trust and act on.
Rejecting p-values entirely is an over-correction. Used properly alongside effect sizes and intervals, they are informative. The problem is not the p-value itself but treating it as the only thing that matters — the 'bright line' of 0.05 as a verdict of truth. The fix is to report and read all three: significance, magnitude, and precision.
Concern about the misuse of p-values grew so widespread that in 2016 the American Statistical Association (ASA) took the unusual step of issuing a formal statement on statistical significance and p-values. It warned, among other things, that a p-value does not measure the probability that a hypothesis is true or the size of an effect, that 'statistical significance' is not the same as scientific or practical importance, and that decisions should not be based on whether a p-value passes a fixed threshold. The statement, and the broader 'new statistics' movement led by researchers such as Geoff Cumming (and earlier by Jacob Cohen, who long championed effect sizes), urged reporting effect sizes and confidence intervals to convey magnitude and uncertainty. The episode shows a field openly correcting a pervasive bad habit — a healthy example of science policing its own methods.
For a result you read, ask: (1) is it statistically significant, (2) how big is the effect, and (3) how precise (how wide is the interval)? Which question does the headline usually skip?
Think Like a Maester: Never let 'significant' be the last word — ask how big the effect is and how sure we can be.
A p-value shows whether an effect is surprising under chance, but not its size or importance — and it depends heavily on sample size. Effect sizes report how big an effect is; confidence intervals report how precise the estimate is. Reading all three together prevents mistaking statistical significance for practical importance. The ASA's 2016 statement and the 'new statistics' movement formalised this correction, urging magnitude and uncertainty over a bright-line p-value.
Mark this lesson complete to track your progress.