Reading Polls and Elections
How polls work: sampling, margin of error, likely-voter models, and why a poll is a snapshot, not a prediction.
Statistics for Everyday Life · Lesson 2
How polls work: sampling, margin of error, likely-voter models, and why a poll is a snapshot, not a prediction.
Polls shape how we see elections, public opinion, and each other, and they are quoted with a confidence the numbers rarely earn. A single point of movement becomes a headline; a lead inside the margin of error is called a "surge." Reading polls well means knowing what they can and cannot tell you.
The core idea is genuinely surprising: a carefully chosen sample of roughly a thousand people can estimate the views of millions to within a few percentage points. But that near-magic depends entirely on the sample being representative. When it is not, size does not rescue it — a lesson the polling industry learned the hard way and relearns every few years.
A poll works by drawing a sample that resembles the population in the ways that matter — age, region, education, and so on. If every member of the population has a known, roughly equal chance of being included, the sample's numbers track the whole. Representativeness, not raw size, is what makes a poll trustworthy. A biased sample of millions is worse than a representative sample of a thousand, because its errors do not cancel out; they compound.
Pollsters report a margin of error — often about plus or minus 3 points for a sample near 1,000 — meaning that if the poll were repeated many times, the estimate would usually land within that range of the true value. Two crucial caveats: the margin applies to each candidate's number, so the margin on the gap between two candidates is larger; and the margin captures only random sampling error. It says nothing about bias from who refused to answer, how questions were worded, or who was reached in the first place.
Not everyone who is polled will vote, so pollsters build likely-voter models — screens and weights that estimate who will actually turn out. These models rest on assumptions about enthusiasm and past behaviour, and reasonable pollsters make different choices. Much of the spread between polls of the same race comes from these hidden modelling decisions, not from real swings in opinion.
A poll measures opinion at the moment it was taken. It is a photograph, not a forecast. People change their minds, undecided voters break late, and turnout varies. Treating a poll as a prediction of a future election ignores everything that can happen between the fieldwork and the vote — and ignores the poll's own stated uncertainty.
A poll of 1,000 voters reports Candidate A at 48 percent and Candidate B at 45 percent, margin of error plus or minus 3 points. Is A winning? Each figure could plausibly sit 3 points either way, so A's true support might be 45 and B's 48. The 3-point lead is within the margin — statistically a tie. Because the margin on the difference between two numbers is larger than the margin on either one alone, the honest reading is "too close to call," not "A leads." Reporting it as a lead overstates what 1,000 interviews can show.
A poll being "wrong" does not prove it was bad. In 2016, national US polls showed Hillary Clinton ahead by about 3 points on average; she won the national popular vote by about 2 points, comfortably within normal error. The widespread sense that "the polls failed" came from treating national numbers as state-level predictions, from ignoring the real chance the polls themselves assigned to a Trump win, and from reading a snapshot as a certainty. The national polls were roughly right about opinion; the reading of them was wrong.
In 1936 the American magazine The Literary Digest ran a mail poll of unprecedented scale, sending out about 10 million ballots and tallying roughly 2.4 million returns. On that mountain of data it predicted that Republican Alf Landon would beat President Franklin D. Roosevelt, about 57 to 43 percent. The election was a Roosevelt landslide: he won about 61 percent of the popular vote and every state but two. The Digest's sample was drawn largely from telephone directories, magazine subscribers, and automobile-registration lists — during the Depression, these skewed toward wealthier Americans more likely to favour Landon. Non-response made it worse, since those who bothered to return ballots differed from those who did not. Meanwhile George Gallup, using a far smaller but deliberately representative sample of around 50,000 people, correctly predicted Roosevelt's victory — and even predicted, within a couple of points, the wrong answer the Digest would publish. The Digest folded soon after; Gallup's approach helped found modern scientific polling. The enduring lesson: representativeness beats raw size.
Find a recent poll reported in the news. Locate four things the headline probably omitted: the sample size, the margin of error, the dates of the fieldwork, and how "likely voters" were defined. Then decide whether the headline's claim still survives once you know them.
Think Like a Maester: Ask not how many people a poll reached, but whether the ones it reached resemble everyone it left out.
A poll is an estimate of opinion built from a sample, and its value rests almost entirely on that sample being representative — not on how large it is. Margin of error describes only random sampling error, is wider for the gap between two candidates than for either alone, and says nothing about bias, wording, or non-response. Likely-voter models add another layer of assumption, and much of the visible disagreement between polls comes from these choices rather than from real shifts. Above all, a poll is a snapshot of the present, not a prediction of the future. The 1936 contest between the Literary Digest's millions and George Gallup's representative thousands remains the clearest proof that in sampling, who you ask matters more than how many.
Mark this lesson complete to track your progress.