The Field Guide · No. 45

How to read an election poll

An election poll is an estimate, with a margin of error around each candidate's share, and the gap between two candidates is roughly twice as uncertain as either share alone.

Updated

An election poll interviews a sample of voters and reports what share back each candidate. A sample is not everyone, so the result carries a margin of sampling error, and that margin is the first number to check before believing a headline that says one candidate is ahead. The American Association for Public Opinion Research (AAPOR), the professional association for US pollsters, calls it the price you pay for not talking to everyone in the population you are targeting.

A typical poll of about 1,000 people has a margin of roughly plus or minus 3 percentage points at the 95 percent confidence level. If Candidate A sits at 48 percent, the poll is consistent with a true figure anywhere from about 45 to 51. The catch is that the margin describes each candidate's share on its own. A lead is the difference between two shares that move in opposite directions, so its margin is about twice as large. Pew Research Center works the arithmetic for a 3-point margin and gets roughly 6 points for the gap. A 4-point lead in that poll could be anywhere from about 2 points behind to 10 points ahead.

Two traps follow. First, the margin covers only the luck of the draw in who got sampled. It says nothing about who was missed, who declined to answer, who will actually vote, or how the pollster adjusted the data, and Pew notes that recent studies put a poll's average total error at closer to twice what the reported margin implies. Second, a poll is a snapshot, not a forecast. AAPOR's guide for journalists says election polls are not meant to be predictive of an outcome, which is one reason two polls of the same race can disagree.

So read a poll in a fixed order: who was sampled and how many, when they were interviewed, what the margin covers, and whether this is one poll or a pattern across many. A lead inside the margin for the gap is unresolved, which is different from a tie and different from a call either way. Look at an average of polls before any single one, and look for the methodology statement that lets you check the rest.

What the margin of error covers

A margin of plus or minus 3 points at the 95 percent confidence level means that if the same survey were run 100 times, the result would land within 3 points of the true value about 95 of those times, in Pew's wording. AAPOR adds the other half: about five times in a hundred it will not, which is one reason even the best poll should be read cautiously, particularly when it differs sharply from similar polls fielded at about the same time.

Worked example (hypothetical): 1,000 likely voters, Candidate A at 48 percent and Candidate B at 44 percent
What is measuredPoll resultMargin of errorRange consistent with the poll
Candidate A's share48%about 3 points45% to 51%
Candidate B's share44%about 3 points41% to 47%
The lead (A minus B)A ahead by 4about 6 pointsB ahead by 2 to A ahead by 10
Assumes a simple random sample of 1,000 at 95 percent confidence. The margin for the lead works out to about 5.9 points, rounded to 6. Weighting widens both margins.

The table works one example with made-up numbers: a poll of 1,000 likely voters in which Candidate A has 48 percent and Candidate B has 44 percent. Each share gets the usual 3-point margin. The gap does not. If A's share is too high by chance, B's is likely too low, and the two errors stack, which is why Pew puts the margin for the difference at about twice that for one candidate.

The range for the lead runs from B ahead by about 2 to A ahead by about 10, so it includes zero. Two more limits apply. The margin belongs to the whole sample, and a subgroup has fewer people and a wider margin: in AAPOR's example, 200 people out of 1,000 carry a margin of plus or minus 6.9 points. And a margin of sampling error applies only to probability samples, where everyone has a known chance of being picked. For opt-in online polls AAPOR says it does not apply, and the plus or minus printed on one is a modeled credibility interval that AAPOR urges caution with.

Sample size, and why more helps less each time

The margin shrinks as the sample grows, but slowly. The American Statistical Association (ASA) gives the benchmarks for a poll estimating a proportion: a sample of 100 gives a margin of no more than about 10 points, 500 about 4.5, and 1,000 about 3. Getting down to 1.5 points would take a sample of well over 4,000. AAPOR makes the same point from the other end: doubling a sample from 1,000 to 2,000 trims the margin by only about a point.

Margin of error by sample size, for one candidate's share and for the lead between two candidates
RespondentsOne candidate's shareThe lead (about double)
100plus or minus 10 pointsplus or minus 20 points
500plus or minus 4.4 pointsplus or minus 9 points
1,000plus or minus 3.1 pointsplus or minus 6 points
2,000plus or minus 2.2 pointsplus or minus 4 points
4,000plus or minus 1.5 pointsplus or minus 3 points
Simple random sample, 95 percent confidence, share near 50 percent: 1.96 times the square root of 0.25 divided by the sample size. The lead column doubles the first and rounds. The ASA's published figures for 100, 500 and 1,000 people (about 10, 4.5 and 3 points) are rounded versions of the same formula.

Size is also not quality. AAPOR's journalist guide says a larger sample is not necessarily better, because other factors may matter more. A huge sample drawn from the wrong people still lands on the wrong answer, and no margin of error warns you.

Registered voters, likely voters and the screen

Election polls face a problem that ordinary opinion polls do not. Pew's methods team describes it as producing a model of a population that does not yet exist at the time the poll is conducted: the future electorate. People who say they will vote often do not, and some who say they will not end up voting. A poll of registered voters counts everyone on the rolls. A poll of likely voters tries to count only those who will turn out, so the two can report different numbers from the same race.

The sorting is done with a likely-voter screen. Pew reports that most pollsters ask several questions that together estimate each person's chance of voting: intention to vote, past voting, knowledge of the voting process and interest in the campaign. A deterministic screen applies a cutoff, so each respondent is a likely voter or is not. A probabilistic screen gives each respondent a probability of voting and weights answers by it. In Pew's comparison every method forecast better than counting all registered voters or all who intend to vote, though some did better than others. Pew also notes that pollsters often rely on past turnout and on current measures of enthusiasm, which are educated guesses about a crowd no one has seen yet. When you read a poll, check which group it reports and whether a screen is described.

Weighting, and what it costs

Weighting is a statistical adjustment that makes the sample match the population on traits such as age, race and education. Pew's example: if a survey has too many college graduates compared with their share of the population, people without a degree are weighted up to the proper share. AAPOR's guide adds that weighting should use outside benchmarks, such as Census Bureau data, that fit the group sampled.

Pew's researchers report a growing view that weighting on only a few variables is not enough, and that adjusting on more of them has produced more accurate results in Pew's own studies. Weighting also has a cost. Pew's Andrew Mercer calls it a crucial step for avoiding biased results that nonetheless makes the margin of error larger, an increase statisticians call the design effect. A pollster who reports the simple margin without that adjustment is claiming more precision than the survey earned. As an illustration, with a design effect of 1.5 the plus or minus 3.1 for 1,000 people grows to about 3.8. Members of AAPOR's Transparency Initiative must disclose how they weighted and whether the margin accounts for it.

Nonresponse

Nonresponse means someone sampled for a survey does not take part. The ASA calls it nearly inevitable, because some people will refuse despite every reasonable effort. It becomes a problem when who answers differs from who does not. Pew's example is that college graduates are more likely to participate, so a raw sample can contain too many of them.

Weighting can correct the traits a pollster can measure and compare with an outside benchmark, such as education. Where responders differ from nonresponders in ways no benchmark captures, the bias stays, which is why nonresponse sits among the errors the margin leaves out. Pew separates it from noncoverage, where part of the target population had no chance of being sampled, and from measurement error, where people misunderstand a question or misreport an answer.

Why two polls of the same race differ

Start with luck. Each poll has its own sampling error, and the difference between two polls carries more uncertainty than either alone. Pew's example is two polls with a 3-point margin each: a lead that moves from 5 points to 8 has a margin of plus or minus 8 on that 3-point change, so two surveys alone cannot reliably separate real change from noise. The table shows two hypothetical polls of 900 people each. Their leads differ by 4 points, and the margin on that difference is about plus or minus 9, so sampling luck alone could produce the gap.

Two hypothetical polls of the same race, each with 900 respondents
FeaturePoll 1Poll 2
Who was sampledRegistered votersLikely voters
How they were reachedTelephone and textOnline panel
Field datesOct. 6 to 9Oct. 10 to 14
ResultA 47, B 45 (A ahead by 2)A 46, B 48 (B ahead by 2)
Margin for the leadabout 6.5 pointsabout 6.5 points
The two leads differ by 4 points. The margin on the difference between two leads is about 1.4 times 6.5, or about 9 points, so sampling error alone can explain the gap before any method difference is counted.

Method differences stack on top. Polls can sample different groups, interview by different modes, ask different questions, field on different dates and weight differently. A consistent lean attributable to a pollster's procedures is called a house effect. Stanford political scientist Simon Jackman defined it as bias specific to a polling organisation, and in his study of the 2004 Australian election polls found that a good portion of the differences across polls came from differences in methodology rather than from movements in voter sentiment. A house effect describes a procedure's tendency, and it does not say which poll is closer to the truth.

Averages beat single polls, with a catch

Pew's rule of thumb is that trends across a number of polls give more confidence than one or two. A single poll can also be an outlier by chance. Pew cited a tracker that listed 590 national polls for one past presidential race and noted that at 95 percent confidence about 5 percent of them, around 30, would be expected to miss the true value by more than the margin. Those outliers draw attention because they imply a dramatic change, which is why Pew advises waiting to see whether a surprising result shows up in other surveys.

Pooling raises precision, and the catch is bias. Jackman writes that pooling always enhances precision, but that the validity of the pooled estimate rests on the assumption that the polls are unbiased, which is often not the case. An average of many polls that share the same blind spot inherits it. This is the same logic as pooling studies in a meta-analysis, and the same limits.

What a lead inside the margin does and does not mean

It means the poll cannot rule out, at the usual 95 percent confidence, that the trailing candidate is ahead among all voters. Pew's test is stricter than many headlines: being ahead by more than the single-candidate margin is not enough. With a 3-point margin, a candidate needs a lead of about 6 points before sampling error is an unlikely explanation. AAPOR's guide puts it as ahead by 1.5 to 2 times the margin.

It does not mean a tie. The estimate still points toward the leader, and the interval is a range around it, so values near the observed lead are more plausible than values at the edge. It does not mean the poll is wrong, and it does not mean the result will repeat. The reverse also holds: a lead outside the margin is no guarantee, because the margin covers only sampling error and about 1 poll in 20 misses by more than it. As with any confidence interval, the useful move is to read the whole range.

How to find a poll's methodology statement

A credible poll publishes how it was done. Pew lists the essentials: the sponsor, the data collection firm, where and how participants were selected, the mode of interview, field dates, sample size, question wording and weighting procedures. AAPOR's guide asks the same questions in groups: what was asked, who conducted and paid for it, when it was fielded, and how. It also asks whether the respondents were real people.

Look first on the pollster's own page for the release: a methodology note, a toplines or crosstabs file, or a footnote under the chart. AAPOR runs a Transparency Initiative, and a member organization commits to release these disclosure elements when it releases findings and to provide more detail within 30 days of a request. AAPOR keeps a list of certified organizations on its website. Membership is not a grade. AAPOR says it makes no judgment about the quality of the methods disclosed, and Pew calls participation a positive signal but not a guarantee of rigor. A poll that gives no sample size, no dates and no sponsor has not given you enough to check, and it earns less weight.

What to remember

From the record

To determine whether or not the race is too close to call, we need to calculate a new margin of error for the difference between the two candidates' levels of support. The size of this margin is generally about twice that of the margin for an individual candidate.

Andrew Mercer principal methodologist, Pew Research Center 5 key things to know about the margin of error in election polls, 2016

Asked often

What does the margin of error in a poll mean?

It describes how far a poll's result is likely to fall from the true value because only a sample was interviewed. A margin of plus or minus 3 points at the 95 percent confidence level means that if the poll were repeated 100 times, the result would land within 3 points of the truth about 95 times. It applies to each candidate's share and covers sampling luck only.

If a candidate leads by less than the margin of error, is the race tied?

No. The lead is still the best estimate, but the poll cannot rule out that the other candidate is ahead. The margin for the gap is about twice the reported margin, so with a 3-point margin a lead needs to be about 6 points before sampling error alone is an unlikely explanation.

Why do two polls of the same race show different results?

Sampling luck, for one: the difference between two polls is more uncertain than either poll. Beyond that, polls differ in who they sample (registered or likely voters), how they interview, when they field, how they weight and how they screen for likely voters. A consistent lean traceable to a pollster's methods is called a house effect.

Read the news better, every morning.

We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.