The Field Guide · No. 08
What a p-value really tells you (and what it doesn't)
A p-value measures how surprising the data would be if there were no real effect, not the probability that a finding is true.
Updated
The p-value is one of the most cited and most misunderstood numbers in science. In plain terms, it measures how surprising your data would be if there were actually no real effect at all. A small p-value means the data would be unlikely to look this way by pure chance, so researchers take that as a hint that something real might be going on.
The common cutoff is 0.05, and results below it get called "statistically significant." But that threshold is a convention, not a law of nature, and this is where people go wrong. A p-value is not the probability that the finding is true, and it is not the probability that the result is due to chance. It only speaks to how compatible the data are with the idea of no effect.
Two traps follow from this. First, a significant p-value says nothing about how big or important an effect is. A trivial difference can be significant in a large study. Second, a p-value just above 0.05 does not mean nothing is there, and one just below it does not mean something definitely is. Treating 0.05 as a magic on-off switch is exactly the mistake the statisticians who invented the tool warn against.
So when a headline leans on the word significant, remember what it does and does not mean. It suggests the pattern is probably not pure luck. It does not tell you the finding is certain, large, or important. Pair the p-value with the effect size and whether the result has been replicated, and you will read studies far more accurately than the headlines do.
What counts as a significant p-value
A p-value runs from 0 to 1. By convention, a result is called statistically significant when its p-value is below 0.05, which means data this extreme would show up less than 1 time in 20 if there were no real effect. Some fields and analyses use a stricter line such as 0.01.
The word only reports which side of a line the result fell on. A p of 0.049 and a p of 0.051 are nearly the same evidence, and one gets called significant and the other does not.
Where 0.05 came from
Ronald Fisher’s 1925 book Statistical Methods for Research Workers put the cutoff into circulation. He wrote that the value for which P = .05, or 1 in 20, is 1.96 or nearly 2, and that “it is convenient to take this point as a limit in judging whether a deviation is to be considered significant or not.” Statistician Lee Kennedy-Shaffer, tracing the history, points to the arithmetic convenience: 0.05 is roughly two standard deviations from the mean of a normal distribution, which mattered before computers, when researchers looked results up in printed tables. Fisher’s tables almost always included 0.05 and 0.01 columns.
Even then it was called a convention. L. H. C. Tippett wrote in 1931 that the 0.05 threshold was “quite arbitrary” but “in common use,” and Fisher himself later wrote that “no scientific worker has a fixed level of significance at which from year to year, and in all circumstances, he rejects hypotheses.”
What significance does not tell you
In 2016 the American Statistical Association issued a statement with six principles. Three of them: P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold. And a p-value, or statistical significance, does not measure the size of an effect or the importance of a result.
Its executive director, Ron Wasserstein, put it this way: “The p-value was never intended to be a substitute for scientific reasoning.” A significant result can still be tiny, and a result that misses the line can still be real but under-tested.
Nominal p-values
A nominal p-value is the one reported straight from a single test, at the usual 0.05 line, with no adjustment for how many other tests were run. Papers use the word to mark the difference from an adjusted p-value.
The adjustment matters because chances add up. Two independent outcomes each tested at the nominal 0.05 level give a 0.098 probability of at least one false positive, according to a statistical review of trials with several primary outcomes. Across 100 tests of things that are all null, about 5 will come out significant by luck alone. That is why a lone nominal p-value from one of many comparisons deserves less trust than the same number from the single test a study was planned around.
The stars: *, ** and ***
Asterisks are shorthand for p-value bands in tables and charts. A common set, used by the statistics program GraphPad Prism, is ns for p above 0.05, * for 0.05 or less, ** for 0.01 or less, and *** for 0.001 or less. Some settings add **** for 0.0001 or less, and others stop at three stars.
More stars mean a smaller p-value, which means the data would be more surprising if there were no effect. They do not mean a bigger effect. Definitions vary, so check the figure legend or table note.
What to read alongside it
The ASA notes that statisticians often supplement or replace p-values with methods that emphasize estimation over testing, such as confidence, credibility or prediction intervals. In practice, look for the size of the effect and the interval around it, then ask whether the finding has been replicated and whether the analysis was planned in advance.
What to remember
- A small p-value means the data would be surprising if there were no real effect.
- It is NOT the probability that the finding is true or that the result is due to chance.
- "Significant" (p under 0.05) is a convention, not proof, and says nothing about how big the effect is.
- By convention, p under 0.05 is called significant. Ronald Fisher suggested that line in 1925 as a convenience, and statisticians have called it arbitrary since.
- A nominal p-value has not been adjusted for how many things were tested. Test enough things and some will cross 0.05 by luck.
- Stars next to a p-value are a code: commonly * is p of 0.05 or less, ** is 0.01 or less, *** is 0.001 or less. Check the figure legend.
From the record
Informally, a p-value is the probability under a specified statistical model that a statistical summary of the data (e.g., the sample mean difference between two compared groups) would be equal to or more extreme than its observed value.
Asked often
Does a p-value below 0.05 mean the finding is true?
No. It means the data would be fairly unlikely if there were no real effect. That is a hint, not proof. The p-value does not give the probability that the hypothesis is true, and 0.05 is just a convention.
Does a small p-value mean a big effect?
No. A p-value only speaks to how compatible the data are with 'no effect.' A tiny, unimportant difference can be statistically significant in a large study, which is why you also need the effect size.
What p-value is significant?
By convention, a p-value below 0.05 is called statistically significant, and some fields use a stricter line such as 0.01. The line is a convenience Ronald Fisher suggested in 1925, not a law of nature, and the ASA advises against basing conclusions only on whether a p-value passes a threshold.
What is a nominal p-value?
It is the p-value from a single test as reported, without adjustment for the number of tests run. When many outcomes are tested, nominal p-values overstate the evidence: two outcomes tested at the nominal 0.05 level carry a 0.098 chance of at least one false positive.
What do the stars *, ** and *** mean next to a p-value?
They are a code for how small the p-value is. A common convention is * for p of 0.05 or less, ** for 0.01 or less, and *** for 0.001 or less, and some software adds **** for 0.0001 or less. The figure legend defines the stars for that paper.
Why is 0.05 the cutoff for significance?
Ronald Fisher suggested it in 1925 because P = .05 is about two standard deviations from the mean, which made it easy to use with printed tables. Later statisticians called it arbitrary but in common use.
Further reading
Go deeper
Paid link. If you buy a book through this link, Bookshop.org pays us a small commission and sends a share of the sale to independent bookstores.
-
Statistics Done Wrong (opens Bookshop.org)
Alex Reinhart · 2015
Statistician Alex Reinhart explains in short chapters what p-values do and do not mean, and the common errors that let unreliable results into published research.
Read the news better, every morning.
We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.
Free forever, no spam. One click to unsubscribe, and we never sell your email.