The Field Guide · No. 28
Statistical power: why a small study can miss a real effect
Statistical power is the probability that a study will detect a real effect if one truly exists, and a study with low power can miss it entirely.
Updated
Statistical power is the probability that a study will detect a real effect, assuming one actually exists. A study with high power is likely to find a true effect if it is there. A study with low power can easily come back empty-handed even when the effect it was looking for is real, simply because it did not have enough participants or observations to see it clearly through the noise.
Power depends mainly on three things: how big the true effect is, how much the measurements vary from person to person, and how many people or observations the study includes. Small samples are the most common reason a study ends up underpowered, since researchers can rarely make an effect bigger or make measurements less noisy, but they can, within budget and time limits, recruit more participants.
Two traps follow. First, a small study that finds no significant effect has not shown the effect is absent, only that the study lacked the power to detect it, a distinction headlines routinely blur. Second, when a small, underpowered study does report a significant result, that result tends to overstate the true effect size, because only an unusually large, and partly lucky, measured effect could clear the bar for significance in such a small sample.
When a study reports a finding, especially a surprising one from a small sample, check whether it states its power or its sample size relative to similar studies in the field. A field with a history of underpowered studies, such as early brain imaging research, tends to produce results that later fail to replicate at the same size. Pair a striking small-sample finding with a look at whether a larger, better powered study has tried to confirm it.
What to remember
- Statistical power is the probability a study detects a real effect, given that the effect actually exists.
- A small sample is the most common reason a study ends up underpowered and prone to missing a true effect.
- A significant result from a small, underpowered study tends to overstate how large the true effect really is.
From the record
Power is the probability of a study to make correct decisions or detect an effect when one exists.
Asked often
Does a study with no significant result prove there is no effect?
No. A study that finds no significant effect may simply have lacked the statistical power to detect a real one, especially if its sample was small. Absence of evidence in an underpowered study is not the same as evidence of absence, which is why sample size and power matter as much as the reported result itself.
Why do small studies sometimes report surprisingly large effects?
In a small, underpowered study, only an unusually large measured effect, partly driven by chance, is big enough to reach statistical significance. This tends to inflate the reported effect size compared with what a larger, better powered study of the same question would find.
Further reading
Go deeper
Paid link. If you buy a book through this link, Bookshop.org pays us a small commission and sends a share of the sale to independent bookstores.
-
Statistics Done Wrong (opens Bookshop.org)
Alex Reinhart · 2015
Alex Reinhart devotes a chapter to underpowered studies, showing how small samples miss real effects and exaggerate the ones they do find.
Read the news better, every morning.
We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.
Free forever, no spam. One click to unsubscribe, and we never sell your email.