The Field Guide · No. 29

Multiple comparisons: how testing enough things finds something

The multiple comparisons problem is that running enough separate statistical tests will turn up some significant results by chance alone, even in pure noise.

Updated

The multiple comparisons problem is that testing many things at once raises the odds that at least one test will look statistically significant purely by chance, even if nothing real is going on. A single test run at the conventional 5 percent significance threshold has a 1 in 20 chance of a false positive. Run 20 such tests on random noise and, on average, one will come back looking significant anyway.

The field's most memorable illustration came from neuroscientists who placed a dead Atlantic salmon, bought whole from a supermarket, into a functional MRI scanner and showed it photographs of people in social situations. Scanning roughly 130,000 tiny brain regions, or voxels, without correcting for the sheer number of tests produced a cluster of voxels in the dead fish that appeared to respond to the images. Applying the standard corrections for multiple testing made that apparent brain activity disappear completely.

Two traps follow. First, an analysis that scans many outcomes, subgroups, or brain regions for anything significant, without adjusting the threshold for how many tests were run, is set up to find noise and report it as a discovery. Second, correcting properly for multiple comparisons is a known, standard step in fields like brain imaging and genetics, so a striking result that skipped it should be treated as unproven rather than simply cautious.

When a study reports a significant finding among many things it measured, such as one gene among thousands or one brain region among many, check whether it corrected its statistical threshold for the number of tests run. A result that survives a proper multiple comparisons correction is far more trustworthy than one plucked from a long list of comparisons. Pair any single striking finding with a look at how many other comparisons the same study quietly ran and did not report.

What to remember

From the record

Statistics that were uncorrected for multiple comparisons showed active voxel clusters in the salmon's brain cavity and spinal column. Statistics controlling for the family-wise error rate (FWER) and false discovery rate (FDR) both indicated that no active voxels were present, even at relaxed statistical thresholds.

Craig M. Bennett, Abigail A. Baird, Michael B. Miller, George L. Wolford Neural correlates of interspecies perspective taking in the post-mortem Atlantic Salmon: an argument for proper multiple comparisons correction, 2010

Asked often

What was the dead salmon fMRI study actually testing?

Neuroscientists placed a dead Atlantic salmon in a functional MRI scanner and showed it photographs of people, then ran the standard analysis used in brain imaging research across roughly 130,000 brain regions. Without correcting for the number of tests, the analysis found a cluster that appeared to respond to the images, purely from noise, since the salmon was dead throughout.

How do researchers correct for the multiple comparisons problem?

Common approaches include the Bonferroni correction, which tightens the significance threshold in proportion to how many tests were run, and false discovery rate methods, which control the expected share of false positives among the results called significant. Applying either method to the salmon data eliminated the apparent brain activity entirely.

Further reading

Go deeper

Paid link. If you buy a book through this link, Bookshop.org pays us a small commission and sends a share of the sale to independent bookstores.

  1. How Not to Be Wrong (opens Bookshop.org)

    Jordan Ellenberg · 2014

    Jordan Ellenberg shows how testing many possibilities turns up impressive-looking coincidences, with examples ranging from a stock-picking scam to research that finds patterns in noise.

Read the news better, every morning.

We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.