The Field Guide · No. 36
Sensitivity and specificity: why a good test can still mislead
Sensitivity and specificity describe how well a test tells sick from healthy, but neither one tells you the chance a positive result is real, which also depends on how common the condition is.
Updated
Sensitivity and specificity describe how good a test is at telling sick from healthy, and each looks at only one side of that question. Sensitivity is the share of people who truly have a condition that the test correctly flags as positive. Specificity is the share of people who truly do not have it that the test correctly clears as negative. A good screening test needs both numbers to be high, but even a test that scores well on both can behave very differently depending on who takes it.
Mammography is a well documented case. A set of statistics used to train gynecologists put the sensitivity of mammography at 90 percent, meaning 9 out of 10 women with breast cancer test positive, and the false-positive rate at 9 percent, meaning specificity of about 91 percent. Numbers like these sound reassuring, and most people, including most doctors, assume a positive result from a test this accurate must mean the disease is very likely. That assumption is where the confusion starts.
Two traps follow. First, sensitivity and specificity do not answer the question a patient actually has, which is: given my positive result, what is the chance I have the disease? That figure, the positive predictive value, needs a third fact neither statistic supplies: how common the condition is in the group being tested, its base rate. Second, when a condition is rare, even a good test produces mostly false alarms. Out of 1,000 women screened at the 1 percent prevalence used in the mammography training, about 10 actually have breast cancer and 9 of them test positive. Of the 990 who do not, about 89 test positive anyway, so of the 98 positive results, only 9 are real, about 1 in 10.
So when a headline touts a test's sensitivity or specificity, ask the question those numbers cannot answer on their own: how common is the condition in the people being tested? A test can be excellent by both measures and still be wrong most of the time when it is used to screen a population where the condition is rare, which is exactly why deciding whether to screen a whole population for something uncommon is such a hard call for medical guideline bodies. Pair sensitivity and specificity with the base rate, and count people out of 1,000 rather than trusting a percentage alone.
What to remember
- Sensitivity is the share of people with a condition a test correctly flags; specificity is the share without it the test correctly clears.
- Neither number is the answer a patient wants: the chance a positive result is real also depends on how common the condition is.
- At 1% prevalence, a test with 90% sensitivity and 91% specificity can still mean only about 1 in 10 positive results is a true case.
From the record
Only about 1 out of every 10 women who test positive in screening actually has breast cancer.
Asked often
If a test is 90% sensitive and 91% specific, does a positive result mean I probably have the disease?
Not necessarily. Those two numbers alone cannot answer that question; you also need to know how common the disease is in people like you. In a mammography training example built on real screening statistics, a 1% cancer prevalence combined with 90% sensitivity and 91% specificity meant that only about 1 in 10 women with a positive result actually had breast cancer.
Why do good tests still produce so many false positives for rare conditions?
Because a specificity of 91% still lets 9% of everyone without the disease test positive, and when almost everyone being tested does not have the disease, that 9% adds up to more false alarms than true cases. Out of 1,000 women screened at 1% prevalence, about 89 healthy women test positive against only 9 women who truly have cancer.
Further reading
Go deeper
Paid link. If you buy a book through this link, Bookshop.org pays us a small commission and sends a share of the sale to independent bookstores.
-
Risk Savvy (opens Bookshop.org)
Gerd Gigerenzer · 2014
Gerd Gigerenzer shows with screening examples why a positive result from an accurate test is often wrong, and how natural frequencies make sensitivity and false positives clear.
Read the news better, every morning.
We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.
Free forever, no spam. One click to unsubscribe, and we never sell your email.