The Field Guide · No. 37

Difference-in-differences: comparing two trends, not two totals

Difference-in-differences compares the change over time in a group exposed to a policy against the change in a similar unexposed group, canceling out trends both shared anyway.

Updated

Difference-in-differences is a method for estimating the effect of a policy or event by comparing two changes rather than two totals: how much an outcome moved in a group exposed to the change, minus how much it moved over the same period in a similar group that was not exposed. Subtracting the second change from the first cancels out whatever else was happening to both groups at the same time, a recession, a season, a nationwide trend, leaving an estimate closer to the effect of the policy itself.

The method's best-known application compared fast-food employment in New Jersey and Pennsylvania after New Jersey raised its minimum wage from $4.25 to $5.05 an hour on April 1, 1992, while Pennsylvania's stayed fixed. Economists David Card and Alan Krueger surveyed 410 restaurants on both sides of the state line before and after the increase. Pennsylvania's average full-time-equivalent staffing fell by 2.16 employees per store over that period; New Jersey's rose by 0.59. The difference between those two changes, 2.76 FTE employees, about 13%, is the difference-in-differences estimate, and its direction ran opposite to what a simple supply-and-demand model predicted: employment did not fall where the minimum wage rose.

Two traps follow. First, the method only works if the comparison group would have moved in step with the treated group had nothing changed, an assumption called parallel trends, which cannot be proven directly, only argued for by showing the two groups tracked each other before the policy hit. Second, the estimate is only as good as what feeds it. A later reanalysis of the same New Jersey increase, built from the restaurants' actual payroll records rather than a phone survey, reported a fall in employment instead, by economists David Neumark and William Wascher, reigniting a debate over which measure of employment the comparison was actually differencing.

So when a headline credits or blames a policy for a change in some outcome, look for the comparison group the change is being measured against, not just the before-and-after numbers for the place that changed. Ask whether the treated and comparison groups were tracking similarly before the policy hit, and whether the underlying data is solid enough to trust the difference. Pair a difference-in-differences claim with a look at its parallel-trends check, the same way a natural experiment's strength rests on how well its 'as-if random' assignment holds up.

What to remember

From the record

As noted in Table 2, New Jersey stores were initially smaller than their Pennsylvania counterparts, but grew relative to Pennsylvania stores after the rise in the minimum wage. The relative gain (the "difference in differences" of the changes in employment) is 2.76 FTE employees (or 13 percent), with a t-statistic of 2.03.

David Card and Alan B. Krueger Minimum Wages and Employment: A Case Study of the Fast Food Industry in New Jersey and Pennsylvania, 1993

Asked often

What is the classic example of difference-in-differences?

Economists David Card and Alan Krueger's 1992 study of fast-food restaurants in New Jersey and Pennsylvania after New Jersey raised its minimum wage. Pennsylvania's average staffing per store fell by 2.16 full-time-equivalent employees while New Jersey's rose by 0.59, a difference-in-differences of 2.76 FTE, about 13%, running opposite to the simple prediction that a higher minimum wage would cost jobs.

What has to be true for a difference-in-differences estimate to be trustworthy?

The comparison group has to be a fair stand-in for what would have happened to the treated group without the policy, usually checked by confirming the two groups moved together beforehand, the parallel-trends assumption. The estimate also depends on the underlying data: a later reanalysis of the same New Jersey case using payroll records instead of a phone survey reported the opposite result, showing how much rides on data quality.

Further reading

Go deeper

Paid link. If you buy a book through this link, Bookshop.org pays us a small commission and sends a share of the sale to independent bookstores.

  1. Mastering 'Metrics (opens Bookshop.org)

    Joshua D. Angrist and Jorn-Steffen Pischke · 2014

    Angrist and Pischke devote a chapter to comparing changes over time between treated and untreated groups, with policy examples such as banking during the Great Depression.

Read the news better, every morning.

We send you the day's real progress in science, medicine and beyond, told plainly and sourced. Free, forever.