Hello there, it’s time for a new post :)
Apologies for the delay, I hope my lack of posting haven’t set you back a few months hehe
Today we will be reviewing three papers:
Difference-in-differences with “bad controls”, by Carolina Caetano, Brantly Callaway, Stroud Payne, and Hugo Sant’Anna
Causal Graphs for Conditional Parallel Trends, by Michael C. Knaus and Henri Pfleiderer
An Evaluation of Difference-in-Differences Methods Using Placebo Event Studies, by John Coglianese and Jade A. Fang
Difference-in-differences with “bad controls”
TL;DR: in DiD, a covariate can be necessary for conditional parallel trends and still be a “bad control” if treatment changes it, so both including it and dropping it can bias the ATT. The paper shows when conditioning only on the pre-treatment value is enough, and when researchers instead need to recover how the covariate would have evolved without treatment.
What is this paper about?
In Mostly Harmless Econometrics, bad controls1 are “variables that are themselves outcome variables in the notional experiment at hand”, and the advice that Angrist and Pischke give us is: control only for variables that are not themselves caused by the treatment. With this in mind, the authors in this paper study what happens to such advice in DiD when three things hold: 1) parallel trends holds only conditional on covariates, 2) some of those covariates vary over time, and 3) some of them could be affected by the treatment. The classic labour examples are a worker’s occupation, industry or union status, and the issue we face is that parallel trends is more plausible when we compare workers in similar jobs, yet once treatment can change someone’s job controlling for the occupation we observe is substantially different that controlling for the occupation they would have had without treatment. The authors agree that putting the observed post-treatment covariate straight into the analysis biases the estimate. Their additional point is that simply dropping it does not solve the problem either. When we drop the covariate, we are changing the identification strategy: a parallel trends assumption that works without the covariate is a different assumption from the one we started with - and if that different assumption holds, the covariate was never needed as a control in the first place. The paper offers two solutions. Under stronger conditions, conditioning on the pre-treatment value of the bad control is enough. Under weaker conditions, researchers instead have to recover how the bad control would have evolved without treatment and incorporate that counterfactual into the estimation.
What do the authors do?
First they characterise the bias of the two traditional approaches: using the observed covariate as a control or dropping it. Their parallel trends assumption conditions on the value the covariate would have taken without treatment (untreated potential occupation), and neither traditional approach conditions on that value. Using the observed covariate conditions on the wrong value because, for treated workers, the post-treatment occupation may itself have been changed by treatment, whereas dropping the covariate conditions on… pretty much nothing about occupation at all. It is a different identification strategy from the one we started with.
The first approach gives conditions under which conditioning on the covariate’s pre-treatment value is sufficient. When those conditions hold, no new estimator is needed because the Callaway and Sant’Anna (2021) estimator with pre-treatment covariates already does this - in the labour example we condition on pre-treatment occupation, industry and union status. The second approach works when a covariate unconfoundedness condition holds. It has two steps: first you recover how the covariate would have evolved without treatment, and then you use that counterfactual distribution to recover the untreated outcome path implied by conditional parallel trends. For estimation the paper provides an imputation estimator and a Neyman-orthogonal one2, and the nuisance pieces can be fitted with ML.
Both approaches also work under staggered adoption. The paper ends with an application to job displacement where a worker’s occupation score is the bad control. They find that displacement lowers occupation scores on average, which confirms that the covariate is affected by treatment. Their baseline estimate implies an earnings loss of about 7%, compared with close to 10% under the traditional TWFE specifications.
Why is this important?
On the theory side, the paper exposes a gap in conditional parallel trends when one of the controls is itself affected by treatment. The assumption really needs the value that the covariate would have taken without treatment. But for treated units after treatment, we do not observe that value. So conditional parallel trends alone is not enough to identify the ATT: we also need an assumption about how the covariate would have evolved without treatment. The paper makes those extra assumptions more visible through simple covariate unconfoundedness and bad-control redundancy, and shows when they are enough for identification. The bad-control problem survives the paper, and the gain is of a different kind: we now know which extra assumptions carry the identification and we can judge for ourselves whether they are plausible. On the practice side, this is potentially a very common problem. The controls we often want for conditional parallel trends are also variables that policies and shocks may change. The paper also weakens a familiar robustness check: estimating the model both with and without the suspect control. Similar estimates do not necessarily mean that the bad-control problem has gone away since both specifications can be wrong for different reasons.
Who should care?
This paper is for applied researchers who need a time-varying covariate for their conditional parallel trends assumption and have reason to think the treatment changes that covariate. In labour economics the usual examples are occupation, industry, union status and employment itself. By the same logic, similar issues could arise in health studies with insurance status or health behaviours, or in firm-level studies with firm size or inputs. It’s also useful for anyone who reports results with and without such a covariate as a robustness check (the paper shows that both specifications can be biased at the same time so their agreement says less than we tend to assume).
Do we have code?
Yes. The authors provide an R package, badcontrols, which implements all the estimators proposed in the paper, and the package website has a conceptual overview of the method. The proofs and additional results are in the Supplementary Appendix.
In summary, this paper takes a piece of advice most of us learned from Mostly Harmless Econometrics and asks what it means for DiD when parallel trends depends on a covariate the treatment can change. The authors show that the two usual responses go wrong in different ways: including the observed covariate conditions on a value that may itself have been changed by treatment, while dropping it swaps the original parallel trends assumption for a different one. They then give two ways to keep the covariate as a control: one conditions on its pre-treatment value, which under their conditions is the Callaway and Sant’Anna estimator many of us already use, and the other recovers how the covariate would have evolved without treatment and uses that counterfactual to construct the untreated outcome path. For applied people, the lesson we can take home is having to decide whether to control for a variable like occupation means deciding which parallel trends assumption we are making.
Causal Graphs for Conditional Parallel Trends
TL;DR: many DiD papers assume parallel trends conditional on covariates and justify their choice of covariates only with pre-trends. This paper builds ∆-SWIGs, a version of causal graphs for conditional parallel trends, so we can read off which covariates, measured in which periods, make the assumption hold. When time-varying covariates affect the outcome, their post-treatment values have to be controlled for, which is only safe if treatment doesn’t change them, and pre-trends speak to only some of the assumptions behind post-treatment effects.
What is this paper about?
Many DiD applications assume parallel trends only after conditioning on covariates, and this is known as conditional parallel trends. In practice, parallel pre-trends are often used to justify the final conditioning strategy. In this paper the authors show how common conditional parallel trends is in the AER: of the papers it published in 2024 and 2025, 19 explicitly assume some form of parallel trends, 14 impose the conditional version and 13 use time-varying controls. None of these papers gives a formal justification for its choice of conditioning variables beyond showing parallel pre-trends. Unconditional and conditional parallel trends are different assumptions, so the choice of covariates needs an argument about the causal structure, and a pre-trends plot can only support part of it. Under selection on observables (unconfoundedness), we already make that argument with causal graphs3, the DAGs many of us met in Prof Scott’s Causal Inference: The Mixtape, which let us draw what causes what and read off which variables to condition on. Bringing that tool over to DiD takes some work - standard DiD identification is typically justified by a functional-form restriction4: time-invariant unobserved confounders enter untreated outcomes additively, so differencing removes them. The authors solve this by introducing ∆-SWIGs5 (transformed Single World Intervention Graphs) that represent differences in untreated potential outcomes. They then use these graphs to work out which covariates, measured in which periods, make conditional parallel trends hold for a given causal structure.
What do the authors do?
Using ∆-SWIGs, the authors work out which conditioning strategies make conditional parallel trends valid in the standard 2x2 case and under staggered adoption. They show first that outcome dynamics (e.g., past outcomes affecting later outcomes, treatment take-up or future covariates) generally prevent conditional parallel trends from being justified by the causal structure alone. They then turn to time-varying covariates. If these covariates affect outcomes, valid DiD specifications may need to condition on values measured after treatment has begun. That is fine when treatment does not itself change the covariate, but when there is treatment–covariate feedback, the relevant object is the covariate value that would have been observed without treatment. For treated units that value is unobserved, therefore dynamic treatment effects are no longer identified without additional assumptions. Without them, using only pre-treatment controls leaves confounding bias, while using the observed post-treatment value creates what the authors call “wrong world control bias”.
This also explains one of the paper’s main practical results: parallel pre-trends can hold even when later treatment-effect estimates are biased because treatment–covariate feedback only appears after treatment and therefore cannot be diagnosed from pre-treatment data.
Why is this important?
The paper gives conditional parallel trends a graphical method for choosing covariates, like the one unconfoundedness has had for decades. We can go from our assumptions to the valid controls, or from a published specification back to the assumptions that would justify it. The one functional form assumption left outside the graph is that time-invariant confounders enter the untreated outcome additively. Pre-trends support only some of the assumptions needed for post-treatment effects, so the rest has to come from what we know about the setting. In practice, many of us show unconditional results first and add covariates afterwards, but each specification’s pre-trends test a different assumption. The authors suggest testing pre-trends conditional on all pre-treatment covariates and comparing three specifications: without time-varying covariates, with the covariates up to the outcome period, and with their values just before treatment and in the outcome period.
Who should care?
Mostly applied people using DiD with covariates, particularly if you:
use covariates that vary over time, or have staggered treatment timing;
aren’t sure which periods of a covariate belong in your conditioning set;
rely on pre-trends to justify a conditional specification;
think the treatment may itself affect your controls;
already use Callaway and Sant’Anna-style estimators.
It’s also useful for those reviewing or reading applied DiD. Instead of asking “does the pre-trends look flat?”, we can ask “which assumptions make this conditioning strategy valid” and “which of them the pre-treatment data can speak to?”.
Do we have code?
This paper is mainly about identification, so no. You can draw the causal structure for your setting then use the authors’ results to work out which covariates belong in the conditioning set, and then estimate the DiD as you would usually do. In staggered settings, that can be done with the did package or another Callaway and Sant’Anna implementation. The simulation from the introduction is described in Appendix G if you want to reproduce it. The authors explicitly say that they abstract from estimation and testing issues.
In summary, the authors provide us with a graphical way to decide which covariates to use in conditional DiD. They show that with time-varying covariates, post-treatment values may be needed, but this becomes a problem if treatment itself changes those covariates. They also show that flat pre-trends only tell us about part of the identification argument, so we would still need to justify our choice of controls using the causal structure of the setting.
An Evaluation of Difference-in-Differences Methods Using Placebo Event Studies
(A good day for the Wooldridge school of “perhaps TWFE was not the problem”)
TL;DR: Coglianese and Fang run more than 134,000 placebo event studies on US state data to compare 13 DiD estimators when a single state is treated. The newer staggered estimators give the same answer as TWFE. Synthetic control-type methods can be much more or much less precise, depending on the outcome, and which state is treated affects precision as much as the choice of estimator.
What is this paper about?
After the last few years of DiD methods papers we have many estimators to choose from, including two-way fixed effects (TWFE), Callaway and Sant’Anna, imputation, synthetic control and synthetic DiD, and most of the evidence on how they compare comes from econometric theory or from Monte Carlo simulations with a data generating process the authors chose. This paper compares them on real data instead, using placebo event studies6. The setting is the one-state event study common in applied micro, where one state changes a policy and is compared with others, as in the local labour market studies that go back to Card and Krueger’s minimum wage paper. The authors ask a simple question: if we apply each estimator to a state and a date where nothing happened, how spread out are the “effects” it finds?
What do the authors do?
They build a monthly panel of US states covering 1990-2024, with five outcomes used in local labour market and regional studies: the unemployment rate, labour force participation, two employment growth series and house price growth. For each of the 50 states + DC they draw one random month per year between 1992-2022 as a fake event, then estimate an event study from 12 months before to 12 months after it and repeat this for 13 estimators - with and without a covariate. That adds up to more than 134,000 placebo event studies. Since the true effect is zero, the mean squared error of each estimator’s placebo estimates shows how far it tends to land from zero, and every estimator is compared with TWFE on that basis. The TWFE-like estimators (e.g., Callaway and Sant’Anna, imputation, local projections) perform exactly like TWFE: with a single treated unit they all reduce to TWFE. Matching estimators perform akin to TWFE - about 10% better to about 20% worse depending on the outcome. Synthetic-control-like estimators vary the most with some combinations of estimator and outcome showing 30 to 60% lower variance than TWFE and others showing several times its variance, and there’s no consistent ranking among them - matrix completion, for example, does best for three outcomes and worse than the others for the remaining two. The authors read this as a bias-variance trade-off, where fitting pre-treatment outcomes closely sometimes captures real differences across states and sometimes overfits. They then look at choices beyond the estimator. Both normalising estimates to the last pre-treatment period and seasonally adjusting the outcome reduce the spread of placebo estimates, while the choice of covariates makes little difference. Which state is “treated” affects performance at least as much as which estimator is used: large, diverse states like Texas, Pennsylvania and Ohio give tighter placebo distributions than idiosyncratic ones like Nevada, Alaska and Hawaii, which the authors think is because large states look more like the rest of the country.
Why is this important?
Most of the knowledge we have about how DiD estimators compare comes from theory or from simulations with a data generating process the authors from the papers chose. In this paper, on the other hand, the authors add evidence from real data. They show that, for a single-state event studies, the choice among the newer staggered-adoption estimators makes pretty much difference since they all reduce to TWFE when only one unit is treated. The actual choice is between TWFE, matching and synthetic-control-like methods. Now, which one does best depends on the outcome. When it comes to precision, it depends at least as much on which state is treated, and two simple steps - normalising to the last pre-treatment period and seasonally adjusting the outcome - reduce the spread of the estimates. The authors’ recommendation is that we run placebo event studies on our own data before choosing an estimator. Is worth keeping in mind, though, that their evidence comes from US states with one treated unit at a time, and they themselves say that staggered designs might show bigger differences between estimators. If your setting has a long pre-treatment period, a short note by de Chaisemartin (2026) proposes using the pre-treatment DiDs themselves as the reference distribution for inference on the post-treatment DiD7. His test replaces parallel trends with the assumption that the size of the differential trend between the two groups has a stable distribution over time, so the groups can drift apart as long as the size of that drift keeps its usual distribution.
Who should care?
Anyone running an event study where a single state, region or city is treated, which covers a lot of local labour market work, like the minimum wage studies that go back to the OG Card and Krueger. Also if you’re:
stuck between choosing TWFE, matching and synthetic-control-like methods for one treated unit;
studying a small or unusual treated unit;
working with monthly outcomes that could be seasonally adjusted, or haven’t normalised your estimates to the last pre-treatment period;
dealing with a referee who’s asked you to switch to one of the newer staggered-adoption estimators, since with one treated unit they give the same answer as TWFE.
Do we have code?
No, unfortunately. I couldn’t find a replication package on the paper or on the authors’ pages. The data are all public (the CPS, LAUS, QCEW and the FHFA house price index), so you can run the same exercise on your own data. Pick states and dates where nothing happened, run your estimator on each one, and see how spread out the estimates are. Section 3 of the paper explains how they set up each of the 13 estimators.
In summary, this paper compares 13 DiD estimators on real data, using more than 134,000 placebo event studies in which one US state is “treated” at a date when nothing happened. When only one state is treated, the newer staggered estimators give the same estimates as TWFE. Matching and synthetic control-type methods give different answers, and which one does best depends on the outcome. Normalising to the last pre-period and using seasonally adjusted data make the estimates more precise. Adding covariates changes very little. The authors recommend running placebo event studies on your own data before you choose an estimator. Their evidence comes from designs with one treated state.
From the “Bad Control” section of Mostly Harmless Econometrics: bad controls are “variables that are themselves outcome variables in the notional experiment at hand”, while good controls can be thought of as fixed at the time the regressor of interest was determined. Their example is the effect of a college degree on earnings, with occupation as the tempting control: college opens the door to white-collar jobs, therefore comparing wages within an occupation stops being an apples-to-apples comparison even if the degree itself were randomly assigned - a subtle version of selection bias. The book’s discussion concerns regression models, while this paper gives the term a formal DiD definition with two conditions that both have to hold: the covariate affects the path of untreated outcomes, which is what makes it a covariate at all, and treatment affects the covariate, which is what makes it bad.
Double (or debiased) machine learning is a way of using flexible machine-learning models inside a causal estimate without inheriting their biases. The ML models handle only the nuisance parts (aka the prediction tasks around the estimate), like modelling the covariate’s evolution and the treatment. The estimate itself comes from a formula built to be insensitive to small errors in those predictions, and this insensitivity is the property called Neyman orthogonality, which is why estimators with this structure get called Neyman-orthogonal.
A causal graph - or directed acyclic graph (DAG) - is a diagram with an arrow from A to B whenever we believe A directly causes B. The useful part is a mechanical rule called d-separation which reads off the statistical independencies those beliefs imply so the choice of conditioning variables follows from the assumptions we drew. The framework comes from Judea Pearl’s work in computer science, epidemiology adopted it early, and a good entry point is the primer by Pearl, Glymour and Jewell (2016).
The textbook version of this restriction writes the untreated outcome as a unit effect that never changes plus a time effect common to everyone plus a shock (Y(0) = α + λ + ε). Taking differences over time removes the unit effect, which is why confounders hidden in it drop out of a DiD comparison, as long as they stay fixed over time and enter this way. The authors use a generalised version that restricts only the untreated potential outcome so the treated outcomes and individual treatment effects stay unrestricted. It’s also the one assumption their graphs can’t represent. In their conclusion they note it has to be defended outside the graph, and that a transparent tool for reasoning about it is still missing.
A single world intervention graph (SWIG) is a causal graph redrawn so that potential outcomes appear on it directly. The treatment node is split in two - one part for the treatment the unit actually received and one for the treatment we set in the thought experiment - and every variable downstream of the treatment becomes its potential version. The graph then describes the same objects as our estimand. The idea comes from Richardson and Robins (2013), who built it to connect the graph approach with the potential outcomes approach most economists use. The ∆ version is the authors’: parallel trends is an assumption about changes over time, and they transform the graph to describe differences in untreated potential outcomes, proving that the independencies read off it by d-separation imply conditional parallel trends.
A placebo event study picks a fake event, say Nebraska in October 1998 (the authors’ own example), estimates an event study as if a policy had started there, and repeats this for many randomly drawn states and dates. Because the events are random, the estimates should be centred on zero, and their spread shows how precise each estimator is in that setting. The authors describe this as a form of randomization inference, and it builds on earlier placebo exercises like Bertrand, Duflo and Mullainathan (2004), who used placebo laws to show how often standard DiD inference finds effects that aren’t there.
Conformal inference builds prediction intervals with very few distributional assumptions. The idea is that if a new observation behaves like the observations in a reference sample, its rank among them should be unremarkable, so we keep the candidate values that rank unremarkably and reject the rest. It comes from machine learning (Vovk, Gammerman and Shafer, 2005), and Chernozhukov, Wüthrich and Zhu (2021) brought it to counterfactual and synthetic control inference, including DiD. Their DiD version works with outcome levels and implies parallel trends, while de Chaisemartin’s works with outcome changes, which is why it needs a stable distribution of differential trends instead.





Regarding the third one. Am I understanding correctly that the authors built placebos based on a single treated unit? Why put CS and the others against that? The placebo exercise for staggered is interesting, but we already know the limitations will be smaller as the size of the never treated grows.