One number doing too much work
On distributional and mediation parameters in DiD
Hi there! Some good news: I submitted my thesis, so I should now have more time to focus on this newsletter. The bad news is that I need a job, hehe :) If you know of anything, please let me know, preferably industry roles., and I now have an official job1 :) I’ve been to China (again, yes, and many more times hopefully) for 10 days and had a great time, will write about it soon. I’ll be in Denmark on Friday Aug 7 for the Emergent Ventures European Conference, if you’re around say hello! Other conferences coming up are the EEA-ESEM in two weeks, which will happen at UCD. I volunteered to help, so if you’re coming to Dublin to present let me know!
Another piece of good news is that Prof Yiqing has a Substack! It’s called Event Study and you can subscribe here. I’m a big fan of his academic work and of the effort he puts into making his research reproducible. I’m now looking forward to seeing him write for a broader audience.
Also a bit of house-cleaning:
Prof Paul GP has been more active on Substack, with great posts recently
JAMA’s Guide to Statistics and Methods has a special DiD article
A new NBER guide “for economists writing a literature review who want to summarize evidence, incorporate covariates, and adjust for selectivity”
Prof Clément just solved all your problems with his ChatDiD! Brilliant and a lot of fun
You can (should!) pre-order Prof Scott’s new book “Causal Inference: The Remix” here! Out August 25 :)
My friend Aniket has started a newsletter on AI and agentic coding filtered for economists! Subscribe here
Prof Damian released the book “Applied Microeconomics”, which is paired with a great companion site
“A curated list of AI tools, libraries, and resources for economics research, teaching, and policy analysis” by Lu Han (BoC)
Today's post is about papers that go beyond our usual ATT, either replacing it with distributional parameters or decomposing it into direct and indirect effects:
A simple distributional difference-in-differences estimator for univariate and bivariate outcomes, by Aico van Vuuren, Iván Fernández‐Val, Francis Vella and Jonas Meier
Difference-in-differences for mediation analysis using double machine learning, by Martin Huber and Sarina Joy Oberhänsli
Difference-in-differences with a mediator, by Yuhao Deng, Haoyu Wei and Zhongzhe Ouyang
The next papers that will be covered in upcoming posts are listed in the footnotes2 :)
A simple distributional difference-in-differences estimator for univariate and bivariate outcomes
TL;DR: most DiD applications stop at the ATT, but the same design can tell you where in the distribution the effects happened. The authors run a small DiD for the probability of being below each threshold, on a logit/probit scale, and stitch the results into a counterfactual distribution for the treated group; an extension covers the joint distribution of two outcomes. The identifying restriction is a PT condition on a transformed probability scale, so the link function is part of the assumption. In the C&K application, full-time employment gains are concentrated in larger establishments, with evidence consistent with substitution between full-time and part-time workers.
What is this paper about?
In most DiD settings we are constrained, for many research-design-related reasons, to asking what a treatment did to the average outcome of the treated group: the ATT. Questions like “did a minimum-wage increase raise or reduce average employment?”, “did a school reform increase average test scores?” or even “did a health policy change average healthcare use?” are common and very useful, but we know that the average can conceal much of what has happened.
So let’s suppose a policy raises an outcome of interest by ten units for people near the bottom of the distribution and lowers it by ten units for people near the top. The average effect could then be close to zero even though the policy produced substantial changes for both groups. From this, we infer that a “zero” ATT doesn’t necessarily mean that nobody was affected, but that the same average effect can also describe very different distributions. A policy that gives every treated person a small gain could produce the same ATT as one that gives a large gain to a small group and leaves everyone else unchanged. Those two results have different implications for important measures like inequality, targeting and welfare, even though a conventional DiD estimate may report the same number.
In this paper the authors ask what we can learn when we replace the ATT with the whole outcome distribution3. Instead of estimating only whether the mean changed, they estimate the treated group’s counterfactual distribution, meaning what the full distribution would’ve looked like after the intervention if treatment hadn’t occurred. They then compare it with the distribution observed under treatment.
This means we can examine where in the distribution the effects occurred. We can ask whether treatment changed the probability of falling below a specific threshold, which gives us a distributional treatment effect, or whether the effect differed at the bottom, middle or top of the distribution, which is captured by quantile treatment effects. The paper also considers cases in which treatment affects two related outcomes. This goes further than estimating a separate DiD for each one. Imagine that a minimum-wage increase affects both full-time and part-time employment. Separate estimates could show what happened to each outcome on its own, while the joint distribution can show whether firms changed the combination of full-time and part-time workers they employed. The paper therefore expands the object studied in a DiD design in two directions: from average effects to distributional effects and from one outcome to the joint behaviour of two outcomes.
What do the authors do?
The authors begin with the canonical DiD setting: two groups and two periods, with treatment received only by the treated group in the second period. We therefore have the familiar four cells, the treated and control groups before and after treatment. For the treated group after the intervention, we observe the distribution under treatment. What we don’t observe is the distribution that this same group would’ve had at the same time without treatment. This is the missing counterfactual distribution that the authors need to recover. A conventional DiD does this calculation for the mean: it uses changes in the control group to estimate how the treated group’s average outcome would’ve evolved without treatment. The authors apply the same logic at every possible value of the outcome.
This is where distribution regression4 comes in. Suppose our outcome is the number of workers employed by a restaurant. The authors take a threshold, such as ten workers, and create an indicator equal to one when a restaurant employs ten workers or fewer. They then estimate a logit or probit model for the probability of being below that threshold, using group and time indicators together with their interaction. Repeating the exercise for many thresholds, five workers, six workers and so on, traces out the outcome distribution. We can think of this as running a sequence of small DiD exercises, one for each threshold. At every point, the question is the same: how did the probability of being below this value change in the treated group relative to the control group?
The observed post-treatment distribution for the treated group comes directly from the data. To estimate its counterfactual distribution without treatment, the authors combine information from the other three group-by-time cells and impose their main identifying restriction, which they call the no-interaction assumption. The name is a little opaque, so it helps to translate it. For every possible outcome threshold, the authors assume that, without treatment, the treated and control groups would’ve experienced the same time change in the probability of falling below that threshold, after the probability has been transformed using the selected link function.
This is related to PT, but it is a different restriction, and one can hold when the other fails. Standard PT concerns the evolution of the average untreated outcome. Here, the restriction concerns the evolution of the entire untreated distribution, one threshold at a time, on a transformed probability scale. The reason for the transformation is more on the practical side. If we used a linear probability model, the DiD calculation could produce counterfactual probabilities < 0 or > 1, mainly near the tails of the distribution. Logit and probit transformations keep the estimated probabilities within their proper range.
The choice of link function still carries an assumption. Parallel changes on the logit scale aren’t the same as parallel changes on the probit scale, so the method doesn’t remove the need to justify the counterfactual. It shifts the identifying claim from trends in average outcomes to trends in transformed probabilities across the distribution.
The authors also introduce covariates. In that version, the no-interaction condition must hold after conditioning on the observed characteristics included in the model. This can help when the composition of the treated and control groups changes over time, although it requires sufficient overlap in those characteristics across all four cells. Once the two potential-outcome distributions have been estimated, the rest follows from comparing them. At each fixed threshold, their difference gives a distributional treatment effect. Researchers can also invert the distributions and compare their quantiles, giving quantile treatment effects. The bivariate extension follows the same idea with two outcomes. The authors first estimate the marginal counterfactual distribution of each outcome, and then use bivariate distribution regression to estimate how the two outcomes are related at different points of their joint distribution.
This dependence component is what we’d miss by running two separate DiD models. Suppose a minimum-wage increase raises full-time employment and reduces part-time employment. Separate regressions can estimate each change, while the joint model can also show whether restaurants increasingly substitute one type of worker for the other. Identification is more demanding in this case because the no-interaction condition now has to hold for both marginal distributions and for the dependence between the two outcomes5.
From the estimated joint distributions, researchers can calculate joint probabilities and measures of dependence such as Kendall’s tau and Spearman’s rank correlation. This provides a way to study whether treatment changes how two outcomes move together. For inference, the paper proposes a weighted bootstrap. In panel data, all observations belonging to the same unit receive the same bootstrap weight, preserving the dependence between that unit’s observations over time. The estimator can also be used with repeated cross-sectional data.
To illustrate the method, the authors return to C&K’s study of New Jersey’s 1992 minimum-wage increase. They examine the distributions of full-time and part-time employment, along with full-time-equivalent employment, in fast-food restaurants, using restaurants in Pennsylvania as the control group. The distributional results indicate that the estimated effects vary with establishment size. The increase in full-time-equivalent employment is mainly driven by gains in full-time employment, with larger gains towards the upper part of the distribution. The estimates for part-time employment are generally zero or negative, although the relatively small sample produces wide confidence intervals. The bivariate analysis adds another piece of information: the estimated relationship between full-time and part-time employment becomes more negative after the minimum-wage increase. The change in Spearman’s rank correlation is statistically significant at the 10 per cent level, which the authors interpret as consistent with greater substitution between full-time and part-time workers.
The authors also conduct a specification check based on whether the estimated counterfactual distribution is properly increasing. They don’t reject monotonicity in their application, which is compatible with the model. This check can detect certain failures of the specification, though it can’t establish that the no-interaction assumption is true.
Why is this important?
A large share of policy questions are distributional at their core. Evaluations of minimum wages or school reforms often need to say who gained and who lost, and a single average can’t answer that. This paper makes the distributional version of the question answerable with tools most applied researchers already have, since the estimation reduces to a sequence of logit or probit regressions followed by a bootstrap.
The identification lesson is worth the read on its own. Justifying PT in means doesn’t justify a counterfactual distribution. The restriction the authors need operates threshold by threshold on a transformed probability scale, and the choice of link function is part of it. Researchers who report quantile treatment effects from DiD designs should be able to say which restriction their numbers rely on, and this paper states it in a very clean way.
The bivariate extension adds a question that single-outcome methods can’t answer: whether treatment changed how two outcomes move together. In the minimum-wage application, that is the difference between knowing that full-time employment rose while part-time employment fell, and knowing that restaurants substituted one type of worker for the other.
Who should care?
You should read this paper if you work on policies where the distribution of the outcome is the question, such as effects on inequality or on the share of units below a threshold; report quantile treatment effects from a DiD design and want the identifying restriction explicitly stated; study two related outcomes and suspect the interesting response is in their combination; and want distributional estimates without leaving standard logit and probit machinery.
Do we have code?
I couldn’t find a package or a link to code in the paper. The estimation itself stays close to standard practice, a set of logit or probit regressions plus a weighted bootstrap, so implementing it from scratch looks feasible for anyone comfortable with distribution regression. The paper is an SSRN preprint and hasn’t been peer reviewed yet, so a package may still appear.
In summary, the paper changes the estimand of a DiD while keeping its logic. We still impute a counterfactual for the treated group from the other three cells, and we still need an untestable assumption to do it. What changes is the object, now a whole distribution or the joint distribution of two outcomes, and the assumption, which lives on a transformed probability scale and includes the link function. If the distribution is what your research question is about, this is a practical way to estimate it and to be explicit about what the estimate relies on.
Difference-in-differences for mediation analysis using double machine learning
TL;DR: mediation analysis usually requires an unconfounded treatment, which is precisely what DiD users don’t want to assume. The authors split the total effect on the treated into direct and indirect channels using PT assumptions imposed across treatment-mediator combinations, with DML handling the covariates. The price is that PT must now hold across groups defined partly by mediator behaviour, which is stronger than the standard version. In their NLSY97 application, health care coverage shows no significant short-term effect on general health, through routine checkups or otherwise.
What is this paper about?
The previous paper changed the object we estimate, whereas this one asks where the effect comes from. A DiD design gives us the total effect of a treatment on the treated, but applied questions often go one step further: how much of that effect “travels” through a specific channel? In the authors’ application, health care coverage may improve general health partly because insured people go for routine checkups. The checkup is a mediator (a variable sitting on the causal path between treatment and outcome) and the goal is to split the total effect into an indirect effect operating through the mediator and a direct effect capturing everything else.
The standard toolkit for this split - causal mediation analysis6 - usually assumes that both the treatment and the mediator are as good as randomly assigned once we condition on observed covariates. That is exactly the assumption that researchers who want to use DiD don’t want to make about the treatment; if we believed it, we wouldn’t be running a DiD in the first place. So the paper asks whether mediation analysis can be carried out inside a DiD framework, replacing unconfoundedness of the treatment with PT conditions.
Their framework is broad on purpose. Treatments and mediators can be binary, multivalued or continuous, and the paper covers several targets. Some fix the mediator directly (joint effects of setting treatment and mediator together, and controlled direct effects that hold the mediator at a chosen value), while others (the natural direct and indirect effects7) let the mediator take whatever value it takes with and without treatment. The distinction is important because, as we’ll see, the two “families” require different assumptions.
What do the authors do?
They define groups by combinations of treatment and mediator values. With two periods, this recreates the four-cell logic of canonical DiD, except that the “treated” cell is now a treatment-mediator pair - say insured people who went for a checkup - and the comparison cell is another pair - say uninsured people who didn’t.
The central identifying assumption is conditional PT across these treatment-mediator combinations: given covariates, the potential outcomes of both groups under the comparison state would have followed the same average trend. This is stronger than PT across treatment groups alone because the groups are also defined by how the mediator responds. In the binary case, it amounts to assuming PT across people who take up the mediator when treated and people who don’t take it up when untreated, and these are groups with different underlying mediator behaviour. The authors are open about this strength and point to the pay-off: unlike earlier DiD-mediation work, they need no monotonicity of the mediator in the treatment, and non-binary treatments and mediators are covered.
To this they add no anticipation, common support (comparable covariate values must exist in all four cells) and the requirement that covariates are unaffected by treatment or mediator. The last condition deserves attention in repeated cross sections where covariates are often measured after treatment.
These assumptions identify the joint effects and the controlled direct effects. Natural effects need more, and the authors offer two routes. The first adds PT across treatment groups, which fails if unobserved factors drive both treatment take-up and the mediator. The second instead adds a distributional PT assumption on the mediator itself, and it requires the mediator to be observed in the pre-treatment period. Neither route is “free”, and which one is more credible depends on the application.
The estimation belongs in the DML framework. The authors derive doubly robust score functions that satisfy Neyman orthogonality, which in practice means two things: covariates can be controlled for with ML (the authors use LASSO), because first-order estimation errors in those auxiliary models don’t contaminate the effect estimate; and cross-fitting (estimating the auxiliary models and the effects on separate folds of the data) guards against overfitting. The estimators are asymptotically normal at the usual rate for discrete treatments and mediators; continuous ones bring kernel smoothing and slower convergence. Simulations at sample sizes of 2000 and 8000 show small bias, with confidence intervals slightly too narrow in the smaller samples.
The empirical application revisits Farbmacher and co-authors’ 2022 study of the effect of obtaining health care coverage in 2006 on self-reported general health in 2008 among NLSY978 respondents, with routine checkups in 2007 as the mediator. The DiD design changes the sample: respondents must report health in 2005 and 2008 and must have had no coverage before 2006, which shrinks the sample from 7486 to 1020. No estimated effect, total, direct or indirect, is statistically significant, although the point estimates of the total and direct effects lean towards better health, broadly in line with the original study’s direction.
Why is this important?
A reform’s evaluation often needs to say through which channel the effect arrived because the channel determines which part of the policy is worth keeping or adjusting. Until now, answering that question usually meant abandoning DiD’s main advantage - its tolerance of unmeasured confounding between treatment and outcome. This paper keeps the DiD logic and prices the mediation question in PT currency instead.
The price we gotta pay is visible, and to me that’s a virtue. The PT conditions are stated cell by cell, so a referee, or the researcher, can interrogate each one. The paper also draws a useful bridge to the staggered adoption literature: its key assumption is of the same type as those used with multiple treatment periods, with early adoption in the role of the treatment and later adoption in the role of the mediator. If you already accept those assumptions in staggered designs, you have a head start on judging these.
The ML component earns its place too. Conditional PT is easier to defend with a generous covariate set, and the DML machinery makes a large set usable without hand-picking controls.
Who should care?
You should read this paper if you use DiD and face questions about mechanisms, such as which channel carried a policy’s effect; work with multivalued or continuous treatments or mediators; have many candidate control variables and want data-driven covariate adjustment with valid inference; and follow the growing DiD-mediation literature and want the assumption trade-offs mapped out. If your mediator is plausibly randomised, standard mediation tools with selection on observables may serve you more simply.
Do we have code?
The paper states that the estimator is implemented in R, with four-fold cross-fitting and LASSO-based nuisance estimation, but I couldn’t find a link to a package or repository in the paper. I’d expect an implementation to show up, given how much of the authors’ related work has been packaged, but I couldn’t verify one at this time.
In summary, this paper translates mediation questions into DiD language. Instead of assuming an unconfounded treatment, you assume PT across treatment-mediator cells and you add further conditions if you want natural rather than controlled effects. The estimands deserve as much attention as the estimator: controlled and natural effects answer different questions and carry different assumptions, and the paper is clear about which is which. For applied researchers with a credible DiD and a mechanism question, the decomposition is now feasible; the assumptions still need arguing one cell at a time.
Difference-in-differences with a mediator
TL;DR: earlier Job Corps studies found no short-run earnings effect of job training. The authors build mediation into a two-period DiD and show that the flat total effect hides a significant positive direct effect on earning capacity, offset by an insignificant negative effect from reduced work time while training. Unmeasured confounding between treatment and outcomes is fine, since differencing removes it, but the mediator must be unconfounded, and that assumption can’t be tested from data. The estimators are multiply robust: consistent if any two of three working models are correct.
What is this paper about?
This paper starts from a substantive puzzle in the Job Corps Study, a large US training programme for disadvantaged youths9: earlier studies found earnings gains only in the long run, and the flat short-run effect may be mechanical (time spent in training is time not spent working, so weekly earnings can fall even while earning capacity rises).
Sorting this out requires a mediation analysis, with the share of weeks employed as the mediator, in a setting where treatment take-up is confounded with outcomes. The paper builds that analysis inside a two-period DiD. The key idea is a division of labour between assumptions: unmeasured confounding between treatment and outcomes is allowed because differencing outcomes over time removes it, while the mediator has to be clean, meaning free of unmeasured confounding with the treatment and with the outcome trend.
What do the authors do?
The target parameters are defined for the treated group: the total effect of training, which splits into a natural indirect effect operating through work time and a natural direct effect capturing the remaining channels.
The identification starts from two familiar DiD ingredients. The first is no anticipation, so that baseline earnings don’t respond to training that hasn’t happened yet, and the second is a PTA, adjusted for the mediator: conditional on covariates and on the value the mediator takes without treatment, untreated potential outcomes follow the same trend in both groups. To these the authors add sequential ignorability - adapted to the DiD setting, which requires that treatment take-up is unconfounded with the mediator given covariates and that the observed mediator doesn’t modify the underlying trend. The usual overlap and consistency conditions complete the set, which ensures that the groups contain comparable observations and that observed outcomes correspond to the relevant potential outcomes.
A design detail I like: the PT condition deliberately avoids conditioning on the observed post-treatment mediator because conditioning on a post-treatment variable can create collider bias (a spurious association produced by conditioning on a common consequence). The authors contrast their approach with related work on exactly this point, including the Huber and Oberhänsli paper above.
For estimation, the authors derive efficient influence functions10 from semiparametric theory and build estimators that are multiply robust. Three working models are involved: one for the propensity score, one for the mediator’s distribution and one for the change in outcomes. The estimator remains consistent if any two of the three are correctly specified, and it reaches the nonparametric efficiency bound when all are. After that a practical section shows how to implement everything with generalised linear models, including a trick that avoids estimating the mediator’s density directly. The paper also treats controlled direct effects - where a continuous mediator creates a technical obstacle (the estimand stops being pathwise differentiable) - handled with kernel smoothing at the cost of a slower convergence rate.
The application returns to Job Corps with actual enrolment as the treatment, log earnings as the outcome and the share of weeks employed as the mediator. In the shorter-run analysis, the total effect on earnings is small and insignificant, and the decomposition explains why: a positive, statistically significant natural direct effect of about 0.11 log points11 is offset by a negative, insignificant indirect effect running through reduced work time. A parallel analysis one year later (using training in the second year and the change in earnings from the second to the third year) finds all effects positive and significant with a direct effect of about 0.38 log points. A sensitivity analysis coding employment in four levels reaches the same conclusions, and the estimated controlled-direct-effect curve is significantly positive across a broad range of employment levels in the later year.
Why is this important?
The substantive lesson pairs well with the methodological one. Earlier readings of Job Corps concluded that training doesn’t help earnings in the short run. Their decomposition suggests the short-run total effect mixes two opposing forces, higher earning capacity and fewer weeks worked, and the “no effect” verdict conceals the first. That is the same theme running through this whole post: a single summary number can hide the part of the answer the policy question really needs.
On the methods side, the paper brings efficiency theory and multiple robustness to DiD mediation. Multiple robustness is practical insurance since you don’t have to bet the analysis on one correctly specified working model. The paper is also honest about its limits. Sequential ignorability can’t be verified from data, and the discussion flags open problems, including the difficulty of modelling the mediator’s distribution and the treatment of zero earnings.
Who should care?
You should read this paper if you evaluate programmes where the treatment absorbs participants’ time (e.g., training or further education) and short-run outcomes look flat; want mediation analysis in settings with non-random take-up; care about multiply robust estimation and are comfortable with (or curious about!) influence-function-based methods; and work with continuous mediators and want effect curves across mediator values instead of a single number.
Do we have code?
I couldn’t find a package or repository linked in the paper. The data are publicly available through the R package causalweight, and the supplementary material spells out a parametric estimation strategy in enough detail to implement, but the implementation work is currently left to the reader.
In summary, the paper shifts the confounding burden of mediation analysis from the treatment, where DiD differencing can deal with it, to the mediator, where only argument can. When that argument is credible, as the authors maintain for work time in Job Corps, the reward is a decomposition that changes the substantive reading of a famous programme: training builds earning capacity from the start and the short-run earnings dip mostly reflects the time participants spend in training.
I’m not moving to London yet, but will be sharing my time here and there.
Three’s a crowd: Identification challenges in the triple difference model with spillover effects, by Silvia De Nicolò, Beatrice Biondi and Mario Mazzocchi
Causal Graphs for Conditional Parallel Trends, by Michael C. Knaus and Henri Pfleiderer
Stacked Triple Differences, by Meng Hsuan Hsieh
Doubly robust local projections difference-in-differences, by Daniel de Abreu Pereira Uhr and Guilherme Valle Moura
When are time series predictions causal? The potential system and dynamic causal effects, by Jacob Carlson and Neil Shephard
Difference-in-Differences using Double Negative Controls and Graph Neural Networks for Unmeasured Network Confounding, by Zihan Zhang, Lianyan Fu and Dehui Wang
Debiased Difference-in-Differences Estimation and Inference of Causal Effects Under Adaptive Controls, by Zhiyuan Tang, Yining Wang and Sentao Miao
It’s About Time (Series): A Simple Correction For Difference-in-Differences Estimators, by Gary Cornwall and Scott Wentland
Using Pre-Trends for Inference in Difference-in-Differences, by Clément de Chaisemartin
Efficient difference-in-differences estimation under partial interference with incremental propensity score policies, by Junjie Li and Yukitoshi Matsushita
Finite-Population Inference for Heterogeneity in Many-Group Synthetic Difference-in-Differences, by Takahiro Hoshino and Makoto Nakakita
Group-Level Treatment Effect Heterogeneity in Difference-in-Differences: A Balanced Approach, by Nora Bearth, Nadja van 't Hoff and Torben S. D. Johansen
Choosing A Headline Estimand from Matching, DID, and Hybrid Designs: A Minimax-Regret Approach, by Yechan Park and Yuya Sasaki
Semiparametric Difference-in-Differences Estimation With Missing Not at Random Data: A Shadow Variable Approach, by Junjie Li and Dongyuan Mu
When Do Treatment Changes Identify Causal Effects?, by Martin Huber
Identification and Estimation of Staggered Difference-in-Differences with Network Spillovers, by Hayato Tagawa
Adaptive Causal Inference for Higher-Order Difference-in-Differences Designs, by Jangsu Yoon
An Evaluation of Difference-in-Differences Methods Using Placebo Event Studies, by John M. Coglianese and Jade A. Fang
Rolling Difference-in-Differences with Reversible Treatment Paths and Moderating Effects, by Soo Jeong Lee
Choosing a Difference-in-Differences Estimand and Estimator for Heavy-Tailed Outcomes: A Practitioner’s Companion, by Daniel Winkler, Christian Hotz-Behofsits and Nils Wlömert
Difference-in-Differences with Spillovers: An Imputation Approach, by Soumen Banerjee and Jianguo Wang
Difference-in-Differences with multiple Treatments under Control, by Daniel Steinberg and Marcus Roller
A bit of background: earlier work extended DiD beyond average effects in several directions. Athey and Imbens (2006) introduced changes-in-changes to recover counterfactual outcome distributions, while later work examined quantile treatment effects (Callaway and Li, 2019) and treatment effects at different points of the outcome distribution using conventional DiD (Dube, 2019; Goodman-Bacon, 2021; Goodman-Bacon and Schmidt, 2020). Other approaches include IPW (Kim and Wooldridge, 2025), distribution regression (Biewen, Rümmele and Fitzenberger, 2022) and a multivariate extension of changes-in-changes (Torous, Gunsilius and Rigollet, 2021). Their paper builds on the distribution regression approach, using nonlinear link functions such as probit or logit instead of linear probability models and providing the identifying conditions required for this version of DR-DiD.
Distribution regression models the whole conditional distribution of an outcome by estimating a separate binary-choice model for the probability of being at or below each threshold. The idea goes back to Foresi and Peracchi (1995), and Chernozhukov, Fernández-Val and Melly (2013) developed its use for counterfactual distributions.
In plain English, without treatment, the relationship between the outcomes must also have evolved in the same way for the treated and control groups on the scale used by the model.
Mediation analysis asks how a treatment produces its effect, by splitting the total effect into two parts: an indirect effect that travels through an intermediate variable (the mediator) and a direct effect that covers everything else. Take job training and earnings, for example: training changes how many weeks people work, and work time changes earnings, so part of the earnings effect runs through work time and part comes from other channels, like higher productivity. The catch is that the mediator is usually not randomised, even when the treatment is, since people who work more weeks (or go for checkups) differ from those who don’t. That’s why the classic versions assume both treatment and mediator are unconfounded given covariates, and why building mediation into DiD - where the treatment is exactly the confounded part - takes some work.
The natural/controlled terminology comes from Robins and Greenland (1992) and Pearl (2001). A controlled direct effect fixes the mediator at the same value for everyone; natural effects let each person’s mediator take the value it would take under a given treatment status. Huber (2021) surveys causal mediation analysis.
The National Longitudinal Survey of Youth 1997, a nationally representative US sample of people aged 12 to 17 at their first interview in 1997.
There are two motivations - one substantive and one methodological - for this paper, and the authors are quite explicit about both. The substantive one is the Job Corps puzzle. Assignment to the programme was randomised, but actual enrolment wasn’t - people chose whether to take up training, and those with more pessimistic labour-market expectations were more likely to join. A series of earlier studies (using IV, principal stratification, and bounds) all reached the same conclusion: a significant long-term earnings effect, and nothing detectable in the short term. The authors’ suspicion is that this “nothing” is partly an artefact: training eats up a large share of participants’ time so they work fewer weeks, and weekly earnings fall “mechanically” when you can’t work full-time. None of the earlier analyses separated “training didn’t build skills yet” from “trainees were busy training”. Doing that separation requires treating work time as a mediator, and doing it in a setting where the treatment is self-selected. The methodological one is that no existing tool fit that problem. Their introduction walks through the alternatives and finds each one wanting for this setting:
Classical mediation analysis requires the treatment, mediator and outcome to all be unconfounded, which is hopeless with self-selected enrolment.
Deuchert, Huber and Schelker (2019) do DiD mediation but need a binary mediator, monotonicity and PT within principal strata.
Huber et al. (2022) use changes-in-changes, with distributional restrictions and binary mediators.
Hsia et al. (2025) put mediation into panel data via linear regressions, which can’t accommodate interactions.
Blackwell et al. (2025) handle controlled effects but need a randomised treatment and a discrete mediator.
Each of these covers part of the territory. The combination this paper goes after - natural direct and indirect effects in a standard two-period DiD - with self-selected treatment and a possibly continuous mediator, wasn’t available before.
An influence function describes how much an estimator moves when the data distribution is perturbed slightly. Building the estimator from the efficient influence function is what delivers both the efficiency bound and the multiple-robustness property.
Roughly, a 0.11 log-point effect corresponds to an 11 to 12 per cent increase in weekly earnings, and 0.38 to about 46 per cent.


