<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[DiD Digest]]></title><description><![CDATA[A curated digest of the latest developments in Difference-in-Differences methodology, featuring new papers, applications and insights from leading researchers. Suggestions welcome!]]></description><link>https://www.diddigest.xyz</link><image><url>https://substackcdn.com/image/fetch/$s_!qf1q!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cb58cbd-6c9d-47bc-9e14-ed1a2a92c226_538x538.png</url><title>DiD Digest</title><link>https://www.diddigest.xyz</link></image><generator>Substack</generator><lastBuildDate>Mon, 31 Aug 2026 05:30:10 GMT</lastBuildDate><atom:link href="https://www.diddigest.xyz/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Beatriz]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[diddigest@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[diddigest@substack.com]]></itunes:email><itunes:name><![CDATA[Beatriz Gietner]]></itunes:name></itunes:owner><itunes:author><![CDATA[Beatriz Gietner]]></itunes:author><googleplay:owner><![CDATA[diddigest@substack.com]]></googleplay:owner><googleplay:email><![CDATA[diddigest@substack.com]]></googleplay:email><googleplay:author><![CDATA[Beatriz Gietner]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[A Field Guide to Missing Data (Part 2)]]></title><description><![CDATA[Continuing our exploration into what economists do when the data they need aren&#8217;t there.]]></description><link>https://www.diddigest.xyz/p/a-field-guide-to-missing-data-part-50f</link><guid isPermaLink="false">https://www.diddigest.xyz/p/a-field-guide-to-missing-data-part-50f</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 24 Aug 2026 22:04:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QkIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QkIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QkIg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 424w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 848w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 1272w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QkIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png" width="1055" height="1491" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1491,&quot;width&quot;:1055,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3226218,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/212152914?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QkIg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 424w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 848w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 1272w, https://substackcdn.com/image/fetch/$s_!QkIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde55bb7e-1d4d-4f49-a7ba-df25a3ea1eab_1055x1491.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>The link to the .pdf is <a href="https://beagietner.github.io/webpage/missing_data_field_guide.pdf">here</a>, as promised.</em></p><p>Hello hello. </p><p>So in <a href="https://www.diddigest.xyz/p/a-field-guide-to-missing-data-part">Part 1</a> we mostly covered the absences inside a single dataset: the unobserved confounder, proxies and measurement error, values that go missing or disappear through attrition, and the counterfactual itself. This time the problems get a bit more <em>structural</em>: trying to get to inference when the data are few or the instrument is weak, attempting to combine datasets that were never meant to be combined (at least initially), the world of official stats with its changing classifications, censoring, publication lags and revisions, and the design complications we kept promising, timing and anticipation, spillovers, and whether our variable measures a concept at all.</p><p>In today&#8217;s post we cover:</p><ul><li><p>P11. Definitions or classifications change</p></li><li><p>P12. The sample is small or the event is rare</p></li><li><p>P13. There is little overlap or little identifying variation</p></li><li><p>P14. Records describing the same units can&#8217;t be linked exactly</p></li><li><p>P15. Variables sit in different datasets</p></li><li><p>P16. Values are censored, top-coded or altered for disclosure</p></li><li><p>P17. The official outcome arrives too late</p></li><li><p>P18. The study population differs from the target</p></li><li><p>P19. Treatment timing is uncertain or anticipated</p></li><li><p>P20. One unit&#8217;s treatment affects another unit</p></li><li><p>P21. The available measure doesn&#8217;t settle the concept</p></li><li><p>P22. Data are revised after first release</p></li></ul><div><hr></div><h4>P11. Definitions or classifications change</h4><p>Going back to our training study for this. Suppose we can now track earnings by occupation for twenty years and halfway through the occupation classification gets revised, e.g. &#8220;web developer&#8221; didn&#8217;t exist in the old system, &#8220;typist&#8221; is practically gone in the new one, and a category like &#8220;clerical worker&#8221; was split into others that now occupy different places. The missing object is comparability - over time or across systems. Nothing was mismeasured and nobody dropped out, but the ruler itself changed. Readers from other fields will recognise the situation because it&#8217;s the same one that affects trade data when product codes are revised and health data when ICD versions turn over: it&#8217;s the problem from P07, where this year&#8217;s laptop was not last year&#8217;s laptop, except now it&#8217;s the categories that are never quite the same twice (Heraclitus vibes).</p><p>We have three direct ways to connect the classifications: a crosswalk, a splice built from an overlap, and a backcast. Chain-linking is a related time-series operation: it joins adjacent price or volume comparisons, but it doesn&#8217;t by itself translate an old category into a new one. A crosswalk<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> builds the bridge from concordance information, meaning documented links between the two systems <a href="https://doi.org/10.3233/JEM-2012-0351">(Pierce and Schott 2012)</a>. A splice builds it from an overlap period - a stretch where both systems ran side by side - so we can see directly how they line up. Backcasting applies an existing bridge backwards, rewriting the earlier data in the new system&#8217;s terms, so the clerical workers of 1985 get re-expressed as today&#8217;s categories. And chain linking replaces one long bridge with many short ones: it compares each year only with the next one, inside whichever system covers both, and multiplies those short comparisons into one long series <a href="https://doi.org/10.5089/9781475589870.069">(International Monetary Fund 2018)</a>. The shared assumption is the same for all the first three: a bridge is still meaningful only as far as the observations used to build it, and the employment shares we measured in the overlap years, applied to 1985, assume the occupation&#8217;s composition hadn&#8217;t changed by then, which is exactly the kind of thing that does change.</p><p>These exercises can produce a usable series, but no bridge can erase a substantive change in what a category means. If the new &#8220;software occupations&#8221; group covers different work than the old &#8220;computer operators&#8221; did, a perfectly smooth spliced series is describing a category whose meaning moved beyond it, and the smoothness is the disguise rather than the reassurance. What we then need to do is to disclose clearly what we did since we can&#8217;t expect the reader to audit a bridged series without knowing the many-to-many mappings, the break dates, the overlap years, and how entrants and exits were handled.  What helps is to create a revision triangle<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, which describes how a statistic&#8217;s estimates change across successive releases, with one row per reference period and one column per release date, so reading along a row shows how the estimate for that period changed as new information arrived. It&#8217;s a diagnostic for how the data arrive, but it has nothing to say about a classification break. Revisions are their own absence with their own exercises (which we will see in a second).</p><div><hr></div><h4>P12. The sample is small or the event is rare</h4><p>Suppose our training programme was only a pilot: six sites, two hundred people, and the outcome we care most about - moving into self-employment - happened eleven times. Two separate problems come from this situation, and they need different tools to be handled. The first is that p-values and confidence intervals depend on large-sample approximations, and with six sites and eleven events, the reference distribution behind them can be a poor description of the uncertainty we face in reality. The second is that the model itself can become unstable or separated<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, and with something like eleven events we are certainly in that territory.</p><p>For the first problem we have three standard solutions - randomisation inference, permutation tests, and the wild-cluster bootstrap<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> - and each one builds its comparison world from a different source. Randomisation inference uses the assignment mechanism we are familiar with: if we know exactly how the six sites were chosen for treatment, we can then re-run that assignment thousands of times (on our software of choice), and under the null of no effect for anyone, we can fill in what each site would have shown so the p-value counts how unusual our real allocation looks among all the allocations that could have happened (<a href="https://doi.org/10.1016/bs.hefe.2016.10.003">Athey and Imbens 2017</a>). A permutation test shuffles labels instead, and it needs an exchangeability argument - a reason to believe the shuffled versions are statistically interchangeable with the real one under the null (<a href="https://doi.org/10.1214/13-AOS1090">Chung and Romano 2013</a>). The wild-cluster bootstrap resamples the data in a particular way, and the quality of what it produces depends on the details we have: how much leverage each cluster has, how treatment is spread across clusters, whether the null is imposed when resampling, the choice of weights, and the statistic being bootstrapped. Their shared commonality is that each of these builds its reference world from one source, a known mechanism, an invariance argument, or a resampling construction, and the result is only as good as that source. There&#8217;s also a fourth, albeit less glamorous, family that adjusts the cluster-robust variance and the degrees of freedom directly for having few clusters (<a href="https://doi.org/10.3368/jhr.50.2.317">Cameron and Miller 2015</a>). These are corrections to an approximation rather than new reference worlds, and how well they do still depends on leverage, on balance, and on how many clusters actually got treated.</p><p>For the second problem the tools work on the estimate rather than on the reference world - Firth correction, rare-events corrections, and Bayesian partial pooling. <em>Firth correction</em> is a penalised version of maximum likelihood that reduces small-sample bias and keeps the estimate finite when the data are separated (<a href="https://doi.org/10.1093/biomet/80.1.27">Firth 1993</a>)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>, whereas <em>rare-events corrections</em> deal with the small-sample bias of logistic models when the outcome is uncommon (<a href="https://doi.org/10.1093/oxfordjournals.pan.a004868">King and Zeng 2001</a>). The right correction depends on why the events are rare and on how the sample was drawn, and recovering the population probabilities can require extra information like the population prevalence or the sampling fractions. <em>Bayesian partial pooling</em> shrinks noisy site-level estimates towards each other, trading them against a hierarchical model and a prior (<a href="https://doi.org/10.1214/06-BA117A">Gelman 2006</a>). All three estimate or regularise, and that solves the second problem only. The first one is still standing: before we can report a p-value or a confidence interval, we still need a reference distribution to compare the estimate against, so the tools in this paragraph and the tools in the previous one do different jobs, and a small-sample study often needs one from each.</p><p>Even the best reference world is still small here. With six sites split three against three, there are only twenty ways treatment could have been assigned, so the smallest p-value randomisation inference can produce is 1/20 = 0.05, and that&#8217;s for a one-sided test where our real allocation is the most extreme of them all. With the usual two-sided statistic the floor is generally 2/20 = 0.10 because every allocation has a mirror allocation with the opposite sign. A placebo-style rank deals with the same arithmetic: standing out among a handful of alternatives is a comparison, and it becomes a formal p-value only with an assignment or invariance argument behind it. None of the tools in this entry escapes this, and none of them repairs a weak design, poor measurement, or a target population the sample doesn&#8217;t support. What they do is make the most of six sites while being clear about how much that is - and with six sites, saying which population our estimate speaks for is often the hardest sentence in a study.</p><div><hr></div><h4>P13. There is little overlap or little identifying variation</h4><p>Back to the observational version of our training study where people chose whether to enrol. Suppose the choosing was &#8220;lopsided&#8221;: almost everyone with a university degree signed up, and almost nobody over sixty did. For those two groups the data contain no comparison because there are no untreated graduates and no treated sixty-somethings to compare against, and no amount of regression adjustment would conjure that. The missing object is the support for the original comparison: treated and untreated people who resemble each other across the whole range of characteristics our question covers.</p><p>We have two standard responses - trimming and overlap weighting<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> - and they share one consequence: both change who the estimate is about. Trimming is about dropping the observations where a comparison is unsupported, e.g. the graduates and the over-sixties, and then estimating the effect for the people who remain (<a href="https://doi.org/10.1093/biomet/asn055">Crump et al. 2009</a>). Overlap weighting keeps everyone but counts people more the more &#8220;contested&#8221; their enrolment was so the estimate concentrates on the middle ground where similar people ended up on both sides (<a href="https://doi.org/10.1080/01621459.2016.1260466">Li, Morgan and Zaslavsky 2018</a>). Neither recovers the effect for the full original population, and the recommendation that follows is the same as in P11: explicitly write/say what the new estimand is and who was excluded since &#8220;the effect of training&#8221; and &#8220;the effect of training for people whose enrolment could have gone either way&#8221; are very different sentences.</p><p>Little identifying variation is the same absence in IV form. Suppose we use distance to the nearest training centre as an instrument for enrolment, and distance turns out to barely move anyone&#8217;s decision: the first stage (the relationship between the instrument and the treatment) is then fragile, and everything built on it inherits the fragility. The diagnostics, like the familiar first-stage F statistic, describe that fragility without fixing it (<a href="https://doi.org/10.2307/2171753">Staiger and Stock 1997</a>). Weak-instrument-robust procedures<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>, like the Anderson-Rubin test, change how the uncertainty is computed so that the reported intervals stay valid however weak the first stage is, and that is all they change: they don&#8217;t strengthen the instrument, and they don&#8217;t validate it (<a href="https://doi.org/10.1146/annurev-economics-080218-025643">Andrews, Stock and Sun 2019</a>). Their guarantees also come from a specific setting, and the classical exactness results don&#8217;t carry over on their own to clustered or serially dependent data, so the version we run has to match the dependence structure we have.</p><p>The two halves of this entry are the same absence in two forms: the data hold little information where our question needs it. The defensible responses either shrink the question to where the information is, and are explicit about it, or keep the question and report uncertainty in a way that respects how little there is. Nothing here manufactures variation, and an estimate that ignores that will look precise for the wrong reason. </p><div><hr></div><h4>P14. Records describing the same units can&#8217;t be linked exactly</h4><p>Our training survey has two hundred people, and the tax records that would tell us their true earnings sit in another file with no shared identifier. Linking means deciding which tax record belongs to which respondent using name, birth date and address, and all three can disagree for the same person: names get typed differently, people move houses, and some change their name at marriage. The missing object is unit identity across files - the fact of the matter about who is who.</p><p>Probabilistic linkage<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> turns this decision into a calculation. Each candidate pair of records gets compared field by field, the pattern of agreements and disagreements goes into a comparison model, and the model returns a match probability<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> - a pair that agrees on birth date and address but differs on one letter of the surname is probably the same person, and a pair that agrees only on being called &#8220;J. Silva&#8221; probably isn&#8217;t (<a href="https://doi.org/10.1080/01621459.1969.10501049">Fellegi and Sunter 1969</a>). For this to work we need three things: 1. linking fields that carry information, meaning fields that distinguish people rather than describe half the country, 2. a comparison model we can defend, and 3. a manageable set of candidate pairs, because comparing two hundred respondents against every tax record in the country is neither feasible nor necessary.</p><p>The linkage will make two kinds of error, and they do different levels of damage. A false link attaches someone else&#8217;s tax record to our respondent, so their &#8220;true earnings&#8221; are now a stranger&#8217;s. A missed match loses the record that was there to be found, and the people it loses are not random since the hardest people to link are the ones who - in our example - have moved, divorced or changed names. What this does to our estimate depends on the full structure of those errors, and it need not be a simple attenuation (<a href="https://doi.org/10.1198/016214504000001277">Lahiri and Larsen 2005</a>), so the defensible practice is to carry the linkage uncertainty into the analysis, either by keeping the match probabilities and letting them propagate, or by modelling the linkage and the analysis as one joint problem. And when the files can only match one-to-one (each person has at most one tax record), the candidate links stop being independent choices, so the model should enforce that structure and the joint versions can then average the analysis over several complete, mutually compatible linkages.</p><p>The diagnostics assess the process, and the process is not the population. Clerical review (humans checking a sample of borderline pairs) and quality metrics like match rates tell us how the linkage went for the records we tried to link. A high match score doesn&#8217;t establish that the linked sample represents the target population: if the people we failed to link are the movers and the name-changers, our linked sample is simply a sample of stable lives, and that is a selection problem in a data-management costume.</p><div><hr></div><h4>P15. Variables sit in different datasets</h4><p>This is the sibling of P14 with one twist now: the files describe different people. Our training survey has enrolment and hours for two hundred respondents, and earnings are in a labour-force survey of other people entirely, so there is no pair of records to link, no matter how clever we get. The missing object is the cross-file association: nobody, anywhere, observed training hours and earnings together, and that association is exactly what our question needs.</p><p>One way out is to give every respondent a statistical twin. We look in the other file for someone with the same age, education and from the same region, we copy that person&#8217;s earnings over, and we end up with a completed dataset where everyone has both variables. This exercise is called statistical matching<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>, and the donated association comes from an assumption, usually conditional independence: given the shared variables, hours and earnings are assumed unrelated, so the shared variables are assumed to carry all the connection there is (<a href="https://www.ajs.or.at/index.php/ajs/article/view/vol33%2C%20no1%262%20-%209">R&#228;ssler 2004</a>). The completed file can&#8217;t test this because the association it contains is the one the assumption put there. A completed synthetic microdata file is evidence that an algorithm worked, but we can&#8217;t say that the missing association was recovered.</p><p>The other way out skips completed files: instead of building records that have both variables, we take summary pieces from each file (like the relationship between distance and enrolment from one survey and between distance and earnings from the other) and combine the pieces into the estimate we want. The exercises in this family are moment combination, two-sample IV, and two-sample 2SLS<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>, and they are separate exercises: each uses different pieces from each file, and each needs its own conditions (<a href="https://doi.org/10.1016/S1573-4412(07)06075-8">Ridder and Moffitt 2007</a>; <a href="https://doi.org/10.1162/REST_a_00011">Inoue and Solon 2010</a>). The shared requirements are a) compatible variable definitions, b) the same population relationships holding in both samples, and c) variance conditions that differ by estimator - and these sit on top of the ordinary IV requirements<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>. When those conditions ask more than we can defend, we can report a set instead - a range of answers consistent with what each file separately shows, instead of a fused point estimate carrying an unverified association inside it.</p><p>Whatever we choose, the useful discipline is to name - again - the bridge: say whether the information is being combined across files, across instruments, or across populations because each link has its own weak point. For a statistical match it is the conditional-independence assumption; for a two-sample IV it is two samples that don&#8217;t come from the same population; and neither problem is revealed automatically in the output, which looks complete either way.</p><div><hr></div><h4>P16. Values are censored, top-coded or altered for disclosure</h4><p>One more visit to our training survey :) This time the data arrive incomplete on purpose. Earnings above 150,000 show up as &#8220;150,000+&#8221;, some respondents answered with a bracket (&#8220;between 30,000 and 40,000&#8221;) - not their fault! - and in the public state tables, cells with too few people are left blank. The missing object depends on the release mechanism (the rule the publisher applied before letting the data out), and the mechanisms are a family with true differences<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a>: a top-code replaces values above a limit with the limit itself, a bracket gives a range instead of a number, censoring cuts off at a known point, swapping exchanges values between similar records, rounding coarsens them, noise addition perturbs them, and suppression removes cells outright. Lots of problems here because each of these actions throws away different information, therefore&#8230; yeah, you guessed, each asks for a different exercise.</p><p>For the two oldest mechanisms, we can name the exercises. When the data give us ranges, we can fit a regression that uses the known endpoints of each range instead of a single value, and that is known as interval regression (<a href="https://doi.org/10.2307/2297773">Stewart 1983</a>). When values pile up at a known limit, we can write a model for the outcome we would have seen without the limit and estimate it from the pile-up and the rest, and that is the Tobit model (<a href="https://doi.org/10.2307/1907382">Tobin 1958</a>)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>. Both assume the mechanism is exactly what the model says, a known endpoint or a known limit, and neither is a general answer to confidentiality treatment: a swapped or noise-injected value is not a censored one, and a model built for limits has nothing to say about it.</p><p>The first exercise is&#8230; actually reading (yourself or your agents - helpful if it&#8217;s in a language you don&#8217;t speak), and it comes before any model: the agency documents its release rules, and the documentation says which mechanism affected which variable in which year (<a href="https://www2.census.gov/programs-surveys/cps/techdocs/cpsmar23.pdf">U.S. Census Bureau 2023</a>). For top-coded earnings, the workable repair is a tail model (a distribution fitted to the upper part of the data and used to fill in above the limit) and it earns its keep only when its family, its threshold and its fit are defended for the application at hand; multiple imputation above the top-code, for example, produces answers that are conditional on the tail model chosen, and different defensible tails give different inequality estimates (<a href="https://doi.org/10.1111/j.1467-985X.2010.00655.x">Jenkins et al. 2011</a>). You will meet more general recipes in the wild (e.g. fit a Pareto tail above any top-code, or recover suppressed cells from the published margins with optimisation) and this is the part where I should stop short of recommending them: the case-by-case defence is the method, and the generic version I can give would simply skip it.</p><p>In this entry, the mechanism decides. Treating a top-code as if it were censoring, or noise-injected values as if they were true ones, mixes two release rules into one analysis, and the output comes without a warning. The exercise that protects us costs nothing but time: match what we do to what the documentation says was done.</p><div><hr></div><h4>P17. The official outcome arrives too late</h4><p>Our training subsidy is up for renewal and the state has to decide this quarter whether it worked, but the official earnings series for this quarter will only be published next year. The missing object is the current value at the decision date: the number exists in the world, in payrolls already paid, and the statistical system hasn&#8217;t finished counting it. Meanwhile the timely data pile up (e.g. unemployment claims, job postings, payroll-processor records) and the exercise of estimating the official number from them before it&#8217;s released is called nowcasting (<a href="https://doi.org/10.1016/j.jmoneco.2008.05.010">Giannone, Reichlin and Small 2008</a>).</p><p>The standard tools to solve this problem are bridge equations, factor models and mixed-frequency models<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>, and they have one thing in common in order to make them work: the information set must be dated. A bridge equation links the quarterly target to monthly indicators through a regression, a factor model compresses many indicators into a few common movements and reads the target off them, and mixed-frequency models let monthly and quarterly data live in one equation without forcing them to the same calendar (<a href="https://doi.org/10.1016/B978-0-444-53683-9.00004-9">Ba&#324;bura et al. 2013</a>). The shared &#8220;discipline&#8221; concerns what the model is allowed to see: for any date we nowcast, only the data releases that existed on that date can enter. Later revisions and indicators published afterwards were not available to the decision we&#8217;re trying to inform.</p><p>We need to be careful because that discipline is also where we can get the evaluation wrong. If we test our nowcasting model against history using today&#8217;s data, we commit a crime in two types of hindsight: revised values of the indicators - which are cleaner than what was on the screen at the time - and series that hadn&#8217;t been released yet on the dates we&#8217;re pretending to nowcast. A fair historical test recreates what was known at each date by using real-time vintages, meaning the data exactly as they stood on each past day<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a> (<a href="https://doi.org/10.1016/S0304-4076(01)00072-0">Croushore and Stark 2001</a>). A model that looks sharp against final data and ordinary against real-time data isn&#8217;t being treated unfairly by the second test; the second test is actually the true one.</p><p>Nowcasting also has a neighbour it gets confused with. A nowcast estimates a current value nobody has measured yet, while <em>backcasting</em> (see previous entries) builds a historically comparable series for periods already measured under other rules; a model can do one well and the other badly because the two jobs need different things from it. And forecast accuracy is a statement about prediction only: an indicator can predict earnings movements while measuring nothing we would interpret as the effect of training, so a model that nowcasts well tells us what the number will be, and says nothing on why.</p><div><hr></div><h4>P18. The study population differs from the target</h4><p>Our pilot went well and the state now wants to roll the programme out to everyone. The missing object is the target-population effect: the effect in the population the decision is about, which is <em>not</em> the population we studied. The two can differ because effects differ across people (e.g. training may help younger workers more than older ones) and our six pilot sites had their own mix of ages, industries and cities, while the state has another. A variable like age here is called an effect moderator, meaning a characteristic that changes the size of the effect, and the pilot&#8217;s mix of moderators is baked into its headline number. </p><p>We have two<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a> ways to move from the study to the target - standardisation and weighting - and they  have four requirements in common. Standardisation computes the effect within subgroups (e.g. for younger and older workers separately) and then re-averages those effects using the target population&#8217;s subgroup shares instead of the study&#8217;s (<a href="https://doi.org/10.1093/aje/kwq084">Cole and Stuart 2010</a>). Weighting goes the other way around: it reweights the study so it resembles the target, giving more weight to the kinds of people the study under-represents, with weights built from the odds of being in the study rather than the target; these are inverse odds of sampling weights<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a> (<a href="https://doi.org/10.1093/aje/kwx164">Westreich et al. 2017</a>). Both need the study effect itself to be right (internal validity), the variables to mean the same thing in both datasets, the effect to carry across given the shared variables (conditional effect exchangeability), and every kind of person in the target to have some counterpart in the study (positivity, the same requirement we met with the missing respondents). The last one is the hard limit: if the pilot enrolled nobody over sixty, no weighting scheme produces the over-sixty effect because there is nothing to reweight.</p><p>The moderators we didn&#8217;t measure get a different exercise. If we suspect the pilot sites differed from the state in something unrecorded (how motivated the local caseworkers were, for example) we can vary that unmeasured moderator in a sensitivity analysis and report how far the target effect could move. The calculation of <a href="https://doi.org/10.1016/j.econlet.2019.02.020">Andrews and Oster (2019)</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a> belongs here: it uses how the observed characteristics relate to participation to calibrate how much the unobserved ones could move the effect, under its own selection model and a small-selection approximation, and it is a stress test under those terms, not an automatic bound, correction, or transport estimator.</p><p>A study that sits inside its target - our pilot counties are part of the state - and a target that is a different population altogether - another state asking whether our results apply there - are different problems with different names: generalisability for the first, transport for the second, and a study should say which one it is doing. And the weights themselves invite one very specific question. Two weighting schemes can produce similar-looking columns of numbers while resting on different assumptions, one about how the study was sampled and one about how effects carry across populations, so the thing to ask of any weight is which assumption it encodes because the numbers look the same either way.</p><div><hr></div><h4>P19. Treatment timing is uncertain or anticipated</h4><p>When we set up the subsidy example, we assumed nobody responded before the policy took effect. This entry is for when that assumption is the problem. The subsidy legally starts in January, but the bill passed in October, the local news covered it, and training providers began advertising in November, so by the time January arrives, some of the response has already happened. The missing object is the behavioural onset of treatment: the date behaviour started responding which is not required to equal the legal date, and which no dataset column announces.</p><p>When the institutional record leaves several defensible onset dates (the bill&#8217;s passage, the news coverage, the legal start, etc), the exercise is to run the analysis under each coding and report the set. Each coding changes real things: who counts as treated from when (the cohorts), which periods serve as untreated comparison, and the estimand itself, since &#8220;the effect of the subsidy from January&#8221; and &#8220;the effect of the subsidy from October&#8221; are different questions. This exercise is a stress test tailored to our design, and there is no general correction behind it: not knowing when treatment began is not a statistical problem with a standard fix, but rather it is a piece of missing institutional knowledge, and the codings are how we show what it costs us.</p><p>Does all of this ring a bell? Anticipation has its own tools when the early response is real and we are not talking about a dating error. People and firms respond before formal adoption whenever they can see the policy coming and acting early pays, and a period we labelled &#8220;untreated&#8221; then already contains part of the effect<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> (<a href="https://doi.org/10.1016/j.jpubeco.2015.01.001">Malani and Reif 2015</a>). The group-time estimators most of us already use can account for this to some extent: we declare an &#8220;anticipation horizon&#8221; (a stated number of periods before adoption in which effects are allowed to exist), and the estimator moves its baseline back accordingly<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a> (<a href="https://doi.org/10.1016/j.jeconom.2020.12.001">Callaway and Sant&#8217;Anna 2021</a>). We &#8220;declare&#8221; the horizon and we try our best to defend it with institutional facts, such as when the bill passed or when coverage began, rather than estimated from the data since the data can&#8217;t distinguish early effects from differential trends on their own. We should do that.</p><p>There&#8217;s one more common thing we can do: we drop the periods right around adoption and compare &#8220;cleanly before&#8221; with &#8220;cleanly after&#8221;, which is called a donut. A donut can be informative, and it has two costs we should account for: it discards the transition data, and it can change the target contrast because an effect measured between &#8220;well before&#8221; and &#8220;well after&#8221; does not say much about the adoption period itself. A donut is a <em>disclosed design choice</em>, and should not be considered a universal repair for ambiguous timing. Our duty, present throughout this entry, is the same one: date the treatment with institutional evidence, show the codings, and treat the calendar as part of the design rather than as a setting nobody chose.</p><div><hr></div><h4>P20. One unit&#8217;s treatment affects another unit</h4><p>Our subsidy trains thousands of workers, and trained workers compete for jobs with untrained ones, so an untrained worker in a treated state is not untouched by the programme: some of what happens to their wages happens because their competitors got trained. Across the border the story continues. Firms in the comparison state recruit workers from ours, so the &#8220;untreated&#8221; comparison may contain a &#8220;diluted dose&#8221; of the treatment. In every entry so far we assumed that each person&#8217;s outcome depends only on their own treatment. In this entry, such assumption holds no more. The missing object is the exposure mapping, which is the rule that says who can affect whom, over what distance, and through which channel. The data don&#8217;t record that rule, so the mapping has to come from an <em>argument</em> about how the market works.</p><p>The designs that handle interference all work by restricting it to somewhere we can see<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a>. One restriction allows effects to travel within groups but not between them (within a local labour market but not across markets). This is called partial interference<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a> (<a href="https://doi.org/10.1198/016214508000000292">Hudgens and Halloran 2008</a>). A more general version states outright which other units count for my outcome and how (such as workers in my commuting zone, weighted by how much we compete), and this is an exposure mapping in the formal sense. Once we declare one, contrasts such as &#8220;treated versus untreated, holding neighbours&#8217; exposure fixed&#8221; become estimable under the design (<a href="https://doi.org/10.1214/16-AOAS1005">Aronow and Samii 2017</a>). This mapping actually does a bit more than just enable estimation. It defines the estimand: the effect of the subsidy now means &#8220;the effect given who counts as exposed to whom, under this specific mapping&#8221;.</p><p>A common check re-runs the analysis with different reaches (spillovers allowed within 10 km, then 20, then 50). This kind of radius sensitivity is often called a ring analysis. What we get from it is sensitivity, meaning how much the estimates move as the assumed reach changes. What we don&#8217;t get is the true spillover radius. Each radius is a different mapping, each mapping may define a different estimand, so the rings compare answers to related questions instead of closing in on one answer (<a href="https://doi.org/10.1093/biomet/asad019">S&#228;vje 2024</a>). And when the mapping is misspecified, we end up with two errors at once<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a>, each with its own name: a definitional error, because we estimated a contrast other than the one we intended, and an identification error, because even that contrast may be estimated with bias.</p><p>When the network or the geography is too weakly known to defend one mapping, the better report is the set of estimates across the defensible mappings, so the reader can see what the missing knowledge costs us. A single exposure model hides that cost. And the assumption we started this entry with should be written down in our studies like any other. No interference is an assumption like parallel trends or missing at random, it can be wrong in ways that change the answer, and it spent nineteen entries being assumed without anyone saying so - which is exactly how an assumption ends up treated as a harmless default.</p><div><hr></div><h4>P21. The available measure doesn&#8217;t settle the concept</h4><p>Suppose the state asks whether our training programme improved workers&#8217; economic security, and we answer them with earnings. Earnings is a real variable, well measured for once, but it is still not the concept: economic security has something to do with stability, benefits, and how easily a job is lost, and a higher wage on a precarious contract may not be more of it. The missing object here is a defensible operationalisation, meaning a way of turning the concept we care about into a variable we can measure, plus the argument that the two line up. Back in P02 the variable existed and was poorly measured. In this entry the measurement can be perfect and the question still remains because what&#8217;s in doubt is whether we (well, others) actually measured the right thing.</p><p>The main exercise in this entry is to check the measure against what theory expects from the concept. If we have prior evidence that earnings really tracks economic security, it should move with the things security moves with (savings, reported financial stress, job tenure) and it should not move one-for-one with things that are clearly not security (like a temporary overtime spike). This is called construct validation (<a href="https://doi.org/10.1037/h0040957">Cronbach and Meehl 1955</a>). A stronger version of it measures several concepts with several methods at once and checks the full pattern: different measures of the same concept should agree, and measures of different concepts should disagree enough to count as <em>different</em>. This is the multitrait-multimethod comparison<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a> (<a href="https://doi.org/10.1037/h0046016">Campbell and Fiske 1959</a>). Both exercises can show us where a measure behaves as the concept should. Neither can prove the indicator is the concept, because behaving as expected is evidence, and identity is not something evidence of this kind can establish.</p><p>Economics runs on examples of this gap. Patent counts are an indicator related to innovation, publication measures indicate research output or visibility, employment is one observable labour-market outcome, and nominal expenditure is a monetary input measure. These are illustrations of common practice, and none of them is a validated proxy relationship. Whether any of them measures the intended concept depends on the question and on domain evidence, which is an argument each study has to make for its own case rather than inherit from the literature.</p><p>When the validation exercise does not support the measure, we have two options, and the less obvious one is often better. We can a) hunt for a better measure of the same concept, or b) we can revise the concept, and ask a question the evidence can answer (not &#8220;did training improve economic security&#8221; but &#8220;did training raise earnings&#8221;, stated as exactly that<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a>) (<a href="https://doi.org/10.1017/S0003055401003100">Adcock and Collier 2001</a>; <a href="https://doi.org/10.1037/0033-2909.110.2.305">Bollen and Lennox 1991</a>). A narrower question answered with a variable we understand beats a grand question answered with a variable we can only dream about.</p><div><hr></div><h4>P22. Data are revised after first release</h4><p>This is the shortest entry in the guide because we&#8217;ve already met its tools. The revision triangle appeared with the classification breaks, and the vintages and the pseudo-real-time evaluation appeared with the nowcasts. What&#8217;s left is the rule that ties them all together. The missing object is the vintage appropriate to our question, and &#8220;appropriate&#8221; changes with the question. A real-time policy analysis should use the latest data available at the decision date while a historical measurement exercise can use first releases, later releases, or one fixed vintage, as long as it says which. The final vintage is not automatic ground truth for a real-time question because the decision-maker we&#8217;re studying never actually saw it (<a href="https://doi.org/10.1016/S0304-4076(01)00072-0">Croushore and Stark 2001</a>).</p><p>The diagnostics here are the exercises we already know. Vintage panels and revision triangles show how the estimates changed as information arrived, and a pseudo-real-time evaluation reruns our whole procedure with each date&#8217;s information set (<a href="https://doi.org/10.5089/9781475589870.069">International Monetary Fund 2018</a>). What these expose is revision risk, meaning how much our answer depends on which release we used. What they can&#8217;t do is decide which vintage is the right one for every estimand, because that choice belongs to the question, and the question is ours to state.</p><div><hr></div><h4>The eight verbs</h4><p>After 22 problems, you might have noticed a pattern. The exercises we walked through are variations on a small set of responses, and the easiest way to see this is to write each response as a verb. We end up with eight of them. A single study usually does several at once, so treat the verbs as a way of reading applied work rather than boxes to sort papers into, and when a procedure combines them, we pull it apart and name the verb for each piece.</p><p><strong>Bound it.</strong> If the data cannot give us one answer, report the range of answers they can support. The width of that range tells us something too: the wider it is, the more we would have to assume to get to a precise answer (<a href="https://doi.org/10.1146/annurev.economics.050708.143401">Tamer 2010</a>).</p><p><strong>Stress it.</strong> Change an assumption and see when the conclusion changes. The useful result is not that the estimate moved, but how far we had to push the assumption before it did (<a href="https://doi.org/10.1016/j.jeconom.2013.08.016">Lu and White 2014</a>).</p><p><strong>Reconstruct it.</strong> Sometimes we can use other information to fill in what we cannot observe directly. The important question is then whether the link we used to fill the gap is credible. A completed dataset does not make the missing information observed.</p><p><strong>Combine, or transport it.</strong> Sometimes the information exists, just somewhere else. We can bring together datasets or populations, but doing so requires a reason to believe that what connects them is comparable. Successfully merging two files is not evidence that this assumption holds.</p><p><strong>Generate reference worlds.</strong> Sometimes what is missing is the comparison itself: what would our statistic look like under some alternative world? We can generate that comparison in different ways, but what we learn from it depends on how that alternative world was constructed.</p><p><strong>Change the question.</strong> Sometimes the original question asks more than the evidence can answer. In that case, the best option may be to ask a narrower question and say clearly that the target has changed.</p><p><strong>Estimate, or regularise it.</strong> Sometimes the problem is that there is too little information to estimate the model reliably. Regularisation can make the estimate more stable, but it does so by adding structure. That structure is part of the answer.</p><p><strong>Report, or audit it.</strong> Sometimes the missing object is the analytical record itself. Better reporting, shared code and computational checks can tell us what was done and whether we can reproduce it (<a href="https://doi.org/10.1257/jel.20171350">Christensen and Miguel 2018</a>; <a href="https://doi.org/10.1257/jep.29.3.61">Olken 2015</a>). They cannot tell us whether the original research design was valid.</p><p>Writing things this way forces us to say what we actually learned. Different exercises produce different kinds of evidence, and they should not all end up in a table labelled &#8220;robustness&#8221;. They also come with a trade-off: if the data cannot answer the original question on their own, we either make an additional assumption, accept a less precise answer, or ask a different question.</p><p>This is where robustness checks can become ritualistic. A useful check starts with a specific problem and tells us something about it. A ritual check starts with a familiar method and works backwards. Some of this is probably driven by publication incentives and convention, but we should be careful about blaming referees for all of it.</p><p>There is good evidence that reasonable analytical choices can produce different results and that selective reporting affects what gets published (<a href="https://doi.org/10.1177/1745691616658637">Steegen et al. 2016</a>; <a href="https://doi.org/10.1038/s41562-020-0912-z">Simonsohn, Simmons and Nelson 2020</a>). We also know something about what editors and referees value and about selection during publication (<a href="https://doi.org/10.1257/jep.31.1.231">Berk, Harvey and Hirshleifer 2017</a>; <a href="https://doi.org/10.1257/aer.20210795">Brodeur et al. 2023</a>). What we do not know is how much of the robustness machinery in published papers is there because referees asked for it. That story is plausible, but the evidence does not get us that far.</p><p>The same point applies to reproducibility. Being able to rerun an analysis is quite useful, but it does not settle whether the design was good in the first place or whether the finding would survive new data (<a href="https://doi.org/10.1561/104.00000053">Chang and Li 2022</a>). Reproducibility can close one gap while leaving the important one untouched.</p><p>So when something we need is missing, start there. Say what is missing and work out what the available evidence can still tell us. Sometimes the answer will be weaker or narrower than the one we wanted. That is fine. The standard robustness table can wait.</p><p>In the end, it is quite a simple rule: say what the exercise does, say what you had to assume to do it, and say what is still missing. A robustness check is useful when it changes what we know, not when it merely gives us another column for the table.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>A <em>concordance</em> is the documented mapping between two classification systems, and the issue here is that it&#8217;s rarely one-to-one in the sense of one old category splitting into several new ones, and several old ones merging into one new one. Allocating a many-to-many mapping needs weights (something like employment, output or trade shares from the overlap period), and those weights are where the assumptions enter, since a share measured in 1997 is being asked to describe 1985. <a href="https://doi.org/10.3233/JEM-2012-0351">Pierce and Schott (2012)</a> built the standard concordance between ten-digit US trade codes and industry classifications and their paper kind of lists everything that can go wrong. The same structure exists in other fields under other names: medicine crosswalks ICD-9 to ICD-10 with official equivalence mappings, and education does it between ISCED versions. Search terms: &#8220;concordance&#8221;, &#8220;crosswalk&#8221;, &#8220;classification change&#8221;, and for the trade version search &#8220;Pierce Schott concordance&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>A <em>revision triangle</em> is a table with one row per reference period (e.g., GDP in 2019Q4) and one column per release date, so reading along a row tells you how the estimate for that quarter changed as later releases arrived. It&#8217;s called a triangle because recent periods have had fewer releases so the table has a staircase edge. Agencies and real-time-data archives publish these, and they answer questions like &#8220;how large is a typical revision?&#8221; and &#8220;do first releases systematically under- or overstate growth?&#8221;. What a triangle can&#8217;t do is repair a classification break because a break is a change in what was being counted, and it shows up in the triangle as a discontinuity that no amount of averaging across releases eliminates. Search terms: &#8220;revision triangle&#8221;, &#8220;real-time data&#8221;, &#8220;data vintages&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><em>Separation</em> is the situation where our predictors split the two outcome groups perfectly: you could draw a line with every self-employment case on one side and every non-case on the other, e.g. all eleven cases sitting in trained sites and none anywhere else, so the variable &#8220;trained site&#8221; sorts the yes&#8217;s from the no&#8217;s on its own. Ordinary maximum likelihood then has no finite best answer because making the coefficient larger always fits a little better, so the estimate runs off towards infinity, the standard errors explode, and some software stops with a warning while some reports enormous numbers. &#8220;Unstable&#8221; means that the estimate stays finite but swings considerably when one or two observations change. The fixes are in the main text (Firth correction, rare-events adjustments); the implementations are <code>logistf</code> in R, <code>firthlogit</code> in Stata, and <code>relogit</code> for the King&#8211;Zeng version. Search terms: &#8220;separation logistic regression&#8221;, &#8220;complete separation&#8221;, &#8220;quasi-separation&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><a href="https://doi.org/10.1162/rest.90.3.414">Cameron, Gelbach and Miller (2008)</a> found substantial improvements in simulations with as few as five to ten clusters, though that&#8217;s no general guarantee because with treatment concentrated in one or two clusters, or badly unbalanced cluster sizes, the bootstrap can still perform poorly. The practitioner&#8217;s guide to all of it is <a href="https://doi.org/10.3368/jhr.50.2.317">Cameron and Miller (2015)</a>. One configuration is tough even for the bootstrap: few <em>treated</em> clusters, e.g. six sites of which only one or two got the programme, where no resampling scheme has much to work with and the randomisation-inference route tends to be the more defensible one. The implementation everyone uses is <code>boottest</code> in Stata (also available in R); search terms: &#8220;wild cluster bootstrap&#8221;, &#8220;few treated clusters&#8221;, &#8220;cluster robust inference&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>To be fully correct, Firth&#8217;s bias-reduction adjustment modifies the likelihood score (<a href="https://doi.org/10.1093/biomet/80.1.27">Firth 1993</a>), and its use to obtain finite logistic estimates under separation was established by <a href="https://doi.org/10.1002/sim.1047">Heinze and Schemper (2002)</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Both tools run on the propensity score, which is the estimated probability of enrolling given a person's observed characteristics. A score near 0 or 1 flags a person whose treatment status was close to decided in advance, and this is exactly where comparisons are unsupported. The practical trimming rule from <a href="https://doi.org/10.1093/biomet/asn055">Crump et al. (2009)</a> is to keep observations with scores between roughly 0.1 and 0.9, and it&#8217;s a rule of thumb rather than a law. Overlap weights, from <a href="https://doi.org/10.1080/01621459.2016.1260466">Li, Morgan and Zaslavsky (2018)</a>, weight each person by the probability of the treatment they didn&#8217;t get, so the weight peaks for people with scores near 0.5 and fades toward the extremes, and the resulting estimand even has its own name, the average treatment effect in the overlap population (ATO). Search terms: &#8220;common support&#8221;, &#8220;propensity score trimming&#8221;, &#8220;overlap weights&#8221;, &#8220;limited overlap&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>The famous F &gt; 10 rule of thumb for the first stage comes from <a href="https://doi.org/10.2307/2171753">Staiger and Stock (1997)</a>, and it was a rough guide from the start; the modern survey treatment of what to do instead is <a href="https://doi.org/10.1146/annurev-economics-080218-025643">Andrews, Stock and Sun (2019)</a>. The robust procedures work by inverting tests rather than trusting the usual standard errors: the Anderson-Rubin test asks, for each candidate effect size, whether the data reject it, and collects the survivors into a confidence set. One feature of these sets is worth knowing before you meet one: when the instrument is weak enough, the set can be wide or even unbounded, and that is the method being truthful about the data rather than the method misbehaving. Anderson-Rubin also has company: conditional likelihood ratio tests can say more in the settings they cover, and the first-stage F we all learned to check has to match the covariance structure of the data, so the usual threshold is not a certificate once heteroskedasticity, clustering or serial dependence enter the picture (<a href="https://doi.org/10.1146/annurev-economics-080218-025643">Andrews, Stock and Sun 2019</a>). The implementations are <code>weakiv</code> and <code>ivreg2</code> in Stata and <code>ivmodel</code> in R. Search terms: &#8220;weak instruments&#8221;, &#8220;Anderson-Rubin test&#8221;, &#8220;conditional likelihood ratio test&#8221;, &#8220;first-stage F statistic&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>The Fellegi-Sunter framework runs on two probabilities per linking field: how often the field agrees when two records truly are the same person (high, but not 1, because of typos), and how often it agrees by coincidence when they aren&#8217;t (low for birth dates, higher for common names). Each field&#8217;s agreement or disagreement contributes evidence, the contributions add up to a score, and two thresholds turn scores into decisions: clear matches above, clear non-matches below, and a middle zone sent to clerical review. The &#8220;manageable candidate pairs&#8221; requirement has its own name: <em>blocking</em>. Rather than comparing everyone with everyone, we compare only within blocks that share something cheap and reliable, like a birth year, and accept that a wrong birth year now means a missed match. Implementations: <code>fastLink</code> in R, <code>dtalink</code> in Stata, and <code>Splink</code> in Python. Search terms: &#8220;probabilistic record linkage&#8221;, &#8220;Fellegi-Sunter&#8221;, &#8220;blocking record linkage&#8221;, &#8220;linkage error bias&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Some later models turn that evidence into an actual match probability, but only by adding assumptions about the linkage process (<a href="https://doi.org/10.1080/01621459.2016.1148612">Sadinle 2017</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>The mechanics are usually a donor pool and a distance rule: the other file&#8217;s records are the donors, we measure similarity on the shared variables, and the closest donor&#8217;s value gets copied over (or several donors get averaged, or a model predicts the value from the shared variables - the flavours differ, the logic doesn&#8217;t). The name for what&#8217;s being assumed is the <em>conditional independence assumption</em>, and we can see how it works in practice in our example: if motivated people train more hours <em>and</em> earn more in ways age, education and region don&#8217;t capture, the matched file will understate the hours-earnings connection, and nothing in the file will show it. Without conditional independence, the data from the two files only bound the association rather than pin it down, which is why the fusion literature increasingly reports uncertainty about the match itself. Search terms: &#8220;statistical matching&#8221;, &#8220;data fusion&#8221;, &#8220;conditional independence assumption&#8221;, and for the bounds version the famous &#8220;Fr&#233;chet bounds&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Two-sample IV runs an IV design with its two halves in different datasets: one file has the instrument and the treatment, so it can estimate the first stage, and the other has the instrument and the outcome, so it can estimate the reduced form, and the ratio of the two estimates the effect - under the usual IV assumptions - even though no single person has all three variables. In our example, one survey shows how distance to a training centre moves enrolment, the other shows how distance relates to earnings. The requirements are that both samples come from the same population and that the variables mean the same thing in both, which sounds ok and reasonable until two surveys define &#8220;earnings&#8221; differently.  Drawing both samples from the same population is the cleanest case; there are methods that let the samples differ (<a href="https://doi.org/10.1214/18-STS692">Zhao et al. 2019</a>), but they work by stating which relationships stay stable across the two, not by waving the difference away. Two-sample IV and two-sample 2SLS are not the same estimator: <a href="https://doi.org/10.1162/REST_a_00011">Inoue and Solon (2010)</a> show they differ in finite samples and in efficiency, so you have to choose between them. Search terms: &#8220;two-sample IV&#8221;, &#8220;two-sample 2SLS&#8221;, &#8220;data combination econometrics&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>Splitting the moments across two files doesn&#8217;t excuse the instrument from its &#8220;day job&#8221;: it still has to move treatment, affect the outcome only through treatment, and be unrelated to the unobserved determinants of earnings.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>The confidentiality mechanisms deserve their own definitions because the words look interchangeable and they aren&#8217;t, actually. Swapping exchanges values, like earnings, between records that look similar on other variables, so each record is real but some are wearing a neighbour&#8217;s number. Noise addition perturbs values with random error whose distribution the agency knows and we usually don&#8217;t. Suppression comes in two rounds: primary suppression blanks the sensitive cell, like the average earnings of the three trainees in a small county, and secondary suppression blanks other cells too since a blank cell inside published row and column totals could otherwise be recovered by subtraction. The umbrella term for all of it is statistical disclosure control, and agencies publish their general approach while keeping some parameters secret <em>on purpose</em>, which is a constraint our analysis has to live with rather than an oversight to complain about - we can still complain tho. Search terms: &#8220;statistical disclosure control&#8221;, &#8220;cell suppression&#8221;, &#8220;data swapping&#8221;, &#8220;top-coding&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>Interval regression and Tobit get confused because both involve limits, and the difference decides the assumptions we get. On one hand, interval regression knows each observation&#8217;s endpoints, so it needs a distributional assumption only to spread probability inside the brackets (kind of a mild request). On the other hand, the Tobit model assumes a full latent outcome (usually ~N) and its estimates lean on such normality harder than we would expect, so a heavy-tailed or skewed truth can move the results. If the question is about conditional quantiles rather than a latent mean, censored quantile regression is the standard option (<a href="https://doi.org/10.1016/0304-4076(86)90016-3">Powell 1986</a>), and it still needs the censoring rule and the quantile restriction to be right, while heavy censoring can leave some quantiles unidentified no matter what we do. There&#8217;s also a third &#8220;animal&#8221;: tail models for top-codes. <a href="https://doi.org/10.1111/j.1467-985X.2010.00655.x">Jenkins et al. (2011)</a> fit a flexible distribution (the generalised beta of the second kind, GB2) to the upper range and impute above the limit, and their multiple-imputation machinery brings the tail uncertainty into the standard errors, which is what makes the exercise conditional-on-model in a way we can inspect. Implementations: <code>intreg</code> and <code>tobit</code> in Stata, <code>survreg</code> in R; search terms: &#8220;interval regression&#8221;, &#8220;Tobit model&#8221;, &#8220;top-coded earnings&#8221;, &#8220;GB2 distribution&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>The simplest one is a bridge equation: you aggregate the monthly indicators to quarterly frequency and regress the target on them, so the &#8220;bridge&#8221; crosses from what&#8217;s published monthly to what&#8217;s published quarterly. A dynamic factor model assumes many indicators move together because a few underlying forces drive them, extracts those forces, and nowcasts from the forces rather than the indicators, which is why it digests dozens of series without falling apart; the version in <a href="https://doi.org/10.1016/j.jmoneco.2008.05.010">Giannone, Reichlin and Small (2008)</a> also updates its nowcast release by release, so you can read how much each morning&#8217;s data changed the estimate, a decomposition the literature calls the &#8220;news&#8221;. Mixed-frequency regressions (search &#8220;MIDAS&#8221;) skip the aggregation step and let a quarterly variable sit on the left with monthly regressors on the right. Many of these models are written in state-space form and updated with a Kalman filter, which is handy when releases arrive on different calendars and the dataset has a ragged edge (<a href="https://doi.org/10.1016/B978-0-444-53683-9.00004-9">Ba&#324;bura et al. 2013</a>). Search terms: &#8220;nowcasting&#8221;, &#8220;bridge equation&#8221;, &#8220;dynamic factor model&#8221;, &#8220;nowcast news decomposition&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>A <em>vintage</em> is a snapshot of a data series as it stood on a given date, and vintage archives exist because researchers kept needing yesterday&#8217;s version of the truth: the Philadelphia Fed&#8217;s Real-Time Data Set, built from the work in <a href="https://doi.org/10.1016/S0304-4076(01)00072-0">Croushore and Stark (2001)</a>, and the St. Louis Fed&#8217;s ALFRED archive are the standard sources for US series. An evaluation that uses them is called a <em>pseudo-real-time evaluation</em>: pseudo because we run it now, real-time because every number the model sees is dated to what existed then. The evaluation has to close two separate gaps, and closing only one still overstates the model&#8217;s accuracy: the indicator values must be the unrevised ones that were on record at the time, and the list of series itself must match what had been published at all on that date. Search terms: &#8220;real-time vintage data&#8221;, &#8220;pseudo real-time evaluation&#8221;. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>If we are being really honest, there&#8217;s a third that combines them: the doubly robust procedures. These stay consistent when either the outcome model or the participation model is right, though only inside the transport assumptions we already signed up for (<a href="https://doi.org/10.1146/annurev-statistics-042522-103837">Degtiar and Rose 2023</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Let&#8217;s think in numbers. Suppose training raises earnings by 800 for under-forties and by 200 for over-forties. Our pilot was 70% under-forty so its headline effect is 0.7 &#215; 800 + 0.3 &#215; 200 = 620. The state is 40% under-forty so the standardised target effect is 0.4 &#215; 800 + 0.6 &#215; 200 = 440. Same study, same subgroup effects, and a headline that drops by a third when the mix changes. This arithmetic we are doing is the entire point of the exercise, and it only needed the subgroup effects and the target&#8217;s shares. Inverse odds weighting reaches the same destination by another road: model each study person&#8217;s probability of being in the study rather than the target given their characteristics, weight by the inverse odds, and the reweighted study behaves like the target on everything the model included. How to choose? Go for standardisation when the moderators are few and the subgroup effects are estimable, and weighting when the characteristics are many. Search terms: &#8220;standardization&#8221;, &#8220;transportability&#8221;, &#8220;inverse odds of sampling weights&#8221;, &#8220;generalizability of trial results&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>The Andrews-Oster logic is: if the people and places that joined the study differ from the target in ways we can see, and those visible differences relate to the effect, then the visible selection gives us a yardstick for how much the invisible selection could move the result, under the assumption that the invisible part behaves like the visible part (the selection model) and that selection into the study is not too strong (the small-selection approximation). The answer moves with both assumptions, which is why the output reads as a stress test: it answers &#8220;how bad could unmeasured selection be if it resembles the measured kind&#8221;, and it doesn&#8217;t answer &#8220;what is the corrected target effect&#8221;. The distinction between generalisability and transport also decides the machinery: when the study is nested in the target, the target&#8217;s covariate distribution is known from the start; when the target is a separate population, its data enter as a second dataset with all of P15&#8217;s compatibility questions attached. Search terms: &#8220;Andrews Oster external validity&#8221;, &#8220;selection on observables and unobservables&#8221;, &#8220;transportability&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p><a href="https://doi.org/10.1016/j.jpubeco.2015.01.001">Malani and Reif (2015)</a> studied tort reform, where physicians responded while reforms were still working through legislatures, and their point cuts two ways. A &#8220;pre-trend&#8221; can be anticipation rather than a violation of comparability: the groups were comparable, and the treated one started responding early. And anticipation contaminates the baseline. If the pre-period already contains part of the effect, the before-after comparison subtracts that early response from the later one. When the anticipation runs in the same direction as the effect, the estimate usually comes out too small, but with other dynamics the bias can go either way (<a href="https://doi.org/10.1016/j.jpubeco.2015.01.001">Malani and Reif 2015</a>). The two readings ask for different responses, which is why the pre-trend discussion from P08&#8217;s diagnostics and this entry&#8217;s dating problem are neighbours rather than the same issue. Search terms: &#8220;anticipation effects&#8221;, &#8220;pre-trends as anticipation&#8221;, &#8220;announcement effects&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>In the <a href="https://doi.org/10.1016/j.jeconom.2020.12.001">Callaway and Sant'Anna (2021)</a> framework, effects are estimated per cohort and period, and the comparison baseline is the cohort&#8217;s last untreated period. Declaring an anticipation horizon of one period tells the estimator that the last <em>untreated</em> period is one earlier than the adoption date, so the baseline moves back and the adoption-adjacent period gets treated as possibly affected. The declaration costs us data since the baseline now sits further from the outcomes it anchors, and that is the price of admitting the response started early. Implementations: the <code>did</code> package in R and <code>csdid</code> in Stata, both with an anticipation argument. Search terms: &#8220;group-time average treatment effects&#8221;, &#8220;anticipation horizon&#8221;, &#8220;Callaway Sant&#8217;Anna&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>In experiments we can build this variation on purpose with a two-stage design, which first randomises the treatment intensity across groups and then randomises treatment within each group (<a href="https://doi.org/10.1198/016214508000000292">Hudgens and Halloran 2008</a>), and that assignment structure is what gives us separate direct and spillover comparisons.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>The partial-interference setting comes from vaccines, where the spillover is entirely the point: your vaccination protects me. <a href="https://doi.org/10.1198/016214508000000292">Hudgens and Halloran (2008)</a> showed that with interference confined to groups, four effects become definable and estimable. The direct effect compares treated with untreated people in the same group. The indirect effect compares untreated people in a heavily treated group with untreated people in an untreated group, and it is the formal version of herd immunity. The total effect combines the two, and the overall effect compares whole groups under different treatment intensities. The same four questions translate to our subsidy study: the indirect effect is what happens to untrained workers&#8217; wages when many colleagues get trained, which for a labour economist is the displacement question. Search terms: &#8220;partial interference&#8221;, &#8220;direct and indirect effects&#8221;, &#8220;spillover effects&#8221;, and the no-interference assumption&#8217;s formal name, &#8220;SUTVA&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>The two-error decomposition is the useful lesson of <a href="https://doi.org/10.1093/biomet/asad019">S&#228;vje (2024)</a>: when the declared mapping is wrong, part of the damage is that our estimand changed, and part is ordinary bias, and the two need different repairs - a better definition for the first, a better design for the second. This is also why ring analyses can disagree without any of them being broken, since each ring answers its own question. When we can&#8217;t defend a single mapping, the practical move is to report estimates under several and to say which institutional knowledge would narrow the set, e.g. commuting-flow data that shows how far workers travel. Search terms: &#8220;exposure mapping misspecification&#8221;, &#8220;ring analysis spillovers&#8221;, &#8220;interference causal inference&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>Our friends at the Psychology department met this problem first. <a href="https://doi.org/10.1037/h0040957">Cronbach and Meehl (1955)</a> had to validate tests for things like anxiety and intelligence, where no direct benchmark exists, so they proposed checking the measure against the relationships theory predicts (if the concept should rise with X and fall with Y, the measure should too). <a href="https://doi.org/10.1037/h0046016">Campbell and Fiske (1959)</a> turned this into a table. The idea is, we measure several concepts with several methods, then ask for two things: different methods measuring the same concept should agree (convergent validity), and the same method measuring different concepts should not agree <em>too much</em> (discriminant validity). The second check catches a failure we in econ rarely test for. Two supposedly different concepts can move together just because the same method produced both, e.g. everything self-reported moving together because the same person answered everything on the same afternoon. Search terms: &#8220;construct validity&#8221;, &#8220;multitrait-multimethod matrix&#8221;, &#8220;convergent and discriminant validity&#8221;, &#8220;common method bias&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p><a href="https://doi.org/10.1017/S0003055401003100">Adcock and Collier (2001)</a> split operationalisation into steps: background concept, systematised concept, indicators, scores. The split is what tells us where the trouble is. Sometimes the indicator is fine and the concept was never pinned down, and no new variable repairs an undefined concept. <a href="https://doi.org/10.1037/0033-2909.110.2.305">Bollen and Lennox (1991)</a> add a distinction that changes what validation means. Reflective indicators are effects of the concept (test answers reflect ability), while formative indicators build it (income, education and occupation together form socioeconomic status). The agreement logic that works for reflective measures is the wrong test for formative ones, whose components can disagree because they measure different parts on purpose. Search terms: &#8220;systematized concept Adcock Collier&#8221;, &#8220;reflective and formative indicators&#8221;, &#8220;measurement model&#8221;.</p></div></div>]]></content:encoded></item><item><title><![CDATA[A Field Guide to Missing Data (Part 1)]]></title><description><![CDATA[What economists do when the data they need aren&#8217;t there.]]></description><link>https://www.diddigest.xyz/p/a-field-guide-to-missing-data-part</link><guid isPermaLink="false">https://www.diddigest.xyz/p/a-field-guide-to-missing-data-part</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 17 Aug 2026 12:48:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q9Hn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q9Hn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png" width="1122" height="1402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1402,&quot;width&quot;:1122,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2657526,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/211036449?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 424w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 848w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 1272w, https://substackcdn.com/image/fetch/$s_!Q9Hn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff21e8df0-b94a-40f5-82d1-c9c4155765a3_1122x1402.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>(This is part 1 of 2 because it was getting way too long, even by my standards. I will format both articles in LaTeX and add the link to the .pdf for download in part 2).</em></p><p>Empirical work often starts with an object we can&#8217;t observe. Sometimes that object is a variable or an observation. Sometimes it is a counterfactual, a comparable definition, a price, a target population, an assignment mechanism, a joint distribution, or a reproducible computational record. I use <em>missing data</em> in this broad sense<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. There&#8217;s a reason I think we should stop and think a bit deeper about these &#8220;absences&#8221;: they require different responses. An economist might report a set instead of a point, stress an estimate under a departure from an assumption, reconstruct an unavailable object, combine information across sources, generate a reference distribution, change the target question, regularise a sparse model, or improve the reporting and audit trail. These are different inferential activities, and because they are different, in nature and in substance, they don&#8217;t become interchangeable when a results table places them under one &#8220;robustness&#8221; heading.</p><p>So to get more out of the imperfect situation we find ourselves in, the first question we should ask is: what exactly is unavailable? This requires a lot of training, actually, and the next ones aren&#8217;t easier either. After we&#8217;ve thought about what is unavailable, one way &#8220;out&#8221; is to ask: what side information is available? Which assumption links it to the missing object? What can the proposed exercise establish under that assumption? What remains unknown afterwards? And has the target estimand (the quantity we&#8217;re trying to estimate) changed along the way?</p><p>With all of this in mind, I tried my best to come up with a field guide organised around 22 recurring problems. Several entries below contain exercises that do different jobs, and I keep them separate whenever the choice between them changes what the analysis can claim: a bound, a reconstruction and a stress test answer to different things, even when they address the same problem.</p><p>Here&#8217;s a list of the entries we will be covering today:</p><ul><li><p>P01. A relevant confounder is unobserved</p></li><li><p>P02. The variable you want is represented by a proxy</p></li><li><p>P03. Some ordinary observations or items are missing</p></li><li><p>P04. Outcomes disappear through attrition or selection</p></li><li><p>P05. A continuous variable is measured with error</p></li><li><p>P06. Either treatment or outcome status is misclassified</p></li><li><p>P07. A price or quality adjustment is unavailable</p></li><li><p>P08. The counterfactual is unavailable</p></li><li><p>P09. Inputs or computational objects are inaccessible</p></li><li><p>P10. Only aggregates are observed</p></li></ul><p>Let&#8217;s start:</p><div><hr></div><h4>P01. A relevant confounder is unobserved</h4><p>We shall make this easier to understand with a familiar setup: we want to know whether a job-training programme raises earnings, and people chose whether to enrol. One worry for us would be &#8220;motivation&#8221;: one would assume more motivated people are more likely to sign up and more likely to earn more anyway, and no .csv has a column called &#8220;motivation&#8221;. That&#8217;s the missing object here: the part of who-gets-treated that&#8217;s related to the outcome but that our controls don&#8217;t capture. Confounding has other responses, and a convincing instrument, a discontinuity or randomised variation can change the design itself (<a href="https://doi.org/10.1257/jep.24.2.3">Angrist and Pischke 2010</a>), but this entry is about what we can learn when the original observational comparison is kept in place.</p><p>Since we can&#8217;t observe motivation directly (and don&#8217;t have a &#8220;suitable&#8221; proxy), the best we can do is ask the &#8220;what if&#8221; question: how strong would the link between motivation, enrolment and earnings have to be before our estimate stopped meaning what we say it means? That&#8217;s what sensitivity calculations do. Robustness values and coefficient-stability exercises (<a href="https://doi.org/10.1111/rssb.12348">Cinelli and Hazlett 2020</a>; <a href="https://doi.org/10.1080/07350015.2016.1227711">Oster 2019</a>) put numbers on that question. Oster&#8217;s version asks it in a more specific way: suppose the unobservables (e.g. the &#8220;motivation&#8221;) are related to enrolment in the same way as the observables we do observe in the data, like education and work history. If controlling for the things we can see barely moves the estimate while those controls explain a good share of the outcome, it would take an implausibly strong hidden factor to overturn it; if the estimate jumps around as we add controls, a hidden factor could plausibly finish the job. Both halves of that condition are needed since stable coefficients alone can be misleading: controls that carry little information also leave the estimate unmoved, without telling us anything about the confounder (<a href="https://doi.org/10.1080/07350015.2018.1462710">Pei, Pischke and Schwandt 2019</a>). To run this we have to declare two inputs: &#948;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, which says how strong the hidden selection is allowed to be relative to the observed selection, and R-max<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, which caps how much of the earnings variation a complete model could ever explain. The calculation won&#8217;t choose those values for us, and what it returns is a threshold that stresses the estimate, not an estimate of the missing confounding itself.</p><p>Rosenbaum sensitivity analysis asks a similar question from the assignment side, and it&#8217;s an easy one to over-read. Take two people who look identical on every variable we observe. Its &#915; parameter<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> says how much more likely one of them could be to enrol because of the hidden factor: &#915; of 2 means one could face twice the odds of treatment. The procedure then reports whether our conclusion would survive that much hidden imbalance. So what it bounds is the inference under that specified hidden-bias model rather than a causal parameter in general (<a href="https://doi.org/10.1257/000282803321946921">Imbens 2003</a>).</p><p>A negative control borrows a habit from epidemiology: find an outcome the programme couldn&#8217;t possibly affect, and check whether we &#8220;find&#8221; an effect there anyway. Earnings in the years before the programme existed are a natural candidate. If our design shows the training &#8220;raising&#8221; earnings that were paid out before anyone was trained, we haven&#8217;t discovered time travel; we&#8217;ve diagnosed a bias pattern, most likely the motivated people pulling ahead in any period (<a href="https://doi.org/10.1097/EDE.0b013e3181d61eeb">Lipsitch, Tchetgen Tchetgen and Cohen 2010</a>). The catch is that the diagnosis is only as good as our claim that the control outcome really is untouchable by the treatment. And going further, actually using negative controls to identify the causal effect, is a different exercise with its own requirements: proxy conditions, conditional independence, and rank or completeness conditions (<a href="https://doi.org/10.1093/biomet/asy038">Miao, Geng and Tchetgen Tchetgen 2018</a>).</p><p>This family also has two more members worth talking about. When economic knowledge supports a sign or relative-magnitude restriction on the omitted bias, the output can be an identified set rather than a threshold (<a href="https://doi.org/10.3982/QE689">Chalak 2019</a>; <a href="https://doi.org/10.1515/jem-2013-0013">Krauth 2016</a>). And for ratio measures, epidemiology has what they call the E-value: it asks how strongly an unmeasured confounder would have to be associated with both treatment and outcome to explain away the observed association, which is one more threshold, and not evidence that no such confounder exists (<a href="https://doi.org/10.7326/M16-2607">VanderWeele and Ding 2017</a>).</p><p>Whichever of these we run, none of them observes motivation for us, and none of them proves exchangeability, meaning that enrolled and non-enrolled people are comparable once we condition on the controls. They tell us how fragile our conclusion is to the confounder we&#8217;re worried about, which is valuable, but different from removing the worry.</p><div><hr></div><h4>P02. The variable you want is represented by a proxy</h4><p>This time the missing object is the variable itself. Staying with our training study: suppose earnings are self-reported, and what we actually want is what employers paid. People round, forget, and sometimes flatter themselves, so the column called &#8220;earnings&#8221; in our dataset is really a proxy for earnings. The same situation shows up everywhere once we look for it: test scores standing in for learning, self-reported health standing in for health, patents standing in for innovation, etc.</p><p>The standard way out is a validation sample: a dataset where we observe both the proxy and a better benchmark for the same people - like a subset of our survey linked to tax records. A replicate measure or a reliability study can recover parts of the noise, while a linked benchmark can also reveal systematic bias, so these tools give us different kinds of side information rather than substitutes for one another<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. With one of these in hand, we can estimate how the proxy relates to the target, which is our model of the error process<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>, and then use that estimated relationship to correct the analysis in the main sample.</p><p>This move rests on an assumption we should say out loud: it needs both the measurement process and the relationship we estimated to transport from the validation setting to ours, meaning they hold in our sample too (<a href="https://doi.org/10.1016/S1573-4412(01)05012-7">Bound, Brown and Mathiowetz 2001</a>). If people misreport differently in the validation study than in our survey (because the years, the interview and/or the people themselves differ), then the correction we imported was calibrated for a different problem, and the apparent fix has just moved the problem from one dataset to the other. This is not ideal. One practical detail follows: for a regression correction, the validation records need to hold the proxy, the benchmark, the outcome and the relevant covariates together, and if the benchmark itself is imperfect, its error has to be modelled too, with the uncertainty from the calibration step carried into the final interval (<a href="https://doi.org/10.1016/S1573-4412(01)05012-7">Bound, Brown and Mathiowetz 2001</a>).</p><p>We should also be clear about what the correction returns when it does work: a regression parameter under the assumed error model. We can recover, let&#8217;s say, the average effect of training on true earnings, and still have no idea what any particular person truly earned. And none of this tells us whether the proxy measures the concept we care about in the first place: a validation sample can tell us how noisy our earnings measure is, but it can&#8217;t tell us whether earnings were the right thing to measure.</p><div><hr></div><h4>P03. Some ordinary observations or items are missing</h4><p>Back to our training survey: some people answered everything except the earnings question, and a few interviews never happened at all. The missing object here is ordinary, a value that should have been sitting in our analysis sample, and the response to it is where most of us first met the phrase &#8220;missing data&#8221;.</p><p>There are three standard routes here: imputation, weighting and observed-data likelihood, and the first two work in opposite directions. Multiple imputation fills each &#8220;empty space&#8221;, several times, with plausible values drawn from a model of the data<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>. Inverse-probability weighting doesn&#8217;t fill anything: it makes the people we did observe count for more so each respondent also stands in for the similar people who didn&#8217;t respond. The third route, observed-data maximum likelihood, neither fills nor reweights but writes down a likelihood for the data we saw and integrates over the values we didn&#8217;t, and it inherits the same ignorability and modelling assumptions as the other two (<a href="https://doi.org/10.1002/9781119482260">Little and Rubin 2019</a>). (<em>The simplest response of all - dropping the incomplete cases - is transparent but it pays in precision and can easily change the target to the kind of person who has complete records.</em>) All three routes can support valid inference, but only under an assumption called missing at random, or MAR<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>, and the fine print on MAR is where the argument lives: it&#8217;s always relative to the information we observe. Given what we see about a person, their age, education and work history, the chance that their earnings answer is missing can&#8217;t depend on the earnings themselves. Whether that holds can&#8217;t be checked from the observed data alone (<a href="https://doi.org/10.1093/biomet/63.3.581">Rubin 1976</a>; <a href="https://doi.org/10.1080/01621459.1994.10476818">Robins, Rotnitzky and Zhao 1994</a>); adding more variables to the model can make it more plausible, but never verified.</p><p>There&#8217;s also a &#8220;belt-and-braces&#8221; version, augmented inverse-probability weighting, which combines a model of the outcome with the model of who responds and stays consistent when either of the two is correctly specified (<a href="https://doi.org/10.1111/j.1541-0420.2005.00377.x">Bang and Robins 2005</a>). This property, double robustness<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>, is a genuine insurance against picking the wrong model, but the insurance operates inside the maintained MAR assumption: it gives us two routes through model specification, and it doesn&#8217;t repair data that are missing not at random (when the chance of a space depends on the value hiding in it), nor a positivity problem (when some kinds of people never respond at all, so nobody is available to stand in for them).</p><p>Imputation carries one more requirement that sounds bureaucratic and turns out to be important: the imputation model has to be compatible with the analysis we are planning to run<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>. If our analysis studies an interaction (e.g., training effects that differ by gender, and the imputation model doesn&#8217;t include it), we&#8217;ve quietly built the absence of that interaction into the completed data before the analysis even starts. With chained equations - which is the popular variable-by-variable default - keeping the two compatible may need a substantive-model-compatible imputation scheme rather than the out-of-the-box settings (<a href="https://doi.org/10.1177/0962280214521348">Bartlett et al. 2015</a>).</p><p>Diagnostics have limits here too. We can spot completed values that look implausible or weights that &#8220;explode&#8221;, which diagnoses instability or thin support, though it doesn&#8217;t tell us which model or assumption caused it, and no diagnostic computed on the observed data verifies the missingness assumption itself. The assumption stays an argument, which is why the choice of variables in the imputation or response model is a substantive decision and belongs in the paper&#8217;s text rather than in a replication file.</p><div><hr></div><h4>P04. Outcomes disappear through attrition or selection</h4><p>Our training study has one more trick to play on us. At the follow-up interview, some people have left the study, and wages exist only for people who are working. This looks like P03, but it&#8217;s a different thing because now the treatment itself is &#8220;editing&#8221; the sample: training changes who finds a job, and having a job determines whose wage we get to see. If we naively compare the wages of trained workers with the wages of untrained workers, we&#8217;re mixing two things: the effect of training on wages AND the effect of training on who shows up in the wage column at all.</p><p>There are four responses here, and they answer different questions. The first borrows from P03: if dropping out is ignorable given everything we observed along the way, inverse-probability-of-censoring weights let the people who stayed stand in for the people who left, with the same MAR and positivity requirements as before, and no help against dropout that depends on the missing wages themselves (<a href="https://doi.org/10.1080/01621459.1994.10476818">Robins, Rotnitzky and Zhao 1994</a>). The second is to give up the point estimate and report a range. Worst-case bounds fill the spaces with the most extreme plausible values in both directions (<a href="https://doi.org/10.1016/S0304-4076(97)00077-8">Horowitz and Manski 1998</a>), which I think is honest and often uncomfortably wide. Lee trimming narrows the range under one extra assumption (monotone selection), meaning training can move people into employment but never out of it. It also needs the treatment itself to be as good as random, or exogenous given the covariates we use, so it doesn&#8217;t fix the self-selection into training we worried about in P01. The narrower bound comes at a price in interpretation: it describes the effect for the always-observed group, the people who would be working with or without training<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>, and not for everyone originally assigned (<a href="https://doi.org/10.1111/j.1467-937X.2009.00536.x">Lee 2009</a>). Fittingly for our example, Lee&#8217;s original application was a job-training programme, the US Job Corps.</p><p>The third response is to reconstruct the missing wages with a selection model. A Heckman-style model writes down who ends up employed and what they earn as one joint system with related errors, and then uses the system to correct the wage equation (<a href="https://doi.org/10.2307/1912352">Heckman 1979</a>). For this to rest on more than an assumption about the shape of the error distribution, we need an exclusion restriction<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>: some variable that shifts whether a person works but has no direct business in their wage. Without a convincing one, the identification leans heavily on distributional form, which is a polite way of saying the estimate can move a lot when we change assumptions we can&#8217;t test.</p><p>The fourth response is to stress the missingness assumption directly. Pattern-mixture and tipping-point exercises<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> say: suppose the people we lost differ from the people we kept by some amount, and let&#8217;s see how large that amount has to be before our conclusion changes (<a href="https://doi.org/10.1080/01621459.1999.10473862">Scharfstein, Rotnitzky and Robins 1999</a>). If it takes a difference that is implausibly large on a scale we can defend, the result is less fragile under that sensitivity model; if a small plausible difference is enough, we&#8217;ve learned our conclusion leans heavily on the missing-outcome assumption (aka, it&#8217;s held up by hope only).</p><p>So: a reweighting, a set, a reconstruction and a stress test, four different answers to &#8220;what happened to the people we can&#8217;t see?&#8221;. And we shouldn&#8217;t expect the data to referee between them because model fit can&#8217;t do it: every missing-not-at-random model has a missing-at-random counterpart that fits the observed data exactly as well (<a href="https://doi.org/10.1111/j.1467-9868.2007.00640.x">Molenberghs et al. 2008</a>). The choice between stories is an argument about behaviour, and not a statistic.</p><div><hr></div><h4>P05. A continuous variable is measured with error</h4><p>Let&#8217;s change what we&#8217;re asking of the training study. Suppose we now want to know how earnings gains relate to the hours of training a person actually attended, and hours are self-reported. You see where I&#8217;m going with this, right? Nobody remembers precisely; people round to whole weeks and guess. The missing object is the error-free value - or really - the regression relationship we would run if we had it.</p><p>Most of us were taught one fact about this situation: noise in a regressor pulls the estimated coefficient towards zero, so our estimate is <em>at least </em>conservative. That fact is real but much narrower than its (big) reputation<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>. It holds for classical error, meaning noise unrelated to the true value and to everything else, in a simple linear regression with one mismeasured variable. Add more regressors, let the errors correlate with the truth (remember from P02 that real survey errors do exactly that), make the model nonlinear, or put the error in the outcome instead, and even the direction of the bias can change (<a href="https://doi.org/10.1257/jep.15.4.57">Hausman 2001</a>). &#8220;Our estimate is biased towards zero, so the true effect is bigger&#8221; is a sentence that needs its assumptions checked before it leaves the seminar room.</p><p>The tools for repairing this all work by learning something about the error process, and each one assumes something different about it. A validation sample, as in P02, estimates the error process directly. Repeated measures use a second noisy report of the same variable, and one clean use is as an instrument: the second measure is correlated with the true hours but, if its error is independent of the first measure&#8217;s error and of the disturbance in our earnings equation, not with the parts that cause the trouble<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a><sup> </sup>(<a href="https://doi.org/10.1016/0304-4076(86)90058-8">Griliches and Hausman 1986</a>). Regression calibration replaces the noisy value with our best estimate of the true one given everything observed, and SIMEX runs the regression several times with extra noise deliberately added, watches how the coefficient degrades, and projects that path backwards to the no-noise case<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a><sup> </sup>(<a href="https://doi.org/10.1080/01621459.1994.10476871">Cook and Stefanski 1994</a>). Likelihood methods write the error process into the model and estimate everything jointly<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>. </p><p>What all of these can deliver (when their assumptions hold!) is a corrected population parameter: the average relationship between true hours and earnings. What they don&#8217;t deliver is each person&#8217;s true hours. The correction works at the level of the regression, and that&#8217;s usually what we need, but it&#8217;s worth saying out loud, because a corrected coefficient sitting in a clean table makes it easy to believe the underlying variable was fixed too, and it wasn&#8217;t.</p><div><hr></div><h4>P06. Either treatment or outcome status is misclassified</h4><p>Sometimes the mismeasured variable isn&#8217;t a number but a yes-or-no, and that changes the problem enough to deserve its own entry. In our training study: some people say they attended when they didn&#8217;t, some attended and don&#8217;t say so, and on the outcome side &#8220;employed&#8221; versus &#8220;not employed&#8221; has its own wrong labels. False positives and false negatives behave differently from the continuous noise of P05 because a binary variable can only be wrong in two directions, and the mix of the two directions is what drives the damage.</p><p>The first question is what we&#8217;re actually trying to recover, because there are three different targets hiding here: the true prevalence (what share of people really attended training), a regression coefficient (how attendance relates to earnings), or a causal effect. Each needs its own correction, and a fix for one is not automatically a fix for the next. A corrected prevalence doesn&#8217;t by itself correct a regression coefficient, and a corrected coefficient isn&#8217;t automatically a causal effect (<a href="https://doi.org/10.1016/S0304-4076(98)00015-3">Hausman, Abrevaya and Scott-Morton 1998</a>). This chain is easy to skip in practice: a paper corrects the share and quietly treats the regression as corrected too.</p><p>What we can do depends on what we know about the error rates, which readers from epidemiology will recognise as sensitivity and specificity<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a>. If we know them, or can estimate them from a validation subsample, a correction can be built for our specific target and error model, and <a href="https://doi.org/10.1016/0304-4076(73)90005-5">Aigner (1973)</a> shows what that full knowledge buys us in his one-sided model for a binary regressor. If we only have partial information, like a plausible range for the error rates rather than their values, the defensible output is a bound (<a href="https://doi.org/10.1016/S0304-4076(95)01730-5">Bollinger 1996</a>): a range of coefficients consistent with what we know instead of a corrected point. Between the two sits a third option: when the error rates are uncertain rather than known, probabilistic sensitivity analysis feeds defensible distributions for sensitivity and specificity through the correction and reports the resulting distribution of corrected results, though it still doesn&#8217;t learn those distributions from our own data (<a href="https://doi.org/10.1093/ije/dyi184">Fox, Lash and Greenland 2005</a>).</p><p>The assumptions doing the work are about the error rates themselves: which ones we know, whether the validation setting is comparable to ours (the transport question from P02 again), and above all whether the rates vary with treatment, outcome or covariates. When they do vary, the misclassification is called differential<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a>, and it needs its own model, because the convenient folk theorem that misclassification biases us towards zero belongs to the non-differential case, and not always even there.</p><p>And when there&#8217;s no gold standard at all (no validation sample, no error rates from anywhere), latent-class methods<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> can seem to rescue us: give the model several imperfect measures of the same status and let it estimate the error rates and the truth together. The rescue is real but runs on assumptions we should lay out, usually that the measures err independently given the true status and that the classes mean the same thing across populations. The output can look impressively precise, and that precision can be fragile to exactly those assumptions.</p><div><hr></div><h4>P07. A price or quality adjustment is unavailable</h4><p>Let&#8217;s forget our training study example for a second because this problem sits in a different corner of economics. Suppose we want to compare the cost of living across regions (spatial variation), or track inflation over time (time variation), and the goods people buy differ in quality (they&#8217;re heterogeneous): rice in one market is not the rice in another, this year&#8217;s laptop is not last year&#8217;s laptop. The missing object may be a comparable price, a quality adjustment, or the weights an index needs. What we observe instead are transactions for goods that are never quite the same twice.</p><p>Household surveys tempt us with a shortcut called the unit value: divide what the household spent on rice by the kilos it bought, and call that the price of rice. The trouble is that a unit value mixes three things: the price the household faced, the quality it chose, and the composition of its purchases. One could say that richer households buy better rice, so their unit values are higher even at identical prices, and if we treat unit values as prices we build the household&#8217;s own choices into the &#8220;price&#8221; we then use to explain those choices (<a href="https://www.jstor.org/stable/1809142">Deaton 1988</a>). Unit values can be validated and adjusted into something useful, but they don&#8217;t start out as prices. Deaton&#8217;s own method is a reconstruction in this article&#8217;s sense: it needs geographically clustered households, a model of quality choice and a measurement-error correction, so dividing expenditure by quantity is where the exercise starts rather than what it is.</p><p>For goods that change quality over time, hedonic methods<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a> price the characteristics instead of the good: regress observed prices on attributes (for a laptop, memory, speed, screen), and use the fitted relationship to ask what the new model would have cost with the old characteristics (<a href="https://doi.org/10.1086/260169">Rosen 1974</a>). The comparison this gives us depends on the characteristics we observed and the functional form we chose, which is why official price indices don&#8217;t stop at a regression: the statistical manuals spell out explicit rules for matching items over time, weighting them, handling a product&#8217;s replacement and linking new items in (<a href="https://doi.org/10.5089/9781484354841.069">Intersecretariat Working Group on Price Statistics 2020</a>). The rules don&#8217;t remove the judgement, but they put it where we can inspect it. Two named index families do much of this work: matched-model indexes compare identical items over time, and multilateral methods pool comparisons across several periods, which also reduces the chain drift we will meet below, though both still depend on the matching rule, the product universe and how entry and exit are handled (<a href="https://doi.org/10.1016/j.jeconom.2010.09.004">De Haan and van der Grient 2011</a>).</p><p>There&#8217;s also a difference between calculating a missing price and inferring one, and it decides how much the number can be trusted. Where a fully specified accounting identity pins the number down, total expenditure, quantities and the other components all observed and compatible, we can calculate the residual and carry the components&#8217; uncertainty through to it. A price backed out of a demand curve, a production function or any other behavioural relationship is a different object: a model-based scenario whose value changes with the model. There is no general &#8220;residual price&#8221; correction. And the diagnostics here have the usual limit. Chain-drift checks<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a> can tell us a chained index has &#8220;wandered&#8221; in a way that pure price change can&#8217;t explain, but they can&#8217;t tell us whether the items being linked were economically comparable in the first place.</p><div><hr></div><h4>P08. The counterfactual is unavailable</h4><p>Readers of this newsletter are more than familiar with this problem :) This is THE central absence in causal inference, the one every other previous entry has been circling around. Let&#8217;s take our example a level up: a state introduces a training subsidy and we want to know what happened to earnings <em>because</em> of it. The missing object is what earnings in that state would have been without the subsidy. No dataset ever contains that column so everything in this entry is about constructing a stand-in for it and then interrogating (a lot) the construction.</p><p>Design, diagnostics, falsification and sensitivity analysis constantly get blurred under the word &#8220;robustness&#8221;, and it helps to keep the four apart. Our chosen design constructs the comparison: these units, these periods, this assumption about why the comparison is fair. The diagnostics check parts of that design, while falsification tests look for effects in places where the design says there should be none, and sensitivity analysis varies an identifying restriction to report which conclusions remain. The four are related, and they are not the same task, so a paper that runs one of them has not automatically done the others. For our subsidy question, the two designs we would most likely use are synthetic control and difference-in-differences<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a>. </p><p>With synthetic control, we build the counterfactual ourselves: the untreated states get combined into a weighted mix, and the weights are chosen so that the mix tracks the treated state&#8217;s earnings path before the subsidy (<a href="https://doi.org/10.1198/jasa.2009.ap08746">Abadie, Diamond and Hainmueller 2010</a>). The comparison is only as good as its conditions, good pre-treatment fit and a donor pool of comparable, untreated states. The usual check is a placebo run<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a>: pretend each donor state passed the policy and see whether our real state&#8217;s post-policy gap stands out among the pretend ones. That comparison produces a rank, and a rank is not a p-value by itself (<a href="https://doi.org/10.1257/jel.20191450">Abadie 2021</a>); it earns a probability interpretation only with an argument about how treatment was assigned. There&#8217;s also an in-time placebo: move the intervention to a date before the real one and check whether the same procedure finds an effect where there shouldn&#8217;t be one, and here a check that finds one counts against the design, while a clean one stays just a diagnostic (<a href="https://doi.org/10.1257/jel.20191450">Abadie 2021</a>).</p><p>With difference-in-differences, the counterfactual comes from assumptions rather than weights, and two of them work together. They are a) parallel trends, which in our example would be &#8220;without the subsidy, the treated state&#8217;s earnings would have moved like the comparison states&#8217; did&#8221;, and b) no anticipation, in our example &#8220;earnings didn&#8217;t start responding before the subsidy took effect because people saw it coming&#8221;. <em>(Anticipation and treatment timing get their own entry in upcoming P19; here we stay with parallel trends, which is where - one can say - the diagnostic action is - wink).</em> The standard diagnostic is a pre-trend check, and it&#8217;s a bit weaker than its popularity suggests<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a>: the test can have little power against the violations that would hurt us most, so passing it doesn&#8217;t validate the assumption (<a href="https://doi.org/10.1257/aeri.20210236">Roth 2022</a>). The newer response is to stop testing and start stressing: the Honest DiD approach<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a> asks what we can conclude if post-treatment violations of parallel trends are allowed to be, say, no larger than the worst pre-treatment one, and reports an interval or a set for the effect under that restriction (<a href="https://doi.org/10.1093/restud/rdad018">Rambachan and Roth 2023</a>). The restriction, and not the method&#8217;s reassuring name, is what does the identifying work.</p><div><hr></div><h4>P09. Inputs or computational objects are inaccessible</h4><p>This entry is about a different kind of disappearance, and the example works best told forward. Ten years from now, someone questions our training-subsidy paper. The tax microdata we used were confidential and our access agreement has expired, the person who wrote the estimation code has left, and what survives is the published paper with its tables (this hits way too close to home). The missing objects are the inputs of the analysis itself: microdata, code, the fitted model, the intermediate files that turned one into the other. Every other entry in this guide worried about data on the world; this one is about the data trail of the research.</p><p>The first thing to keep in mind is what question each exercise answers here, because the vocabulary gets used loosely<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-27" href="#footnote-27" target="_self">27</a>. Reproduction asks whether the same data and the same instructions regenerate the published number (<a href="https://www.jstor.org/stable/1806061">Dewald, Thursby and Anderson 1986</a>; <a href="https://doi.org/10.1561/104.00000053">Chang and Li 2022</a>). That is a question about the record, and it&#8217;s separate from whether the design identifies anything, from whether the finding is true about the world, and from whether a new study with new data would find it again. A number can reproduce and still be wrong, and a paper whose files were lost can still have been right. We could never know.</p><p>Some questions need more than the published record to answer. If we want to know whether one state, or a handful of unusual people, drove our earnings estimate, we need influence or leave-one-out calculations<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-28" href="#footnote-28" target="_self">28</a>, and those run on the objects themselves: unit-level contributions, the fitted model, score contributions, or something equivalent that survived. When those objects are gone, they are gone. The defensible report in that situation is a documented limit on reproducibility, saying what was checked, what could not be, and why; an influence analysis can&#8217;t be recreated from objects that no longer exist, and writing one up as if it could is a reconstruction wearing a &#8220;diagnostic&#8217;s clothes&#8221;.</p><p>Before settling for that report, there are moves we haven&#8217;t used yet, and they&#8217;re the same verbs from the beginning of this article. The first one is to reconstruct the inputs from their sources instead of from the paper: the tax microdata still exist at the agency even though our agreement expired, agreements can be renewed, and if we documented our extraction criteria, a new extract plus a reimplementation of the described pipeline can regenerate the analysis. The result needs plain labelling, because it&#8217;s a new extract of possibly revised data (this is P22&#8217;s problem, arriving a bit early), so a difference from the published number can come from the data as well as from the code.</p><p>The published record itself can also be checked without touching any lost object, because a table has arithmetic obligations: the coefficients have to cohere with their standard errors and t-statistics, the sample sizes have to agree across tables, and the subgroups have to add up to their totals. This is diagnosis rather than regeneration: it can reveal contradictions in the surviving report, and it doesn&#8217;t regenerate the missing analysis (psychology built automated versions of these checks; search &#8220;statcheck&#8221; and &#8220;GRIM test&#8221;).</p><p>There is a narrow path between giving up (and crying on the floor) and pretending. On that path, a scenario range or a set of bounds can stand in for the lost analysis, but only after we define which missing inputs we consider compatible with the published record and justify that set: &#8220;the result holds under any reasonable version of the lost code&#8221; is an empty sentence until &#8220;reasonable&#8221; has content. Even then, the published totals on their own rarely settle it because a table of coefficients doesn&#8217;t reveal how each unavailable component, a weight here, a sample restriction there, moved the original estimate. When none of this recovers the analysis, the remaining move is to change the question: a replication with new data stops asking whether the published number was computed correctly and asks whether the claim is true, which is usually the thing we cared about in the first place.</p><div><hr></div><h4>P10. Only aggregates are observed</h4><p>Suppose the individual data behind our training question were never available to us in the first place, and all we have is a table by state: the share of workers who went through training and the average earnings. The missing object is the joint distribution, which is the individual relationship living inside those margins. The temptation we face is to correlate the two columns and say that trained workers earn more, and this temptation has a name and a famous cautionary tale<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-29" href="#footnote-29" target="_self">29</a>: the correlation between two aggregates can differ in size - and even in sign - from the correlation between the same two things across people.</p><p><span>Once we stop trusting the two-column correlation, the correct move is to ask less of the aggregates we have, and ecological bounds are the exercise built for that: they describe the set of individual relationships consistent with the margins we observe, and the set is often wide, because the margins don&#8217;t determine what happens inside them (</span><a href="https://doi.org/10.1093/oxfordhb/9780199286546.003.0024"><span>Cho and Manski 2008</span></a><span>; open-access chapter </span><a href="https://tam.cas.vanderbilt.edu/papers/chomanski.pdf"><span>here</span></a><span>). </span>Sadly, the bounds don&#8217;t recover the micro relationship, and more aggregate margins alone won&#8217;t either; a narrower set or a point estimate comes only from additional restrictions, and then the restrictions carry the conclusion. King&#8217;s ecological-inference model is the best-known point-estimation route, and it works exactly that way: it selects an answer from what the margins allow by adding distributional and homogeneity assumptions that the margins themselves can&#8217;t test (<a href="https://doi.org/10.1515/9781400849208">King 1997</a>). <span>The available diagnostics tell us more about the aggregation than about individuals: re-estimating across plausible geographies (like counties instead of states, if possible) shows how sensitive the aggregate association is to where the (sometimes literal) lines were drawn, a dependence with its own name (the modifiable areal unit problem, or MAUP), and comparing the aggregate analysis with an individual-level one, wherever even a small individual sample exists, can expose the contextual and ecological biases in the aggregate estimate</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-30" href="#footnote-30" target="_self">30</a><span> (</span><a href="https://doi.org/10.1093/ije/30.6.1343"><span>Greenland 2001</span></a><span>).</span></p><p>Multilevel and within-between models sound like the modern solution, and they are, but only for the data they actually need: relevant lower-level observations, or repeated hierarchical information such as many periods per state. Given one aggregate cross-section, a multilevel model has nothing to work with underneath the aggregates; it can&#8217;t manufacture individual data from a table of forty-something rows.</p><p>Shift-share and component decompositions<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-31" href="#footnote-31" target="_self">31</a> can also mislead here, because they look like they answer the individual question and they don&#8217;t. A decomposition tells us how the aggregate change adds up, how much came from composition shifting across groups and how much from changes within groups, and that&#8217;s a descriptive answer to a descriptive question. It isn&#8217;t a solution to ecological identification, because both pieces of the sum are themselves aggregates.</p><p>The useful discipline is to state which margins we observe, which associations are missing, and what extra structure would be needed to connect the two. Papers that do this read differently, because the reader can see where the aggregates stop and the assumptions start.</p><div><hr></div><p><em>That&#8217;s all for today. We will be back next week. Any comments/suggestions/corrections please send me an email or leave them below.</em></p><p><em>Fable-5 helped with the literature review and gpt-5.6-sol audited the text (twice).</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;m using missing data more broadly than usual, although the basic idea has precedents. <a href="https://doi.org/10.1007/s40471-015-0050-8">Howe, Cain and Hogan (2015)</a> treat confounding, selection and measurement error as missing-data problems; <a href="https://doi.org/10.1093/ije/dyu272">Edwards, Cole and Westreich (2015)</a> describe measurement error as missing information within the potential-outcomes framework; and <a href="https://doi.org/10.1214/18-STS645">Ding and Li (2018)</a> develop causal inference as a missing-potential-outcomes problem. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>&#948; is the coefficient of proportionality: a number that says how strong selection on unobservables is allowed to be relative to selection on observables. If &#948; = 1, the factors we can&#8217;t see (like motivation) are related to programme take-up just as strongly as the factors we included in the regression (like education and work history); &#948; = 2 would mean twice as strongly, and &#948; = 0 takes us back to assuming no hidden selection at all. The idea comes from </span><a href="https://doi.org/10.1086/426036">Altonji, Elder and Taber (2005)</a><span>, who were studying Catholic schools and argued that if the variables we observe are effectively a random subset of all the factors that drive selection, then how much the observables shift our estimate is informative about how much the unobservables could. </span><a href="https://doi.org/10.1080/07350015.2016.1227711">Oster (2019)</a><span> turned this into the calculation applied papers now report, where the usual output is the &#948; that would drive the estimate to zero: if it would take hidden selection stronger than everything we measured combined (&#948; &gt; 1), many authors read the result as safer, though that reading still depends on how good our observables were in the first place. A search tip: the useful terms are &#8220;proportional selection&#8221;, &#8220;coefficient stability&#8221;, and the command names, </span><code>psacalc</code><span> in Stata and </span><code>robomit</code><span> in R.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>R-max is the R-squared of a hypothetical regression we can never run: outcome on treatment plus every relevant control, the observed ones and the unobserved ones together. It answers the question &#8220;how much of this outcome could we ever explain if we had all the right variables?&#8221;, and the sensitivity calculation needs it because the room left between our actual R-squared and R-max is exactly the room where a hidden confounder could live. It can&#8217;t be lower than the R-squared we actually got, and for most outcomes it shouldn&#8217;t be 1 either, since part of the variation in something like earnings is noise, luck and measurement error that no variable will ever explain. </span><a href="https://doi.org/10.1080/07350015.2016.1227711">Oster (2019)</a><span> suggests a default of 1.3 times the observed R-squared, calibrated so that the method would validate the majority of a sample of results from randomised experiments, but that default is a starting point for an argument, and the right move is to justify a value for your own outcome (or show a range) rather than copy the 1.3. When searching, write it out as &#8220;Rmax&#8221; together with &#8220;Oster&#8221;, or look for &#8220;maximum R-squared&#8221; and the &#8220;1.3 rule&#8221; for the discussions around the default.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><span>&#915; is Rosenbaum&#8217;s sensitivity parameter, and the easiest way to read it is through a matched pair. Take two people who look identical on every covariate we observe. If assignment were as good as random given those covariates, both would have the same odds of ending up treated, and &#915; measures how far from that ideal we&#8217;re willing to imagine: &#915; = 1 means no hidden bias at all, while &#915; = 2 allows the hidden factor to double the odds of enrolment for one of them relative to the other (odds, and not probability, is the technically right word here). The procedure then asks how our p-value would look in the worst case allowed by each value of &#915;, raising &#915; until </span>the worst-case p-value crosses our significance level, and that value of &#915; is what gets reported:<span> a study that is &#8220;insensitive up to &#915; = 1.8&#8221; keeps its conclusion as long as no hidden factor comes close to doubling the odds of treatment. The method comes from </span><a href="https://doi.org/10.1093/biomet/74.1.13">Rosenbaum (1987)</a><span> and is developed in his book </span><em><a href="https://doi.org/10.1007/978-1-4757-3692-2">Observational Studies</a></em><span> (2002, second edition), whose &#8220;Sensitivity to Hidden Bias&#8221; chapter is the standard introduction. Two reading habits are worth keeping: the bound applies to the inference (the p-value or CI) under this specific model of hidden bias, and a large &#915; doesn&#8217;t certify the design, it just says the rejection survives that much imagined imbalance. When searching, the terms that work are &#8220;Rosenbaum bounds&#8221;, and the implementations are </span><code>rbounds</code><span> in R and Stata plus </span><code>sensitivitymv</code><span> and </span><code>sensitivitymw</code><span> in R for matched sets.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><span>A validation sample, a linked benchmark, a replicate measure and a reliability study sound interchangeable, but they observe different things, and it&#8217;s worth keeping them apart. A </span><em>validation sample</em><span> observes the proxy and a better benchmark for the same people, so it can estimate both the noise and any systematic gap between them. A </span><em>linked benchmark</em><span> is the administrative version of the same idea, like survey answers linked to tax or payroll records. A </span><em>replicate measure</em><span> is the same instrument applied twice, asking the same person their earnings in two interviews: comparing the answers tells us how noisy the instrument is, but if both answers share the same bias (everyone rounding their salary up), the replicate can&#8217;t see it, because the bias cancels in the comparison. A </span><em>reliability study </em>is what statistical agencies run when they re-interview a subsample and publish the resulting reliability ratio - which is the share of the measure&#8217;s variance that is signal rather than noise. The practical rule: replicates and reliability studies can tell us about noise, while only a benchmark can tell us about bias.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>T<span>he textbook error model is </span><em>classical measurement error</em><span>: the gap between the proxy and the truth is random noise, unrelated to the true value and to everything else in the analysis. It&#8217;s the assumption behind the familiar claim that mismeasurement pulls estimates toward zero (more on when that claim holds in P05). The reason to be suspicious of it is that when economists actually checked, real survey errors turned out not to be classical. </span><a href="https://doi.org/10.1086/298256">Bound and Krueger (1991)</a><span> compared what people told the Current Population Survey with their Social Security payroll records and found errors that were </span><em>mean-reverting</em><span>: high earners tend to underreport and low earners to overreport, so the error is correlated with the truth, which is exactly what the classical model rules out. </span><a href="https://ideas.repec.org/a/ucp/jlabec/v12y1994i3p345-68.html">Bound, Brown, Duncan and Rodgers (1994)</a><span> found the same pattern in the PSID Validation Study, a survey run on workers of a firm that shared its payroll records. This is why the transport assumption in the main text is doing real work: a correction built on the wrong error model can move an estimate in the wrong direction, not just by the wrong amount. When searching, the useful terms are &#8220;errors-in-variables&#8221;, &#8220;classical measurement error&#8221;, &#8220;mean-reverting measurement error&#8221;, &#8220;reliability ratio&#8221; and &#8220;PSID Validation Study&#8221;, and the survey covering all of it </span><a href="https://doi.org/10.1016/S1573-4412(01)05012-7">Bound, Brown and Mathiowetz (2001)</a><span>.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><span>Why fill each space several times instead of only once? Because filling once with our best guess pretends we knew the value all along, and the analysis then reports standard errors as if there had been no space, which overstates our certainty. Rubin&#8217;s solution was to draw several plausible values from a model that reflects how uncertain we are, produce several completed datasets, run the analysis on each, and combine the results with a set of formulas known as Rubin&#8217;s rules, so the disagreement across the completed datasets flows into the final standard errors. The foundations are in </span><a href="https://doi.org/10.1093/biomet/63.3.581">Rubin (1976)</a><span> and his 1987 book </span><em><a href="https://doi.org/10.1002/9780470316696">Multiple Imputation for Nonresponse in Surveys</a></em><span>. The implementations most people use are </span><code>mice</code><span> in R and </span><code>mi impute</code><span> in Stata, and those are also the best search terms together with &#8220;multiple imputation&#8221; and &#8220;Rubin&#8217;s rules&#8221;.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p><span>The taxonomy comes from </span><a href="https://doi.org/10.1093/biomet/63.3.581">Rubin (1976)</a><span>, and the names are famously confusing, so here they are in plain words. MCAR, or missing completely at random: the spaces are unrelated to anything, as when a page of questionnaires is lost in the office. MAR, or missing at random: given the variables we observe, the chance of a space doesn&#8217;t depend on the missing value itself, so younger respondents may skip the earnings question more often, but among people with the same observed characteristics, skipping is unrelated to earnings. MNAR, or missing not at random: the space depends on the value hiding in it even after we condition on everything observed, as when high earners &#8220;prefer not to say&#8221;. The names mislead because MAR doesn&#8217;t mean &#8220;missing randomly&#8221;; it means &#8220;random enough once we condition on what we see&#8221;. Search terms that work: &#8220;missing at random&#8221;, &#8220;ignorability&#8221;, and for the MNAR case &#8220;nonignorable nonresponse&#8221;.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>The definition is worth quoting, because the original authors say it more cleanly than most textbooks I&#8217;ve checked: &#8220;an estimator is doubly robust (DR) or doubly protected if it remains consistent when either a model for the missingness mechanism or a model for the distribution of the complete data is correctly specified. Because of the frequency and near inevitability of model misspecification, double robustness is a highly desirable property&#8221; (</span><a href="https://doi.org/10.1111/j.1541-0420.2005.00377.x">Bang and Robins 2005</a><span>). The augmented weighting estimator itself goes back to </span><a href="https://doi.org/10.1080/01621459.1994.10476818">Robins, Rotnitzky and Zhao (1994)</a><span>. Positivity, the other requirement in the main text, is the condition that every kind of person in the analysis has some chance of being observed; weighting can stretch respondents to cover similar non-respondents, but it can&#8217;t conjure information about groups with no respondents at all. Search terms: &#8220;AIPW&#8221;, &#8220;doubly robust estimation&#8221;, &#8220;positivity&#8221; or &#8220;overlap&#8221;.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p><span>The technical term for this compatibility is </span><em>congeniality</em><span>, coined by </span><a href="https://doi.org/10.1214/ss/1177010269">Meng (1994)</a><span>, and the concern is practical rather than pedantic. Imputation is often done by one person (or one agency) and analysis by another, and if the imputer&#8217;s model is poorer than the analyst&#8217;s (let&#8217;s say it leaves out interactions, nonlinearities or subgroup structure that the analysis cares about), the completed datasets systematically understate exactly the patterns being studied. The working rule that falls out: make the imputation model at least as rich as the analysis model, and when in doubt, richer. Search &#8220;congeniality multiple imputation&#8221; or &#8220;uncongenial imputation&#8221; for discussions.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p><span>The formal name for this group is a </span><em>principal stratum</em><span>, and the idea is worth thirty seconds because it explains who the estimate is really about. Sort people by how their employment would respond to the programme: some would work whether or not they were trained (the always-employed), some would work only if trained, some wouldn&#8217;t work either way. We can&#8217;t tell person by person who belongs where since we only ever see each person under one treatment, but the monotone-selection assumption rules out the awkward group that training pushes </span><em>out</em><span> of work, and that&#8217;s what lets the trimming produce a bound. The bound then belongs to the always-employed stratum: a policy claim like &#8220;training raised wages&#8221; silently becomes &#8220;training raised wages for the kind of person who&#8217;d be employed regardless&#8221;. The framework is from </span><a href="https://doi.org/10.1111/j.0006-341X.2002.00021.x">Frangakis and Rubin (2002)</a><span>. Search terms that work: &#8220;principal stratification&#8221;, &#8220;Lee bounds&#8221;, and the Stata command </span><code>leebounds</code><span>.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>An <em>exclusion restriction</em> is a variable that moves the selection (whether we observe the outcome) without moving the outcome itself. In the classic female labour-supply applications, the number of young children was argued to shift whether a woman worked but not her offered wage; whether the exclusion is believable is always the fight, and it should be. What happens without one is specific: the Heckman model is still estimable, because the assumed joint normality of the errors gives the correction term (the inverse Mills ratio) a particular nonlinear shape, and that shape alone identifies the model. That means the estimate is resting on a distributional assumption we can&#8217;t check, and modest changes to it can move the results a lot. The original is <a href="https://doi.org/10.2307/1912352">Heckman (1979)</a>, the Stata command is <code>heckman</code>, and the useful search terms are &#8220;Heckman selection model&#8221;, &#8220;exclusion restriction&#8221; and &#8220;inverse Mills ratio&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>A <em>pattern-mixture model</em> splits the sample by response pattern, the people we kept and the people we lost, models the ones we can see, and then states openly how the lost ones are assumed to differ, usually through an offset: &#8220;dropouts earn 10% less than observationally similar completers&#8221;. That factorisation by response pattern comes from <a href="https://doi.org/10.1080/01621459.1993.10594302">Little (1993)</a>. The <em>tipping-point</em> exercise repeats the analysis over a range of offsets and reports the value where the conclusion changes (here &#8220;changes&#8221; means whatever we declared as the criterion: usually that the effect is no longer statistically distinguishable from zero, sometimes that the estimate changes sign). Instead of defending one assumption, we show the whole range and discuss whether that value is realistic. <a href="https://doi.org/10.1080/01621459.1999.10473862">Scharfstein, Rotnitzky and Robins (1999)</a> is the standard reference for the related selection-model version, where a selection-bias parameter plays the role the offset plays here. One naming warning: this literature (especially in clinical trials where regulators routinely ask for these analyses) calls the offset a &#8220;delta adjustment&#8221;, and that delta has nothing to do with Oster's &#948; from P01, an unlucky collision of Greek letters. Maybe we should start using Chinese characters instead? :) Search terms: &#8220;pattern-mixture model&#8221;, &#8220;tipping point analysis&#8221;, &#8220;delta adjustment&#8221;, &#8220;nonignorable dropout&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>The mechanics are worth going over once. A regression coefficient is a covariance divided by a variance. Classical noise in the regressor leaves the covariance with the outcome alone but inflates the regressor&#8217;s variance, so the ratio shrinks: that&#8217;s attenuation, and the shrinkage factor is the reliability ratio from P02&#8217;s footnote, the share of the measure&#8217;s variance that is signal. The textbook result stops there, and the trouble starts where the textbook stops hehe. With several regressors, the bias from one mismeasured variable spreads to the coefficients of the correctly measured ones, in directions that depend on how everything correlates. If the error is mean-reverting rather than classical, the covariance is contaminated too, and the estimate can be biased away from zero. In nonlinear models there is no single shrinkage factor at all. And error in the <em>outcome</em> behaves differently again: if classical, it inflates standard errors without biasing the slope, but if it correlates with the truth (as in the earnings validation studies), it biases the slope as well. Search terms: &#8220;attenuation bias&#8221;, &#8220;errors-in-variables&#8221;, &#8220;reliability ratio&#8221;, and for the multi-regressor case is &#8220;measurement error contamination&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>We have two noisy reports of the same true variable, say self-reported hours and hours from the training provider&#8217;s records, which are both imperfect. Using one as an instrument for the other works because a valid instrument needs two things: it must be correlated with the variable we&#8217;re instrumenting (which may hold if the shared signal is strong enough to give a relevant first stage), and it must be unrelated to the errors that cause the bias (which holds only if the two errors are independent of each other and of the earnings equation&#8217;s disturbance). That second condition is the one to argue about: if both reports come from the same person&#8217;s memory, their errors probably share a component, and the instrument inherits the &#8220;disease&#8221; it was meant to cure. <a href="https://doi.org/10.1016/0304-4076(86)90058-8">Griliches and Hausman (1986)</a> developed this logic for panel data, where differencing can amplify measurement-error bias, and their paper is also the standard warning about that amplification. Search terms: &#8220;instrumental variables measurement error&#8221;, &#8220;repeated measures instrument&#8221;, &#8220;errors in variables panel data&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>I&#8217;d say these two have the most personality of the correction methods. <em>Regression calibration</em> asks: given the noisy report and everything else we observe, what&#8217;s our best estimate of the person&#8217;s true hours? It substitutes that estimate (a conditional expectation) into the regression and then adjusts the standard errors to account for the substitution, which makes it simple and popular in epidemiology. <em>SIMEX</em>, simulation-extrapolation, is the counterintuitive one: since we can&#8217;t remove noise, just <em>add</em> more of it, in several known amounts, and rerun the regression each time. The coefficient degrades along a path as noise grows, and fitting that path lets us project it back to the point of zero noise (<a href="https://doi.org/10.1080/01621459.1994.10476871">Cook and Stefanski 1994</a>). The projection step is the assumption: we never observe the zero-noise end, so the answer depends on the fitted shape of the path. Implementations: <code>simex</code> in R and Stata&#8217;s <code>simex</code> command; search &#8220;regression calibration&#8221; and &#8220;simulation extrapolation SIMEX&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>In the simple classical linear case, we also have two older tools: a known reliability ratio can correct the naive slope directly, and running the regression in both directions (earnings on hours, then hours on earnings, inverted) brackets a positive slope between the two estimates, though neither trick survives nonclassical error or richer models (<a href="https://doi.org/10.1257/jep.15.4.57">Hausman 2001</a>). One caution on regression calibration: it&#8217;s exact for a linear mean model under its assumptions and generally an approximation in nonlinear ones (<a href="https://doi.org/10.1201/9781420010138">Carroll et al. 2006</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p><em>Sensitivity</em> is the probability that a true yes is recorded as a yes (a real trainee reports attending); <em>specificity</em> is the probability that a true no is recorded as a no. Together they pin down the misclassification process for a binary variable, and they make the prevalence correction almost arithmetic: the observed rate of yeses mixes true yeses caught by the test with false positives, so observed rate = sensitivity x true rate + (1 &#8722; specificity) x (1 &#8722; true rate), and knowing the two error rates lets us solve for the true rate. The same logic extends (with more algebra) to regression coefficients. The reason economists cite <a href="https://doi.org/10.1016/0304-4076(73)90005-5">Aigner (1973)</a> for the known-rates correction and <a href="https://doi.org/10.1016/S0304-4076(95)01730-5">Bollinger (1996)</a> for the partial-information bounds is that they answered the two different situations: Aigner shows what full knowledge of the error rates buys, Bollinger shows what can still be said when we only know something. Search terms: &#8220;sensitivity specificity&#8221;, &#8220;misclassification correction&#8221;, &#8220;misclassified binary regressor&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>Misclassification is <em>non-differential</em> when the error rates are the same for everyone regardless of their treatment, outcome or characteristics, and <em>differential</em> when they aren&#8217;t. The distinction decides how much trouble we&#8217;re in. A classic example from epidemiology is recall bias in case-control studies: people who got sick search their memory for exposures harder than healthy controls do, so the exposure&#8217;s error rate depends on the outcome, and the bias can go in any direction. In our training example, people who benefited from the programme may be more likely to remember and report attending, which is the same disease. Even the comfortable non-differential case has fine print: the towards-zero result is for a binary exposure in simple settings, and <a href="https://doi.org/10.1016/S0304-4076(98)00015-3">Hausman, Abrevaya and Scott-Morton (1998)</a> show that misclassification in a discrete <em>outcome</em> biases coefficients in probit and logit models even when the errors are non-differential, which is the nonlinear-model warning from P05 wearing a different coat. Search terms: &#8220;differential misclassification&#8221;, &#8220;non-differential misclassification bias&#8221;, &#8220;recall bias&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>The best-known design is due to Hui and Walter: with two imperfect tests applied in two populations whose true prevalences differ, the error rates of both tests and both prevalences become estimable at once, with no gold standard anywhere. It feels like magic, and the magic has a price list. The measures must err independently given the true status, which is exactly what stops holding when both measures come from the same interview or the same administrative pipeline, and the true classes must be the same objects across the populations, and the design adds two requirements of its own: the two populations need different true prevalences, and each test&#8217;s sensitivity and specificity must stay constant across them (<a href="https://doi.org/10.2307/2530508">Hui and Walter 1980</a>). When any of these is wrong, the model may still converge and print tight standard errors (<a href="https://doi.org/10.1111/j.0006-341X.2004.00187.x">Albert and Dodd 2004</a>), which is why an apparently precise answer can be fragile: the precision is conditional on assumptions the data can&#8217;t check. Search terms: &#8220;latent class analysis without gold standard&#8221;, &#8220;Hui-Walter design&#8221;, &#8220;conditional dependence diagnostic tests&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>The hedonic idea is that a differentiated good is a bundle of characteristics, and the market implicitly prices the characteristics. Practically it&#8217;s a regression of price on attributes, and statistical agencies use exactly this to quality-adjust fast-changing goods like computers and housing. <a href="https://doi.org/10.1086/260169">Rosen (1974)</a> is the theory paper behind the practice, and it comes with a warning that often gets lost: the regression coefficients are equilibrium implicit prices, formed where buyers&#8217; preferences meet sellers&#8217; costs, and not demand parameters or willingness to pay. Two things limit any hedonic adjustment: the characteristics we didn&#8217;t record (a camera improves in ways no spec sheet captures) and the functional form we imposed. Search terms: &#8220;hedonic regression&#8221;, &#8220;hedonic quality adjustment&#8221;, &#8220;implicit prices&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>A chained index links many short comparisons: it compares this month with last month, then multiplies that by the comparison of last month with the month before, and so on. Chain drift is when this multiplication goes wrong in a specific way: prices and quantities come back exactly to where they started, but the index doesn&#8217;t. One well-documented cause is supermarket sales in scanner data (<a href="https://doi.org/10.1016/j.jeconom.2010.09.004">De Haan and van der Grient 2011</a>). During a sale the price is low and households buy a lot, so the price cut enters the chain with a big weight. When the sale ends the price goes back up, but by then purchases have returned to normal, so the price increase enters with a smaller weight. The cut and the increase don&#8217;t cancel out, the index ends the round trip below where it began, and over many sales the gap keeps growing. A chain-drift check tells us this is happening, and the multilateral methods built to avoid it can reduce the drift under a stated product universe, window and formula. What they can&#8217;t answer is whether a discounted, end-of-line item is still the same good as its full-price predecessor: that is a judgement about comparability, and - sadly - no statistic settles it. Search terms: &#8220;chain drift&#8221;, &#8220;scanner data price index&#8221;, &#8220;GEKS&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>As readers of this newsletter can remember, both designs have newer hybrids: augmented synthetic control brings in outcome modelling to patch residual pre-treatment imbalance (<a href="https://doi.org/10.1080/01621459.2021.1929245">Ben-Michael, Feller and Rothstein 2021</a>), and synthetic difference-in-differences combines unit and time weighting with a DiD adjustment (<a href="https://doi.org/10.1257/aer.20190159">Arkhangelsky et al. 2021</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>A step by step of the placebo logic: run the entire synthetic-control procedure once per donor state, each time pretending that donor passed the policy, and collect the resulting post-&#8220;policy&#8221; gaps. If our real state&#8217;s gap is the largest of twenty (or something) such runs, we can say its rank is 1 out of 20, and it&#8217;s tempting to call that p = 0.05, but we should be really careful because that number is a real randomisation p-value only if the policy was as good as randomly assigned among the states, and nothing in the procedure makes that true since states pass subsidies for n different reasons. Without that argument, the rank is still informative, our state stands out, but it&#8217;s a comparative statement instead of a significance level. <a href="https://doi.org/10.1257/jel.20191450">Abadie (2021)</a> is the survey treatment of this and of the feasibility conditions, and it&#8217;s the right first read before using the method. Search terms: &#8220;synthetic control method&#8221;, &#8220;placebo test synthetic control&#8221;, &#8220;in-space placebo&#8221;; implementations are <code>synth</code> (Stata, R) and <code>tidysynth</code> (R).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>A pre-trend check regresses the outcome on time relative to treatment and asks whether the treated and comparison groups were already diverging before the policy. <a href="https://doi.org/10.1257/aeri.20210236">Roth (2022)</a> documents two problems with it. The first one is low power: when the data are noisy, a violation large enough to bias our estimate badly can still pass the test, so a finding of &#8220;no significant pre-trend&#8221; often means we couldn&#8217;t detect the divergence rather than that there was none. The second problem comes from the pre-testing itself: if we only carry on with the analysis when the pre-test looks clean, then the estimates that end up published are a selected sample, and that selection introduces its own bias. None of this makes pre-trend plots useless since they still show what the two groups did before the policy; the trouble starts when we treat a passed test as a validated assumption. Search terms: &#8220;pre-trends test power&#8221;, &#8220;pretesting parallel trends&#8221;, and Prof Roth&#8217;s <code>pretrends</code> R package.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p>The Honest DiD idea is to replace &#8220;parallel trends holds&#8221; with &#8220;parallel trends can be violated, but only by this much&#8221;, and then to state the &#8220;this much&#8221; as a number. One version allows the post-treatment violation to be at most M times the largest pre-treatment violation, so M = 1 means the assumption can be broken after the policy by as much as it ever was before. Another version restricts how fast the violation can change from period to period (a smoothness restriction). For each choice of restriction, the method reports the range of effects consistent with the data, and the usual summary is the breakdown value: the M at which our conclusion changes, which in most applications means the M where the reported interval starts to include zero (<a href="https://doi.org/10.1093/restud/rdad018">Rambachan and Roth 2023</a>). Reading a result then becomes an argument about whether violations of that size are plausible in our application, and that argument is on econ rather than on stats. The implementation is the <code>HonestDiD</code> package (R and Stata); search terms: &#8220;Honest DiD&#8221;, &#8220;Rambachan Roth&#8221;, &#8220;relative magnitudes restriction&#8221;, &#8220;breakdown value&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-27" href="#footnote-anchor-27" class="footnote-number" contenteditable="false" target="_self">27</a><div class="footnote-content"><p>Discussing the taxonomy is important here since the words get swapped quite frequently in seminars: <em>reproduction</em> (or computational reproducibility) means running the same data through the same code and getting the same number; <em>replication</em> means asking the question again with new data or a new design; and both are different from <em>identification</em> (does the design isolate the effect?) and from empirical validity (is the claim true about the world?). The reason economists cite <a href="https://www.jstor.org/stable/1806061">Dewald, Thursby and Anderson (1986)</a> is the story: the <em>Journal of Money, Credit and Banking</em> asked authors for their data and programs, and most published results could not be regenerated, often for silly reasons like lost files and undocumented steps. <a href="https://doi.org/10.1561/104.00000053">Chang and Li (2022)</a> repeated the exercise decades later (sixty papers from thirteen journals) and their title answers the question: &#8220;often not&#8221;. The silly causes are the point of this entry: in those studies, results became irreproducible through ordinary decay, lost files, expiring agreements and departing coauthors, which doesn&#8217;t rule out misconduct elsewhere, but does show how far ordinary entropy gets on its own. Search terms: &#8220;computational reproducibility economics&#8221;, &#8220;replication in economics&#8221;, &#8220;data availability policy&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-28" href="#footnote-anchor-28" class="footnote-number" contenteditable="false" target="_self">28</a><div class="footnote-content"><p>An influence analysis asks how much each observation contributed to the final estimate, and a leave-one-out analysis answers it by brute force: drop one unit, re-estimate, repeat. Both need - at least - the per-unit material, the microdata or the fitted model&#8217;s unit-level pieces (in regression these are the score contributions, each observation&#8217;s push on the coefficient). This is why they can&#8217;t be run from a published table: a coefficient and a standard error are sums over units, and a sum doesn&#8217;t remember its parts. The practical lesson runs in both directions. Looking backward, if the objects are gone, the analysis is gone. Looking forward, archiving the fitted model and unit contributions alongside the code is what keeps this door open for our own papers. Search terms: &#8220;influence function&#8221;, &#8220;leave-one-out&#8221;, &#8220;dfbeta&#8221;, and for the recent automated version &#8220;finite-sample robustness metric&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-29" href="#footnote-anchor-29" class="footnote-number" contenteditable="false" target="_self">29</a><div class="footnote-content"><p>The tale is <a href="https://doi.org/10.2307/2087176">Robinson (1950)</a>, and it&#8217;s worth knowing in its original form. Across US states in 1930, the share of foreign-born residents correlated <em>positively</em> with English literacy rates: states with more immigrants had higher literacy. Across individuals, the correlation ran the other way since immigrants were on average less literate in English than the native-born. Both facts are true at once because immigrants had settled in states with high native literacy. Reading the state-level correlation as a statement about individuals became known as the <em>ecological fallacy</em>, and Robinson&#8217;s paper is the reason every methods course warns about it. Search terms: &#8220;ecological fallacy&#8221;, &#8220;ecological inference&#8221;, &#8220;cross-level inference&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-30" href="#footnote-anchor-30" class="footnote-number" contenteditable="false" target="_self">30</a><div class="footnote-content"><p>The aggregate association mixes up to three things, and <a href="https://doi.org/10.1093/ije/30.6.1343">Greenland (2001)</a> is the careful accounting of them. There is the individual effect we wanted (does training raise your earnings?), a possible contextual effect (does living in a high-training state raise your earnings through the labour market, whatever you did yourself?), and confounding at the group level (high-training states differ in other ways). The contextual effect deserves respect rather than deletion because sometimes it&#8217;s the actual question since spillovers from a trained workforce are a real policy object; the fallacy here is in reading an aggregate coefficient as an individual one without saying which of the three ingredients it contains. Search terms: &#8220;contextual effects&#8221;, &#8220;ecological bias&#8221;, &#8220;multilevel analysis&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-31" href="#footnote-anchor-31" class="footnote-number" contenteditable="false" target="_self">31</a><div class="footnote-content"><p>A shift-share decomposition splits an aggregate change into parts using an accounting identity: the change in a state&#8217;s average earnings, for example, into the part from employment shifting between industries (the composition part) and the part from earnings changing within industries. The arithmetic is exact and the output is descriptive. One naming warning because economics reused the phrase: &#8220;shift-share&#8221; also names the Bartik instrument literature where industry composition is used as an identification strategy, and that is a different object with its own assumptions rather than an accounting identity. A reader searching should use &#8220;shift-share decomposition&#8221; for this footnote&#8217;s tool and &#8220;shift-share instrument&#8221; or &#8220;Bartik instrument&#8221; for the other one. Related decompositions worth knowing by name: &#8220;within-between decomposition&#8221; and &#8220;Oaxaca-Blinder&#8221;.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[One number doing too much work]]></title><description><![CDATA[On distributional and mediation parameters in DiD]]></description><link>https://www.diddigest.xyz/p/one-number-doing-too-much-work</link><guid isPermaLink="false">https://www.diddigest.xyz/p/one-number-doing-too-much-work</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Tue, 04 Aug 2026 14:00:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IaPj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IaPj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IaPj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IaPj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2200520,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/196523680?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IaPj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!IaPj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F733f79b4-0e4a-433c-ad7a-b3d2ecaec4de_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! Some good news: I submitted my thesis, <s>so I should now have more time to focus on this newsletter. The bad news is that I need a job, hehe :) If you know of anything, please let me know, preferably industry roles.</s>, and I now have an <a href="https://science.works/about/">official job</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> :) I&#8217;ve been to <a href="https://www.linkedin.com/posts/beatrizgietner_just-got-back-from-beijing-the-second-time-ugcPost-7477686602417381378-tGyU/?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAACTHmhoBpQRxz0hUy8fhU1GuRQuRyKVjdd0">China</a> (again, yes, and many more times hopefully) for 10 days and had a great time, will write about it soon. I&#8217;ll be in Denmark on Friday Aug 7 for the Emergent Ventures European Conference, if you&#8217;re around say hello! Other conferences coming up are the <a href="https://eeassoc.org/events/eea-esem-dublin-2026">EEA-ESEM</a> in two weeks, which will happen at UCD. I volunteered to help, so if you&#8217;re coming to Dublin to present let me know!</p><p>Another piece of good news is that Prof <a href="https://yiqingxu.org/">Yiqing</a> has a Substack! It&#8217;s called Event Study and you can subscribe <a href="https://yiqingxu.substack.com/">here</a>. I&#8217;m a big fan of his academic work and of the effort he puts into making his research reproducible. I&#8217;m now looking forward to seeing him write for a broader audience.</p><p>Also a bit of house-cleaning:</p><ul><li><p>Prof Paul GP has been more active on <a href="https://paulgp.substack.com/">Substack</a>, with great posts recently</p></li><li><p>JAMA&#8217;s Guide to Statistics and Methods has a special DiD <a href="https://jamanetwork.com/journals/jama/fullarticle/2851729?guestAccessKey=38dd8365-5c2f-47c1-afb2-5623852ef29d&amp;utm_source=linkedin_company&amp;utm_medium=social_jama&amp;utm_term=21121981182&amp;utm_campaign=article_alert&amp;linkId=981460181">article</a> </p></li><li><p>A new NBER <a href="https://www.nber.org/papers/w35403">guide</a> &#8220;for economists writing a literature review who want to summarize evidence, incorporate covariates, and adjust for selectivity&#8221;</p></li><li><p>Prof Cl&#233;ment just solved all your problems with his <a href="https://chatgpt.com/g/g-6a20206ca0508191bcbb14e28b35a5b6-chatdid">ChatDiD</a>! Brilliant and a lot of fun</p></li><li><p>You can (should!) pre-order Prof Scott&#8217;s new book &#8220;Causal Inference: The Remix&#8221; <a href="https://www.amazon.com/Causal-Inference-Remix-Scott-Cunningham/dp/0300272448/ref=mp_s_a_1_15?crid=2G9KHBOU61E14&amp;dib=eyJ2IjoiMSJ9.k3Y8aN0W1Bk8pivsUu0K6Czgqxx4O8UF6jWGBSsqfTcKQoWJGLnUpUyQmtWOVDraPOsBJnjxRWcAFGcMYlqkGJUbocBqu9dcKBLX0kaQcMsMs3lw30tfgmooVNSkT7GmqbcSobJvtcylJk98C5Wgix1KWgDW-TI6RO3yBHDWWDiTmhOy6Yk3qlMJS5zmmYDP_4hIxDRpL3q0znWCRdDeLg.jyXLelqsK-hpU1UqWy35XcCWId_ijoUXKJIWZThHW1g&amp;dib_tag=se&amp;keywords=causal+inference&amp;qid=1779059317&amp;s=books&amp;sprefix=causal+inferenc%2Cbooks%2C390&amp;sr=1-15">here</a>! Out August 25 :)</p></li><li><p>My friend Aniket has started a newsletter on AI and agentic coding filtered for economists! Subscribe <a href="https://aieconomist.io/subscribe">here</a> </p></li><li><p>Prof Damian released the book &#8220;<a href="https://mitpressbookstore.mit.edu/book/9780262053648">Applied Microeconomics</a>&#8221;, which is <span>paired with a great companion </span><a href="https://damianclarke.net/books/microeconometrics/"><span>site</span></a></p></li><li><p>&#8220;A curated <a href="https://github.com/hanlulong/awesome-ai-for-economists">list</a> of AI tools, libraries, and resources for economics research, teaching, and policy analysis&#8221; by Lu Han (BoC)</p></li></ul><p>Today&#8217;s post is about papers that go beyond our usual ATT, either replacing it with distributional parameters or decomposing it into direct and indirect effects:</p><ol><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6386993">A simple distributional difference-in-differences estimator for univariate and bivariate outcomes</a>, by Aico van Vuuren, Iv&#225;n Fern&#225;ndez&#8208;Val, Francis Vella and Jonas Meier</p></li><li><p><a href="http://arxiv.org/abs/2602.23877">Difference-in-differences for mediation analysis using double machine learning</a>, by Martin Huber and Sarina Joy Oberh&#228;nsli</p></li><li><p><a href="https://arxiv.org/abs/2604.24049">Difference-in-differences with a mediator</a>, by Yuhao Deng, Haoyu Wei and Zhongzhe Ouyang</p></li></ol><p>The next papers that will be covered in upcoming posts are listed in the footnotes<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> :)</p><div><hr></div><h3>A simple distributional difference-in-differences estimator for univariate and bivariate outcomes</h3><h5>TL;DR: most DiD applications stop at the ATT, but the same design can tell you where in the distribution the effects happened. The authors run a small DiD for the probability of being below each threshold, on a logit/probit scale, and stitch the results into a counterfactual distribution for the treated group; an extension covers the joint distribution of two outcomes. The identifying restriction is a PT condition on a transformed probability scale, so the link function is part of the assumption. In the C&amp;K application, full-time employment gains are concentrated in larger establishments, with evidence consistent with substitution between full-time and part-time workers.</h5><p><em>What is this paper about?</em></p><p>In most DiD settings we are constrained, for many research-design-related reasons, to asking what a treatment did to the average outcome of the treated group: the ATT. Questions like &#8220;did a minimum-wage increase raise or reduce average employment?&#8221;, &#8220;did a school reform increase average test scores?&#8221; or even &#8220;did a health policy change average healthcare use?&#8221; are common and very useful, but we know that the average can conceal much of what has happened.</p><p>So let&#8217;s suppose a policy raises an outcome of interest by ten units for people near the bottom of the distribution and lowers it by ten units for people near the top. The average effect could then be close to zero even though the policy produced substantial changes for both groups. From this, we infer that a &#8220;zero&#8221; ATT doesn&#8217;t necessarily mean that nobody was affected, but that the same average effect can also describe very different distributions. A policy that gives every treated person a small gain could produce the same ATT as one that gives a large gain to a small group and leaves everyone else unchanged. Those two results have different implications for important measures like inequality, targeting and welfare, even though a conventional DiD estimate may report the same number.</p><p>In this paper the authors ask what we can learn when we replace the ATT with the whole outcome distribution<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. Instead of estimating only whether the mean changed, they estimate the treated group&#8217;s counterfactual distribution, meaning what the full distribution would&#8217;ve looked like after the intervention if treatment hadn&#8217;t occurred. They then compare it with the distribution observed under treatment.</p><p>This means we can examine where in the distribution the effects occurred. We can ask whether treatment changed the probability of falling below a specific threshold, which gives us a distributional treatment effect, or whether the effect differed at the bottom, middle or top of the distribution, which is captured by quantile treatment effects. The paper also considers cases in which treatment affects two related outcomes. This goes further than estimating a separate DiD for each one. Imagine that a minimum-wage increase affects both full-time and part-time employment. Separate estimates could show what happened to each outcome on its own, while the joint distribution can show whether firms changed the combination of full-time and part-time workers they employed. The paper therefore expands the object studied in a DiD design in two directions: from average effects to distributional effects and from one outcome to the joint behaviour of two outcomes.</p><p><em><span>What do the authors do?</span></em></p><p><span>The authors begin with the canonical DiD setting: two groups and two periods, with treatment received only by the treated group in the second period. We therefore have the familiar four cells, the treated and control groups before and after treatment. For the treated group after the intervention, we observe the distribution under treatment. What we don&#8217;t observe is the distribution that this same group would&#8217;ve had at the same time without treatment. This is the missing counterfactual distribution that the authors need to recover. A conventional DiD does this calculation for the mean: it uses changes in the control group to estimate how the treated group&#8217;s average outcome would&#8217;ve evolved without treatment. The authors apply the same logic at every possible value of the outcome.</span></p><p><span>This is where distribution regression</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span> comes in. Suppose our outcome is the number of workers employed by a restaurant. The authors take a threshold, such as ten workers, and create an indicator equal to one when a restaurant employs ten workers or fewer. They then estimate a logit or probit model for the probability of being below that threshold, using group and time indicators together with their interaction. Repeating the exercise for many thresholds, five workers, six workers and so on, traces out the outcome distribution. We can think of this as running a sequence of small DiD exercises, one for each threshold. At every point, the question is the same: how did the probability of being below this value change in the treated group relative to the control group?</span></p><p><span>The observed post-treatment distribution for the treated group comes directly from the data. To estimate its counterfactual distribution without treatment, the authors combine information from the other three group-by-time cells and impose their main identifying restriction, which they call the no-interaction assumption. The name is a little opaque, so it helps to translate it. For every possible outcome threshold, the authors assume that, without treatment, the treated and control groups would&#8217;ve experienced the same time change in the probability of falling below that threshold, after the probability has been transformed using the selected link function.</span></p><p><span>This is related to PT, but it is a different restriction, and one can hold when the other fails. Standard PT concerns the evolution of the average untreated outcome. Here, the restriction concerns the evolution of the entire untreated distribution, one threshold at a time, on a transformed probability scale. The reason for the transformation is more on the practical side. If we used a linear probability model, the DiD calculation could produce counterfactual probabilities &lt; 0 or &gt; 1, mainly near the tails of the distribution. Logit and probit transformations keep the estimated probabilities within their proper range.</span></p><p><span>The choice of link function still carries an assumption. Parallel changes on the logit scale aren&#8217;t the same as parallel changes on the probit scale, so the method doesn&#8217;t remove the need to justify the counterfactual. It shifts the identifying claim from trends in average outcomes to trends in transformed probabilities across the distribution. </span></p><p><span>The authors also introduce covariates. In that version, the no-interaction condition must hold after conditioning on the observed characteristics included in the model. This can help when the composition of the treated and control groups changes over time, although it requires sufficient overlap in those characteristics across all four cells. Once the two potential-outcome distributions have been estimated, the rest follows from comparing them. At each fixed threshold, their difference gives a distributional treatment effect. Researchers can also invert the distributions and compare their quantiles, giving quantile treatment effects. The bivariate extension follows the same idea with two outcomes. The authors first estimate the marginal counterfactual distribution of each outcome, and then use bivariate distribution regression to estimate how the two outcomes are related at different points of their joint distribution.</span></p><p><span>This dependence component is what we&#8217;d miss by running two separate DiD models. Suppose a minimum-wage increase raises full-time employment and reduces part-time employment. Separate regressions can estimate each change, while the joint model can also show whether restaurants increasingly substitute one type of worker for the other. Identification is more demanding in this case because the no-interaction condition now has to hold for both marginal distributions and for the dependence between the two outcomes</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a><span>.</span></p><p><span>From the estimated joint distributions, researchers can calculate joint probabilities and measures of dependence such as Kendall&#8217;s tau and Spearman&#8217;s rank correlation. This provides a way to study whether treatment changes how two outcomes move together. For inference, the paper proposes a weighted bootstrap. In panel data, all observations belonging to the same unit receive the same bootstrap weight, preserving the dependence between that unit&#8217;s observations over time. The estimator can also be used with repeated cross-sectional data.</span></p><p><span>To illustrate the method, the authors return to C&amp;K&#8217;s study of New Jersey&#8217;s 1992 minimum-wage increase. They examine the distributions of full-time and part-time employment, along with full-time-equivalent employment, in fast-food restaurants, using restaurants in Pennsylvania as the control group. The distributional results indicate that the estimated effects vary with establishment size. The increase in full-time-equivalent employment is mainly driven by gains in full-time employment, with larger gains towards the upper part of the distribution. The estimates for part-time employment are generally zero or negative, although the relatively small sample produces wide confidence intervals. The bivariate analysis adds another piece of information: the estimated relationship between full-time and part-time employment becomes more negative after the minimum-wage increase. The change in Spearman&#8217;s rank correlation is statistically significant at the 10 per cent level, which the authors interpret as consistent with greater substitution between full-time and part-time workers.</span></p><p><span>The authors also conduct a specification check based on whether the estimated counterfactual distribution is properly increasing. They don&#8217;t reject monotonicity in their application, which is compatible with the model. This check can detect certain failures of the specification, though it can&#8217;t establish that the no-interaction assumption is true.</span></p><p><em><span>Why is this important?</span></em></p><p><span>A large share of policy questions are distributional at their core. Evaluations of minimum wages or school reforms often need to say who gained and who lost, and a single average can&#8217;t answer that. This paper makes the distributional version of the question answerable with tools most applied researchers already have, since the estimation reduces to a sequence of logit or probit regressions followed by a bootstrap.</span></p><p><span>The identification lesson is worth the read on its own. Justifying PT in means doesn&#8217;t justify a counterfactual distribution. The restriction the authors need operates threshold by threshold on a transformed probability scale, and the choice of link function is part of it. Researchers who report quantile treatment effects from DiD designs should be able to say which restriction their numbers rely on, and this paper states it in a very clean way.</span></p><p><span>The bivariate extension adds a question that single-outcome methods can&#8217;t answer: whether treatment changed how two outcomes move together. In the minimum-wage application, that is the difference between knowing that full-time employment rose while part-time employment fell, and knowing that restaurants substituted one type of worker for the other.</span></p><p><em><span>Who should care?</span></em></p><p><span>You should read this paper if you work on policies where the distribution of the outcome is the question, such as effects on inequality or on the share of units below a threshold; report quantile treatment effects from a DiD design and want the identifying restriction </span>explicitly <span>stated; study two related outcomes and suspect the interesting response is in their combination; and want distributional estimates without leaving standard logit and probit machinery.</span></p><p><em><span>Do we have code?</span></em></p><p><span>I couldn&#8217;t find a package or a link to code in the paper. The estimation itself stays close to standard practice, a set of logit or probit regressions plus a weighted bootstrap, so implementing it from scratch looks feasible for anyone comfortable with distribution regression. The paper is an SSRN preprint and hasn&#8217;t been peer reviewed yet, so a package may still appear.</span></p><p><span>In summary, the paper changes the estimand of a DiD while keeping its logic. We still impute a counterfactual for the treated group from the other three cells, and we still need an untestable assumption to do it. What changes is the object, now a whole distribution or the joint distribution of two outcomes, and the assumption, which lives on a transformed probability scale and includes the link function. If the distribution is what your research question is about, this is a practical way to estimate it and to be explicit about what the estimate relies on.</span></p><h3>Difference-in-differences for mediation analysis using double machine learning</h3><h5>TL;DR: mediation analysis usually requires an unconfounded treatment, which is precisely what DiD users don&#8217;t want to assume. The authors split the total effect on the treated into direct and indirect channels using PT assumptions imposed across treatment-mediator combinations, with DML handling the covariates. The price is that PT must now hold across groups defined partly by mediator behaviour, which is stronger than the standard version. In their NLSY97 application, health care coverage shows no significant short-term effect on general health, through routine checkups or otherwise.</h5><p><em><span>What is this paper about?</span></em></p><p><span>The previous paper changed the object we estimate, whereas this one asks where the effect comes from. A DiD design gives us the total effect of a treatment on the treated, but applied questions often go one step further: how much of that effect &#8220;travels&#8221; through a specific channel? In the authors&#8217; application, health care coverage may improve general health partly because insured people go for routine checkups. The checkup is a mediator (a variable sitting on the causal path between treatment and outcome) and the goal is to split the total effect into an indirect effect operating through the mediator and a direct effect capturing everything else.</span></p><p><span>The standard toolkit for this split - causal mediation analysis</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a><span> - usually assumes that both the treatment and the mediator are as good as randomly assigned once we condition on observed covariates. That is exactly the assumption that researchers who want to use DiD don&#8217;t want to make about the treatment; if we believed it, we wouldn&#8217;t be running a DiD in the first place. So the paper asks whether mediation analysis can be carried out inside a DiD framework, replacing unconfoundedness of the treatment with PT conditions.</span></p><p><span>Their framework is broad on purpose. Treatments and mediators can be binary, multivalued or continuous, and the paper covers several targets. Some fix the mediator directly (joint effects of setting treatment and mediator together, and controlled direct effects that hold the mediator at a chosen value), while others (the natural direct and indirect effects</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a><span>) let the mediator take whatever value it takes with and without treatment. The distinction is important because, as we&#8217;ll see, the two &#8220;families&#8221; require different assumptions.</span></p><p><em><span>What do the authors do?</span></em></p><p><span>They define groups by combinations of treatment and mediator values. With two periods, this recreates the four-cell logic of canonical DiD, except that the &#8220;treated&#8221; cell is now a treatment-mediator pair - say insured people who went for a checkup - and the comparison cell is another pair - say uninsured people who didn&#8217;t.</span></p><p><span>The central identifying assumption is conditional PT across these treatment-mediator combinations: given covariates, the potential outcomes of both groups under the comparison state would have followed the same average trend. This is stronger than PT across treatment groups alone because the groups are also defined by how the mediator responds. In the binary case, it amounts to assuming PT across people who take up the mediator when treated and people who don&#8217;t take it up when untreated, and these are groups with different underlying mediator behaviour. The authors are open about this strength and point to the pay-off: unlike earlier DiD-mediation work, they need no monotonicity of the mediator in the treatment, and non-binary treatments and mediators are covered.</span></p><p><span>To this they add no anticipation, common support (comparable covariate values must exist in all four cells) and the requirement that covariates are unaffected by treatment or mediator. The last condition deserves attention in repeated cross sections where covariates are often measured after treatment.</span></p><p><span>These assumptions identify the joint effects and the controlled direct effects. Natural effects need more, and the authors offer two routes. The first adds PT across treatment groups, which fails if unobserved factors drive both treatment take-up and the mediator. The second instead adds a distributional PT assumption on the mediator itself, and it requires the mediator to be observed in the pre-treatment period. Neither route is &#8220;free&#8221;, and which one is more credible depends on the application.</span></p><p><span>The estimation belongs in the DML framework. The authors derive doubly robust score functions that satisfy Neyman orthogonality, which in practice means two things: covariates can be controlled for with ML (the authors use LASSO), because first-order estimation errors in those auxiliary models don&#8217;t contaminate the effect estimate; and cross-fitting (estimating the auxiliary models and the effects on separate folds of the data) guards against overfitting. The estimators are asymptotically normal at the usual rate for discrete treatments and mediators; continuous ones bring kernel smoothing and slower convergence. Simulations at sample sizes of 2000 and 8000 show small bias, with confidence intervals slightly too narrow in the smaller samples.</span></p><p><span>The empirical application revisits Farbmacher and co-authors&#8217; 2022 </span><a href="https://academic.oup.com/ectj/article-pdf/25/2/277/43772864/utac003.pdf?casa_token=Gu3yTFCqcloAAAAA:QKbhNGqlCfFO9HsPmL3trGEGmZf0HJqhgOztb2fi76W-BsFm2RPAZma0mNmFTz-wrWTHkgLW9l-uxH4"><span>study</span></a><span> of the effect of obtaining health care coverage in 2006 on self-reported general health in 2008 among NLSY97</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a><span> respondents, with routine checkups in 2007 as the mediator. The DiD design changes the sample: respondents must report health in 2005 and 2008 and must have had no coverage before 2006, which shrinks the sample from 7486 to 1020. No estimated effect, total, direct or indirect, is statistically significant, although the point estimates of the total and direct effects lean towards better health, broadly in line with the original study&#8217;s direction.</span></p><p><em><span>Why is this important?</span></em></p><p><span>A reform&#8217;s evaluation often needs to say through which channel the effect arrived because the channel determines which part of the policy is worth keeping or adjusting. Until now, answering that question usually meant abandoning DiD&#8217;s main advantage - its tolerance of unmeasured confounding between treatment and outcome. This paper keeps the DiD logic and prices the mediation question in PT currency instead.</span></p><p><span>The price we gotta pay is visible, and to me that&#8217;s a virtue. The PT conditions are stated cell by cell, so a referee, or the researcher, can interrogate each one. The paper also draws a useful bridge to the staggered adoption literature: its key assumption is of the same type as those used with multiple treatment periods, with early adoption in the role of the treatment and later adoption in the role of the mediator. If you already accept those assumptions in staggered designs, you have a head start on judging these.</span></p><p><span>The ML component earns its place too. Conditional PT is easier to defend with a generous covariate set, and the DML machinery makes a large set usable without hand-picking controls.</span></p><p><em><span>Who should care?</span></em></p><p><span>You should read this paper if you use DiD and face questions about mechanisms, such as which channel carried a policy&#8217;s effect; work with multivalued or continuous treatments or mediators; have many candidate control variables and want data-driven covariate adjustment with valid inference; and follow the growing DiD-mediation literature and want the assumption trade-offs mapped out. If your mediator is plausibly randomised, standard mediation tools with selection on observables may serve you more simply.</span></p><p><em><span>Do we have code?</span></em></p><p><span>The paper states that the estimator is implemented in R, with four-fold cross-fitting and LASSO-based nuisance estimation, but I couldn&#8217;t find a link to a package or repository in the paper. I&#8217;d expect an implementation to show up, given how much of the authors&#8217; related work has been packaged, but I couldn&#8217;t verify one at this time.</span></p><p><span>In summary, this paper translates mediation questions into DiD language. Instead of assuming an unconfounded treatment, you assume PT across treatment-mediator cells and you add further conditions if you want natural rather than controlled effects. The estimands deserve as much attention as the estimator: controlled and natural effects answer different questions and carry different assumptions, and the paper is clear about which is which. For applied researchers with a credible DiD and a mechanism question, the decomposition is now feasible; the assumptions still need arguing one cell at a time.</span></p><h3>Difference-in-differences with a mediator</h3><h5><span>TL;DR: earlier Job Corps studies found no short-run earnings effect of job training. The authors build mediation into a two-period DiD and show that the flat total effect hides a significant positive direct effect on earning capacity, offset by an insignificant negative effect from reduced work time while training. Unmeasured confounding between treatment and outcomes is fine, since differencing removes it, but the mediator must be unconfounded, and that assumption can&#8217;t be tested from data. The estimators are multiply robust: consistent if any two of three working models are correct.</span></h5><p><em><span>What is this paper about?</span></em></p><p><span>This paper starts from a substantive puzzle in the Job Corps Study, a large US training programme for disadvantaged youths</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a><span>: earlier studies found earnings gains only in the long run, and the flat short-run effect may be mechanical (time spent in training is time not spent working, so weekly earnings can fall even while earning capacity rises).</span></p><p><span>Sorting this out requires a mediation analysis, with the share of weeks employed as the mediator, in a setting where treatment take-up is confounded with outcomes. The paper builds that analysis inside a two-period DiD. The key idea is a division of labour between assumptions: unmeasured confounding between treatment and outcomes is allowed because differencing outcomes over time removes it, while the mediator has to be clean, meaning free of unmeasured confounding with the treatment and with the outcome trend.</span></p><p><em><span>What do the authors do?</span></em></p><p>The target parameters are defined for the treated group: the total effect of training, which splits into a natural indirect effect operating through work time and a natural direct effect capturing the remaining channels. </p><p>The identification starts from two familiar DiD ingredients. The first is no anticipation, so that baseline earnings don&#8217;t respond to training that hasn&#8217;t happened yet, and the second is a PTA, adjusted for the mediator: conditional on covariates and on the value the mediator takes without treatment, untreated potential outcomes follow the same trend in both groups. To these the authors add sequential ignorability - adapted to the DiD setting, which requires that treatment take-up is unconfounded with the mediator given covariates and that the observed mediator doesn&#8217;t modify the underlying trend. The usual overlap and consistency conditions complete the set, which ensures that the groups contain comparable observations and that observed outcomes correspond to the relevant potential outcomes.</p><p><span>A design detail I like: the PT condition deliberately avoids conditioning on the observed post-treatment mediator because conditioning on a post-treatment variable can create collider bias (a spurious association produced by conditioning on a common consequence). The authors contrast their approach with related work on exactly this point, including the Huber and Oberh&#228;nsli paper above.</span></p><p><span>For estimation, the authors derive efficient influence functions</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a><span> from semiparametric theory and build estimators that are multiply robust. Three working models are involved: one for the propensity score, one for the mediator&#8217;s distribution and one for the change in outcomes. The estimator remains consistent if any two of the three are correctly specified, and it reaches the nonparametric efficiency bound when all are. After that a practical section shows how to implement everything with generalised linear models, including a trick that avoids estimating the mediator&#8217;s density directly. The paper also treats controlled direct effects - where a continuous mediator creates a technical obstacle (the estimand stops being pathwise differentiable) - handled with kernel smoothing at the cost of a slower convergence rate.</span></p><p><span>The application returns to Job Corps with actual enrolment as the treatment, log earnings as the outcome and the share of weeks employed as the mediator. In the shorter-run analysis, the total effect on earnings is small and insignificant, and the decomposition explains why: a positive, statistically significant natural direct effect of about 0.11 log points</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a><span> is offset by a negative, insignificant indirect effect running through reduced work time. A parallel analysis one year later (using training in the second year and the change in earnings from the second to the third year) finds all effects positive and significant with a direct effect of about 0.38 log points. A sensitivity analysis coding employment in four levels reaches the same conclusions, and the estimated controlled-direct-effect curve is significantly positive across a broad range of employment levels in the later year.</span></p><p><em><span>Why is this important?</span></em></p><p><span>The substantive lesson pairs well with the methodological one. Earlier readings of Job Corps concluded that training doesn&#8217;t help earnings in the short run. Their decomposition suggests the short-run total effect mixes two opposing forces, higher earning capacity and fewer weeks worked, and the &#8220;no effect&#8221; verdict conceals the first. That is the same theme running through this whole post: a single summary number can hide the part of the answer the policy question really needs.</span></p><p><span>On the methods side, the paper brings efficiency theory and multiple robustness to DiD mediation. Multiple robustness is practical insurance since you don&#8217;t have to bet the analysis on one correctly specified working model. The paper is also honest about its limits. Sequential ignorability can&#8217;t be verified from data, and the discussion flags open problems, including the difficulty of modelling the mediator&#8217;s distribution and the treatment of zero earnings.</span></p><p><em><span>Who should care?</span></em></p><p><span>You should read this paper if you evaluate programmes where the treatment absorbs participants&#8217; time (e.g., training or further education) and short-run outcomes look flat; want mediation analysis in settings with non-random take-up; care about multiply robust estimation and are comfortable with (or curious about!) influence-function-based methods; and work with continuous mediators and want effect curves across mediator values instead of a single number.</span></p><p><em><span>Do we have code?</span></em></p><p><span>I couldn&#8217;t find a package or repository linked in the paper. The data are publicly available through the R package </span><a href="https://cran.r-project.org/web/packages/causalweight/index.html"><span>causalweight</span></a><span>, and the supplementary material spells out a parametric estimation strategy in enough detail to implement, but the implementation work is currently left to the reader.</span></p><p><span>In summary, the paper shifts the confounding burden of mediation analysis from the treatment, where DiD differencing can deal with it, to the mediator, where only argument can. When that argument is credible, as the authors maintain for work time in Job Corps, the reward is a decomposition that changes the substantive reading of a famous programme: training builds earning capacity from the start and the short-run earnings dip mostly reflects the time participants spend in training.</span></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;m not moving to London <strong>yet</strong>, but will be sharing my time here and there.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><a href="https://arxiv.org/abs/2608.03881v1">Difference-in-differences with &#8220;bad controls&#8221;</a>, by Carolina Caetano, Brantly Callaway, Stroud Payne and Hugo Sant&#8217;Anna</p><p><a href="https://arxiv.org/abs/2601.15764v1">Three&#8217;s a crowd: Identification challenges in the triple difference model with spillover effects</a>, by Silvia De Nicol&#242;, Beatrice Biondi and Mario Mazzocchi</p><p><a href="https://arxiv.org/abs/2604.12818">Causal Graphs for Conditional Parallel Trends</a>, by Michael C. Knaus and Henri Pfleiderer</p><p><a href="https://arxiv.org/abs/2604.22982">Stacked Triple Differences</a>, by Meng Hsuan Hsieh</p><p><a href="https://arxiv.org/abs/2604.27035">Doubly robust local projections difference-in-differences</a>, by Daniel de Abreu Pereira Uhr and Guilherme Valle Moura</p><p><a href="https://arxiv.org/abs/2603.20394">When are time series predictions causal? The potential system and dynamic causal effects</a>, by Jacob Carlson and Neil Shephard</p><p><a href="https://arxiv.org/abs/2601.00603">Difference-in-Differences using Double Negative Controls and Graph Neural Networks for Unmeasured Network Confounding</a>, by Zihan Zhang, Lianyan Fu and Dehui Wang</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6161168">Debiased Difference-in-Differences Estimation and Inference of Causal Effects Under Adaptive Controls</a>, by Zhiyuan Tang, Yining Wang and Sentao Miao</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5991734">It&#8217;s About Time (Series): A Simple Correction For Difference-in-Differences Estimators</a>, by Gary Cornwall and Scott Wentland</p><p><a href="https://arxiv.org/abs/2607.21312"><span>Using Pre-Trends for Inference in Difference-in-Differences</span></a><span>, by Cl&#233;ment de Chaisemartin</span></p><p><a href="https://arxiv.org/abs/2607.19925"><span>Efficient difference-in-differences estimation under partial interference with incremental propensity score policies</span></a><span>, by Junjie Li and Yukitoshi Matsushita</span></p><p><a href="https://arxiv.org/abs/2607.08324"><span>Finite-Population Inference for Heterogeneity in Many-Group Synthetic Difference-in-Differences</span></a><span>, by Takahiro Hoshino and Makoto Nakakita</span></p><p><a href="https://arxiv.org/abs/2606.24785"><span>Group-Level Treatment Effect Heterogeneity in Difference-in-Differences: A Balanced Approach</span></a><span>, by Nora Bearth, Nadja van 't Hoff and Torben S. D. Johansen</span></p><p><a href="https://arxiv.org/abs/2606.20435"><span>Choosing A Headline Estimand from Matching, DID, and Hybrid Designs: A Minimax-Regret Approach</span></a><span>, by Yechan Park and Yuya Sasaki</span></p><p><a href="https://arxiv.org/abs/2606.08474"><span>Semiparametric Difference-in-Differences Estimation With Missing Not at Random Data: A Shadow Variable Approach</span></a><span>, by Junjie Li and Dongyuan Mu</span></p><p><a href="https://arxiv.org/abs/2606.02234"><span>When Do Treatment Changes Identify Causal Effects?</span></a><span>, by Martin Huber</span></p><p><a href="https://arxiv.org/abs/2605.15119"><span>Identification and Estimation of Staggered Difference-in-Differences with Network Spillovers</span></a><span>, by Hayato Tagawa</span></p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6753318">Adaptive Causal Inference for Higher-Order Difference-in-Differences Designs</a>, by <span>Jangsu Yoon</span></p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7006383">An Evaluation of Difference-in-Differences Methods Using Placebo Event Studies</a>, by John M. Coglianese and Jade A. Fang</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6844479">Rolling Difference-in-Differences with Reversible Treatment Paths and Moderating Effects</a>, by Soo Jeong Lee</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6813701">Choosing a Difference-in-Differences Estimand and Estimator for Heavy-Tailed Outcomes: A Practitioner&#8217;s Companion</a>, by Daniel Winkler, Christian Hotz-Behofsits and Nils Wl&#246;mert</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7157418">Difference-in-Differences with Spillovers: An Imputation Approach</a>, by Soumen Banerjee and Jianguo Wang</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6991840">Di&#64256;erence-in-Di&#64256;erences with multiple Treatments under Control</a>, by Daniel Steinberg and Marcus Roller</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>A bit of background: earlier work extended DiD beyond average effects in several directions. <a href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1468-0262.2006.00668.x?casa_token=mEVjxEyd75EAAAAA:y8K93mnLlaolxH4orFueUjEaXz4fwKyOEuT6LKRQ2Ky75bY6k3D-TRDd4_jp-L_hr5eYs453MBBRdGPY">Athey and Imbens</a> (2006) introduced changes-in-changes to recover counterfactual outcome distributions, while later work examined quantile treatment effects (<a href="https://onlinelibrary.wiley.com/doi/pdf/10.3982/qe935">Callaway and Li, 2019</a>) and treatment effects at different points of the outcome distribution using conventional DiD (<a href="https://assets.aeaweb.org/asset-server/files/8842.pdf">Dube, 2019</a>; <a href="https://www.jstor.org/stable/pdf/27086084.pdf?casa_token=2qi1UeeuZoYAAAAA:6Ok5YlAWZSivOg5n4a_JCYCw4VWXfi1rLwiSEfcyil7WGL26wgzvilUnG9OuqsPs26MGfBg7CQsPPoA-2upxen3sJ_4hYs34xCTOywq4moTv580nJGfr">Goodman-Bacon, 2021</a>; <a href="https://www.sciencedirect.com/science/article/pii/S0047272720300384?casa_token=kTxMK6U4NR4AAAAA:nziaEgYSIGv7bfLr7BNhyjcqpr8EJXYK5PgIEiaZj7RrRS-HxmGnjt3Nql4BTrD4htYqb2ZwcA">Goodman-Bacon and Schmidt, 2020</a>). Other approaches include IPW (<a href="https://www.tandfonline.com/doi/pdf/10.1080/07350015.2024.2388643?casa_token=Y_At_uB228YAAAAA:lR_6Kt9p10aSnFsxYYP5cYMWHqTIQdSOiNCCwp1EDqcBdQjuVCIDmrZvrjpWbuoneI3aqjPUwZAOnQ">Kim and Wooldridge, 2025</a>), distribution regression (<a href="https://www.econstor.eu/bitstream/10419/265755/1/dp15534.pdf">Biewen, R&#252;mmele and Fitzenberger, 2022</a>) and a multivariate extension of changes-in-changes (<a href="https://arxiv.org/pdf/2108.05858">Torous, Gunsilius and Rigollet, 2021</a>). Their paper builds on the distribution regression approach, using nonlinear link functions such as probit or logit instead of linear probability models and providing the identifying conditions required for this version of DR-DiD.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>Distribution regression models the whole conditional distribution of an outcome by estimating a separate binary-choice model for the probability of being at or below each threshold. The idea goes back to <a href="https://www.jstor.org/stable/pdf/2291056.pdf?casa_token=bk_XKyNm1LkAAAAA:2ra5XVL5ooevSxXc5Lee-jtV2H2qWaUHeVEsUAY_kmT9sia4ggRPD1m32NwAa8BtuMv0oaFZU4wqBIdIUkJyQwxeYjt-Gg8i7PITA83iax9eIwvVpPta">Foresi and Peracchi (1995)</a>, and <a href="https://onlinelibrary.wiley.com/doi/pdf/10.3982/ECTA10582?casa_token=4S9K95vVXIsAAAAA:eF0Wt5yvXSLrDzSgQNeVQvTw0oCsmTOw4buYGKwVx5Qq-_vw7rc1kXTuF86AQP8VhMgsMb8JN3ZXtjGG">Chernozhukov, Fern&#225;ndez-Val and Melly (2013)</a> developed its use for counterfactual distributions.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p> In plain English, without treatment, the relationship between the outcomes must also have evolved in the same way for the treated and control groups on the scale used by the model.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Mediation analysis asks how a treatment produces its effect, by splitting the total effect into two parts: an indirect effect that travels through an intermediate variable (the mediator) and a direct effect that covers everything else. Take job training and earnings, for example: training changes how many weeks people work, and work time changes earnings, so part of the earnings effect runs through work time and part comes from other channels, like higher productivity. The catch is that the mediator is usually not randomised, even when the treatment is, since people who work more weeks (or go for checkups) differ from those who don&#8217;t. That&#8217;s why the classic versions assume both treatment and mediator are unconfounded given covariates, and why building mediation into DiD - where the treatment is exactly the confounded part - takes some work.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>The natural/controlled terminology comes from <a href="https://www.ovid.com/jnls/epidem/pdf/00001648-199203000-00013~identifiability-and-exchangeability-for-direct-and-indirect?casa_token=HDSrBATkzn4AAAAA:jbb1I9Qm373LuZ3ZVyLBfPQun0BYTf4soVQ0Z4RUYx8IOMqMCW5kONKxd4yw0-jFZeA873UVfnIE_5Pi4G_crrp27eKQzVwJ71s0">Robins and Greenland (1992)</a> and <a href="https://idp.springer.com/authorize/casa?redirect_uri=https://link.springer.com/content/pdf/10.1023/A:1020315127304.pdf&amp;casa_token=aN38lNEtYFwAAAAA:2SPNyvfgsavqarVMbt1odZZaNe0TjOFWrEVyMQIPBK68TFjoYBCrWIz5G9BSEifT2e35W-I133GzxHHSzg">Pearl (2001)</a>. A controlled direct effect fixes the mediator at the same value for everyone; natural effects let each person&#8217;s mediator take the value it would take under a given treatment status. <a href="https://books.google.com/books?hl=en&amp;lr=&amp;id=GoqZEAAAQBAJ&amp;oi=fnd&amp;pg=PA1&amp;dq=huber+2021+causal+analysis&amp;ots=kGj3isLNpk&amp;sig=5nQLOjHiM9HlbfYEVQCjz1A1c9k">Huber (2021)</a> surveys causal mediation analysis.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>The National Longitudinal Survey of Youth 1997, a nationally representative US sample of people aged 12 to 17 at their first interview in 1997.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>There are two motivations - one substantive and one methodological - for this paper, and the authors are quite explicit about both. The substantive one is the Job Corps puzzle. Assignment to the programme was randomised, but actual enrolment wasn&#8217;t - people chose whether to take up training, and those with more pessimistic labour-market expectations were more likely to join. A series of earlier studies (using IV, principal stratification, and bounds) all reached the same conclusion: a significant long-term earnings effect, and nothing detectable in the short term. The authors&#8217; suspicion is that this &#8220;nothing&#8221; is partly an artefact: training eats up a large share of participants&#8217; time so they work fewer weeks, and weekly earnings fall &#8220;mechanically&#8221; when you can&#8217;t work full-time. None of the earlier analyses separated &#8220;training didn&#8217;t build skills yet&#8221; from &#8220;trainees were busy training&#8221;. Doing that separation requires treating work time as a mediator, and doing it in a setting where the treatment is self-selected. The methodological one is that no existing tool fit that problem. Their introduction walks through the alternatives and finds each one wanting for this setting:</span></p><ul><li><p><span>Classical mediation analysis requires the treatment, mediator and outcome to all be unconfounded, which is hopeless with self-selected enrolment.</span></p></li><li><p><a href="https://www.tandfonline.com/doi/pdf/10.1080/07350015.2017.1419139?casa_token=lvIMP1cckwAAAAAA:V9Yo2J5qo40kBe_-A02mQrKtC4L6madmGcf6RK_rJlFVnl7qCAuqevg-w5rkiBhf1VcvLcDvLIJBag"><span>Deuchert, Huber and Schelker (2019)</span></a><span> do DiD mediation but need a binary mediator, monotonicity and PT within principal strata.</span></p></li><li><p><a href="https://www.tandfonline.com/doi/pdf/10.1080/07350015.2020.1831929?casa_token=Fpy5F0yPxUIAAAAA:reboFcDrnsNOeG_l73isU0FezYn81whlJehDh63vVE0VH_11rYnRPHdFJgHbad0Ey8-1oUtlQjgRZg"><span>Huber et al. (2022)</span></a><span> use changes-in-changes, with distributional restrictions and binary mediators.</span></p></li><li><p><a href="https://www.degruyterbrill.com/document/doi/10.1515/em-2024-0025/html?recommended=sidebar"><span>Hsia et al. (2025)</span></a><span> put mediation into panel data via linear regressions, which can&#8217;t accommodate interactions.</span></p></li><li><p><a href="https://academic.oup.com/jrsssa/article-pdf/189/2/1070/63054224/qnaf042.pdf?casa_token=0mSgzcLq4okAAAAA:aPO6jUSfc-wHeFA4j0_AirMz5I4ZWXp_pMjHrMxhOxjMYKieDXYXW2WkE-Yt2vcS32cooSHvDz_5Uuo"><span>Blackwell et al. (2025)</span></a><span> handle controlled effects but need a randomised treatment and a discrete mediator.</span></p></li></ul><p>Each of these covers part of the territory. The combination this paper goes after - natural direct and indirect effects in a standard two-period DiD - with self-selected treatment and a possibly continuous mediator, wasn&#8217;t available before.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p><span>An influence function describes how much an estimator moves when the data distribution is perturbed slightly. Building the estimator from the efficient influence function is what delivers both the efficiency bound and the multiple-robustness property.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Roughly, a 0.11 log-point effect corresponds to an 11 to 12 per cent increase in weekly earnings, and 0.38 to about 46 per cent.</p></div></div>]]></content:encoded></item><item><title><![CDATA[The estimate is in the design ]]></title><description><![CDATA[On stacking, discrete outcomes, matched inference, and data-driven controls]]></description><link>https://www.diddigest.xyz/p/the-estimate-is-in-the-design</link><guid isPermaLink="false">https://www.diddigest.xyz/p/the-estimate-is-in-the-design</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 30 Mar 2026 14:36:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iVR_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iVR_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iVR_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iVR_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2366185,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/191847251?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iVR_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!iVR_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb923c8e-b365-4284-bda7-3667e3ddffdf_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! Apologies for the disappearance! My thesis is due in a month so I was/am busy with that. There are 10 papers to catch up on, so we will start with four of them today and hopefully I can write another post next week.</p><p>I picked these first four because I see an overall theme in them: DiD looks fairly straightforward in theory, but irl applications tend to make the assumptions do a lot of work, and these papers all show, in different ways, how design, outcome structure, and inference determine what your estimates are identifying in practice. So today&#8217;s post is based on:</p><ol><li><p><a href="https://www.nber.org/papers/w32054">Stacked Difference-in-Differences</a>, by Coady Wing, Seth M. Freedman and Alex Hollingsworth</p></li><li><p><a href="http://arxiv.org/abs/2603.07914">Event-Study Designs for Discrete Outcomes under Transition Independence</a>, by Young Ahn and Hiroyuki Kasahara</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6151287">Inference for Matched Difference-in-Differences</a>, by Mijeong Kim and Mingue Park</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5956311">Parallel Trends Forest: Data-Driven Control Sample Selection in Difference-in-Differences</a>, by Yesol Huh and Matthew Vanderpool Kling</p></li></ol><div><hr></div><h3>Stacked Difference-in-Differences </h3><p><em>( This is for Prof Khoa :) Also there are lots of info in the footnotes because this topic is important - not that the others aren&#8217;t but I feel like applied people may come across this sort of approach/design-based implementation or want to apply it themselves more often than some of the other methods we covered. So just for you to remember, stacked DiD is a regression-based approach used in staggered adoption settings, where different units receive treatment at different times<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. It became popular as researchers moved away from standard two-way fixed effects (TWFE) in settings with heterogeneous treatment effects, since TWFE can compare already-treated units with later-treated units in ways that don&#8217;t recover a clear causal average<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>.)</em></p><h5>TLD;DR: stacked DiD became popular in staggered adoption settings because it avoids the late-versus-early treated comparisons that caused problems for standard TWFE. This paper shows that the usual unweighted stacked regression still doesn&#8217;t have a clear causal interpretation because treatment and control trends are combined with different implicit weights across sub-experiments. The authors define a target parameter - the trimmed aggregate ATT - and show that a weighted stacked estimator lines up with it. The result is a more defensible stacked DiD approach for applied work, with code provided for the weighting procedure.</h5><p><em><strong>What is this paper about?</strong></em></p><p>This paper studies the performance of stacked DiD estimators<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> in staggered adoption settings. The authors take a trimmed aggregate ATT<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> as the target parameter and ask whether stacked regressions recover it. Their main point is that the most basic stacked estimator doesn&#8217;t identify that target because it places different implicit weights on treatment and control trends across sub-experiments. As a result, even if each individual sub-experiment satisfies the standard DiD assumptions, the pooled stacked regression<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> can&#8217;t be given a clean causal interpretation as an ATT. The paper then derives corrective sample weights and proposes a weighted stacked estimator that does identify the trimmed aggregate ATT.</p><p><em><strong>What do the authors do?</strong></em></p><p>The paper moves in a fairly clear sequence. First, the authors define the parameter they want to study: the trimmed aggregate ATT, which is a weighted average of group-time treatment effects over a trimmed set of adoption cohorts. The trimming step is there to keep the composition of the event-study sample stable across the chosen pre- and post-treatment window so that movements in the estimates over event time are not driven by different cohorts entering and leaving the average.</p><p>They then turn to the standard unweighted stacked regression and ask whether it recovers that target parameter. Their answer is no :) The problem is not that the individual sub-experiments are invalid on their own. The problem is that, once they are pooled into one stacked regression, treatment and control trends are combined with different implicit weights across sub-experiments. Because of that the basic stacked estimator doesn&#8217;t identify the trimmed aggregate ATT, or any other average causal effect in general.</p><p>The next step is the main constructive part of the paper. The authors derive corrective sample weights and use them to build a weighted stacked DID estimator. With those weights in place, the stacked regression does recover the trimmed aggregate ATT. They show that this can be implemented through weighted least squares, either in a simple DID setup or in an event-study specification. They also note that the same logic can be adapted if the researcher wants a different aggregate such as a population-weighted version or a sample-share-weighted version.</p><p>After that, the paper looks at stacked fixed-effects specifications that had already been used in applied work. This part is useful because many of you will have seen exactly those regressions in practice. The authors show that these fixed-effects versions run into the same basic problem: in general, the unweighted specification doesn&#8217;t identify the target aggregate either. Once the corrective weights are added, that problem is resolved and the fixed effects themselves become unnecessary for identification. One of the paper&#8217;s broader points is that a simpler weighted event-study specification is enough.</p><p>The final step is inference. Because stacked datasets can reuse the same groups across multiple sub-experiments, the paper discusses how standard errors should be handled and compares clustering approaches. The authors run a small Monte Carlo exercise and find that clustering at the group level and clustering at the group-by-sub-experiment level both perform reasonably well when the number of clusters is not too small. So the paper is not only about identification but also gives us some guidance on implementation, which is always welcome.</p><p><em><strong>Why is this important?</strong></em></p><p>This paper is important because stacked DiD was already being used in applied work yet the literature hadn&#8217;t clearly established what the standard stacked regression identified. That is a problem if researchers are treating the coefficient from a stacked regression as a causal average without being able to say what average it is (isn&#8217;t this always a problem?). One of the paper&#8217;s main contributions is to show that the usual unweighted stacked regression doesn&#8217;t, in general, identify the target aggregate, or any other convex combination of causal effects.</p><p>It&#8217;s also important because the paper deals directly with a common practical use of event studies. We often want to look at treatment effects over time and use the pre-period to check for possible violations of the DiD assumptions. The authors show why that can go wrong if the composition of the sample changes across event time. Their trimmed aggregate ATT is built to avoid that problem so that changes in the event-study line reflect treatment dynamics in the post-period or differential pre-trends in the pre-period, rather than different cohorts entering and leaving the average at different horizons.</p><p>There is also a practical gain for applied researchers. The paper&#8217;s weighted stacked approach is still regression-based, which makes it relatively easy to implement and explain. The authors are quite explicit on this point. They argue that stacked estimation is attractive for applied work because it&#8217;s regression-based, keeps attention on the underlying research design and doesn&#8217;t rely on extra modelling assumptions beyond the standard DiD setup. They also show that the weighted estimator gives coefficients that correspond to a well-defined average of group-time ATT parameters, which gives us a clearer rationale for using the method.</p><p>Finally, this paper is useful because it speaks directly to what people were already doing. The authors discuss earlier applied papers using stacked DiD and show that the stacked fixed-effects versions used there run into the same general identification problem when left unweighted. They show that a simpler weighted event-study specification is sufficient to identify the trimmed aggregate ATT. By contrast, the more complicated fixed-effects setup does not solve the core issue on its own. For readers who may want to use stacked DiD in their own work, that is a very practical takeaway.</p><p><em><strong>Who should care?</strong></em></p><p>Applied researchers may want to use stacked DiD in staggered adoption settings, where different units are treated at different times and the aim is to compare each treated cohort with clean controls that remain untreated over the relevant event window. This can come up in labour, health, public, education and other applied fields where policies are rolled out in stages across places or organisations.</p><p><em><strong>Do we have code?</strong></em></p><p>Yes. The authors provide a <a href="https://github.com/hollina/stacked-did-weights">GitHub repo</a> with example code for calculating and implementing the weights used in their stacked DiD approach.</p><p>In summary, this paper gives stacked DiD a much clearer foundation. Wing, Freedman and Hollingsworth show that the standard unweighted stacked regression doesn&#8217;t have a clear causal interpretation in staggered adoption settings even when the underlying sub-experiments satisfy the usual DiD assumptions. They then define a target parameter, the trimmed aggregate ATT, and show that a weighted stacked estimator can recover it. For applied researchers, the takeaway is fairly straightforward: stacking the data is only part of the job. You also need to think really carefully about which cohorts enter the sample, which units count as clean controls and how the regression weights are constructed. That makes this a very useful paper for anyone who may want to use stacked DiD in practice, or who wants a clearer link between the regression they run and the parameter they claim to estimate.</p><div><hr></div><h3>Event-Study Designs for Discrete Outcomes under Transition Independence</h3><p><em>(<a href="https://youngahn.com/">Young Ahn </a>is a 5th-year Ph.D. Candidate at UPenn, but he will soon start as a postdoc at Brown. All the best!)</em></p><h5>TL;DR: this paper asks how to do DiD or event-study analysis when the outcome is discrete, such as employment status, complaints, or patenting. The authors argue that standard PT can be a poor fit in these settings, so they propose transition independence and then add a latent-type Markov structure to handle unobserved heterogeneity and short panels. In their applications, this leads to estimates that differ markedly from conventional DiD.</h5><p><em><strong>What is this paper about?</strong></em></p><p>This paper is trying to solve a basic identification problem in DiD with discrete outcomes<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. Standard DiD relies on PT, but that logic can be a poor fit when outcomes are bounded and evolve through transitions across categories<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>. The authors&#8217; point is that this isn&#8217;t a minor technical issue. With discrete outcomes, differences in baseline distributions can generate mean reversion even in the absence of treatment, binary outcomes can imply impossible counterfactual probabilities and for multi-category outcomes the notion of a single trend becomes hard to pin down. The paper proposes transition independence as an alternative identification strategy<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>, and then adds a latent-type Markov structure to deal with unobserved heterogeneity and short panels<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>. The broader aim is to give researchers a way to do event-study or DiD analysis with discrete outcomes without leaning on a PTA that may be &#8220;incoherent&#8221; in this setting.</p><p><em><strong>What do the authors do?</strong></em></p><p>They set up a potential-outcomes framework for discrete outcomes and show how the ATT can be identified under transition independence. The basic idea is to construct the treated group&#8217;s counterfactual using control-group transition probabilities rather than mean outcome trends. They then relate this to conventional DiD, show that transition independence is equivalent to a version of conditional PT based on the full pre-treatment outcome history, and derive the bias of DiD under their framework. To make the approach workable in practice, they introduce a latent-type Markov structure, establish identification of latent-type-specific treatment effects and the overall ATT, and then develop an estimator. Finally, they apply the method to the Dodd-Frank Act, Norway&#8217;s patent reform, and the ADA. In all three cases their results differ substantially from conventional DiD<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>.</p><p><em><strong>Why is this important?</strong></em></p><p>This paper is important because discrete outcomes are everywhere in applied work, yet DiD is still often used as though the usual PT logic carries over without much trouble. The authors show that this can lead to badly misspecified counterfactuals and, in turn, misleading treatment effects. Their alternative gives researchers a way to build counterfactuals that respects the bounded, transition-based nature of discrete outcomes. The applications make the stakes clear: in one case DiD implies complaint rates below zero, in another it overstates a negative effect, and in the ADA application the transition-based framework also shows which channels are driving the employment effect. So the paper is useful both methodologically and empirically, because it presents a different identification strategy for a common class of outcomes and shows that the choice of strategy can change the substantive conclusion.</p><p><em><strong>Who should care?</strong></em></p><p>This paper should be useful for researchers working with outcomes like employment status, unemployment, labour-force participation, occupational choice, complaint filing, patenting, disability status, or other bounded categorical variables. It is directly relevant in settings where the outcome moves across states over time and researchers might still be tempted to use a standard DiD or event-study design on binarised outcomes. </p><p><em><strong>Do we have code?</strong></em></p><p>Yes. The authors provide a standalone R package with replication codes, available as <a href="https://github.com/bayesiahn/ak">bayesiahn/ak</a>.</p><p>In summary, Ahn and Kasahara argue that standard DiD can be a poor fit for discrete outcomes because the counterfactual is built from mean trends in settings where outcomes are bounded and evolve through transitions across states. Their solution is to replace PT with transition independence and then use a latent-type Markov structure to make the framework workable with unobserved heterogeneity and short panels. The paper gives applied researchers a different way to study treatment effects with discrete outcomes, and the empirical applications show that it can lead to very different conclusions from conventional DiD.</p><div><hr></div><h3>Inference for Matched Difference-in-Differences</h3><h5>TL;DR: this paper studies inference in matched DiD and shows that standard variance estimators can be too conservative once matching induces dependence within treated-control pairs. Kim and Park propose a projection-based variance estimator that removes the variation explained by the matching covariates and targets the correct design-based variance instead. In their simulations it performs much better than standard unit-level or pair-level procedures in informative matching designs - but watch out: when matching is uninformative (R&#178;_Z &#8776; 0), the projection step overcorrects and the estimator undershoots, so standard pair-cluster inference is the safer default in that setting.</h5><p><em><strong>What is this paper about?</strong></em></p><p>This paper is about inference in matched DiD<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>. Kim and Park start from a simple point: matching is often used in DiD to improve covariate balance, but once we match treated units to controls, the design itself creates dependence within matched pairs<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>. Standard variance estimators usually ignore that feature, so they target the wrong variance. In this setting, the consequence is systematic overestimation of standard errors<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>. The paper&#8217;s goal is to fix that. More specifically, the authors study the variance implied by the matched design and show why standard unit-level and cluster-level inference can be too conservative once matching is built into the estimator. They then propose a projection-based variance estimator that removes the component of variation explained by the matching covariates and is designed to recover the relevant design-based variance<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>. The broader aim is to show that, in matched DiD, getting the point estimate is only part of the job. You also need a standard error that reflects the dependence structure created by the matching step.</p><p><em><strong>What do the authors do?</strong></em></p><p>Kim and Park begin by setting up a matched DiD framework in which each treated unit is matched one-to-one, without replacement, to a control unit using pre-treatment or time-invariant covariates. They then write the matched DiD estimator in an asymptotic linear form and use that setup to show what variance target the matched design implies. The next step is to explain why standard inference goes wrong. The paper shows that matching creates positive covariance within matched pairs because treated units and their matched controls share a common component linked to the matching covariates. Once the estimator is differenced, that covariance reduces the true variance. Standard unit-level CRVE<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a> ignores this term, so it overstates uncertainty, and pair-level clustering only partly fixes the problem because it still doesn&#8217;t use the covariate structure that generated the dependence.</p><p>Their solution is a projection-based variance estimator. The basic idea is to take the within-pair sum of first-differenced residuals, project it onto the matching covariates and remove the variation explained by those covariates. This yields an adjusted variance estimator designed to recover the design-based variance implied by the matched sample. They also show that the estimator can be written as a scaled version of the pair-cluster variance estimator where the scaling factor depends on how much of the within-pair variation is explained by the matching covariates.</p><p>They then study finite-sample performance in Monte Carlo simulations. The results show that standard unit-level and pair-level variance estimators are too conservative in informative matching designs, while the projection-based estimator comes much closer to the Monte Carlo benchmark.</p><p><em><strong>Why is this important?</strong></em></p><p>This paper is important because matched DiD is widely used to improve comparability between treated and control units, yet the matching step also changes the dependence structure of the estimator. Kim and Park show that standard variance estimators can then become systematically too conservative because they miss the variance reduction created by matching-induced covariance within pairs. It is also important because the paper offers a very concrete correction. The projection-based estimator adjusts inference using the same matching covariates that created the dependence in the first place. The simulation results show that this performs much better than standard unit-level or pair-level procedures in informative matching designs.</p><p><em><strong>Who should care?</strong></em></p><p>This paper should be useful for researchers using matched DiD to improve covariate balance before estimation. It&#8217;s really relevant in settings where one-to-one matching is used and inference is reported with standard cluster-robust standard errors without much discussion of what variance is being estimated<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>. More broadly, it should interest applied researchers who use matching and then run DiD as though the inference problem were unchanged by the design step.</p><p><em><strong>Do we have code?</strong></em></p><p>Not yet. Will update if this changes.</p><p>In summary, Kim and Park show that, in matched DiD, the matching step changes the variance structure of the estimator in a way that standard inference usually ignores. Because matching induces positive covariance within matched pairs, standard unit-level and pair-level variance estimators can become too conservative. Their solution is a projection-based variance estimator that adjusts for the dependence created by the matching covariates. The paper&#8217;s main message is simple: once matching is part of the design, the standard errors need to reflect that design too.</p><div><hr></div><h3>Parallel Trends Forest: Data-Driven Control Sample Selection in Difference-in-Differences</h3><h5>TL;DR: this paper proposes parallel trends forest, a machine-learning method for choosing control units in DiD when treatment assignment has little randomness and there are many plausible controls. Instead of fixing the control sample by hand, it uses pre-treatment data to build a weighted control for each treated unit that moves as closely in parallel as possible.</h5><p><em><strong>What is this paper about?</strong></em></p><p>This paper is trying to solve a very practical problem in DiD: when treatment assignment has little randomness and there are many possible control units, how do we choose the control sample in a disciplined way? The authors&#8217; answer is parallel trends forest, a data-driven method that uses pre-treatment data and machine learning to construct, for each treated unit, an optimal weighted control sample<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a> that moves as closely in parallel as possible before treatment<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a>. The broader concern is familiar. In many DiD applications, we don&#8217;t have random treatment assignment, so we choose a control group that looks &#8220;reasonable&#8221;, inspect the pre-treatment series, and hope the PTA is credible. Huh and Kling are trying to improve on that workflow. Their method uses random forests<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> to search over a large set of covariates and control units rather than relying on researcher judgment alone or on a standard TWFE regression with a fixed control sample<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a>.</p><p>The paper is also aimed at a specific type of setting: relatively long panels with noisy, granular and sometimes non-normal data. The authors argue that existing methods such as SC, synthetic DiD and matrix completion can struggle there, either because they fit poorly, overfit in-sample<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a> or don&#8217;t adapt well to a large and chaotic donor pool<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a> (which is true). Parallel trends forest is meant to work better in exactly those cases. The main goal is to give us a more reliable way to build a control sample before running DiD, especially when the design is useful but far from clean.</p><p><em><strong>What do the authors do?</strong></em></p><p>They start by formalising the problem: for each treated unit, they want to find a weighted combination of control units that would have moved in parallel with it in the absence of treatment. That means the method constructs an optimal weighted control for each treated unit using pre-treatment data.</p><p>They then define a new measure of deviation from parallel trends<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a>. Instead of relying on the usual sum of squared errors, they use a placebo-style criterion based on how large the estimated treatment effect would be if treatment were falsely assigned at different dates in the pre-treatment period. This is meant to work better with noisy, granular and non-normal data.</p><p>The next step is the machine-learning part. They build trees using this deviation measure, split units only along the unit dimension and then combine many trees into a parallel trends forest. From that forest, they recover weights for each treated-control pair based on how often they end up in the same leaf. Those weights are then used to construct the counterfactual for each treated unit and, from there, the ATT.</p><p>After laying out the method, they compare it to SC, synthetic DiD and matrix completion using a placebo exercise and Monte Carlo simulations. In their setting, they find that the parallel trends forest tracks the treated sample more closely and performs better than the existing methods. They then apply the method to the rollout of post-trade transparency in the corporate bond market. They use it to study the effect of TRACE Phase 2 on bond turnover, compare the results to TWFE estimates using several different control samples and then also show how the method can be combined with the Rambachan-Roth framework to allow for bounded deviations from parallel trends<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a>. They also introduce an &#8220;honest&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a> version of the forest to reduce overfitting.</p><p><em><strong>Why is this important?</strong></em></p><p>A few reasons. This paper is important because control-sample choice is often one of the weakest parts of applied DiD. In many settings we don&#8217;t have much randomness in treatment assignment so the credibility of the design depends heavily on which controls are chosen. Huh and Kling turn that choice into an explicit data-driven step rather than leaving it to researcher judgment alone. It is also important because the paper is aimed at a setting where many existing methods struggle: long panels with many candidate controls, noisy outcomes and granular data. The empirical application makes the point clearly. In their TRACE setting, different &#8220;reasonable&#8221; control samples produce quite different TWFE estimates, sometimes even with different signs. Their method gives a smaller estimate of the effect of transparency on turnover and once they allow for bounded deviations from parallel trends, the effect is not statistically significant. </p><p><em><strong>Who should care?</strong></em></p><p>Anyone working in settings where there are many plausible controls, little randomness in treatment assignment and enough pre-treatment data to compare how units move over time. That makes it relevant for work on phased policy rollouts, regulatory changes, market-design reforms and other settings where the main practical difficulty is choosing a credible control sample rather than writing down the DiD regression itself.</p><p><em><strong>Do we have code?</strong></em></p><p>No yet. Will update if anything changes. The method is described well enough to attempt a reconstruction, but not enough to call it fully reproducible out of the box.</p><p>In summary, Huh and Kling take one of the messiest parts of applied DiD - choosing the control sample - and turn it into an explicit data-driven step. Their method is built for settings with long panels, large donor pools and noisy data, where standard tools can fit badly or become very sensitive to which controls the researcher picked. The broader contribution is practical: rather than asking readers to accept a &#8220;reasonable&#8221; control group on faith, the paper offers a structured way to search for one using the pre-treatment data itself.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>In stacked DiD, the we construct a separate sub-experiment for each treatment cohort using that cohort as the treated group and a set of clean controls, then pools those sub-experiments into one stacked dataset for estimation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>The appeal of this approach is that it avoids the late-versus-early treated comparisons that created problems for conventional TWFE in staggered adoption designs under treatment effect heterogeneity. In the paper the authors note that versions of stacked DiD were already being used in applied work before this formal treatment.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Plural because there are 3: the basic unweighted stacked estimator, the weighted stacked estimator they propose, and the stacked fixed-effects specification used in earlier applied work, which they also evaluate.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The trimmed aggregate ATT is a weighted average of cohort-specific treatment effects at each event time in a staggered adoption design, computed after trimming the sample so that the same cohorts contribute across the chosen event window. This keeps the event-study graph from being driven by cohorts entering and leaving the average at different horizons.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>A pooled stacked regression is just the regression you run after you have built all the sub-experiments and stacked them into one dataset, so instead of estimating a separate DiD for each treatment cohort, you create a sub-experiment for each valid adoption cohort, then append those sub-experiments into one stacked file, and finally run one regression on the pooled stacked dataset.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>By &#8220;discrete outcomes&#8221; the authors mean outcomes that take a limited set of categories, such as employed, unemployed or out of the labour force, or which occupation a worker is in. These are settings where we often still use DiD, usually by turning categories into binary indicators.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>The paper argues that PT can fail in discrete settings for three reasons: mean reversion can create different trends when groups start from different baseline distributions, binary outcomes can imply impossible counterfactual probabilities 0 &lt; P or P &gt; 1, and for multi-category outcomes the notion of a single &#8220;trend&#8221; isn&#8217;t well defined.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>Transition independence means that, absent treatment, treated and control units with the same pre-treatment outcome path would have the same transition dynamics. So the counterfactual is built from transition probabilities rather than from average outcome trends.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>To allow for unobserved heterogeneity, the authors assume that units belong to latent types with different transition dynamics, and that outcomes follow type-specific Markov processes. This is what lets them recover both latent-type-specific and aggregate treatment effects from short panels.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>Funnily enough, it&#8217;s all due to the identification strategy :) The two methods are building different counterfactuals, and in these applications the usual DiD counterfactual is a poor fit for the DGP. The paper&#8217;s general argument is that with discrete outcomes, conventional DiD can go wrong for 3 main reasons: baseline differences create mean reversion, binary outcomes can generate impossible probabilities, and multi-category outcomes are driven by transitions across states rather than a single smooth trend. In the 3 applications, each one highlights a different version of that problem. In the Dodd-Frank case, the DiD counterfactual complaint rate falls below zero, so the estimates differ because DiD is extrapolating linearly in a setting where probabilities are bounded. Their transition-based method avoids that by construction since it works with transition probabilities rather than mean trends. In the Norway patent reform case, the treated group starts with a much higher patenting rate than the control group, so there is more room for patenting to fall &#8220;mechanically&#8221;. DiD reads that decline as treatment, while the authors argue that much of it is mean reversion. Once they model the transition dynamics directly, the estimated effect moves much closer to zero. In the ADA case, the issue is that level comparisons hide what is happening in the underlying transitions. The paper says DiD misses statistically significant employment effects because pre-treatment level differences mask the transition dynamics, whereas their method can trace flows across labour-force states and shows that the negative employment effect is operating through specific exit channels (mainly transitions from employment into out-of-labour-force status). The estimates are so different because the paper replaces the identification strategy. DiD uses PT in outcome levels. Their method uses transition independence and constructs counterfactuals from transition probabilities conditional on pre-treatment histories. When the DiD counterfactual is badly misspecified, the resulting treatment effects can move a lot, and sometimes even flip sign.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Matched DiD combines matching with a DiD setup: treated units are first matched to similar control units using observed covariates and the DiD estimator is then computed on the matched sample. The attraction is better covariate balance and - in principle! - a more credible comparison group.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>The key point is that matched treated and control units are deliberately chosen to be similar on the matching covariates, which then creates positive correlation within matched pairs since both units share a common component related to those covariates.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Here &#8220;conservative&#8221; means the standard errors are too large. The paper&#8217;s argument is that standard variance estimators miss the variance reduction coming from the positive covariance within matched pairs, so they systematically overstate uncertainty.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p><a href="https://onlinelibrary.wiley.com/doi/abs/10.3982/ecta6474?casa_token=8mZ0UHMNFJQAAAAA%3AYxdELHf72yEFPIFGKh9i3tDLDsvaGHOSKrpioUwgDzBJpxv1TLFlrSvNTtOwzLxmESXq6l6qmljOT0BG">Abadie and Imbens (2008)</a> identified the same mechanism in paired experiments and proposed a resampling-based correction. Kim and Park&#8217;s contribution is a closed-form fix that targets the same problem specifically in matched DiD designs.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>The proposed estimator works by projecting the within-pair error sum onto the matching covariates and removing the part explained by them. Intuitively, this filters out the variation induced by the matching design and leaves the idiosyncratic component relevant for variance estimation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Standard unit-level CRVE means the usual cluster-robust variance estimator applied at the unit level, treating each unit as independent from the others and allowing correlation only within the unit over time. In this paper the problem is that this estimator ignores the extra dependence created by the matching step, so it misses the covariance within matched pairs and ends up overstating uncertainty.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>The framing here draws on prof Wooldridge (2023)&#8217;s &#8220;<a href="https://id.elsevier.com/as/authorization.oauth2?platSite=SD%2Fscience&amp;additionalPlatSites=GH%2Fgeneralhospital%2CLS%2FLS%2CMDY%2Fmendeley%2CSC%2Fscopus%2CRX%2Freaxys&amp;scope=openid%20email%20profile%20els_auth_info%20els_idp_info%20els_idp_analytics_attrs%20urn%3Acom%3Aelsevier%3Aidp%3Apolicy%3Aproduct%3Ainst_assoc&amp;response_type=code&amp;redirect_uri=https%3A%2F%2Fwww.sciencedirect.com%2Fuser%2Fidentity%2Flanding&amp;authType=SINGLE_SIGN_IN&amp;prompt=none&amp;client_id=SDFE-v4&amp;state=retryCounter%3D0%26csrfToken%3Ddabbbca6-c4bd-4cad-a431-3be2152925b1%26idpPolicy%3Durn%253Acom%253Aelsevier%253Aidp%253Apolicy%253Aproduct%253Ainst_assoc%26returnUrl%3D%252Fscience%252Farticle%252Fabs%252Fpii%252FS0304407623002336%26prompt%3Dnone%26cid%3Datn-1800ecb4-516c-4a72-9a3e-ae52a1bd5076">What is a standard error?</a>&#8221;, which argues that the validity of any standard error depends on correctly specifying the variance target relative to the underlying source of randomness. Worth reading alongside this paper.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Instead of picking one control group for the whole treatment sample, the method builds a weighted control for each treated unit. The weights are based on which control units end up in the same leaves as the treated unit across many trees in the forest.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>Parallel trends forest is a machine-learning method that builds an optimal control sample for each treated unit using only pre-treatment data. It assigns weights to control units so that the weighted control outcome moves as closely in parallel as possible with the treated unit before treatment.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>Random forests let the method work with many covariates and many potential control units. They also automate covariate selection so the algorithm can use the variables that really help predict which units move in parallel instead of relying only on what the researcher guessed would matter.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>The paper is motivated by settings with little or no randomness in treatment assignment. In those cases, we often choose controls by hand, compare pre-treatment trends visually, and then run DiD. Huh and Kling are trying to replace that partly <em>ad hoc</em> step with a data-driven selection rule. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>Synthetic control actually does the opposite of fitting poorly, it overfits in-sample very closely but then performs badly out-of-sample. Synthetic DiD and matrix completion fail for a different reason: they assign near-equal weights across the whole control pool rather than upweighting the units that actually track the treated series.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>In the paper the authors include earlier-treated units - bonds that became transparent before the sample window - in the donor pool. They argue these would organically receive near-zero weight if their trends have already diverged from the newly-treated bonds by the start of the sample period, which serves as a data-driven answer to the Goodman-Bacon contamination concern.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>The paper doesn&#8217;t use the usual sum of squared pre-treatment errors as its objective. Instead, it defines a placebo-style measure based on how large the estimated treatment effect would be if treatment were assigned at different dates in the pre-treatment period.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>The paper doesn&#8217;t derive analytical standard errors for the parallel trends forest estimator, and this is explicitly flagged as outside scope. The convergence analysis (Figure 8 in the paper) shows the weight estimates stabilise quickly across trees, but that addresses estimation variance, not inferential uncertainty more broadly. Do keep this in mind before reaching for the method.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p>The &#8220;honest&#8221; version of the forest follows the logic in Wager and Athey: one sample is used to grow the tree and another is used to estimate the weights. The aim is to reduce overfitting to pre-treatment noise.</p></div></div>]]></content:encoded></item><item><title><![CDATA[The outcomes you measure and the assumptions holding them up]]></title><description><![CDATA[New papers on compositional shifts, coarsened outcomes, pre-trend testing and event-study identification]]></description><link>https://www.diddigest.xyz/p/the-outcomes-you-measure-and-the</link><guid isPermaLink="false">https://www.diddigest.xyz/p/the-outcomes-you-measure-and-the</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 12 Jan 2026 16:16:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kVgJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kVgJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kVgJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kVgJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3291361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/177726641?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kVgJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!kVgJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F500aa814-eb20-4c89-b7ca-d755a20dfe70_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! </p><p>Thank you (all 1553 of you) for being here for one year now :)</p><p>I wanted to close 2025 with a &#8220;banger&#8221; - meaning lots of banger papers - but I got distracted by the only 2 weeks I have of free-time in the year. Apologies. Also I am set to graduate in May while also teaching a course this term, so probably the rate of posts will come down just a bit, but after May it should resume to normal.</p><p>Here&#8217;s to many more posts in 2026 :)</p><p>Also before we get to the papers, I&#8217;d like to thank my friend <a href="https://samenright.substack.com/">Sam Enright</a> who picked 3 posts from this newsletter for <a href="https://www.thefitzwilliam.com/p/join-the-fitzwilliam-reading-group">The Fitzwilliam Reading Group</a> monthly meetup a couple of weeks ago. Special thanks to those who attended and who are reading this now :)</p><p>As always, please let me know if you would like for me to cover a specific paper. I will do my best to go through it.</p><p>Today&#8217;s papers are:</p><ul><li><p><a href="https://arxiv.org/abs/2304.13925">Difference-in-Differences with Compositional Changes</a>, by<strong> </strong>Pedro H. C. Sant&#8217;Anna and Qi Xu</p></li><li><p><a href="https://arxiv.org/abs/2512.08759">Difference-in-Differences with Interval Data</a>, by<strong> </strong>Daisuke Kurisu, Yuta Okamoto and Taisuke Otsu</p></li><li><p><a href="https://arxiv.org/abs/2310.15796">Testing for equivalence of pre-trends in Difference-in-Differences estimation</a>, by<strong> </strong>Holger Dette and Martin Schumann</p></li><li><p><a href="https://www.nber.org/system/files/working_papers/w34550/w34550.pdf">Harvesting Difference-in-Differences and Event-Study evidence</a>, by Alberto Abadie, Joshua Angrist, Brigham Frandsen and J&#246;rn-Steffen Pischke</p></li></ul><p>Honorable mention:</p><ul><li><p><a href="https://arxiv.org/abs/2511.05870">Synthetic Parallel Trends</a>, by Yiqi Liu</p></li></ul><div><hr></div><h3>Difference-in-Differences with Compositional Changes</h3><h5>TL;DR: standard DiD with repeated cross-sections relies on a strong stationarity assumption that often fails in practice. Sant&#8217;Anna and Xu show that the ATT remains identifiable when composition shifts, derive the correct efficiency bound and provide robust estimators and a diagnostic test to assess whether stationarity is empirically relevant. The paper formalises the bias&#8211;variance trade-off and gives applied researchers a clear framework for working with moving samples.</h5><p><em>What is this paper about?</em></p><p>This paper is about a quite strong and often unrealistic (in practice) assumption in DiD with repeated cross-sections<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>: no compositional changes over time, which means assuming that the treated and control groups come from the same underlying population before and after treatment. In many real applications this is straight implausible: people enter and leave samples, early adopters differ from late adopters, firms exit, migrants move, and survey frames rotate. This is the default setting for many labour, health, education and dev applications that rely on repeated survey data.</p><p>A classic example cited by the authors is the study of Napster&#8217;s effect on music sales. Over time, the composition of internet users changed substantially: early adopters were typically younger and wealthier while later adopters were more demographically diverse. If a researcher ignores these shifts, the negative effect of Napster on sales might be overestimated because the &#8220;post-Napster&#8221; group naturally includes more households with lower reservation prices for music, regardless of the technology&#8217;s impact.</p><p>Most DiD methods simply impose stationarity<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> and move on. Sant&#8217;Anna and Xu do the opposite. They study DiD when composition is allowed to change and show that the ATT remains identifiable under conditional PT<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. They develop a general framework that makes clear what is identified, at what cost in efficiency, and how standard approaches break down when compositional changes are ignored.</p><p><em>What do the authors do?</em></p><p>They start by dropping the usual stationarity assumption and allowing the joint distribution of <strong>(D,X)</strong> to change over time. In this setup, they show that the ATT is still identified under conditional PT, no anticipation, and overlap. However, the &#8220;math&#8221; changes: because you can no longer assume the population is stable, you have to track four distinct groups (treated/control in both periods) using a generalized propensity score. </p><p>They derive the efficient influence function (EIF) and efficiency bound<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> (essentially the &#8220;gold standard&#8221; for precision) for the ATT when composition is allowed to shift, and build nonparametric doubly robust<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> estimators that remain valid under these compositional shifts. Unlike classical DR, which relies on getting one of two models &#8220;correct&#8221;, their estimators enjoy a rate doubly robust property. This means your estimate is reliable even if one part of your model (like the outcome regression) is very complex, provided the other part (the propensity score) is simple enough to compensate. They prove that the widely used <a href="https://www.sciencedirect.com/science/article/pii/S0304407620301901">Sant&#8217;Anna&#8211;Zhao (2020)</a> estimator can be biased because it wasn&#8217;t designed to handle these moving populations.</p><p>They then characterize the bias&#8211;variance trade-off. If you wrongly impose stationarity, you can get biased ATT estimates. If you don&#8217;t impose it when it actually holds, you lose efficiency.  You might be attributing a change in the data to the treatment when it&#8217;s actually just a result of the &#8220;mix&#8221; of people in your sample changing. However, robustness isn&#8217;t free. If stationarity really holds but you choose to ignore it by using the robust estimator, you lose efficiency (precision). Proposition 2.2 identifies three factors that determine how much precision you lose: the period ratio (the balance between the size of your pre-treatment and post-treatment sample; the group ratio (the proportion of comparison units relative to treated units; and treatment heterogeneity (aka how much the treatment effect varies across different types of people). In a world where everyone responds to a policy in exactly the same way (homogeneous effects), the robust estimator is just as precise as the standard one. In the real world, where effects usually vary, the test they propose becomes your essential &#8220;GPS&#8221; for navigating this trade-off. </p><p>To operationalise this, they propose a <a href="https://www.jstor.org/stable/1913827">Hausman-type tes</a>t<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> that compares the robust estimator (valid under compositional changes) with the stationarity-based estimator. The idea is simple: if composition doesn&#8217;t matter for the ATT, the two should coincide; if they differ, stationarity is empirically relevant. This concentrates power exactly in the directions that matter for the target parameter.</p><p>Then they show how to implement all of this in practice. They use local polynomial methods for outcome regressions and local multinomial logit for generalized propensity scores, allow for both continuous and discrete covariates, discuss bandwidth selection, extend the framework to cross-fitting and mixed panel/cross-section designs, and illustrate the methods in simulations and an application.</p><p><em>Why is this important?</em></p><p>Many applied DiD papers rely on repeated cross-sections and proceed under an implicit stationarity assumption. This paper shows that this choice has real consequences for identification and precision. When composition changes, standard DiD estimators can be biased even if conditional PT holds. This affects the credibility of causal claims in a wide range of applications, as researchers might mistake a demographic shift for a policy effect.</p><p>The contribution here is diagnostic as much as methodological. Sant&#8217;Anna and Xu show exactly where the bias comes from, how large it can be and which assumptions remove it. This replaces hand-waving about &#8220;changing samples&#8221; with a precise link between assumptions, estimands and estimators.</p><p>The bias&#8211;variance trade-off is very important for practice. Imposing stationarity &#8220;buys&#8221; precision at the cost of bias when the assumption fails. Dropping it &#8220;buys&#8221; robustness at the cost of efficiency when the assumption holds. The paper formalises this trade-off and provides a test to decide which side you&#8217;re on. This then gives applied researchers a way to justify their design choices rather than rely on convention.</p><p>More broadly, the paper aligns DiD theory with how data are generated. Repeated cross-sections aren&#8217;t panels. Populations evolve. Treating moving samples as if they were fixed is convenient, but often inaccurate. By building the theory for the setting we actually face, this paper closes an important gap between methodological assumptions and empirical practice.</p><p><em>Who should care?</em></p><p>Applied researchers using DiD with survey or administrative data should pay attention. This includes work in labour, health, education, development and public policy where repeated cross-sections are the norm. Those using repeated cross-sectional data from major national sources like the Current Population Survey (CPS) or the Consumer Expenditure Survey (CEX). If your identification strategy relies on treating moving samples as stable groups, this paper speaks directly to your design choices.</p><p>Policy evaluators working on national reforms, programme expansions or regulatory changes are another core audience. In many of these settings, the underlying population evolves at the same time as the policy environment. The framework here clarifies when standard DiD logic carries through and when different estimators are needed, which is directly relevant for government analysts and policy units producing causal evidence under real-world constraints.</p><p>Methodologically, this paper will be useful for researchers building or using modern DiD estimators. It tightens the link between assumptions, estimands and efficiency, and shows how robustness to compositional change interacts with precision. If you work with DR methods, semiparametric efficiency or ML in DiD, the technical contributions here are relevant.</p><p>Referees and editors should also care. Many papers implicitly rely on stationarity without stating it. This paper provides language and tools to assess that assumption rather than take it on faith which raises the bar for what we can reasonably claim in applied work.</p><p><em>Do we have code?</em></p><p>Not in the paper itself. While no dedicated software package is listed, the authors provide the exact mathematical estimators in Section 3. They recommend using local polynomial and multinomial logit tools, which are available in general-purpose R packages like <code>np</code> or <code>npsf</code>, and then combining them according to the ATT formula derived in the paper. Check <a href="https://psantanna.com/files/DiD_CC_Appendix.pdf">Supplemental Appendix</a> for implementation, where they give a practical recipe (local polynomial outcome regressions + local multinomial logit generalised propensity scores + cross-validation/bandwidth choices), but you&#8217;d code it yourself (or adapt existing nonparametric/kernel + multinomial logit tooling).</p><p>In summary, this great paper revisits DiD with repeated cross-sections. Sant&#8217;Anna and Xu show that the ATT remains identifiable under conditional PT, but the estimand, efficiency bound and appropriate estimator change once stationarity is dropped. They derive the efficient influence function, build rate DR estimators that remain valid under shifting composition and make the bias&#8211;variance trade-off explicit. The Hausman-type test gives applied researchers a way to assess whether stationarity is empirically relevant rather than assume it. Overall, the paper replaces a convenient convention with a clear design framework and provides tools that align DiD practice with how samples evolve in reality.</p><div><hr></div><h3>Difference-in-Differences with Interval Data</h3><p>(<a href="https://sites.google.com/view/yutaokamoto">Yuta</a> is a 3rd-year PhD student at the Graduate School of Economics, Kyoto University)</p><h5>TL;DR: traditional DiD assumes scalar (exact) outcomes, but real-world data is often &#8220;coarsened&#8221; into intervals through rounding or binned reporting. The authors demonstrate that instead of naively extending &#8220;parallel trends&#8221;, a &#8220;parallel shifts&#8221; assumption is necessary to ensure the resulting bounds are both mathematically valid and intuitively sensible<strong>.</strong> Through their reanalysis of the Card and Krueger (1994) study, they show that this new method provides a more disciplined and informative way to handle rounding and heaping in employment data.</h5><p><em>What is this paper about?</em></p><p>This paper is about a problem almost nobody talks about in DiD: what happens when your outcome isn&#8217;t a number, but an interval. In a lot of survey and administrative data, outcomes are reported in brackets, bins or rounded chunks. Income, e.g., is a range. Hours worked are heaped at multiples of five. Headcounts are approximate. Even when datasets look scalar, they often aren&#8217;t, because rounding, heaping and reporting conventions mean the &#8220;true&#8221; value sits somewhere inside an interval. In this paper, Kurisu, Okamoto, and Otsu tell us to reconsider our choices. They ask: how should we do difference-in-differences when outcomes are interval-valued rather than point-valued? And more importantly: what does &#8220;parallel trends&#8221; even mean in that world?</p><p>Their first result is quite uncomfortable: if you naively apply the PT assumption to interval data, you can get bounds that are either so wide they&#8217;re useless, or that move in directions that make no sense. You can satisfy &#8220;parallel trends&#8221; and still get counterintuitive or uninformative results.</p><p>So the paper&#8217;s core contribution is conceptual. It shows that extending PT to interval outcomes is not without a cost, and that you need a different way of thinking about how treated and control groups should evolve over time. The authors propose an alternative assumption - parallel shifts - that&#8217;s designed to respect how intervals move and scale, and to avoid the pathologies that show up under na&#239;ve extensions.</p><p><em>What do the authors do?</em></p><p>They set up DiD properly for interval data. Instead of observing a single outcome, each unit is observed as a lower and an upper bound. The target is still the ATT, but now both the treated mean and the counterfactual mean are only partially identified.</p><p>They then study three ways of bringing PT into this setting. First, they apply standard PT to the unobserved true outcome. This is the most direct extension of textbook DiD. It performs badly. The bounds are extremely wide and often dominated by worst-case combinations of lower and upper bounds across groups.</p><p>Second, they apply PT directly to the interval bounds, that is, they track how the control group&#8217;s lower and upper bounds change over time and project those changes onto the treated group. This looks intuitive, but it can behave in strange ways. You can end up with treated bounds moving in the opposite direction to control bounds, even though &#8220;parallel trends&#8221; is imposed.</p><p>They then propose a different assumption: parallel shifts. Instead of matching trends over time, it matches shifts between groups. The idea is that whatever mapping takes the control group&#8217;s interval to the treated group&#8217;s interval before treatment is assumed to also apply after treatment. Under this assumption, the bounds move in the same direction across groups, interval widths behave sensibly, and the identified sets remain well-behaved. They derive closed-form bounds for the ATT so the method is easy to implement.</p><p>Finally, they apply the framework to the C&amp;K minimum wage study, treating employment as interval-valued because of rounding and heaping, and show how the different assumptions lead to very different conclusions.</p><p><em>Why is this important?</em></p><p>Because the way outcomes are recorded affects what DiD can identify.</p><p>In many survey and administrative datasets, outcomes are reported in brackets, rounded, or subject to heaping. Income is often binned. Hours worked cluster at focal values. Headcounts are approximate. In these settings, the outcome is a range rather than a point.</p><p>The standard DiD framework treats outcomes as exact and applies PT to their means. This paper shows that once outcomes are interval-valued, that logic does not carry over in a clear way. Some natural ways of extending PT lead to bounds that are very wide or that move in ways that are difficult to justify, which has direct consequences for interpretation. With the same data, different extensions of PT can produce very different identified sets. The identifying assumptions are doing more work than is usually acknowledged.</p><p>The contribution of the paper is to make that structure explicit. It shows that transporting information from control to treated is not neutral when outcomes are coarsened, and that some ways of doing it behave better than others. The parallel shifts assumption is proposed as a disciplined way of mapping intervals across groups and over time.</p><p>More broadly, the paper sits in the partial identification space. Once outcomes are coarsened, point identification is no longer guaranteed. Bounds are often the appropriate object. That is less convenient, but it reflects the information in the data.</p><p><em>Who should care?</em></p><p>Researchers using DiD with survey or administrative data where outcomes are coarsened. This includes work on income, wages, hours worked, employment, firm size, time use, and test scores when these are reported in brackets, rounded, or subject to heaping. It is common in labour, health, education, and public policy applications.</p><p>It is also relevant for people working with repeated cross-sections. In these settings, binning and rounding are standard, and the choice to treat midpoints as exact values is widespread. This paper shows that those choices have identifying consequences.</p><p>More broadly, anyone interested in identification in DiD, rather than only estimation, will find this useful. The paper makes clear that different ways of extending PT are not equivalent once outcomes are interval-valued. It is also relevant for readers working on partial identification and bounded inference. The framework fits well in settings where point identification is not credible.</p><p><em>Do we have code?</em></p><p>The paper derives closed-form bounds, so implementation is should be more or less straightforward. Everything is based on sample means of observed lower and upper bounds, with no optimisation or simulation required. Reproducing the results should be easy in R, Stata, or Python. The authors don&#8217;t provide replication code with the paper. The C&amp;K reanalysis uses standard data and simple transformations, so I think we could reconstruct it. If you are planning to use this in practice, you will need to code it yourself.</p><p>In summary, this paper asks what DiD is really doing when outcomes are interval-valued rather than exact. It shows that standard ways of extending PT can produce wide or poorly behaved bounds. The authors propose an alternative assumption, parallel shifts, that maps intervals across groups in a more disciplined way and leads to well-behaved identified sets. The contribution is conceptual. It makes clear that once outcomes are coarsened, identification depends heavily on how structure is imposed. Different choices lead to different answers. If your outcomes are rounded, binned, or heaped, this paper is directly relevant.</p><div><hr></div><h3>Testing for equivalence of pre-trends in Difference-in-Differences estimation</h3><h5>TL;DR: we test the wrong thing when we check pre-trends in DiD. The usual null is &#8220;no difference&#8221;. Failing to reject is then treated as evidence for PT. The authors say that logic is weak: it often means the data say nothing. In large samples it can reject for tiny, irrelevant differences. This paper proposes equivalence tests instead. You test whether pre-trend differences are small enough to ignore. The null is that they are too large. Rejection gives actual support for the identifying assumption.</h5><p><em>What is this paper about?</em></p><p>This paper is about how we test the PTA in DiD. In applied work, the standard approach is to test whether pre-treatment differences between treated and control groups are statistically zero. If you fail to reject, you proceed. The authors argue this logic is backwards. Failing to reject zero doesn&#8217;t mean trends are similar. It often means the test has low power. In large samples, the same tests can reject for tiny, irrelevant differences. Either way, the usual pre-tests do not answer the question practitioners actually care about.</p><p>The paper proposes replacing &#8220;no difference&#8221; tests with equivalence tests. Instead of asking whether pre-trends are exactly the same, it asks whether any differences are small enough to be considered negligible. The null becomes &#8220;differences are non-negligible&#8221;. Rejection then gives positive evidence in favour of PT.</p><p>The core object here is the size of pre-treatment trend differences, and whether the data support treating them as effectively zero for identification. The paper is about formalising that logic and giving tools to implement it.</p><p><em>What do the authors do?</em></p><p>They formalise pre-trend checking as an equivalence testing problem. Instead of testing whether each pre-treatment coefficient is zero, they define null hypotheses where pre-trend differences are at least as large as a user-chosen threshold. The alternative is that all differences are smaller than that threshold. Rejection then gives evidence that deviations from PT are negligible.</p><p>They propose three ways to summarise pre-trend differences: the maximum deviation across periods, the average deviation, and the root mean square deviation. Each corresponds to a different notion of what &#8220;small enough&#8221; means. The choice is left to the researcher and should be justified in context. They develop test statistics for each case, show their asymptotic properties and discuss implementation. In practice this means estimating the usual pre-treatment coefficients, then testing whether their joint behaviour stays within the chosen bound. They also show how to recover the smallest threshold for which equivalence would be accepted, if the researcher does not want to commit to a bound ex ante.</p><p>They extend the framework to staggered adoption and heterogeneous treatment effects using regression-based DiD setups. The idea is the same: define placebo effects in pre-treatment periods and test whether those are jointly small.</p><p>They back this up with simulations and an empirical illustration. The simulations compare equivalence tests to standard pre-tests under different violations of PT. The application re-examines a well-known DiD study and shows that, under their approach, the data give weak support for treating pre-trends as negligible.</p><p><em>Why is this important?</em></p><p>This is important because the usual pre-trend checks do not answer the identification question we are pretending they answer. In DiD, PT is an assumption, we never observe the counterfactual. Pre-trend tests are meant to give reassurance that the assumption is plausible, but the standard null is &#8220;no difference&#8221;. Failing to reject that null is often read as evidence of similarity. That is a logical error. It usually means the data are uninformative. This then creates two problems. In small samples, you can have meaningful differences in trends that go undetected. In large samples, you can reject for tiny differences that are irrelevant for the treatment effect. In both cases, the test result is hard to interpret for identification.</p><p>Equivalence testing flips the burden of proof. The null is that differences are large enough to worry about. To proceed, the data must show that differences are smaller than a threshold you consider acceptable. That aligns the test with the actual identifying assumption. You are no longer asking whether trends are identical, but rather whether deviations are small enough that treating them as zero is defensible.</p><p>This also forces structure. You have to say what &#8220;small enough&#8221; means in your application. That makes the identifying assumptions explicit rather than implicit. If you can&#8217;t justify any threshold, you can report the smallest one the data would support. That gives a direct sense of how much non-parallelism is compatible with your design. From an applied perspective, this changes interpretation. Instead of &#8220;we don&#8217;t see pre-trends&#8221;, you get &#8220;the data support ruling out pre-trend differences larger than X&#8221;. That is a statement about identification strength, not about p-values.</p><p><em>Who should care?</em></p><p>Everyone. We all do pre-trend checks. They sit at the start of almost every DiD design. They are usually the first thing people look at and the first thing readers use to judge credibility. This paper is about that step. The first one in the default workflow.</p><p>If you run event studies, plot leads, or report placebo regressions to justify PT, this applies to you. The argument is that the usual logic behind those checks is weak and often misleading. Equivalence testing is a way to make that step mean what we pretend it means.</p><p>This includes work with small samples, short panels, survey data, administrative data, and everything in between. It includes clean policy designs and messy real-world ones. If PT is doing identifying work in your paper, this framework is relevant.</p><p><em>Do we have code?</em></p><p>In a narrow sense, yes. The simulations are implemented in R and the empirical illustration is fully reproducible. The paper describes the procedures in enough detail that replication is straightforward, and the authors indicate that code is available. There is no standalone package or user-friendly wrapper. This is not a plug-and-play tool. That is consistent with the paper&#8217;s aim. The contribution is conceptual and inferential. If you&#8217;re comfortable writing your own pre-trend regressions and working with bootstrap routines, you can implement these tests without much friction. If you are not, there is no ready-made function you can drop into your workflow so maybe wait a little?</p><p>In summary, this paper is about aligning pre-trend checks with what PT means for identification. The standard workflow tests whether pre-treatment differences are zero, but that&#8217;s not the question we care about. We care whether differences are small enough to ignore. The authors show how to formalise that directly, using equivalence tests that put the burden of proof where it belongs. Their contribution is a way of being honest about assumptions: you either justify a threshold for negligible differences or you report how large differences could be before the design breaks. In both cases, interpretation improves. If you take PT seriously as an identifying restriction, this is a cleaner way to assess it.</p><div><hr></div><h3>Harvesting Difference-in-Differences and Event-Study evidence</h3><h5>TL;DR: this paper dissects what modern DiD and event-study workflows are really identifying, using a single policy setting to show how normalisation, heterogeneity and design choices shape the estimand. It shows that dynamic paths, exposure coefficients and pre-trend checks carry more identifying content than many of us realise, and that regression output is often easier to compute than to interpret. The message is that standard implementations impose structure that is rarely acknowledged. If you use event studies or exposure designs, this paper sharpens the link between your code and your causal claims.</h5><p><em>What is this paper about?</em></p><p>This paper is a guided walk through modern DiD and event-study practice, it&#8217;s a synthesis and critique of what people are actually doing with staggered adoption, dynamic effects, exposure designs and pre-trend checks.</p><p>The authors use the unilateral divorce reforms in US states as a running example. They use it to show how static DiD, event studies and newer heterogeneity-robust approaches behave in practice. Their focus is on interpretation rather than technique, what is being identified under which assumptions, and how easy it is to get something that looks clean but is hard to interpret.</p><p>Three themes run through the paper. First, normalisation isn&#8217;t harmless in event studies because when and how you anchor the coefficients can change the shape of the estimated dynamics. Second, heterogeneous and time-varying effects complicate both static DiD and event-study estimates. The paper is explicit about where regression averages do and do not correspond to meaningful causal objects. Third, common workflow steps like pre-trend testing and log specifications carry stronger identifying content than many of us realise.</p><p>The paper&#8217;s goal is rather practical. It&#8217;s a great attempt at trying to discipline how we read DiD and event-study output, what object is this coefficient estimating, which assumptions are doing the work, and where the usual visual and statistical checks can mislead.</p><p><em>What do the authors do?</em></p><p>They take a single policy setting and run it through the full modern DiD toolkit: static two-way fixed effects, event studies with leads and lags, alternative normalisations, heterogeneity-robust approaches, exposure designs, pre-trend testing, and logs versus levels. The unilateral divorce reforms are the running example, used to illustrate how identification and interpretation change across specifications.</p><p>They start from the canonical TWFE model and make the identifying assumptions explicit: additive state and time effects +PT in untreated potential outcomes + no anticipation. They then introduce event-time indicators to allow for dynamic effects and show how this immediately creates normalisation problems. With staggered adoption and no never-treated units, at least two coefficients must be set to zero. Which ones you choose changes the shape of the estimated path.</p><p>They then move to treatment effect heterogeneity. First in simple two-period settings, where TWFE still recovers an average effect under clear conditions, then in staggered designs with dynamic effects, where already-treated units act as controls and regression averages can become hard to interpret. Small numerical examples are used to show when static DiD collapses and when event studies produce misleading dynamics.</p><p>For exposure designs, they write down a model where treatment is continuous rather than binary and allow effects to vary with exposure intensity. They show that the usual regression coefficient identifies an average marginal effect, not an average of underlying treatment effects. This distinction is carried through using the potato example.</p><p>They also implement the BJS imputation estimator on the divorce data and compare it to standard event-study regressions. This is used to assess how much cross-sectional heterogeneity actually distorts regression-based dynamics in a real application.</p><p>At the end they examine pre-trend testing and inference, discuss why lead coefficients are a pre-test, why this can create problems for inference and why it is still informative in practice. They also simulate clustered standard errors in staggered event studies to show how leverage and limited treated observations at long horizons can understate uncertainty.</p><p><em>Why is this important?</em></p><p>A large share of applied DiD and event-study work relies on regression output that looks intuitive but is tightly tied to modelling choices that are rarely spelled out, and this paper shows in a very concrete way where structure is being imposed and how that structure shapes what is being identified. Normalisation choices in event studies are not cosmetic. Using already-treated units as controls is not neutral. Exposure designs are not estimating the same object as binary DiD. Lead coefficients are not just informal diagnostics. Each of these decisions changes the estimand, often in ways that are easy to miss when reading a table or a plot.</p><p>The paper is also important because it forces some discipline into how we talk about dynamics and heterogeneity. Many applied papers present event-time paths, average effects and pre-trend plots as if they were direct summaries of the data, when in fact they are products of models that restrict untreated potential outcomes, restrict how treatment effects evolve over time and restrict how heterogeneity enters. Once those restrictions are written down, it becomes much harder to treat the output as self-explanatory.</p><p>The exposure design discussion makes this very clear. A single regression coefficient is often read as &#8220;the effect of exposure&#8221;, but in the presence of heterogeneity it is an average marginal effect, not an average of underlying treatment effects, and those two objects diverge as soon as effects vary with exposure intensity. That is an interpretation problem and it applies directly to a large class of shift-share and intensity-based designs.</p><p>The same logic applies to event studies. When there are no never-treated units, at least two coefficients are pinned down by assumption rather than data, and with heterogeneous effects the estimated dynamics can reflect changing composition of contributing units rather than changes in causal impacts. None of this requires pathological designs. It arises in settings that look standard.</p><p>On pre-trends and inference, the paper is explicit about trade-offs that are often waved away. Lead coefficients are pre-tests. Pre-tests affect post-test interpretation. Clustered standard errors behave poorly when only a few units identify long leads and lags. These are features of the designs we actually use.</p><p>Overall, the value of the paper is that it tightens the link between regression output and causal interpretation in a way that is directly relevant for applied work. It keeps bringing the reader back to the same questions: what is being identified, under which assumptions, and whether that object matches the story being told.</p><p><em>Who should care? </em></p><p>Anyone running DiD or event studies in applied work, which in practice means most people doing policy evaluation with panel or repeated cross-section data. If you use staggered adoption, plot leads and lags, average dynamic effects, or rely on pre-trend checks to support identification, this paper is written for you.</p><p>This includes people working with clean administrative panels and people working with short, noisy surveys. It includes settings with clear policy shocks and settings with fuzzy timing. It includes binary treatments and exposure designs. The issues the paper raises are not niche. They show up in default workflows.</p><p>It is also relevant for readers and referees. If you are judging whether an event-study plot support PT, whether a dynamic path is meaningful, or whether an exposure coefficient has a sensible interpretation, the paper gives you a sharper set of questions to ask. Where is the normalisation? Which units are acting as controls? What object is this coefficient estimating? If DiD is doing identifying work in your paper, or in papers you read and referee, this applies.</p><p><em>Do we have code?</em></p><p>The paper is accompanied by replication code for the divorce application, and the examples are implemented in standard Stata-style workflows. The authors are explicit about how the event-study specifications, alternative normalisations, and the BJS imputation estimator are implemented, which makes it easy to map the discussion to actual code. This paper is not a software contribution and it is not packaged as a toolkit. The value is that the code mirrors what people already do, which makes the points about normalisation, heterogeneity, and inference easy to check in practice. You can reproduce the figures, change the normalisation, drop never-treated units, and see the mechanics for yourself. If you have used event-study code from recent DiD packages, nothing here will look exotic. That is part of the point. </p><p>In summary, this paper is a reality check on modern DiD practice. It does not propose a new estimator or a new identification strategy. It shows how the tools people already use behave once you take their assumptions seriously. The unifying message is simple: DiD and event studies remain useful, but their output is only as interpretable as the assumptions underneath. This paper helps make those assumptions visible, and it shows where common workflows quietly add identifying content. If you want your coefficients to mean what you think they mean, you need to be clear about what is being identified and why.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>A repeated cross-section is when you observe different units at each point in time rather than following the same units over time, which is common in surveys, firm data, and rotating samples, and is precisely why compositional changes become a problem. When you write &#8220;the treated group&#8217;s outcome changed over time&#8221;, what you&#8217;re really saying is &#8220;the average outcome of whoever is in the treated group sample at time 1 is different from the average outcome of whoever was in the treated group sample at time 0&#8221;. <em>Those aren&#8217;t the same individuals</em>. If the composition changes, part of the difference can come from who is being observed, not from the treatment. That is why standard DiD methods impose no compositional changes/stationarity. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>By stationarity, the DiD literature typically means that the joint distribution of treatment status and covariates doesn&#8217;t change over time, i.e., treatment status and covariates are independent of time:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot; (D,X)&#8869;T&quot;,&quot;id&quot;:&quot;FFREWBOWOD&quot;}" data-component-name="LatexBlockToDOM"></div><p>which rules out compositional changes in who is observed before and after treatment, and is different from stationarity in time-series analysis. The kinds of people (or firms, schools, etc.) you observe before treatment and after treatment come from the same population, with the same mix of characteristics and the same treated/control composition. With repeated cross-sections, you aren&#8217;t following the same units over time, so you need something like &#8220;even though we observe different people each period, they are drawn from the same underlying population&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>We spoke about this before. Conditional parallel trends means that, after conditioning on observed covariates <strong>X</strong>, treated and untreated units would have followed the same average outcome trends in the absence of treatment. This allows different groups to have different raw trends, as long as these differences are explained by <strong>X</strong>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>The efficient influence function characterizes the best possible (lowest variance) regular estimator under the assumed model. The associated efficiency bound is the variance lower bound that no estimator can beat under those assumptions.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Here &#8220;doubly robust&#8221; is meant in the rate sense (how fast the estimation error shrinks as sample size grows). This is different from classical DR based on correct specification of parametric models. In DiD (and many causal methods), you usually have to estimate two auxiliary pieces before you get your treatment effect: an outcome model (how outcomes relate to covariates), and a propensity/selection model (how treatment or group membership relates to covariates). These are called nuisance components because they aren&#8217;t your target, but you need them to get the target. Classical DR means: if either the outcome model is correctly specified or the propensity model is correctly specified, the final estimate is still consistent. In other words: you can mess up one, as long as the other is right you still get the right answer. In this paper DR means something different and weaker: the final estimator converges to the truth at the usual speed (root-n) as long as one of the two nuisance pieces is estimated well enough, even if the other is a bit messy or slow. in summary, classical DR &#8594; about correctness of models, whereas rate DR &#8594; about speed of convergence.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>A Hausman-type test is a consistency check between two estimators that rely on different assumptions. One estimator is robust: it remains valid even when the data are messy, but it&#8217;s less precise. The other is more efficient: it has lower variance, but it only works if a strong assumption holds. The test asks whether the two estimates are statistically indistinguishable. If they are, the strong assumption isn&#8217;t empirically relevant and you can safely use the more efficient estimator. If they aren&#8217;t, the efficient estimator is biased and you should stick with the robust one. In this paper, the robust estimator allows for compositional changes, while the efficient one assumes stationarity. The Hausman test tells you whether stationarity matters for the ATT in your data.</p></div></div>]]></content:encoded></item><item><title><![CDATA[On Finding Causes vs. Finding Levers]]></title><description><![CDATA[What are you actually trying to say?]]></description><link>https://www.diddigest.xyz/p/on-finding-causes-vs-finding-levers</link><guid isPermaLink="false">https://www.diddigest.xyz/p/on-finding-causes-vs-finding-levers</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Tue, 18 Nov 2025 14:28:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IbDJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IbDJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IbDJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IbDJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg" width="1456" height="825" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:825,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Best binoculars for stargazing 2025: Spot stars and galaxies | Live Science&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Best binoculars for stargazing 2025: Spot stars and galaxies | Live Science" title="Best binoculars for stargazing 2025: Spot stars and galaxies | Live Science" srcset="https://substackcdn.com/image/fetch/$s_!IbDJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IbDJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d4e7d06-b3f3-4b17-811d-aea9bda59b43_2560x1450.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there!</p><p>This post is a bit different from the others, but I felt compelled to write it as someone who does research on education and runs this newsletter on a specific causal inference technique. Nothing I will say in this post is new in the sense that it hasn&#8217;t been discussed before, but I do hope to provide more perspective to those who are not &#8220;in the trenches&#8221; of these sorts of nuanced topics.</p><p>It all started with a post on Twitter/X about the effects of school spending on educational outcomes (heterogenous treatment effects&#8217; discussion: almost absent, except maybe for <a href="https://x.com/unboxpolitics/status/1990578814645137474">this</a>). On one side, some users pointed to specific RCTs in developing countries as evidence that increasing school funding increases test scores. On the other, some pointed to meta-analyses finding mixed results in developed countries. The biggest point of discussion was on effect size and statistical-significance.</p><p>I&#8217;m not going to weigh in on that specific example since it&#8217;s not the point of this post (also Twitter/X is a better place for this type of discussion). Instead, I want to talk about the &#8220;causal inference revolution&#8221;, how its legacy might be amplified - positively or negatively - by the problems we will (already?) have (for example, when it comes to outsourcing the reviewing process to LLMs<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>), and circle back to the discussion on effect size and statistical-significance. Admittedly, the referee-example is a niche concern directly affecting a small circle of researchers. But since many of us go on to teach the present and next generations (and maybe some are out there reading this!), a nuanced view here could do wonders to spark their interest in what actually drives variation. </p><p>The debate also struck a chord because it touches on something I love to explore (thanks prof Frank for asking me a very relevant question at my first ever PhD seminar): the mechanisms and channels behind a causal claim. They are the pathways through which the cause creates an effect. The &#8220;cause&#8221; (e.g., spending) doesn&#8217;t &#8220;occur&#8221; through the mechanism; it &#8220;operates&#8221; through the mechanism. The mechanism can be of all sorts, like increasing teacher training or providing meals at no direct cost to the students. It&#8217;s one thing to say &#8220;spending is a cause&#8221; (identification). It is entirely another to ask if spending is the primary driver of student success (decomposition). We need to distinguish between identifying a causal link and finding a policy &#8220;lever&#8221;, a factor that really accounts for the variation we see in the world.</p><h3>How &#8220;credibility&#8221; became the new gold-standard</h3><p>I&#8217;m not sure if it was an &#8220;coordinated revolution&#8221; with a defined plan<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, but something undeniably changed in empirical economics in the early 00&#8217;s. The &#8220;credibility revolution&#8221; of the 90&#8217;s and 2000&#8217;s completely shifted our methodological goalposts. I would argue it wasn&#8217;t so much a planned success as it was a fundamental shift in standards.</p><p>Before, running a kitchen-sink OLS<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> and hand-waving about selection bias and omitted variables was common (first thing you teach the undergrads *not* to do). After, the new gold standard became clean identification<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>. The &#8220;Big Four&#8221; methods - RCTs, DiD, IV, and RDD - became the &#8220;gatekeepers&#8221; of truth<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. In this sense, the movement won. It won the argument over what &#8220;credible&#8221; empirical work looks like, and it gave us a powerful, trusted toolkit to find &#8220;cause&#8221;, and therefore has changed the actual output of the profession<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. </p><p>This new toolkit also created a very powerful new incentive structure. We now know that &#8220;causal novelty&#8221; (finding a new relationship) and &#8220;causal narrative complexity&#8221; (using these sophisticated methods) are strong predictors of getting published in a Top 5 journal. But, anecdotally, it goes further down the line. Every week I get a few friends and colleagues asking me &#8220;how can I DiD this?&#8221;, &#8220;what should be my treatment group?&#8221;, &#8220;is it ok if I have around 20 treated units, shifting in and out of treatment, and 400 control units in some period?&#8221;. I find these questions fascinating because they make us all think deeper about the limits of research questions and proper study design. But notice what is missing from these conversations.</p><p>We spend hours debating the feasibility of the identification strategy (&#8220;can we get the standard errors right?&#8221;, &#8220;is this even worth of publishing?&#8221;) but we rarely stop to ask: if this ~does~ work, will the effect size be large enough to matter?</p><p>And that&#8217;s where the problem starts. The revolution succeeded in teaching us how to find a cause. But in our quest to publish novel, statistically significant causal chains, we often forget to ask if we&#8217;ve found a lever<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>.</p><h3>Why a cause is not a lever</h3><p>So we have the toolkit. We can &#8220;confidently&#8221; say that X causes Y. But this brings us back to the Twitter debate and that 0.056 SD effect (a number which I made up, by the way, and got smart pushback on). The problem happens when we assume finding a cause is the same as finding a lever. It isn&#8217;t.</p><p>Prof Cochrane - just a few days ago - cleverly articulated this in his essay &#8220;<a href="https://www.grumpy-economist.com/p/causation-does-not-imply-variation">Causation Does Not Imply Variation</a>&#8221;. From my understanding, his point was that you can rigorously prove a policy causes a change in an outcome, but that cause might account for almost none of the variation between success and failure in the real world. You can prove that school funding causes test scores to move (p &lt; 0.05), but if the effect is small, it explains almost none of the difference between a struggling school and a thriving one.</p><p>Prof Galiani wrote a <a href="https://sebastiangaliani.substack.com/p/causality-causes-and-what-our-methods">post</a> after prof Prof Cochrane&#8217;s in which he takes this further, explaining why our methods fail us here. He reiterates that the &#8220;causal revolution&#8221; tools (DiD, RCTs, etc.) are built to answer a specific question - &#8220;what happens to Y when we induce an exogenous change in X?&#8221; - and that they are ~not~ built to answer THE broader question: &#8220;what are the main causes of variation in Y?&#8221;.</p><p>Prof Galiani points out a couple of limitations in treating these &#8220;causes&#8221; as &#8220;levers&#8221;. In the first one, which is a &#8220;battle&#8221; between &#8220;local vs structural&#8221;, he says that our methods identify a total effect in a specific context rather than a universal structural parameter. This focus on the &#8220;snapshot&#8221; is likely to obscure dynamic levers: a small effect size might look dismissible in a static regression, but if it captures a structural parameter that compounds over time (i.e., a daily learning gain), it can be massive. We learn what happened there, not necessarily how the machine works over time and everywhere. In the second one, characterized as &#8220;the SUTVA trap, he says that these methods rely on SUTVA, which essentially assumes no spillovers. But true &#8220;levers&#8221; - like macro policies or institutional changes - are defined by their equilibrium effects and spillovers. By assuming them away to get a clean estimate, we often strip the &#8220;cause&#8221; of the very mechanism that would make it a &#8220;lever&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>. </p><p>Statisticians and econometricians have been warning about this for decades, often using a vocabulary that is painfully absent from our current seminars. Profs McCloskey and Ziliak and  created the term &#8220;<a href="http://valjhun.fmf.uni-lj.si/~mihael/ef/ps/pdfpapers/mccloskey.pdf">Oomph</a>&#8221; to describe &#8220;the difference a treatment makes&#8221; arguing that our profession suffers from a disproportionate focus on &#8220;precision&#8221; (statistical significance) at the expense of &#8220;Oomph&#8221; (magnitude).</p><p>This outcome shouldn&#8217;t surprise us. Nearly 40 years ago, sociologist Peter Rossi formulated the &#8220;<a href="https://gwern.net/doc/sociology/1987-rossi">Stainless Steel Law of Evaluation</a>&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>: &#8220;The better designed the impact assessment of a social program, the more likely is the resulting estimate of net impact to be zero&#8221;. The credibility revolution gave us the stainless steel tools Rossi predicted; we shouldn&#8217;t be shocked that they are doing exactly what he said they would: revealing that many well-meaning interventions don&#8217;t work.</p><p>Prof Kr&#228;mer formalized this critique. He <a href="https://www.econstor.eu/bitstream/10419/292347/1/schm.131.3.455.pdf">argued</a> that our obsession with precision has birthed a new category of scientific failure: the Type C Error (his nomenclature).</p><p>We are familiar with Type A errors (finding an effect where there is none) and Type B errors (missing a large effect because of noise). But a Type C Error occurs when there is a &#8220;small effect (no &#8220;oomph&#8221;), but due to precision it is highly &#8220;significant&#8221; and therefore taken seriously&#8221;. This is the (abstract) 0.056 SD problem. It is a Type C error masquerading as a scientific victory. We used our sophisticated &#8220;credibility&#8221; toolkit to maximize precision, driving our standard errors down to zero, which allowed us to put three stars next to a number that, for all practical purposes, is zero.</p><p>We found a cause. We did not find a lever. A &#8220;lever&#8221; is a confirmed cause that possesses &#8220;Oomph&#8221;, meaning it explains a substantial portion of the actual variation in the outcome. But we need to be precise. A &#8220;lever&#8221; isn&#8217;t a large coefficient in isolation. As <a href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/joes.12330">Prof Sterck (2019)</a> provides the background for me to argue, a true lever is a variable that drives the variation we see in the real world. It is a factor that explains a meaningful percentage of the deviations in our outcome. So a 0.056 SD effect might be a &#8220;cause&#8221;, but if it explains only 0.1% of the variation in test scores, it is certainly not a lever.</p><p>This distinction seems clear &#8220;enough&#8221; in theory, but in practice - as when we are staring at a regression table - it gets muddy. Where&#8217;s the &#8220;lever&#8221; column? Instead, we fall back on the standard metrics we were trained to trust, a lot of the times without realizing they aren&#8217;t answering the question we think they are.</p><h3>What do we even mean by &#8220;important&#8221;?</h3><p>In the back of our heads there&#8217;s always a voice that goes &#8220;check N, R^2, effect size, and statistical significance&#8221;. I don&#8217;t think that&#8217;s bad, but it is superficial. </p><p>When we find a &#8220;cause&#8221;, how do we decide if it matters? We play around with terms like &#8220;economically significant&#8221; but as Prof. Sterck (2019) argued, the literature is surprisingly vague on what that really means<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>. We usually rely on standard tools - like standardized coefficients (the 0.056 SD!) or R^2 decompositions, but Prof Sterck shows these are often &#8220;flawed and misused&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> (standardized coefficients, for instance, are sensitive to sample variance, making them unreliable for comparing effects across different populations). If our tools for measuring &#8220;importance&#8221; are not ideal, it&#8217;s no wonder we are left to rely on one tool that gives us a clear binary answer and that all scientists (I am being charitable) understand: the p-value.</p><p>But this is where the &#8220;stars&#8221; can &#8220;blind&#8221; us. To see why, look at the fertility case study in Prof Sterck&#8217;s paper. He analyzes the determinants of fertility rates (specifically the number of births per woman) using a massive dataset (N &gt; 490,000). Because the sample is so huge, almost everything is statistically significant. If you look at the regression output, you see stars next to ln(GDP/capita), land quality, distance to water, child mortality, age, and education. If you are hunting for &#8220;causes&#8221; (stars), you get lost in the sauce. The regression output shows that all of the aforementioned variables are &#8220;significant&#8221; causes. This makes it incredibly easy to engage in HARKing (Hypothesizing After the Results are Known). A researcher looking at those stars could pick any variable (e.g., distance to water) and justify a paper on it. They are all &#8220;causes&#8221;. It&#8217;s not even &#8220;data mining&#8221; or &#8220;p-hacking&#8221; in the traditional sense because no manipulation was necessary. The significance was guaranteed by the sample size -  which is a feature, not a bug, as it gives us precise estimates (though, as I&#8217;ve <a href="https://www.diddigest.xyz/p/in-defense-of-machine-learning-in">argued before</a>, modern ML tools are better suited to handle this &#8220;star-gazing&#8221; problem in high-dimensional data than standard OLS); the error is in pretending that significance equals importance. Of course, context dictates the threshold for &#8220;Oomph&#8221;. If the outcome is &#8220;survival rates&#8221;, a 0.056 SD effect *is* a massive victory. But for continuous variables like test scores or wages, where we are trying to explain inequality or performance gaps, a 0.056 SD effect without a cost-benefit argument should be taken with a grain of salt.</p><p>But&#8230; if you shift your thinking to look for &#8220;levers&#8221;, i.e., by asking both &#8220;is it non-zero?&#8221; and &#8220;how much of the deviation does this actually explain?&#8221;, the picture changes from night to day. To go back to his example: ln(GPD/capita) explains only 0.19% of the deviation (a cause, but not a lever), distance to water explains 0.73%, age explains 45.6%, and woman&#8217;s education explains 10.0%. The &#8220;stars&#8221; told us all these variables were valid. But only by asking the right question - &#8220;does this variable actually drive the variation?&#8221; - do we see that GDP is a distraction in this context, while education is a massive policy lever.</p><p>This is the trap we fall into. We use the precision of our causal tools to find the 0.19% drivers, and because they have stars next to them, we treat them as if they matter just as much as the 10% drivers.</p><h3>What we publish vs what we cite</h3><p>This brings us back to the incentives. If we know that &#8220;levers&#8221; (like education in the example above) are what actually drive outcomes, why do we spend so much time hunting for and publishing tiny &#8220;causes&#8221;?</p><p>Because that is what the market *rewards*.</p><p><a href="https://arxiv.org/pdf/2501.06873">Garg and Fetzer (2025)</a> find a stark divergence between what gets published (causes) and what creates impact (levers). Top 5 journals reward &#8220;novelty&#8221;, e.g., finding a new causal link, even a tiny, obscure one&#8230; is the path to acceptance. The long-term impact (citations), tho, comes from engaging with &#8220;central, widely recognized concepts&#8221;. </p><p>We are incentivized to hunt for the 0.056 SD effect because it&#8217;s novel. But we cite the papers that discuss the big, central questions. I think both are somewhat fair ways to advance the science, and not totally exclusionary. We need architects to design the building (structural/levers) and plumbers to fix the pipes (causal/identification). The problem arises when we only reward the plumbers but expect them to explain the architecture. Prof. Megan Stevenson (2023)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> argues that this expectation stems from a mistaken &#8220;Engineer&#8217;s View&#8221; of the world - the belief that society is a machine where we can isolate and pull specific levers for predictable results. But the social world is actually full of &#8220;stabilizers&#8221;, which are forces that push people back onto their original trajectories after a small, limited in scope intervention. A small &#8220;cause&#8221; (like a job training program) rarely triggers a &#8220;cascade&#8221; of success because structural forces dampen the effect. Finding a statistically significant effect in a vacuum ignores the stabilizing forces that make that effect irrelevant in the aggregate.</p><p>And as I briefly mentioned, this problem is likely to get worse. As AI-assisted tools are increasingly used in peer review, we risk amplifying this exact incentive structure. A LLM can be trained to find &#8220;novel causal claims&#8221; and check for the &#8220;right&#8221; methods, but it cannot ~easily~ be trained to understand &#8220;economic importance&#8221; (or can it???? I won&#8217;t bet against it) a concept we humans struggle to agree on sometimes since it&#8217;s highly context-dependent. An AI reviewer seems to be the ultimate &#8220;cause&#8221; hunter, as of today. I think it will likely reward the 0.056 SD effect as long as it&#8217;s novel and significant. It has no intuition for finding the &#8220;lever&#8221;. LLMs are trained on the corpus of published literature - a literature that systematically selects for significance over importance. The AI will likely entrench, rather than correct, this bias.. But I want to be proven wrong.</p><p>As Prof. Kr&#228;mer wrote, &#8220;Cheap t-tests... have in equilibrium a marginal scientific product equal to their cost&#8221;. AI makes those tests almost free. My point isn&#8217;t that we should abandon the p-value, but that we must be far more rigorous about the questions we ask during the design phase. Instead of stopping at &#8220;can I identify this?&#8221; or &#8220;is this novel?&#8221;, we need to ask: &#8220;does this variable explain the variation?&#8221;. If it works, how much of the problem does it actually solve? Practically, this means reporting variance decomposition alongside our regression tables, or justifying small effects with explicit cost-benefit or compounding arguments. If we don&#8217;t start valuing the answers to these questions, we are going to drown in statistically significant noise.</p><p>Further readings:</p><ul><li><p>Angrist, J. D., &amp; Pischke, J. S. (2010). <a href="https://www.researchgate.net/publication/43296056_The_Credibility_Revolution_in_Empirical_Economics_How_Better_Research_Design_Is_Taking_the_Con_out_of_Econometrics">The credibility revolution in empirical economics: How better research design is taking the con out of econometrics.</a> <em>Journal of Economic Perspectives</em>, 24(2), 3-30.</p></li><li><p>Deaton, A. (2010). <a href="https://www.researchgate.net/publication/46554038_Instruments_Randomization_and_Learning_about_Development">Instruments, randomization, and learning about development.</a> <em>Journal of Economic Literature</em>, 48(2), 424-455.</p></li><li><p>Imbens, G. W., &amp; Angrist, J. D. (1994). <a href="https://www.gsb.stanford.edu/faculty-research/publications/identification-estimation-local-average-treatment-effects#:~:text=Angrist,Issue%202%20Pages%20467%E2%80%93475.">Identification and Estimation of Local Average Treatment Effects.</a> <em>Econometrica</em>, 62(2), 467&#8211;475.</p></li></ul><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Already happening in <a href="https://x.com/iclr_conf/status/1990204431959470099">other sciences</a>. Good point <a href="https://x.com/alexolegimas/status/1990424516112056351?s=20">here</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Anyone wanting to discuss this case against Kuhn&#8217;s definition, e-mail me!</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>&#8220;OLS&#8221; (Ordinary Least Squares) is the standard method for drawing a line of best fit through data. A &#8220;kitchen-sink&#8221; regression is when a researcher throws every available variable into the model (control variables) in hopes of isolating an effect. Endogeneity is a fatal flaw here: it happens when there is something else driving your result that you didn&#8217;t (or couldn&#8217;t) put in the &#8220;sink&#8221;. For example, if you find that &#8220;tutoring causes higher test scores&#8221; but you didn&#8217;t control for &#8220;parental motivation&#8221; your result is <a href="https://x.com/Jackbmeyer/status/1989809239842570666/photo/1">endogenous</a>, meaning it&#8217;s biased because motivated parents are likely the ones hiring the tutors.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>In econometrics, &#8220;identification&#8221; refers to the strategy used to try to strip away bias and identify the true causal link. A &#8220;clean&#8221; identification strategy mimics a laboratory experiment. For example, an RCT flips a coin to assign treatment; a Regression Discontinuity (RDD) compares people just above and just below an arbitrary cutoff (like a test score threshold for a scholarship); and DiD compares the change in a treated group to the change in a control group over time.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><a href="https://arxiv.org/pdf/2501.06873">Garg and Fetzer (2025)</a> note that &#8220;leading journals now prioritize studies employing these methods over traditional correlational approaches&#8221;, and they point out that these four specific methods are the ones that have seen &#8220;substantial growth&#8221; as the discipline shifted toward rigorous identification.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><a href="https://arxiv.org/pdf/2501.06873">Garg and Fetzer (2025)</a> show that the share of claims supported by these rigorous causal methods skyrocketed from about 4% in 1990 to nearly 28% in 2020. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>A &#8220;cause&#8221; answers the question &#8220;does X affect Y?&#8221; (precision), a &#8220;lever&#8221; answers the question &#8220;how much of the difference between success and failure does X actually explain?&#8221; (importance). </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>He argues that we need both: the local identification of causal effects <em>and</em> models (comparative statics) to understand the &#8220;causal architecture&#8221; of a phenomenon. He also notes in an addendum that low partial R^2 isn&#8217;t bad per se, provided the intervention is cost-effective - which, by extension, aligns with our definition of a &#8220;lever&#8221;: it has to be worth pulling.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Thank you <a href="https://x.com/lu_sichu/status/1990883038826225987?s=20">Sichu</a> for flagging this.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>Standardized coefficients are tricky (not as straight forward as logs) because they depend heavily on the variance of your sample, making comparisons across studies a literal nightmare (good luck on your meta-analyses).<strong> </strong>R^2 measures (like Shapley values) are often computationally heavy and not intuitive, and they that can attribute importance to irrelevant variables just because they are correlated with relevant ones (that is also one of the criticisms regarding the use of ML in Econ, specifically for large samples)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>And &#8220;with large N you don&#8217;t need randomization&#8221;&#8230; :P</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>Another suggestion by <a href="https://x.com/lu_sichu/status/1990882868600422473?s=20">Sichu</a>.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Hidden Bias, Pie Charts and Spillovers]]></title><description><![CDATA[What to consider when reality has stopped cooperating]]></description><link>https://www.diddigest.xyz/p/hidden-bias-pie-charts-and-spillovers</link><guid isPermaLink="false">https://www.diddigest.xyz/p/hidden-bias-pie-charts-and-spillovers</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 03 Nov 2025 14:18:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!AW3t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AW3t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AW3t!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AW3t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2740786,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.diddigest.xyz/i/176027761?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AW3t!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!AW3t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78ea862e-6215-4617-9953-f2146e10e18a_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hello! Today we have 3 theory papers, 1 &#8220;software&#8221;/guide paper, and 3 applied I found (the Japanese one is behind a paywall, so email me if you would like a copy). All of them are super cool.<br>Buckle up because we have many concepts from areas we might not be entirely familiar with&#8230; check the footnotes!</p><ul><li><p><a href="https://arxiv.org/abs/2510.09064">Sensitivity Analysis for Treatment Effects in Difference-in-Differences Models using Riesz Representation</a>, by Philipp Bach, Sven Klaassen, Jannis Kueck, Mara Mattes and Martin Spindler</p></li><li><p><a href="https://onilboussim.github.io/files/JMP.pdf">Compositional Difference-in-Differences for Categorical Outcomes</a>, by Onil Boussim (on the JM!)</p></li><li><p><a href="https://arxiv.org/abs/2502.03414">Efficient nonparametric estimation with difference-in-differences in the presence of network dependence and interference</a>, by Michael Jetsupphasuk*, Didong Li and Michael G. Hudgens (* on the JM!)</p><div><hr></div></li></ul><p>Software/guide:</p><ul><li><p><a href="https://arxiv.org/pdf/2510.19426">Using did_multiplegt_dyn to Estimate Event-Study Effects in Complex Designs: Overview, and Four Examples Based on Real Datasets</a>, by Cl&#233;ment de Chaisemartin, Diego Ciccia, Felix Knau, M&#233;litine Mal&#233;zieux, Doulo Sow, David Arboleda, Romain Angotti, Xavier D&#8217;Haultf&#339;uille, Bingxue Li, Henri Fabre and Anzony Quispe <em>(De Chaisemartin et al. extend their 2024 event-study estimators into a new command for Stata and R. The command handles on/off, continuous and multivalued treatments, and it replaces did_multiplegt for this use. The new command is faster because variance uses analytic formulas rather than bootstrap which they show how in the paper. It&#8217;s a neat and practical guide, they go through the estimators, validate them in simulations and show four real-data applications).</em></p><div><hr></div></li></ul><p>Applied:</p><ul><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5375017">The Labor Market Effects of Generative AI: A Difference-in-Differences Analysis of AI Exposure</a>, by Andrew C. Johnston and Christos A. Makridis (USA, DiD with continuous treatment intensity in an event study framework) &#8594; <em>I was seriously considering starting a Substack focused on Econ and AI, but I don&#8217;t have time. If you feel like you could do this, you should do it. It&#8217;s an area in development, at its early stages, and we would all benefit from learning more about it. This paper is the first comprehensive study providing economy-wide evidence of real labour market impacts from AI exposure. What I particularly enjoyed was the analysis at both extensive (e.g., whether workers remain employed or lose jobs, whether firms enter or exit sectors, whether establishments expand or contract their workforce) and intensive (e.g., how AI changes the composition of tasks workers perform within their jobs, changes in productivity/output per worker, changes in hours worked or work intensity) margins. They also decompose AI exposure into augmenting (where AI complements human work) versus displacing (where AI substitutes for it) components, finding that augmenting exposure drives employment growth, while displacing exposure causes contraction. Because the analysis was done using administrative data that covered the universe of U.S. employers, they were able to show that high-skilled workers see larger gains, which represents a departure from previous automation waves.</em></p></li><li><p><a href="https://link.springer.com/article/10.1007/s42973-025-00214-8">The role of human interaction in innovation: evidence from the 1918 influenza pandemic in Japan</a>, by Hiroyasu Inoue, Kentaro Nakajima, Tetsuji Okazaki, and Yukiko U. Saito (Japan, tech-class x year DiD with common shock and differential exposure; event-study; PPML for counts) &#8594; <em>I&#8217;m a sucker for natural experiments, even though there were lots of deaths involved. I&#8217;m also a sucker for historical data collection, which I would love to do if I spoke Japanese. This paper mixes both things. The 1918 flu in Japan had three waves: Oct 1918&#8211;Mar 1919, Dec 1919&#8211;Mar 1920, Dec 1920&#8211;Mar 1921, about 23.8 million infections and 390k deaths out of a 55 million population, which is crazy. Transport and factories were disrupted, mortality peaked among working-age adults and face-to-face contact became costly for inventors. This setting allowed the authors to test the claim that tacit knowledge and spillovers travel through direct interaction. We know that collaboration drives the flow of ideas which has a direct effect on the quality of output, therefore the test was straight-forward: if tacit knowledge moves through direct contact, then tech classes that collaborated more before 1918 should stumble more when contact gets costly. The authors proxy collaboration needs with pre-1918 co-inventing shares and track patents by class over 1911&#8211;1930, so you watch the pandemic hit play out. The drop is sharper in collaboration-intensive classes in 1919&#8211;1921, and it comes from fewer new inventors rather than incumbents retreating. The key message then becomes: raise interaction costs and innovation slows most where collaboration is the input. This paper made me think of Covid as the modern analogue: a common shock that raised interaction costs, with differential exposure by how much collaboration must be in-person. Unlike in 1918, we had Zoom and cloud tools which softened the blow, so we should expect the effects to be smaller where remote work was feasible and larger in wet-lab, field or hardware R&amp;D(if anyone wants to write about it).</em></p></li><li><p><a href="https://onlinelibrary.wiley.com/doi/full/10.1002/soej.12793">Forgoing Nuclear: Nuclear Power Plant Closures and Carbon Emissions in the United States</a>, by Luke Petach (USA, TWFE DiD with staggered treatment timing; robustness checks with Callaway-Sant&#8217;Anna and Borusyak et al. estimators) &#8594; <em>We were discussing the other day on X/Twitter about the difficulty of using DiD to quantify the effect of new data centres on energy prices due to several threats to identification: data centre planners select locations based on energy characteristics that independently affect price trends (selection bias), their arrival coincides with other infrastructure investments and policy changes (confounders), energy markets are interconnected across regions (spillovers) and prices may react to announcements before construction begins (anticipation effects). Fundamentally, it&#8217;s hard to find a credible counterfactual for what would have happened to energy prices without the data centre. Luke has addressed lots of those concerns in his paper, but more importantly, a plant closing is more exogenous than a data centre (or plant) opening. The &#8220;treatment&#8221; happens to the state rather than being chosen by strategic actors based on local conditions. Luke finds that nuclear closures increased state-level CO_2 emissions by 6-8%, with the gap filled almost entirely by coal rather than cleaner natural gas. Event studies showed no pre-trends and the results hold when analyzed at the regional electricity market (RTO/ISO) level to account for cross-state spillovers, meaning that the effect is real and it&#8217;s carbon-intensive. This is an important example that showcases how identification strategy matters as much as the question itself. When you can&#8217;t randomize treatment, the next best thing is finding settings where treatment is plausibly exogenous, even if that means studying the problem in reverse. </em></p></li></ul><div><hr></div><h3>Sensitivity Analysis for Treatment Effects in Difference-in-Differences Models using Riesz Representation</h3><h5>TL;DR: this paper shows how to measure how much hidden bias could affect DiD estimates when the PTA might not hold. Using the Riesz representation, the authors express the treatment effect as a weighted average and show how bias from unobserved factors links to familiar quantities. The result is a practical way to report DiD effects with clear bounds instead of relying on untestable assumptions.</h5><p><em>What is this paper about?</em></p><p>When analyzing a treatment in order to quantify a causal effect, we have to consider a lot of things. But we never really &#8220;see&#8221; the full picture. People, firms or regions decide to take up a policy for reasons we can&#8217;t fully observe, and those same reasons often affect how their outcomes change over time. In DiD we &#8220;deal&#8221; (to the best of our intentions) with that by conditioning on covariates and assuming that, after controlling for them, treated and control units would have followed similar trends in the absence of treatment (the PTA). The problem is that this assumption of parallel trends given covariates can still fail if we&#8217;re missing something important that we can&#8217;t quantify or account for.</p><p>This paper is about measuring how much of such missing information &#8220;could&#8221; matter. Instead of treating &#8220;unobserved confounders&#8221; as a black box, the authors show how to express their influence in terms of quantities we can describe and argue about. They build on a mathematical idea called the Riesz representation<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>, which expresses the treatment effect as an average of predicted outcome changes, each multiplied by a specific weight. Once written that way, we can study how much the estimate would move if some unobserved factors were left out of either the prediction or the weighting step.</p><p>The motivation is there: the goal with a DiD design is to estimate the ATT, but this estimate is only as good as the assumption that the PTA holds once we condition on covariates. Since that can never be fully verified, the question becomes how to describe and measure the possible bias rather than ignore it.</p><p><em>What do the authors do?</em></p><p>They start off by rewriting the DiD effect as an average of predicted outcome changes taken with a particular set of weights. That step (which comes from the Riesz representation) makes it easier to see what happens when some parts of the data are missing. Once written that way, the estimate depends on two things: the model used to predict outcome changes and the weights that compare treated and control units. If either of these leaves out important factors, the bias can be expressed in terms of how much those missing variables would help explain the outcome or the likelihood of treatment.</p><p>This is where the paper becomes useful for applied work. It shows that the size of the bias depends on how much the unobserved variables would improve the chosen model&#8217;s fit for outcomes *and* for treatment assignment, and on how correlated those two pieces are. Each of these quantities can be linked back to measures we are familiar with, such as the partial R^2<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. With that the authors turn &#8220;how sensitive is my DiD estimate to hidden bias?&#8221; into something measurable.</p><p>They go to extend the same idea to staggered adoption designs, showing how to bound both point estimates and confidence intervals when some confounders are missing. Instead of producing a single number that might rest on unrealistic assumptions, the paper shows how to report an effect that comes with well-defined bounds, or how large the unobserved bias would need to be to make the result disappear/insignificant. The estimation itself follows a DML approach, where modern ML methods are used to learn both the prediction and weighting steps without having to assume a specific functional form. This keeps the method practical for high-dimensional data while preserving standard large-sample inference.</p><p><em>Why is this important?</em></p><p>Couple of reasons, both related to real-world problems and to the econometrics itself. The first one is straight-forward: the PTA does half of the work in DiD, and it comes directly from the design. Because we never observe the treated group&#8217;s counterfactual path after treatment, we need to assume it would have followed the same trend as the control group once we condition on covariates. The issue is that this path is unobserved by construction, so we can&#8217;t ever prove that the assumption holds. That&#8217;s what makes DiD possible, but also what makes it fragile. In practice we check pre-trends, add controls, and then try to justify the identification assumptions, but there is always a gap between what we observe and what drives selection into treatment and changes in outcomes. This paper gives a way to describe that gap and put numbers on it, so we are not asked to take the key assumption on faith.</p><p>It also changes how we present results. Instead of a single estimate that depends on untestable judgment calls, you can say: here is my best estimate and here is how far it would move if unobserved factors were strong enough to improve prediction by this much, which then makes claims more honest without turning them into hand-waving. It also helps avoid overclaiming and it gives a common language for authors, referees, and seminar audiences to discuss what &#8220;plausible&#8221; unobserved bias means.</p><p>For applied work, the paper ties sensitivity to the tools we already use: pre-trend placebos, leave-one-covariate-out checks and familiar fit measures. If you have multiple pre-periods, you can calibrate scenarios from the size of the placebos. If you don&#8217;t, you can benchmark against observed covariates and ask whether any truly unobserved factor would be as powerful as the ones you already &#8220;see&#8221;. Either way, you get clear bounds for the point estimate and the confidence interval and the same logic carries over to staggered adoption settings.</p><p>There&#8217;s a policy angle too. When a result is sensitive, you can show how and by how much rather than pretending the issue isn&#8217;t there. When it&#8217;s robust, you can document that claim with numbers. In both cases the takeaway becomes easier to trust, which is the point of methods in the first place.</p><p><em>Who should care?</em></p><p>Applied economists using DiD in labour, education, health, development or policy evaluation, anyone who has ever been told by a referee to &#8220;test robustness to hidden confounders&#8221;, researchers who rely on DiD but worry about unobserved bias they can&#8217;t rule out, and methodologists interested in linking ML-based estimation with identification theory.</p><p><em>Do we have code?</em></p><p>There&#8217;s an open-source implementation in <a href="https://docs.doubleml.org/stable/index.html">DoubleML</a> for Python, with a user guide that covers this DiD sensitivity setup (2&#215;2 and staggered), examples, and the Riesz-based formulas. You can run the DML estimator, set sensitivity scenarios (from pre-trends or benchmarking) and get point and CI bounds straight from the package.</p><p>In summary, this paper focuses on the most fragile part of DiD: the fact that the PTA can&#8217;t be proved. By rewriting the effect using the Riesz representation, the authors show how to express bias from unobserved factors in terms of familiar quantities like partial R^2. The estimation uses DML, which keeps it workable when there are many covariates or flexible models. The idea is simple: instead of assuming the PTA holds, measure how far it would have to fail for the result to change. It&#8217;s a way to be clearer about what we can learn from the data and what still depends on judgment.</p><div><hr></div><h3>Compositional Difference-in-Differences for Categorical Outcomes</h3><p><em>(<a href="https://onilboussim.github.io/">Onil</a> is a PhD candidate at PSU, he&#8217;s on the JM this year and this is his JMP! Good luck, Onil :) I also liked this paper because the tables are well-structured, the plots and diagrams are pretty and it&#8217;s well formatted. See, a better world is possible)</em></p><h5>TL;DR: this paper develops a DiD framework for categorical outcomes where shares must stay positive and sum to one. By replacing additive changes with proportional growth and redefining the PTA, it yields valid counterfactuals and interpretable effects for outcomes that are distributions rather than single values.</h5><p><em>What is this paper about?</em></p><p>One of the first things you learn in data analysis is types of data. This is gonna be important here, as we deal with categorical vars<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. These are vars that take on labels instead of numbers (e.g., vote choice, occupation, education level) and what we often analyse are the shares of each category. The issue is that those shares form a composition: they&#8217;re all positive and must add up to one. If one category&#8217;s share increases, at least one other must decrease. We can&#8217;t do this with our standard DiD, which relies on additive changes that can move freely in either direction. In a compositional setting, those additive differences can produce impossible results (negative shares, totals above one, and most importantly changes that have no behavioural meaning). Onil then introduces Compositional DiD (CoDiD), a framework that fixes this by working with proportional (rather than additive) changes.</p><p>The idea is to replace the PTA with a parallel growths assumption: in the absence of treatment, each category&#8217;s share would have grown (or shrunk) at the same *proportional rate* in treated and control groups. This keeps the counterfactual composition valid (shares stay positive and sum to one) and lines up with how we usually mode choices across categories in econ. Onil then uses this framework to define new treatment effects that capture how the overall distribution changes and how probability mass shifts between categories<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>. He extends it further to cases with staggered adoption and to a version that builds synthetic counterfactuals for compositional data (did anyone say synthetic DiD?) that works when the outcome is a distribution instead of a single number.</p><p><em>What does the author do?</em></p><p>Onil does *a lot*. He starts with the simplest possible setup (our canonical the 2&#215;2 case) to lay out the idea of parallel growths and show both its economic and geometric meaning. From there, he generalises the framework to multiple periods, which is what most of us actually deal with in applied work. He then connects this new approach to familiar methods like standard DiD and Synthetic DiD, showing where they overlap and where the compositional logic changes the interpretation.</p><p>To make it concrete, the paper includes two empirical examples: one on how early voting reforms affected turnout and party vote shares and another on how the Regional Greenhouse Gas Initiative (RGGI) shifted the composition of electricity generation. Both illustrate how treating the outcome as a distribution rather than a single variable changes the estimated effect and keeps the counterfactuals within the realm of what&#8217;s possible.</p><p><em>Why is this important?</em></p><p>The core issue is how to think about the PTA when the outcome is categorical. A na&#239;ve way would be to run a standard DiD separately on each category&#8217;s raw share, predict what those shares would have been without treatment and then normalize everything so the shares add up to one. That might sound reasonable, yet it&#8217;s not. Doing this ignores how categories relate to each other (e.g., an increase in one share &#8220;mechanically&#8221; means a decrease in another) and it messes with the link to any behavioural or structural model of how people make choices. In other words, you&#8217;d be imposing linear trends on something that we don&#8217;t know if evolves linearly.</p><p>From an econometrics point of view, this causes several issues: it violates the probability constraint (shares might turn negative or exceed one), it treats each category as independent even though they&#8217;re jointly determined and it produces counterfactuals that have no theoretical justification in how choices across categories actually adjust. The bigger point is that additive DiD logic doesn&#8217;t fit categorical data. What Onil does is rebuild the entire framework so the counterfactual distribution is coherent (i.e., the total probability mass stays fixed, categories remain linked and the treatment effect can be interpreted in economic terms rather than as an artifact of normalization).</p><p><em>Who should care?</em></p><p>Beyond the ones Onil talks about, pretty much anyone studying outcomes where the composition is what you&#8217;re interested in. Migration (domestic, international, return), industry shifts (manufacturing, services, tech), language use (home language shares), transportation modes (car, public, walking) or household spending patterns (food, housing, leisure) all fit this setup. Development work often tracks how populations move across employment types, informal versus formal work, or agricultural crops. If you can think of a pie chart when modeling the var at stake, you should consider CoDiD.</p><p><em>Do we have code?</em></p><p>No, and I tried to think of a way to code it but couldn&#8217;t (#skillissue). If code is released later, I&#8217;ll link it here.</p><p>In summary, CoDiD provides a framework for analysing categorical outcomes within a coherent probabilistic structure. It replaces additive comparisons with proportional growth, ensuring that counterfactuals remain consistent with the underlying composition and that estimated effects reflect economically meaningful reallocation across categories. Onil&#8217;s contribution is both conceptual and practical: it formalises how to conduct DiD when the object of interest is a distribution, preserving internal consistency and interpretability across empirical settings.</p><div><hr></div><h3>Efficient nonparametric estimation with difference-in-differences in the presence of network dependence and interference</h3><p><em>(<a href="http://jetsupphasuk@unc.edu">Michael</a> is a Ph.D. Candidate in UNC&#8217;s Department of Biostats, he&#8217;s on the JM this year and this is his JMP! Good luck!)</em></p><h5>TL;DR: this paper builds a DiD framework for settings where one unit&#8217;s treatment affects others through networks. It replaces the usual no-interference assumption with a setup that measures both direct and spillover effects, using an estimator that stays reliable even when some parts of the model are wrong.</h5><p><em>What is this paper about?</em></p><p>In lots of DiD setups, units are assumed to be &#8220;isolated&#8221; (no contamination, SUTVA anyone?): one unit&#8217;s treatment doesn&#8217;t affect another&#8217;s outcome. That assumption is broken the moment there&#8217;s a network (e.g., people, firms and/or regions connected in ways that let effects spill over). When a factory installs a &#8220;pollution scrubber&#8221;, nearby counties benefit too. When a vaccination campaign starts in one area, infection risk falls for neighbours as well. These are classic cases of interference: what happens to you depends not only on your own treatment but also on others&#8217;.</p><p>This paper deals with that. The authors build a Difference-in-Differences framework that can handle both network dependence (correlated outcomes across connected units) and interference (treatment spillovers). Instead of assuming independent, neatly separated units, they let each unit&#8217;s exposure depend on its *position* in the network and on how treatment spreads through its connections.</p><p>They then show how to estimate average treatment effects in this environment efficiently and without bias, even when nuisance functions like treatment probabilities or outcome models are learned flexibly with ML. This gives us a doubly robust, semiparametric estimator that remains valid under complex dependence structures  (a kind of &#8220;network-aware&#8221; DiD that accounts for who is connected to whom and how that matters for identification and inference).</p><p><em>What do the authors do?</em></p><p>They start by defining what &#8220;treatment&#8221; means when spillovers exist. Each unit isn&#8217;t just treated or untreated, but it also has <em>exposure</em> through its neighbours. The authors formalise this by introducing exposure mappings which describe how each unit&#8217;s treatment and its neighbours&#8217; treatments combine to determine potential outcomes.</p><p>From there, they extend the DiD framework to this network setting. They define a conditional parallel trends assumption: in the absence of treatment, both directly treated and indirectly exposed units would have followed the same <em>expected trend</em>, conditional on their covariates and network position. That assumption plays the same role as PTA in standard DiD but now accounts for dependence across connected units.</p><p>They then build a doubly robust + semiparametric estimator for the average treatment effect under interference. It combines an outcome model (predicting changes) with a treatment and exposure model (predicting who gets treated and how that treatment propagates through the network). Either model can be misspecified, but as long as one is correct, the estimator remains consistent. The authors also derive its asymptotic efficiency bound and show that the estimator achieves it, even when using flexible ML methods to estimate the nuisance functions (we talked all about these terms before).</p><p>Finally, they run simulations to check performance under &#8220;realistic&#8221; (IRL scenarios) network structures and apply the method to US county-level data, studying how the adoption of scrubbers in coal power plants affected cardiovascular mortality; both in treated counties and in neighbouring ones that received indirect benefits.</p><p><em>Why does this matter?</em></p><p>We don&#8217;t exist in isolation, we are all part of  network. Factories share &#8220;air sheds&#8221;, counties share labour markets and hospitals, schools share catchment areas, firms share suppliers. In these settings treatment in one place changes exposure in nearby places. If we ignore that structure, the DiD contrast can mix up direct effects with spillovers and the policy story becomes blurred (and incredibly difficult to justify).</p><p>The paper gives a way to define and estimate effects that respect our reality. By writing outcomes in terms of own treatment and exposure through neighbours, and by stating a parallel-trends condition that conditions on network position, we get an estimand that matches the question policymakers ask: what changed for treated units, and what changed for those connected to them?</p><p>There is also a precision and credibility gain. The estimator is doubly robust (either the outcome model or the treatment/exposure model can be off and consistency is kept) and it reaches the efficiency bound while letting nuisance pieces be learned with modern ML. That combination is super important when treatment is rare, networks are sparse or irregular and/or the signal is relatively small.</p><p>Finally, it changes how we report results. Instead of one average with vague caveats about &#8220;contamination,&#8221; you can report direct and spillover effects with valid inference under dependence. For environmental regulation, vaccination campaigns, transport investments or school reforms, that is the difference between a credible evaluation and one that misses how benefits spread across the network.</p><p><em>Who should care?</em></p><p>Environmental and health economists studying pollution or disease spread across regions, labour and urban economists dealing with commuting zones and shared markets, education researchers looking at peer or school-network effects, and development economists evaluating geographically clustered programmes. It also matters for applied researchers who suspect &#8220;contamination&#8221; but don&#8217;t want to drop affected observations or pretend it away. If treatments diffuse through networks (physical, social or economic) this framework gives a way to formalise that dependence and still estimate interpretable causal effects. And for methodologists, it extends the efficiency theory of DiD to a frontier problem: identification and inference under interference.<em><br>Do we have code?</em></p><p>No, the paper does not appear to provide a public code repository or package for the estimator itself. The authors mention using other R packages in their application (such as <code>disperseR</code> to calculate the interference matrices and various ML packages like BART and HAL for the nuisance functions) but they do not link to their own code that implements the proposed network-based doubly robust DiD estimator. I will update the post if anything changes.</p><p>In summary, this paper extends our DiD to the networked world. It formalises what &#8220;treatment&#8221; means when outcomes depend on neighbours&#8217; exposure, defines a conditional PTA that accounts for those links and derives an efficient, DR estimator for both direct and spillover effects. This results in a DiD framework that stays valid when SUTVA is frail. It treats interference as part of the design rather than a violation, which lets us measure how policies spread through connected units instead of assuming an isolation that rarely exists.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Some of the papers in this newsletter can be summarized as &#8220;econometricians writing papers using ML methods based on obscure Maths from the early 1900s&#8221;. The Riesz representation is a concept from Maths that says you can always rewrite a linear estimate (e.g., a treatment effect) as an average of predictions with the &#8220;right&#8221; weights.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Partial R^2 measures how much additional variation a variable explains once other covariates are already included in the model.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Not all outcomes behave like &#8220;heights&#8221; or &#8220;incomes&#8221;. Some are categories. If we&#8217;re looking at vote choice, we don&#8217;t measure a number for each person; we record a label (Democrat, Republican, Independent). At the group level we talk about shares across those labels. Those shares are probabilities, they&#8217;re all non-negative and they must add up to one. Think of a fixed pie cut into slices: make one slice bigger and at least one other slice must get smaller. That constraint is always there. There&#8217;s also a difference between categories and ordered categories. Letter grades are ordered (A above B above C), but the jump from B to A is not the same thing as the jump from C to B. If we map A=3, B=2, C=1, we&#8217;re pretending the steps are evenly spaced when they aren&#8217;t. That can push a linear DiD to say things that don&#8217;t make sense for the underlying scale. Standard DiD is built for additive, nearly continuous outcomes where plus/minus changes behave well. With categorical or compositional outcomes, additive changes can mess with the basic rules (negative &#8220;probabilities&#8221;, totals above one) or hide the fact that gains in one category must come from losses elsewhere. The fix is to work with the whole distribution and with changes that respect the &#8220;pie-chart constraint&#8221;. That&#8217;s the problem this paper takes on: how to do DiD when the outcome is a set of category shares rather than a single continuous number.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>As Onil says, his framework is particularly suited to settings with discrete, unordered outcomes (e.g., employment status, voting choice, other health categories) where the policy question goes beyond an average estimate, where the entire composition of categories was reshaped. We can think of labour-market reforms that shift people between employment, unemployment and out-of-labour-force states; or health interventions that move patients across diagnostic categories rather than changing a single health index. Even education and migration policies often work this way: they change who ends up where. In all these cases, what matters is the redistribution of probability mass across categories, i.e., how treatment changes the mix of outcomes beyond their mean. </p></div></div>]]></content:encoded></item><item><title><![CDATA[Missing and Messy]]></title><description><![CDATA[A DiD horror story?]]></description><link>https://www.diddigest.xyz/p/missing-and-messy</link><guid isPermaLink="false">https://www.diddigest.xyz/p/missing-and-messy</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Wed, 08 Oct 2025 11:55:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-7yH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-7yH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-7yH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-7yH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!-7yH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!-7yH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F984fdc28-9877-4b11-90a7-9c461b3db8b8_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hello! We have three new papers to cover :)</p><p>Here they are:</p><ol><li><p><a href="https://arxiv.org/pdf/2509.25009">Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random</a>, by Lorenzo Testa, Edward H. Kennedy and Matthew Reimherr</p></li><li><p><a href="https://arxiv.org/pdf/2508.21536">Triply Robust Panel Estimators</a>, by Susan Athey, Guido Imbens, Zhaonan Qu and  Davide Viviano</p></li><li><p><a href="https://arxiv.org/abs/2509.01829">Cohort-Anchored Robust Inference for Event-Study with Staggered Adoption</a>, by Ziyi Liu</p></li></ol><div><hr></div><p>Before we jump to those papers, I will recommend some DiD applied papers I found (actually went looking for them) online. </p><p><strong>&#8594; In the next post I want to focus on DiD papers by </strong><em><strong>JOB MARKET CANDIDATES</strong></em><strong>, both theory and applied. If you know of someone&#8217;s paper (or you wrote one yourself), please let me know (send me an email at b.gietner@gmail.com) :) &#8592; </strong></p><ul><li><p>Let&#8217;s begin with these ones, but there are many more on NBER:<br><a href="https://www.nber.org/papers/w34304">Maternity Leave Extensions and Gender Gaps: Evidence from an Online Job Platform</a>, by Hanming Fang, Jiayin Hu and Miao Yu (DDD, China, unintended consequences of maternity leave extension on gender gaps in the labour market)</p></li><li><p><a href="https://www.nber.org/papers/w34303">Price and Volume Divergence in China&#8217;s Real Estate Markets: The Role of Local Governments</a>, by Jeffery (Jinfan) Chang, Yuheng Wang, and Wei Xiong (dynamic DiD, China, divergence between price and volume in residential land and new housing transactions across Chinese cities during the Covid-19 pandemic)</p></li><li><p><a href="https://www.nber.org/papers/w34270">Moving for Good: Educational Gains from Leaving Violence Behind</a>, by Mar&#237;a Padilla-Romo and Cecilia Peluffo (stacked DiD, Mexico, effects of moving away from violent environments into safer areas on migrants&#8217; academic achievement)</p></li></ul><div><hr></div><h3>Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random</h3><p><em>(There&#8217;s a considerable amount of stats lingo in this paper, which I tried to explain in the footnotes, but would recommend having this <a href="https://www.stat.berkeley.edu/~stark/SticiGui/Text/gloss.htm">Glossary</a> saved somewhere. Would appreciate if someone could provide a &#8220;Stats lingo for economists&#8221; *wink*)</em></p><h5>TL;DR: when outcomes are missing in panel data, simply discarding those cases can result in a biased sample. This paper sets out a framework for DiD under missing-at-random assumptions and develops estimators that achieve efficiency while remaining robust to some model misspecification.</h5><p><em>What is this paper about?</em></p><p>This paper speaks to all of us who ever had to drop observations in longitudinal/panel data because of incomplete information on covariates and/or attrition<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. This practice, in one way or another, introduces selection bias<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>. How, then, should we conduct DiD analysis when pre-treatment outcome data are missing for a subset of the sample? This paper provides a formal framework for identifying and efficiently estimating the ATT when pre-treatment outcomes are missing at random (MAR)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. In the Appendix, they also extend it to two scenarios: one where all pre-treatment outcomes are observed, but post-treatment outcomes can be MAR, and the other in which outcomes can be missing both before and after treatment.</p><p><em>What do the authors do?</em></p><p>They start with setting up the problem by assuming a 2x2 DiD<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>, then provide sets of assumptions for identifiability (including two different sets of assumptions for the missingness mechanism in the pretreatment outcome<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>, one where baseline outcomes are missing conditional on covariates and treatment, and another where missingness can also depend on the post-treatment outcome) and semiparametric efficiency bounds<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>, which set the benchmark for the lowest possible estimation variance.</p><p>Building on this, they then construct estimators based on influence functions<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> and cross-fitting<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>. These estimators achieve the efficiency bound and are multiply robust, which means consistency is guaranteed if *at least one* of several combinations of nuisance functions<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> is correctly specified. They also develop a way to estimate the &#8220;nested regression&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> needed under the more complex MAR assumption using regression or conditional density methods and show that augmented approaches can recover efficiency even with misspecification.</p><p>Finally, they validate their estimators through a very extensive simulation study. By systematically varying whether nuisance models are correctly or incorrectly specified, they demonstrate that the estimators perform as theory predicts: bias and RMSE<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> are negligible when the robustness conditions are met, and performance deteriorates only when the assumptions are violated.</p><p><em>Why is this important?</em></p><p>Dropping observations/units with missing baseline outcomes is common practice, but it introduces selection bias and can distort treatment effect estimates. This paper shows how to formally identify and estimate the ATT even when pre-treatment or post-treatment outcomes are missing at random. By deriving efficiency bounds, the authors also tell us how well any estimator could possibly do. Their proposed estimators are not only efficient but also multiply robust, meaning they remain consistent as long as certain subsets of nuisance models are specified correctly. This is a strong safeguard for applied work, where model misspecification is almost inevitable</p><p><em>Who should care?</em></p><p>Anyone using DiD with survey or administrative data where attrition, item nonresponse or incomplete records are common, which then includes applied researchers in labour economics, education, and health, as well as methodologists developing new DiD estimators.</p><p><em>Do we have code?</em></p><p>The paper itself does not release replication code but the estimators are influence-function based and can be implemented with standard tools for cross-fitting and doubly robust estimation in R or Python. The design closely parallels DR-Learner and targeted ML routines therefore adapting existing code should be straightforward. The Monte Carlo simulations in the paper provide guidance on implementation.</p><p>In summary, the authors extend the DiD framework to situations where pre or post-treatment outcomes are not fully observed. They derive efficiency bounds that show the best precision researchers can hope for and propose estimators that reach these bounds while retaining robustness if certain models are misspecified. This gives applied researchers a principled way to handle incomplete outcome data, a problem that arises frequently in labour, education and health applications.</p><div><hr></div><h3>Triply Robust Panel Estimators</h3><p><em>(This one is also about panel data, but you should remember some concepts from linear algebra and matrix theory</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a><em>)</em></p><h5>TL;DR: when evaluating policies with panel data, choosing the wrong estimation method can completely change your answer, but there&#8217;s no way to test which method is right. This paper introduces TROP, an estimator that protects you from choosing wrong by combining unit matching, time weighting and outcome modeling, so if any one approach is valid, you get the right answer. It consistently outperforms traditional methods across diverse real-world settings.</h5><p><em>What is this paper about?</em></p><p>The core problem in causal panel data estimation is estimating what would have happened (the counterfactual) to a treated unit had it not been treated, given that unobserved factors likely drive both the outcome and the treatment decision. The challenge with panel data and a binary intervention is not necessarily the lack of causal inference methods but more like the abundance of them. There&#8217;s a myriad to choose from: DiD TWFE, SC, MC, or a hybrid like SDiD, and each comes with its own identifying story (PT, factor models, or pre-treatment fit). Because of this, none of these assumptions can be tested against one another. Choosing between them often feels arbitrary.</p><p>Even when you settle on a method, their weighting &#8220;rules&#8221; raise problems. DiD spreads weight evenly across all periods and units. SC fixes a single set of unit weights and applies them to every post-treatment period. Both approaches ignore that some pre-treatment periods are more predictive than others. Also, SC itself is not built for modern settings where multiple units receive treatment at different times or where treatment can switch on and off. Its balancing logic does not extend naturally to these cases. Messy landscape much?</p><p>The paper steps in with TROP (Triply RObust Panel), an estimator designed to work across these situations rather than being tied to one fragile set of assumptions. Think of it as a general, unifying framework, from which all the others are &#8220;special cases&#8221;.  </p><p><em>What do the authors do?</em></p><p>The key idea of their estimator is to unify existing approaches rather than picking one. TROP builds counterfactuals for treated units by combining three things:</p><ol><li><p>A flexible outcome model &#8211; a low-rank factor structure<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> layered on top of unit and time fixed effects &#8594; this captures broad common trends and latent dynamics that DiD alone would miss.</p></li><li><p>Unit weights &#8211; like in SC, treated units are matched with controls that had similar pre-treatment trajectories, but the weights are &#8220;learned&#8221; from the data and can vary.</p></li><li><p>Time weights &#8211; unlike DiD or SC, TROP doesn&#8217;t treat all pre-periods as equally useful. It can down-weight very old data and put more weight on periods closer to treatment.</p></li></ol><p>All three components are tuned jointly using cross-validation, so the estimator learns from the data which combination best predicts untreated outcomes. Because of this structure, TROP nests existing methods: if you shut down the factors, or set all weights equal, you recover DiD, SC, SDID, or MC as special cases.</p><p>More importantly, this combination gives TROP its superior theoretical guarantee: the property of triple robustness. TROP is asymptotically unbiased<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> if any one of three conditions holds: a) perfect balance over unit loadings, b) perfect balance over time factors, or c) correct specification of the regression adjustment. This structure means the estimator&#8217;s final bias is tightly bounded by the product of the errors in these three components, making it more robust than any existing doubly-robust estimator.</p><p>The authors then put TROP through a series of semi-synthetic simulations calibrated to classic datasets (minimum wage, Basque GDP, German reunification, smoking, boatlift). Across 21 designs, TROP outperforms existing methods in 20 of them. In the one exception (Boatlift treated unit case), standard SC performs better by 24%. They also show why: the ability to combine unit weights, time weights and factor modeling makes it more robust when PT fail, when pre-treatment fit is imperfect or when the assignment is more complex.</p><p><em>Why is this important?</em></p><p>Those of us doing applied work are often stuck with panel data that looks nothing like the textbook case, e.g. interventions don&#8217;t arrive all at once, different units get treated at different times, and PT are hard to justify. In those settings, the choice of estimator can make us confuse, and this conundrum is laid out in the paper. The authors show that DiD might work well in one dataset and fail miserably in another, while SC can look great in some cases and terrible in others. There&#8217;s no single safe bet.</p><p>This is exactly the situation policy evaluations run into. A government wants to know whether to keep, expand, or cut back a programme. The answer depends on estimates that are only as good as the assumptions behind them. If you can&#8217;t be sure whether PT hold, or whether a factor model really captures the dynamics, you&#8217;re taking a gamble.</p><p>Here&#8217;s why TROP is so useful: it doesn&#8217;t force you to pick one story. It combines unit weights, time weights and an outcome model in a way that means if one assumption fails but another holds, the estimator still works. For applied economists, that triple robustness acts as a safeguard when the data don&#8217;t really fit into any single framework<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>.</p><p><em>Who should care?</em></p><p>Applied econometricians and empirical people group should care most, as TROP directly addresses the daily challenges with data complexity. More specifically, anyone working with panel data (for estimating treatment effects, e.g., policy impact, firm-level interventions) with staggered treatment timing or multiple treated units, where traditional DiD doesn&#8217;t work really well. Other researchers who have to do interactive confounding, since TROP provides a defense against the interactive fixed effects (unobserved trends that affect units heterogeneously) that the authors find are both common in real data and fatal to the simpler DiD model. Other individuals involved in high-stakes policy analysis should care (but I don&#8217;t believe this paper will reach them) since TROP &#8220;protects&#8221; against uncertain assumptions. Also data scientists, since they often work with large panel or time-series cross-section data in tech, finance or marketing.</p><p><em>Do we have code?</em></p><p>The paper provides detailed algorithms (Algorithm 1, 2, 3) with step-by-step procedures. While there may not be a published software package, the algorithms are complete enough that implementation is straightforward.</p><p>In summary, TROP is best thought of as a unifying method for causal inference with panel data. Instead of choosing between DiD, SC, or MC, it combines their core ideas: balance units, weight time and model outcomes flexibly. That design gives it a &#8220;triple robustness&#8221;, meaning that if any one of these channels is valid, the estimator delivers. The simulations make clear how unstable results can be if you commit to a single method. For applied economists, the value of TROP is that it consistently performs well across very different settings. It&#8217;s a practical safeguard when you face uncertainty about which assumptions your data can plausibly support.</p><div><hr></div><h3>Cohort-Anchored Robust Inference for Event-Study with Staggered Adoption</h3><p><em>(<a href="https://x.com/liuziyi233">Ziyi</a> is a third year PhD student at UC Berkley, Haas School of Business. Keep an eye out for him! This paper is quite long but it&#8217;s well-written and Ziyi explains the questions really well)</em></p><h5>TL;DR: Ziyi introduces a cohort-anchored framework for robust inference in staggered DiD event-studies, using block bias to replace unreliable aggregated pre-trend checks and yielding confidence sets that stay valid, and more informative, when cohorts differ.</h5><p><em>What is this paper about?</em><br>Event studies with staggered adoption lean on the PTA. Since PT cannot be tested directly, we usually eyeball the pre-trend plot and/or test whether pre-treatment coefficients are jointly zero. But these approaches are quite limited: the tests have low power, and conditioning on them can distort inference. <a href="https://academic.oup.com/restud/article-pdf/90/5/2555/51356029/rdad018.pdf">Rambachan and Roth (2023)</a> proposed a more robust approach: rather than treating PT as all-or-nothing, use the observed pre-trend to bound how much post-treatment bias could plausibly be. This would deliver confidence sets that remain valid even when PT doesn&#8217;t hold &#8220;exactly&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a>.</p><p>That works fine when treatment happens at one point in time. With staggered adoption, things get messier because inference usually relies on event-study coefficients aggregated across cohorts. This gives birth to three issues: 1) relative-period coefficients from TWFE can be contaminated by weighted averages of heterogeneous effects, sometimes with negative weights<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>; 2) cohort composition shifts over relative time, so pre and post-trends are not directly comparable<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a>; and 3) for estimators that use not-yet-treated controls, the control group itself changes as cohorts adopt<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a>.</p><p>Let me illustrate the issue. Let&#8217;s imagine two states raising the minimum wage, but at different times. State A does it in 2005, State B waits until 2007. A standard event-study lines them up by &#8220;time since adoption&#8221; and averages the results. That looks ok on a plot, but it hides a problem: the early pre-period is based only on State B, while the long post-period is based only on State A . If A and B had different pre-trends, then the &#8220;average pre-trend&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> you&#8217;re using as a benchmark has nothing to do with the post-treatment cohorts you care about.</p><p>This is where Ziyi&#8217;s paper comes in. Instead of averaging across everyone, it anchors inference at the cohort level<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a>. Each group of adopters is always compared to the same set of controls (here defined as the units that were untreated when that cohort first adopted)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a>. Ziyi calls the resulting comparison a block<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a> , and the difference in trends between the cohort and its fixed control group is the block bias. Because the block bias has the same definition before and after treatment, the pre-period provides a credible benchmark<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a> for what could happen post-treatment (meaning it has a consistent interpretation across all periods)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a>. The paper shows then how to use this setup to build robust inference<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a> in event-studies even when treatment timing is staggered and cohorts look different from each other.</p><p><em>What does the author do?</em></p><p>Ziyi does a lot. As I said, he develops a cohort-anchored robust inference framework that operates at the cohort&#8211;period level. His framework is designed to address the three aforementioned methodological complications that arise under staggered adoption: negative weighting from TWFE event-studies, shifting cohort composition across relative periods and changing control groups in estimators that use not-yet-treated units. </p><p>The central building block is the block bias (&#916;), which is the difference in trends between a treated cohort and its fixed initial control group. Because the block bias is defined consistently before and after treatment, the observed pre-treatment block biases provide a valid benchmark for bounding the unobserved post-treatment ones. The author then establishes a bias decomposition. The overall estimation bias (&#948;) contaminating any cohort-period treatment effect estimate can be written as an invertible linear transformation of block biases &#948;=W&#916;. Here, each cohort&#8217;s overall bias equals its own block bias plus a weighted sum of block biases from cohorts that adopt later, with the weights determined by relative cohort sizes and adoption timing. This let us impose transparent restrictions directly on the interpretable block biases and then translate those restrictions into bounds on the overall bias. The result is a valid confidence set for the treatment effect.</p><p>He then goes to implement two specific types of restrictions (adapted from Rambachan and Roth, 2023): relative magnitudes (RM), which bounds how much a cohort&#8217;s block bias can change from one period to the next, with the benchmark coming from pre-treatment variation; and second differences (SD), which bounds the change in the slope of the block bias path (it&#8217;s well-suited for settings where pre-trends look approximately linear). These restrictions can be applied globally (using the largest observed pre-trend variation across cohorts) or cohort-specifically (using each cohort&#8217;s own pre-trends).</p><p>His framework is illustrated in two simulation exercises with heterogeneous pre-trends. Compared to the aggregated approach, the cohort-anchored method delivers confidence sets that are better centered on the truth and, in some cases, narrower. He then finishes it by revisiting Callaway and Sant&#8217;Anna&#8217;s (2021) study of minimum wages and teen employment. Under the cohort-anchored framework, the confidence set remains centered well below zero even after accounting for cohort-specific linear pre-trends (which means the negative employment effect is robust). The aggregated approach, by contrast, centers around zero and kinda obscures this conclusion.</p><p><em>Why is this important?</em></p><p>Ziyi&#8217;s paper addresses the issues left unresolved by the major innovations in the DiD literature between 2018 and 2021. Traditional TWFE estimators are invalid under heterogeneous treatment effects. New HTE-robust estimators were developed to fix this, but they introduced complications for the robust inference framework of Rambachan and Roth (2023), which aggregates results across cohorts. The key problems are dynamic treated composition (the average pre-trend is irrelevant for the cohorts driving the post-trends) and dynamic control groups (the definition of parallel trends shifts as adoption unfolds).</p><p>By introducing the concept of block bias, Ziyi provides a coherent way to conduct robust inference in this setting. Block bias solves the dynamic control group problem by anchoring comparisons to a fixed baseline, and the cohort-anchored framework demonstrates that the aggregated approach can distort results when cohorts have heterogeneous pre-trends. The cohort-anchored method instead produces confidence sets that are better centered on the true effect.</p><p><em>Who should care?</em></p><p>Anyone running event-studies with staggered adoption, which includes researchers studying policies that roll out at different times across states, firms or schools, as well as applied economists working with treatment programs that expand gradually. If you rely on modern HTE-robust estimators and want to report confidence sets that remain valid without assuming exact PT, check this paper out. It is especially relevant when pre-trends differ across cohorts since the standard aggregated approach can distort inference in that case.</p><p><em>Do we have code?</em></p><p>Ziyi told me he&#8217;s working on the package that accompanies the paper, and that he hopes he can make it public soon.</p><p>In summary, this paper pushes robust inference in DiD one step further. Rambachan and Roth (2023) showed how to move beyond binary PT by using pre-trends to bound post-treatment violations, but their framework was limited when treatment timing was staggered. Ziyi builds on that insight by introducing the concept of block bias and anchoring inference at the cohort&#8211;period level. The result is a framework that works with modern HTE-robust estimators, avoids the distortions of aggregated event-studies and produces confidence sets that remain valid even when cohorts look very different from each other. The simulations and the minimum wage reanalysis show clearly that this matters: once you anchor inference properly, robust negative employment effects emerge that the standard approach washes out.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For an interesting discussion on this, see <a href="https://www.sciencedirect.com/science/article/pii/S0304407624001295">Bell&#233;go, Benatia and Dortet-Bernadet</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>As the authors point out, &#8220;if the mechanism that causes the data to be missing is related to treatment or other characteristics that influence the outcome, the remaining sample is no longer representative of the population of interest and the resulting estimates will be biased&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Ok, some house cleaning first. This classification was first proposed by <a href="https://academic.oup.com/biomet/article-pdf/63/3/581/756166/63-3-581.pdf">Rubin in 1976</a>. He formalized the three categories (MCAR, MAR, and MNAR) and provided the statistical framework for understanding how different missing data mechanisms affect inference. This work built on earlier ideas but Rubin really crystallized the taxonomy and its implications.</p><ol><li><p>Missing Completely At Random (MCAR): the missingness here has no relationship to any variables in the dataset (observed or unobserved, e.g., a survey page that randomly fails to load for some users due to some technical glitch). It is the least problematic type because the missing data is essentially a random sample of all data.</p></li><li><p>Missing At Random (MAR): here the missingness is related to observed variables, but not to the missing values themselves (e.g., younger people might be less likely to report their income, but among people of the same age, whether income is missing is random). It is &#8220;random&#8221; CONDITIONAL on other observed variables. Not so problematic seeing that most statistical methods can handle MAR if you account for the related variables.</p></li><li><p>Missing Not At Random (MNAR): the missingness is related to the unobserved (missing) values themselves (e.g., people with very high incomes are less likely to report their income specifically because their income is high). MNAR is the most problematic type because the missing data mechanism is related to what you&#8217;re trying to measure, and thus requires special modeling approaches or sensitivity analyses. </p></li></ol></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;We want to emphasize that our framework differs from both the balanced panel data and the repeated cross-section data analyzed by <a href="https://www.sciencedirect.com/science/article/pii/S0304407620301901">Sant&#8217;Anna and Zhao (2020)</a>. The former postulates that each sample is observed both before and after treatment: the latter assumes that each sample is observed either before or after treatment. Instead, our setup mirrors the so-called unbalanced, or partially missing, panel data framework, where the outcomes of some samples are observed both before and after the treatment, while some other samples miss pre-treatment outcomes.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#8220;These assumptions are novel to our setting and are equivalent to a MAR missingness design. Notice that we let the missingness pattern depend on *both* covariates X and the treatment A, and eventually on the post-treatment outcome. In other words, we admit the possibility that pre-treatment outcomes can be missing due to covariates, the treatment that will be administered, and the post-treatment outcome values.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>The (theoretical) lowest possible variance that an estimator can achieve in a given model. It&#8217;s a benchmark to judge whether an estimator is statistically optimal. If your estimator hits this bound, you cannot do better in large samples.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>A mathematical way of describing how each individual observation affects an estimator. Think of it as a linear &#8220;approximation&#8221; that shows the sensitivity of your estimate to each data point. We don&#8217;t need the formula, just the idea that it underpins efficiency and robustness proofs.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>A technique where the sample is split into folds (subsets of the data, e.g. 1,000 observations split into 5 folds gives 5 groups of 200). Nuisance functions (like propensity scores) are estimated on some folds and then evaluated on a different fold. Each observation is used for training and evaluation, but never at the same time. This rotation avoids overfitting and preserves efficiency.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>These are not the main parameter of interest (the ATT), but rather functions that must be estimated from the data to construct the efficient estimator and guarantee its desirable properties. The primary nuisance functions in the paper are:</p><ol><li><p>Regression Functions (&#956;&#8727;): Models for the expected values of the outcomes conditional on other variables (e.g., the expected post-treatment outcome for the control group, E[Y1&#8203;&#8739;X,A=0]).</p></li><li><p>Propensity Score (&#960;&#8727;): The probability of a unit being in the treatment group conditional on their covariates (P[A=1&#8739;X]).</p></li><li><p>Missingness Probability (&#947;&#8727;): The probability of an outcome being observed (not missing) conditional on other observed variables.</p></li><li><p>Nested Regression (&#951;&#8727;): A specialized conditional expectation needed under the most complex missing-at-random assumption.</p></li></ol><p>Correctly modeling just a subset of these functions is what gives the estimators their multiple robustness property.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>Nested Regression (&#951;0&#8727;&#8203;(X,0)): it&#8217;s a specific and complex type of conditional expectation. It&#8217;s called &#8220;nested&#8221; because it involves taking an expected value of an expected value. It&#8217;s needed for the efficient estimator when missingness depends on the post-treatment outcome (Assumption 2.4).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>Root Mean Squared Error, the standard measure combining bias and variance into a single metric. Lower RMSE means better estimator performance.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>The key person bridging advanced matrix theory (specifically the low-rank assumption and PCA) into modern econometrics and setting the stage for its application in both macro (dynamic factor models) and micro (interactive fixed effects) panel settings is Jushan Bai (who&#8217;s now at Columbia). While the foundation for factor models was laid in macro, the move to rigorous estimation of coefficients in panel data - which is the context of this paper - is his central contribution.  </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>The low-rank factor structure is a statistical modeling assumption which says that the observed outcomes of many units over time can be approximated by a small number of unobserved, common factors or trends. It&#8217;s used to simplify and understand complex, high-dimensional data, such as the panel data (many units observed over many time periods) studied in the paper.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>An estimator is asymptotically unbiased if, as the sample size grows infinitely large, the expected value of the estimator converges to the true value of the parameter it is trying to estimate. Unbiased means the estimator&#8217;s average guess is exactly the true value, regardless of sample size, and asymptotically means it might be biased in small samples, but that bias completely vanishes as you get more and more data (i.e., as the number of units N and/or the number of time periods T goes to infinity).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p> However&#8230; researchers should still verify performance in their specific context as the one exception in the simulations shows that no single method dominates universally.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Or, in the authors words, this allows the average treatment effect to be set-identified, and one can construct a confidence set for it.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>Also known as heterogenous treatment effect contamination. This problem can be addressed by HTE-robust estimators (e.g., <a href="https://www.sciencedirect.com/science/article/pii/S030440762030378X">Sun and Abraham, 2021</a>; <a href="https://academic.oup.com/restud/article-pdf/91/6/3253/60441633/rdae007.pdf">Borusyak et al., 2024</a>), which estimate cohort-period level effects. <a href="https://www.aeaweb.org/conference/2024/program/paper/YebdZQ3S">Borusyak and Hull (2024)</a> is a good reading if you&#8217;re worried.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>Even with HTE-robust estimators, the set of treated cohorts contributing to the aggregated coefficients changes across relative periods, and a a result, aggregated pre-treatment and post-treatment coefficients are not directly comparable, since they are based on different treated cohort compositions.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>Aka, dynamic control group problem. This issue arises for estimators using not-yet-treated cohorts as controls (like what Ziyi refers to as the <a href="https://www.sciencedirect.com/science/article/pii/S0304407620303948">CS-NYT estimator</a>) . A shifting control group means the underlying definition of the PTA changes, making it difficult to impose credible restrictions.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>Aggregation can mask cohort-level heterogeneity and produce a distorted benchmark; differences then reflect shifts in cohort composition.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p>The cohort-anchored framework bases inference, as the name suggests, on cohort-period level coefficients.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p>The block bias compares a treated cohort to its fixed initial control group (units untreated when the cohort first adopts).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p>The comparison operates within a block-adoption structure, hence &#8220;block bias&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p>Block biases are &#8220;anchored&#8221; to a fixed control group, enabling observable pre-treatment block biases to serve as valid benchmarks for their post-treatment counterparts.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p>The block bias concept enables robust inference through a bias decomposition. This decomposition shows that the overall estimation bias (&#948;) can be expressed as an invertible linear transformation of the block biases (&#948;=W&#916;). By imposing restrictions on the consistent block bias (&#916;), the framework can translate those bounds to the overall estimation bias (&#948;) and successfully construct a valid confidence set.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p>The entire paper proposes a cohort-anchored framework for robust inference, solving issues like dynamic control group and cohort heterogeneity.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Where Standard Assumptions Go To Die ]]></title><description><![CDATA[Universal treatments without clean controls, continuous policies without untreated groups, reversible reforms over time and statistical inference with handful of cases]]></description><link>https://www.diddigest.xyz/p/where-standard-assumptions-go-to</link><guid isPermaLink="false">https://www.diddigest.xyz/p/where-standard-assumptions-go-to</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 22 Sep 2025 12:47:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kpuj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kpuj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kpuj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kpuj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png" width="1024" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!kpuj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 424w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 848w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!kpuj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1229a10d-d55b-4217-9fdb-e268ed5ad8cf_1024x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! I didn&#8217;t forget about you folks, but things have been pretty quiet so I was pilling up the papers as to have a meaningful post.<br>Let&#8217;s then check the latest ones:</p><p><a href="https://arxiv.org/abs/2407.11937">Factorial Difference-in-Differences</a>, by Yiqing Xu, Anqi Zhao and Peng Ding <em>(professor Yiqing told me this is not new - it was originally uploaded last year - but we will hear about it in soon so I thought about including it here. He was also very kind and shared <a href="https://yiqingxu.org/papers/2025_fdid/fdid_slides.pdf">slides</a> so we could understand the paper better. Thank you, prof! There&#8217;s also a talk available <a href="https://www.youtube.com/watch?v=NzFFlXxkloE">here</a>)</em></p><p><a href="https://arxiv.org/pdf/2405.04465">Difference-in-Differences Estimators When No Unit Remains Untreated</a>, by Cl&#233;ment de Chaisemartin, Diego Ciccia, Xavier D&#8217;Haultf&#339;uille and Felix Knau <em>(this was also first uploaded last year)</em></p><p><a href="https://arxiv.org/abs/2508.07808">Treatment-Effect Estimation in Complex Designs under a Parallel-trends Assumption</a>, by Cl&#233;ment de Chaisemartin and Xavier D&#8217;Haultf&#339;uille <em>(if you&#8217;re in Europe and want learn more DiD from prof D&#8217;Haultf&#339;uille, he will be at the ISEG Winter School in January. You can sign up for it <a href="https://www.iseg.ulisboa.pt/en/event/iseg-winter-school-2026/">here</a>)</em></p><p><a href="https://arxiv.org/pdf/2504.19841">Inference With Few Treated Units</a>, by Luis Alvarez, Bruno Ferman, and Kaspar W&#252;thrich<em> (first draft: April 26, 2025; this draft: June 26, 2025)</em></p><h3>Factorial Difference-in-Differences</h3><h5>TL;DR: what should we do when there&#8217;s no untreated group? Events like famines, wars, or nationwide reforms hit everyone at once, leaving us without the clean control group that canonical DiD relies on. Xu, Zhao and Ding formalise what applied researchers have long been doing in these cases: using baseline differences across units to recover interpretable estimates. They call this Factorial DiD (FDiD). The framework clarifies what the DiD estimator identifies, what extra assumptions are needed to make causal claims, and how canonical DiD emerges as a special case.</h5><p><em>What is this paper about?</em></p><p>This paper is about a situation many of us applied people face: what to do when a big event affects everyone? In canonical DiD we need a treated group and a clean control group. But in practice there are cases where no such control exists (e.g., a famine, a war, or a nationwide policy reform) and no one escapes exposure. What we often do is still run a DiD regression but instead of comparing treated and untreated units, the idea is to interact the event with some baseline factor, which is a characteristic that units already have before the event happens and that doesn&#8217;t change because of it<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. The authors call this approach Factorial DiD<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>, and the name reflects the fact that the research design involves two factors: the event itself (before versus after) and the baseline factor (e.g., high versus low social capital). </p><p>The same DiD estimator is used, but the <em>research design</em> is different. The estimator no longer identifies the standard treatment effect, but instead it captures how the event&#8217;s impact differs across groups with different baseline factors, a quantity they call effect modification. The authors go further and show that moving from this descriptive contrast to a causal claim about the baseline factor (what they call causal moderation) requires stronger assumptions. They also show that canonical DiD is a special case of FDiD, but only if you impose an exclusion restriction assuming some group is exposed yet unaffected by the event. This paper gives a language, a framework and clear identification results for a class of designs that are already widely used but rarely well defined in the methodological literature.</p><p><em>What do the authors do?</em></p><p>They begin with the simplest possible case: 2 groups, 2 periods + one event that hits everyone. They show that with universal exposure + no anticipation + the usual PT, the DiD estimator doesn&#8217;t give us the &#8220;standard&#8221; treatment effect but instead identifies effect modification, which is the descriptive contrast in how much the event changes outcomes for high-baseline versus low-baseline units (similar to the idea of heterogeneous treatment effects). They then introduce three related quantities<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>: </p><p>1) Causal moderation: the stronger claim that the baseline factor itself changes how the event matters (e.g., that social capital causes the famine to be less deadly);</p><p>2) An ATT analogue: the average effect of the event on the high-baseline group (which looks like the &#8220;treated group effect&#8221; in canonical DiD);</p><p>3) And the baseline factor&#8217;s effect given exposure: how the baseline characteristic matters once the event has happened. </p><p>The key contribution is to line up which assumptions are needed for each one. With only canonical DiD assumptions, you can get effect modification. To move to causal moderation, you need their new factorial PTA. Canonical DiD then emerges as a special case if you add an exclusion restriction that one group is exposed but unaffected. A different exclusion restriction, assuming the baseline factor has no impact in the absence of the event, links the design to the baseline-given-exposure quantity.</p><p>After working through these four cases, they move on to show how the framework can be stretched further. The authors show how to add baseline covariates, work with repeated cross-sections rather than panels and extend to continuous baseline factors. They also clarify what this means for practice: when TWFE regressions with interactions are &#8220;coherent&#8221; and when they are not. The paper closes by revisiting the famine and social capital example, showing how the framework applies to a real case that has already been studied with this kind of design.</p><p><em>Why is this important?</em></p><p>I think giving names to things matter. We&#8217;ve been doing FDiD, we just don&#8217;t call it that. We run a TWFE regression with an interaction term, call it DiD and then interpret the coefficient as if it were the standard ATT. The problem is that the target estimand is *not* the ATT. </p><p>This paper makes the implicit explicit. It 1) names the design and distinguishes it from canonical DiD, 2) spells out what the estimator is actually identifying under different assumptions, and 3) clarifies what extra assumptions we need if we want to go beyond descriptive contrasts. This obviously matters in practice. It gives us a diagnostic tool by stating plainly which estimand we are after and which assumptions justify it. I think it also sort of prevents over-claiming. With only canonical PT, we&#8217;re talking about effect modification. To get causal moderation, we need factorial PT. To recover an ATT-like interpretation, we need an exclusion restriction. </p><p>This paper also ties regression practice back to design. TWFE with interactions is ok if the assumptions are &#8220;right&#8221;; if not, we are not estimating what we think we are. And because the authors extend the framework to covariates, repeated cross-sections and continuous baseline factors, their approach covers the kinds of data we actually use.</p><p>At a broader level, the contribution is about transparency. FDiD gives us language that makes clear to authors, referees (and seminar audiences) what is being identified  and under what conditions. It &#8220;builds a bridge&#8221; between factorial designs in stats and DiD in econ, and in doing so makes a very common empirical strategy much easier to defend (and critique).</p><p><em>Who should care?</em></p><p>Anyone in the social sciences doing causal work. Applied people<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> working on labour, development, education, political economy, trade or even finance will recognise the setup (wars, famines, nationwide reforms, regulatory changes or financial crises all fit the bill). PolSci dealing with regime shifts, revolutions and/or national elections face the same issue. I think it&#8217;s also useful for both referees, editors, authors and us, graduate students, learning causal inference.</p><p><em>Do we have code?</em></p><p>No, for many reasons. FDiD is not a new estimator, but a design. The DiD estimator is standard, what changes is the estimand and assumptions. Implementation here is context-specific. Which quantity you estimate (effect modification, causal moderation, ATT-analogue, baseline-given-exposure) depends on which assumptions you&#8217;re willing to make. A &#8220;one-click&#8221; function would hide that choice. The same regression (TWFE with an interaction or &#8710;Y on G) can map to different targets under different designs. Packaging it would risk misinterpretation. Extensions vary (covariates, repeated cross-sections, continuous baseline factor). Too many branches for a single simple API. If you want a practical recipe, it&#8217;s just: 1) difference outcomes, 2) run &#8710;Y on baseline-group indicator(s) with covariates as needed, and 3) interpret the coefficient as effect modification under standard PT (add stronger assumptions if you want causal moderation or the other targets).</p><p>In summary, FDiD gives a name and a framework to something we were already doing when there was no untreated group. The estimator is the same as in canonical DiD but the design is different (and so is the interpretation). What you get by default is effect modification, which is a descriptive contrast in how the event matters across baseline groups. To turn that into causal moderation, or to recover an ATT-like interpretation, you need extra assumptions. The paper&#8217;s value is in making these distinctions clear, mapping assumptions to estimands and also showing how regression practice fits within the framework. For anyone who has ever interacted an event with a baseline factor and called it DiD, please take a moment and think about rewriting it a bit, hehe.</p><h3>Difference-in-Differences Estimators When No Unit Remains Untreated</h3><h5>TL;DR: what if a policy affects everyone, but by different amounts? Minimum wage hikes, China&#8217;s WTO entry, Medicare Part D, etc, no group is left fully untreated. The authors call this a heterogeneous adoption design with no untreated group. Standard DiD fails here and TWFE with treatment intensity can mislead. The paper shows how to recover effects if quasi-untreated units exist, proves impossibility without them, and offers minimal assumptions for partial identification. It also gives tests for when TWFE is &#8220;defensible&#8221;, with code in Stata and R.</h5><p><em>What is this paper about?</em></p><p>What happens when a policy affects everyone, but not in the same way? A <a href="https://academic.oup.com/qje/advance-article-pdf/doi/10.1093/qje/qjab028/41929667/qjab028.pdf">nationwide minimum wage hike</a> raises pay everywhere, yet the bite is bigger in some regions than others. <a href="https://www.nber.org/system/files/working_papers/w18655/w18655.pdf">China&#8217;s entry into the WTO</a> reduced tariff uncertainty for every U.S. industry, but some faced larger potential tariff spikes. <a href="https://www.nber.org/system/files/working_papers/w13917/w13917.pdf">Medicare Part D</a> applied to all drugs, but exposure varied with each drug&#8217;s Medicare market share. </p><p>In each of these examples, there&#8217;s no group left *entirely* untouched by the policy, every unit is affected to some degree. The authors call this setup a heterogeneous adoption design (HAD) with no untreated group. It turns out this situation is not rare at all. Many policy changes are universal by design and even when a small untreated group exists we often exclude it (drop if) because it looks too different from the rest of the sample. The problem is that standard DiD relies on having a control group that stays at zero treatment. If everyone is affected, that assumption collapses. And the default fix in applied work (a TWFE regression with treatment intensity as the regressor) doesn&#8217;t solve the problem. As the authors show, TWFE can give misleading results in this setting when treatment effects vary across units. </p><p>This paper then &#8220;attacks&#8221; the issue directly. It shows when effects are identified and when they are not. With quasi-untreated units (exposure near zero), they identify a weighted average of slopes (WAS) using a local-linear, optimal-bandwidth estimator adapted from RDD, with valid bias-corrected inference. When no quasi-untreated units exist, they prove an impossibility result for WAS without extra assumptions; they then give minimal assumptions to learn the sign of the effect or an alternative parameter relative to the lowest observed dose. Because TWFE is widely used, they also supply tests: a tuning-free test for the existence of quasi-untreated units and a nonparametric Stute test of the TWFE linearity restriction, alongside pre-trend checks.</p><p><em>What do the authors do?</em></p><p>First the authors set up the framework. They define what they call a heterogeneous adoption design (HAD) with no untreated group: every unit is untreated in period one and in period two everyone gets some positive treatment dose, but the size of that dose differs across units. They show why TWFE regression, which regresses outcome changes on treatment intensity, can be misleading in this setup when effects differ across units.</p><p>They then ask: what can we still learn? Their first target is a parameter they call the weighted average of slopes (WAS), which weights each unit&#8217;s slope (change in potential outcomes between zero and actual treatment) by the unit&#8217;s treatment dose relative to the average dose. If there is a set of quasi-untreated units (those with exposure close to zero), WAS can be identified by comparing outcome changes in the full sample to outcome changes in this group. To estimate it, they import tools from the RD literature (local linear regression at the boundary + optimal bandwidth choice + bias-corrected confidence intervals). </p><p>But what if even quasi-untreated units don&#8217;t exist? Here the paper proves a complete impossibility result: without extra assumptions, WAS has an identification set that is the entire real line - meaning the data can rationalize any value of WAS. To make progress, the authors then propose 2 minimal routes forward. One option is to assume enough structure to at least pin down the sign of WAS, by comparing the least-treated units to the rest. Another is to define a slightly different parameter, WAS_d&#8203;, which compares outcomes to the lowest observed dose rather than to zero. Under an additional assumption, that parameter can be identified and estimated.</p><p>Because the whole strategy depends on whether a quasi-untreated group exists, they also build a tuning-free statistical test for its presence based on simple order statistics of the treatment variable.</p><p>After laying out these new estimators, they circle back to TWFE. Since we often use it anyway, they ask: under what conditions could TWFE actually be valid here? They show that if treatment effects are homogeneous and linear, then TWFE delivers the average slope. This leads to a key insight: TWFE implies that the conditional mean of outcome changes is linear in treatment dose, and this is testable. They adapt a nonparametric specification test (the Stute test) to check this linearity. Combined with standard pre-trend tests and their quasi-untreated test, this gives a three-part testing procedure: if all tests pass, TWFE may be valid. For very large samples, they also provide a faster alternative (the yatchew_test).</p><p>Finally, they show the methods in action. In the case of bonus depreciation in the U.S. (2002), they find positive employment effects (often larger than the original TWFE estimates) with pre-trends largely consistent. In the case of China&#8217;s PNTR entry, their nonparametric estimates are noisy and mostly insignificant. However, when they control for industry-specific linear trends, the pre-trend and linearity tests are no longer rejected, making TWFE with trends defensible (which points to negative employment effects).</p><p><em>Why is this important?</em></p><p>Many real-world policies look exactly like this. Minimum wage hikes, trade reforms, new health programmes&#8230; these are designed to be universal, and they end up affecting everyone to some extent. Untreated groups either don&#8217;t exist or they&#8217;re so small and different that we usually drop them.</p><p>That means the standard DiD design doesn&#8217;t apply. Without a zero-treatment group, the whole logic of comparing treated and untreated units makes no sense. Yet in practice, we often still run a TWFE regression with treatment intensity as the regressor, hoping it will stand in for DiD. The problem is that with heterogeneous effects, TWFE can produce misleading results.</p><p>This paper matters because it tells us exactly what is and isn&#8217;t identifiable in these designs. If quasi-untreated units exist, treatment effects can still be recovered in a transparent way. If they don&#8217;t, the paper shows that the data itself can&#8217;t tell you the answer (an impossibility result that sets a clear boundary for empirical work).</p><p>The authors also bring new tools to the table. Their nonparametric estimators adapt ideas from RD to this setting, and their tuning-free tests give us a way to check assumptions rather than relying on hope (:P). At the same time, instead of throwing TWFE in the bin, they provide guidance on when it might still be valid, and a concrete test for its key assumption.</p><p>Also the parameters they focus on (WAS and WAS_s) have direct policy meaning. They tell us whether a universal reform was beneficial compared to no reform, or compared to the minimal observed treatment, which ties the methods back to cost&#8211;benefit analysis.</p><p><em>Who should care?</em></p><p>First and foremost, applied microeconomists who study policies that are universal but uneven in their bite (labour market reforms, trade liberalisation, health programmes, education policies). These are exactly the settings where untreated groups don&#8217;t exist, yet exposure varies. Policy evaluators and government analysts should also pay attention. When reforms roll out nationwide, the usual evaluation strategies no longer apply. Having a framework that makes clear what can and can&#8217;t be identified is very important if you&#8217;re tasked with producing credible estimates for policy decisions. Researchers who rely on TWFE in continuous-treatment settings also have a lot at stake. The paper doesn&#8217;t say &#8220;don&#8217;t use TWFE&#8221;. The authors rather show when it is misleading and when, with the right tests, it might still be defensible. That&#8217;s a message many of us need to hear. And also methodologists and students who are expanding the DiD toolkit will find this work useful. </p><p><em>Do we have code?</em></p><p>Yes, lots. The did_had package (available in both <a href="https://ideas.repec.org/c/boc/bocode/s459331.html">Stata</a> and <a href="https://www.rdocumentation.org/packages/DIDHAD/versions/1.0.1/topics/did_had">R</a>) estimates the WAS and can also generate pre-trend and event-study estimators. It implements the nonparametric local-linear estimator with optimal bandwidth selection and bias-corrected confidence intervals, borrowing routines from the <a href="https://cran.r-project.org/web/packages/nprobust/index.html">nprobust</a> package. For checking the assumptions behind TWFE, they provide the stute_test package (also in <a href="https://ideas.repec.org/c/boc/bocode/s459349.html">Stata</a> and <a href="https://cran.r-project.org/web/packages/StuteTest/StuteTest.pdf">R</a>). This runs the nonparametric specification test of whether outcome changes are linear in treatment dose using a wild-bootstrap implementation. It works fast for moderate datasets, though memory limits kick in as the sample size grows large. For very large datasets, the authors also offer <a href="https://github.com/chaisemartinPackages/yatchew_test">yatchew_test</a>, a faster alternative that scales up to tens of millions of observations.</p><p>In summary, this paper shows what to do when every unit is treated. With quasi-untreated units, treatment effects can still be identified using a new nonparametric estimator. Without them, effects are not identified unless you add assumptions (an impossibility result that sets clear limits). The authors also develop tests to decide when TWFE is valid and provide ready-to-use code. For universal reforms, this is now the reference point for what we can and can&#8217;t infer and learn.</p><h3>Treatment-Effect Estimation in Complex Designs under a Parallel-trends Assumption</h3><h5>TL;DR: what to do when treatments are non-binary, reversible or vary at the baseline? This paper introduces AVSQ event-study effects, which compare actual outcomes to a status-quo where each unit keeps its initial treatment. The framework makes it possible to evaluate dynamic and multi-dose policies, while offering a cost&#8211;benefit measure (ACE), and shows why distributed-lag TWFE regressions can mislead. The authors also provide software (<code>did_multiplegt_dyn</code>) and illustrate the method with US newspapers and turnout.</h5><p><em>What is this paper about?</em></p><p>In this paper the authors look at how to estimate treatment effects with DiD when the treatment is more complex than a one-time, binary switch. Many policies can vary in intensity, be reversed, or differ in their baseline levels across units. The usual staggered adoption design doesn&#8217;t capture these cases, even though they are common in applied work. The authors propose a way forward by comparing outcomes under the actual treatment path to a counterfactual &#8220;status-quo path&#8221; where each unit simply keeps its initial treatment. The resulting actual-versus-status-quo (AVSQ) effects generalise the ATT from staggered designs and provide a framework for analysing dynamic treatments beyond the standard setup.</p><p><em>What do the authors do?</em></p><p>The authors begin by setting up a framework for complex DiD designs, where treatments are not just binary switches but can vary in intensity, be reversed, or differ across units at baseline. To deal with this, they propose comparing each unit&#8217;s actual treatment path to a counterfactual &#8220;status-quo path&#8221; where the unit simply keeps its initial treatment level. The difference between these two paths defines what they call AVSQ event-study effects, which extend the familiar ATT from staggered adoption designs.</p><p>They then show how these AVSQ effects can be identified under two assumptions: no anticipation and PT conditional on baseline treatment. This conditional PTA is very important because it requires that units with the same initial treatment level have similar outcome trends in the absence of treatment changes, but allows units with different baseline treatments to have different trends. Because raw AVSQ effects can be difficult to interpret, they also introduce a normalised version that rescales by the change in treatment dose, turning them into weighted averages of marginal effects of current and lagged treatments. The AVSQ framework also includes a &#8220;no-crossing&#8221; condition to ensure interpretable results (this means a unit&#8217;s treatment path doesn&#8217;t zigzag above and below its initial level, which would make it unclear whether treatment increases or decreases are driving the effects). On top of this, they develop a statistic called the Average Cumulative Effect (ACE), which aggregates AVSQ effects into a cost&#8211;benefit style measure of whether the realised treatment path improved outcomes per unit of treatment compared to the status quo.</p><p>Beyond defining these new estimands, the authors take a closer look at what happens when researchers run distributed-lag TWFE regressions in these settings. They show that TWFE coefficients do not reliably estimate causal effects when treatment effects are heterogeneous: they end up mixing together different lags and in some cases can even reverse sign. To provide a cleaner alternative, they introduce a random-coefficients distributed-lag model which assumes that treatment effects are constant over time but can vary across units. Under this model they show how to identify and estimate average effects of current and lagged treatments. Estimation is straightforward when treatments take on discrete values, but becomes more complex with continuous treatments which require truncation procedures to handle cases where some units have extreme influence on the estimates.</p><p>At the end, they bring these ideas to an empirical application. Revisiting <a href="https://www.nber.org/system/files/working_papers/w15544/w15544.pdf">Gentzkow, Shapiro, and Sinkinson</a>&#8217;s study of newspapers and voter turnout in the United States, they extend the original analysis by allowing for dynamic effects of current and past newspaper exposure. Their AVSQ estimates show that newspapers increased turnout and that the effect of current exposure was larger than the effect of lagged exposure. A conventional TWFE regression suggests the opposite, which illustrates the practical importance of their framework.</p><p><em>Why is this important?</em></p><p>Many policies that we care about don&#8217;t look like simple one-off treatments. In Kenya, whether a district shares the president&#8217;s ethnicity changes with each election, so exposure to favouritism in public spending can <a href="https://www.nber.org/system/files/working_papers/w19398/w19398.pdf">switch on and off</a>. In the US, states deregulated financial markets during the 1990&#8217;s at <a href="https://www.jstor.org/stable/pdf/43495408.pdf">different times and with different intensities</a> and some even made multiple changes. In Germany, municipalities have local business tax rates that <a href="https://www.jstor.org/stable/pdf/26527909.pdf">vary across space from the start and move up or down over time</a>. Even the number of newspapers in circulation across US counties (which is a continuous measure) has long been used to study voter turnout.</p><p>None of these cases fit perfectly into a staggered binary design. If we force them into that framework or fall back on TWFE regressions, we risk estimates that don&#8217;t mean what we think they mean. This paper is important because it provides a framework that lets us evaluate these policies as they actually happened. The AVSQ approach measures the effect of departing from the status quo baseline, the normalisation step makes the results interpretable in terms of current and lagged effects, and the ACE parameter links the findings to a cost&#8211;benefit view of policy.</p><p>It also matters that the paper shows how and why distributed-lag TWFE can give the wrong answer in these settings. It&#8217;s more than just a theoretical point: in their application to newspapers and turnout, the AVSQ framework points to strong effects from current exposure, while TWFE suggests the opposite. For applied researchers, the message is clear: dynamic and non-binary treatments need methods designed for them, not workarounds.</p><p><em>Who should care?</em></p><p>Anyone working with DiD in situations where treatment is not a clean one-off adoption, which includes researchers studying policies that can be introduced and repealed, vary in strength or differ across units from the very beginning. As the authors point out, examples range from political economy questions like ethnic alignment with leaders, to local fiscal policy where tax rates shift up and down, to historical work on newspapers and turnout. It also matters for people still relying on distributed-lag TWFE regressions to study dynamic effects. The paper shows those regressions can give sign-reversed estimates when effects differ across units. If your design involves intensity, reversals, or lagged exposure, this framework offers more credible alternatives.</p><p>In short, applied researchers in labour, public finance, political economy and economic history all have reasons to pay attention (basically anywhere policies don&#8217;t come as simple binary treatments).</p><p><em>Do we have code?</em><br>We have commands for both Stata and R, called <a href="https://github.com/chaisemartinPackages/did_multiplegt_dyn">did_multiplegt_dyn</a>, which compute the AVSQ event-study effects and related estimators. These commands also implement the placebo tests and the diagnostics for dynamic effects. The paper itself is theory-heavy, but the software is designed to be practical for applied work.</p><p>In summary, this paper shows how to use DiD when treatments don&#8217;t look like neat one-time switches. The AVSQ approach compares what actually happened to a status-quo where nothing changed and the normalised version tells us how current and past exposure matter. The ACE parameter turns this into a cost-benefit measure. The authors also explain why standard distributed-lag TWFE regressions can mislead and give a simple alternative. In their application to newspapers and turnout, the new method finds current exposure matters more, while TWFE says the opposite.</p><h3>Inference With Few Treated Units</h3><h5>TL;DR: what if only one or a few units get treated, like a single country&#8217;s reform or a couple of states changing policy? Clustered or robust errors break down because they assume many treated clusters and badly understate uncertainty. This paper surveys the fixes: borrowing from controls, pre-trends, or randomisation inference. Each comes with trade-offs. The key point is that the count of treated units drives inference, and much of what we see in practice is invalid.</h5><p><em>What is this paper about?</em></p><p>This paper is about a situation many applied researchers find themselves in: what if your treatment only applies to one unit, or just a handful of them? Think of German reunification (only one country &#8220;treated&#8221;) or a state-level reform adopted by 2 or 3 states. In these cases, the usual inference machinery is a bit more tricky and difficult to justify. We know from practice that standard cluster-robust or heteroskedasticity-robust errors rely on having many treated clusters. With only a few, the treated units have outsized leverage, standard errors are biased down and rejection rates can be way off. What the authors do is to step back and survey the growing literature on this problem. They classify methods into two big &#8220;families&#8221;:</p><ul><li><p>Model-based inference, which treats uncertainty as coming from potential outcomes (sampling from a super-population); and</p></li><li><p>Design-based inference, which treats uncertainty as coming from random assignment of treatment.</p></li></ul><p>They then go through the different approaches available when treated units are scarce. Some methods lean on cross-sectional variation (borrowing information from controls), others on time-series variation (borrowing from pre-treatment periods). Each approach makes trade-offs: cross-sectional methods require many controls but allow unrestricted serial correlation; time-series methods require many untreated periods but allow cross-sectional dependence. They also examine extreme cases, like when there&#8217;s *literally* only one treated unit and one treated period. Here, inference is only possible under strong assumptions, like assuming treatment effects are homogeneous or restricting how errors behave. As more treated units, periods, or within-cluster observations become available, assumptions can be relaxed somewhat, but the problem never fully goes away.</p><p>The main message is that the inferential challenge is about the number of treated units. Even if you have a thousand controls, one treated unit means asymptotics won&#8217;t save you. The survey pulls together existing fixes (wild bootstrap variants, randomization inference, conformal methods, etc.), shows how some heuristics can be justified theoretically and proposes small modifications to improve finite-sample performance.</p><p>This is a good paper that provides us with a map of what can and can&#8217;t (shouldn&#8217;t?) be done when &#8220;few treated&#8221; is the reality, and links econometrics back to the practical settings (DiD, SC, matching) where these problems show up most often.</p><p><em>What do the authors do?</em></p><p>They begin with a simple example: estimating a treatment effect when only one unit is treated. With just one treated unit, heteroskedasticity-robust standard errors collapse (what variance would be left to estimate anyway?) and cluster-robust errors in a DiD setting massively understate uncertainty. They use this to show why conventional inference fails when the number of treated units is small (regardless of how many controls you have).</p><p>From there, they organise the literature into categories:</p><ol><li><p>Cross-sectional approaches: methods like <a href="https://direct.mit.edu/rest/article-pdf/93/1/113/1919185/rest_a_00049.pdf">Conley and Taber (2011)</a>, which assume the error distribution for treated and control units is comparable and typically require homogeneous treatment effects or strong distributional assumptions. With many controls you can use the control residuals to approximate the treated distribution. Other refinements handle heteroskedasticity (e.g. state size differences) or variation in treatment timing.</p></li><li><p>Time-series approaches: methods that rely on long pre-treatment periods and require stationarity and weak dependence assumptions on the errors. Here the idea is to compare the treated unit&#8217;s residuals after treatment to its residuals before treatment, testing whether the last observation looks unusually large. These require assumptions like stationarity and weak dependence but allow for cross-sectional correlation.</p></li><li><p>Allowing for stochastic treatment effect heterogeneity: they discuss when inference can be framed around testing sharp nulls (no effect whatsoever) or around realised treatment effects conditional on shocks. This shifts the interpretation of what we&#8217;re learning but opens up additional strategies.</p></li><li><p>Design-based approaches: when treatment is randomly assigned but only to a handful of units, the focus shifts. They show how randomisation inference can still be valid, but imbalances between treated and controls become much more likely with small treated groups.</p></li><li><p>Sharp null testing: many methods test whether there is no effect whatsoever (sharp nulls) rather than testing average treatment effects, which changes how results should be interpreted.</p></li></ol><p>Along the way, they show connections between methods (e.g., how a wild bootstrap with the null imposed can be asymptotically equivalent to a randomisation test) and they propose tweaks that improve finite-sample performance.</p><p>The unifying theme is that each method trades one type of assumption for another. Cross-sectional methods need many controls and comparability. Time-series methods need long pre-trends and stationarity. Design-based methods hinge on how treatment was assigned. The survey doesn&#8217;t offer a &#8220;best&#8221; choice, but it does clarify what&#8217;s available, what each method buys you and where each one fails. </p><p><em>Why is this important?</em></p><p>I&#8217;m preaching to the choir here, but situations like these are everywhere. Many influential DiD and SC papers study events like a single country&#8217;s reform, one state-level policy or a handful of treated clusters. Standard inference procedures underestimate uncertainty because they rely on asymptotic theory that fails when the number of treated units is small, regardless of the total sample size. They underestimate uncertainty and can lead to massive over-rejection.</p><p>This paper gives us both structure and clarity on a messy corner of applied practice. Instead of a patchwork of one-off fixes, it provides a taxonomy: which methods work when you only have one treated unit, when you have a few, when you have many controls or when you have long pre-treatment histories. It also explains the assumptions you&#8217;re implicitly making when you pick one approach over another.</p><p>There&#8217;s also a bigger point: forget about the share of treated units, consider  the count. Even if 1% of units are treated out of a thousand, that&#8217;s still just ten treated units. Asymptotics won&#8217;t serve much. Knowing this helps us avoid the false comfort of &#8220;large N&#8221; when what actually matters is the number of treated clusters.</p><p>The survey also connects different strands of work to demonstrate that many of the inferential problems are variations of the same underlying issue. For practitioners, that means better tools and clearer diagnostics. For theorists, it ties together a set of results that were previously all over the place. It also reveals that methods designed for single treated units may be more powerful than methods requiring multiple treated units, even when more treated units are available, which is an important practical consideration for researchers.</p><p><em>Who should care?</em></p><p>Applied researchers working with one or a few treated units (country case studies, state-level reforms or costly RCTs). Referees and editors since many papers still report invalid clustered errors in these settings.</p><p><em>Do we have code?</em></p><p>Code doesn&#8217;t make sense here.</p><p>In summary, this paper reviews what to do when only a handful of units are treated. Standard inference fails in these cases, and the number of treated units - not their share - drives the problem. The authors map out model-based, time-series, cross-sectional, and design-based approaches, show how existing heuristics can be justified and suggest tweaks to improve small-sample performance. </p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>It took me a good half an hour to understand this because the justification for what classifies as a baseline factor is context-dependent. How do we know if something is a baseline factor? Apparently the simplest rules are: 1) it must already be in place before the event, and 2) it must not *plausibly* (oh no my identification assumption!) be changed by the event. We should be able to look at it and say: this is a characteristic of the unit that existed beforehand and won&#8217;t flip suddenly just because the event occurred. From what I understand, there are three practical checks and they relate to timing, persistence and isolation (to a certain extent). Timing refers to: was the factor measured before the event? If not, it can&#8217;t be baseline. Persistence relates to the factor being something like slow-moving or fixed (e.g., geography, long-standing institutions or cultural traits). And finally, isolation: can the event directly or indirectly alter it in the short run? If yes, it&#8217;s not safe to treat it as baseline. Going back to what the authors use as an example, social capital measured in the 1950s works in a famine study because kinship networks don&#8217;t suddenly disappear during a few &#8220;bad years&#8221;. But soil quality would not work if the &#8220;event&#8221; were a soil restoration reform, since the policy directly targets the factor itself.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>&#8220;Factorial&#8221; comes from stats. <a href="https://home.iitk.ac.in/~shalab/anova/DOE-RAF.pdf">Fisher</a> (also <a href="https://www.tandfonline.com/doi/abs/10.1080/00031305.1980.10482701">here</a>) and <a href="https://repository.rothamsted.ac.uk/item/98765/the-design-and-analysis-of-factorial-experiments">Yates</a> developed factorial experiments to study the main effects of two factors and their interaction. FDiD borrows from this logic: treat the event (pre/post) and a baseline factor (e.g., high/low social capital) as the two factors, then interpret what DiD identifies under stated assumptions. For background, the authors recommend these: <a href="https://journals.lww.com/epidem/fulltext/2009/11000/on_the_distinction_between_interaction_and_effect.16.aspx">VanderWeele (2009)</a>; <a href="https://ideas.repec.org/a/bla/jorssb/v77y2015i4p727-753.html">Dasgupta et al (2015)</a>; <a href="https://academic.oup.com/jrsssa/article/184/1/65/7056364">Bansak (2020)</a>; <a href="https://arxiv.org/abs/2101.02400">Zhao and Ding (2021)</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>Effect modification: identified with the usual DiD assumptions: universal exposure, no anticipation, and canonical parallel trends. This is descriptive, not causal. Causal moderation: to interpret the contrast as causal, you need a stronger assumption, which the authors call <em>factorial PT</em>. ATT analogue: if you add an exclusion restriction that one group is exposed but unaffected, then the design &#8220;collapses&#8221; to the familiar ATT interpretation from canonical DiD. Baseline effect given exposure: if instead you assume the baseline factor has no effect in the absence of the event, then the estimator captures how much the baseline factor matters once the event has happened.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>In the Concluding Remarks, the authors make 3 recommendations for applied researchers:</p><ol><li><p>Be explicit about the design. When using FDID, say so. Don&#8217;t call it a standard DiD, because the estimand is different.</p></li><li><p>State clearly what assumptions justify your claim. With only canonical PT, you can talk about effect modification; to claim causal moderation, you need factorial PT; for ATT-like interpretations, you need an exclusion restriction.</p></li><li><p>Check robustness with alternative assumptions and specifications. For example, examine whether your baseline factor is truly unaffected by the event, and test sensitivity to different ways of grouping or measuring it.</p></li></ol><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Parallel Trends… Conditionally: Covariates, Violations, and More Robust DiD Methods]]></title><description><![CDATA[The devil in the ~DiD details~]]></description><link>https://www.diddigest.xyz/p/parallel-trends-conditionally-covariates</link><guid isPermaLink="false">https://www.diddigest.xyz/p/parallel-trends-conditionally-covariates</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 11 Aug 2025 16:33:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!frLA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!frLA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!frLA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!frLA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!frLA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!frLA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!frLA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!frLA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!frLA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!frLA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!frLA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fe1bc44-5d96-4641-926b-02df80700db9_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! Following a friend&#8217;s suggestion, I decided to add some quick explanations of concepts that eventually show up in the main text. I don&#8217;t want to overcrowd it, so they will appear as footnotes whenever it&#8217;s necessary. I hope this will make things clearer :) For example, the first paper talks about the covariate balancing propensity score, which is a method for estimating propensity scores in a way that automatically balances covariates between treated and control groups (rather than estimating the score first and then checking balance later). I&#8217;ll no longer assume everyone is already familiar with these &#8220;specifics&#8221;. Any other suggestions are welcome, just send me an email (b[dot]gietner[at]gmail[dot]com) or comment in the comment box at the end of the post.</p><p>Let&#8217;s goooo then. Here are the papers:</p><ul><li><p><a href="https://arxiv.org/abs/2508.02097">A difference-in-differences estimator by covariate balancing propensity score</a>, by Junjie Li and Yukitoshi Matsushita </p></li><li><p><a href="https://arxiv.org/abs/2412.14447">Good Controls Gone Bad: Difference-in-Differences with Covariates</a>, by Sunny Karim and Matthew D. Webb</p></li><li><p><a href="https://arxiv.org/abs/2508.02970">Bayesian Sensitivity Analyses for Policy Evaluation with Difference-in-Differences under Violations of Parallel Trends</a>, by Seong Woo Han, Nandita Mitra, Gary Hettinger and Arman Oganisian</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5375017">The Labor Market Effects of Generative AI: A Difference-in-Differences Analysis of AI Exposure</a>, by Andrew Johnston and Christos Makridis <em>(this is an applied paper that I judged to be super relevant to the current discussion we are having regarding the role of AI on the labour market: is it good, is it bad, does it generate consumer surplus, who&#8217;s benefitting from it, who&#8217;s being left out, is the generative AI revolution fundamentally different from past technological shifts (like the industrial revolution or the rise of computers), if generative AI is so powerful, why aren't we seeing a massive, economy-wide productivity boom? etc. The narratives are swinging between utopian optimism and dystopian fear, and it doesn&#8217;t need to be like that. I think you should check it out for a couple of reasons: it provides evidence that AI appears to be both creating and destroying opportunities, but in different parts of the labour market; and it directly tackles the central question of whether AI will help or replace human workers. I&#8217;m not telling you more, go read it :))</em></p></li></ul><h3>A difference-in-differences estimator by covariate balancing propensity score</h3><h5>TL;DR: Professors Li and Matsushita adapt the Covariate Balancing Propensity Score (CBPS) to estimate treatment effects in DiD designs. Their proposed CBPS-DiD estimator enforces covariate balance while offering double robustness for both the point estimate and its statistical inference. Tested in simulations and on the classic LaLonde dataset, it often delivers lower bias and more reliable confidence intervals than standard estimators like regression adjustment, IPW or AIPW. Its advantage is most pronounced when the underlying models are slightly misspecified (a common challenge in applied policy evaluation).</h5><p><em>What is this paper about?</em></p><p>This paper by Li and Matsushita takes the Covariate Balancing Propensity Score (CBPS)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> method introduced by <a href="https://academic.oup.com/jrsssb/article-abstract/76/1/243/7075938">Imai and Ratkovic (2014)</a> and adapts it to the DiD framework. Their focus is on estimating the ATT in situations where the PTA is unlikely to hold unconditionally but may be more plausible after conditioning on covariates<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>.</p><p>The authors&#8217; contribution is to integrate CBPS into a DiD setup by showing that their &#8220;CBPS-DiD&#8221; estimator inherits many &#8220;desirable&#8221; properties. When both the outcome model and the propensity score model are specified correctly, the estimator approaches the semiparametric efficiency bound<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> (meaning it makes full use of the available information in the data). When only one of the models is correct, the estimator remains consistent (doubly-robust<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>), and unlike the augmented inverse probability weighting (AIPW) DiD estimator, it is also doubly robust in terms of *inference* (so confidence intervals remain valid even if one model is wrong and you&#8217;re not being misled about how precise your estimates are). </p><p>In the paper&#8217;s simulation exercises they also find that CBPS-DiD converges faster to the true ATT than AIPW-DiD when both models are only &#8220;slightly&#8221; misspecified (it&#8217;s a subtle but practically relevant advantage since exact model correctness is very rare in applied work). </p><p><em>What do the authors do?</em></p><p>To examine how their CBPS-DiD estimator performs, the authors perform Monte Carlo simulations that let them control the &#8220;truth&#8221; and test the method under different conditions. They construct artificial datasets in which they vary whether the outcome model and the propensity score model are both correct, only one is correct, both are wrong or both are &#8220;slightly&#8221; wrong. These scenarios are important because they mimic the realities of applied work: rarely do we specify every model perfectly and small deviations from the truth are very common even with good theory AND data. For each of the proposed scenarios they compare CBPS-DiD to standard alternatives (outcome regression, IPW and AIPW) tracking bias, coverage probability and confidence interval length (checking if point estimates are close to the truth and whether the measures of uncertainty are reliable).</p><p>They then move from simulations to an empirical illustration using LaLonde&#8217;s (1986) dataset, which is a &#8220;benchmark&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> in program evaluation where the true average treatment effect is known to be zero. It&#8217;s like a tough test: the dataset is well known for &#8220;tripping up&#8221; estimators, often producing spurious effects due to covariate imbalance. By applying CBPS-DiD alongside the other estimators to this data they can see how much bias each method produces when estimating an effect that should, IN THEORY, be absent. They run these comparisons both with a small covariate set and a larger one, displaying how performance changes as the dimensionality of the adjustment set grows.</p><p><em>Why is this important?</em></p><p>The PTA is one of the pillars of DiD, but in many applications it is unlikely to hold unless you adjust for pre-treatment differences in observed characteristics. The problem is that some traditional approaches (like regression adjustment or IPW) often struggle here for a variety of reasons, such as poor covariate balance or relying too much on the model being specified correctly. CBPS-DiD addresses both of these problems by a) building covariate balance into the estimation process and b) keeping the &#8220;protection&#8221; of double robustness. It also offers double robustness for inference, meaning confidence intervals remain valid even when one of the models is wrong. This matters because in observational data the covariates that matter for trends are often measured with error or modeled imperfectly. If balance is poor the ATT estimate can be biased. If inference is wrong you can be misled about the effect&#8217;s precision. CBPS-DiD reduces the risk of both problems without requiring perfect knowledge of the true data-generating process<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. It is designed for the realities of applied policy evaluation where theory guides model choice but can&#8217;t guarantee correctness.</p><p><em>Who should care?</em></p><p>Applied researchers who rely on DiD in settings where the PTA is unlikely to hold without covariate adjustment and where model misspecification is a concern. It is especially relevant if you work with observational data where balance between treated and control units is hard to achieve or if you want more reliable inference when model assumptions are only partly correct. Policy evaluators, labour economists, education researchers and public health analysts working with staggered or selective treatment adoption can all benefit from a method that strengthens both balance and inference without requiring perfect models. </p><p><em>Do we have code?</em></p><p>No, the paper doesn&#8217;t provide replication files or a package info. If you want to try CBPS-DiD you would need to adapt the existing CBPS framework from Imai and Ratkovic (2014) (available in the <code>CBPS</code> package for R) and embed it into a DiD setup following the steps in the paper. The package itself only covers cross-sectional and basic panel weighting so this is not a &#8220;plug and play&#8221; situation for DiD. It&#8217;s doable if you&#8217;re comfortable working with both propensity score weighting and DiD estimation but there&#8217;s no one-line function for it yet.</p><p>In summary, this paper extends the CBPS to the DiD framework, producing an estimator that binds covariate balance into the weighting stage while maintaining double robustness, and even extends it to inference. It&#8217;s a method built for real-world policy evaluation where covariates matter for trends, models are rarely perfectly specified and you want both credible point estimates and trustworthy CIs. The simulations and the LaLonde application show it can outperform standard regression adjustment, IPW and AIPW in both bias and precision, with the advantage being more pronounced when model misspecification is mild but inevitable.</p><h3>Good Controls Gone Bad: Difference-in-Differences with Covariates</h3><p><em>(We have plots, DAGs AND flowcharts in this paper, which I love. Also lots of things in this paper follow from the previous one - check the footnotes)</em></p><h5>TL;DR; many DiD studies add time-varying covariates assuming they help, but this only works if the CCC assumption holds (meaning covariate effects are stable across groups and time). When CCC fails, popular estimators like TWFE, CS-DiD, imputation and FLEX can produce biased treatment effects, even if PT hold without covariates. This paper formalizes CCC, shows how violations cause bias and introduces DiD-INT, a new estimator that remains unbiased under CCC violations and can recover PT that are hidden by misspecified covariate adjustments.</h5><p><em>What is this paper about?</em></p><p>This paper is about what happens when the relationship between your covariates and your outcome isn&#8217;t stable across time or across groups, a situation the authors formalize as the Common Causal Covariates (CCC) assumption<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>.</p><p>In most DiD studies, we add covariates to make the PTA more plausible or to control for other influences, but when the effect of such covariates changes over time or differs between treated and control groups, &#8220;standard&#8221; estimators (from the familiar TWFE model to modern options like Callaway&#8211;Sant&#8217;Anna, imputation and FLEX) can produce biased treatment effect estimates. Karim and Webb show that this problem is widespread, formalize three versions of the CCC assumption and demonstrate how violations cause bias in widely used estimators. They also introduce a new estimator, the Intersection Difference-in-Differences (DID-INT) which remains unbiased even when CCC fails and works in settings with staggered treatment rollout and heterogeneous treatment effects.</p><p><em>What do the authors do?</em></p><p>They do lots of stuff, so I separated the contributions into four subsections.</p><ol><li><p>First, they formalize the CCC assumption</p><p>They introduce three explicit versions of the CCC assumption: two-way, state-invariant, and time-invariant, and explain how each shapes the way covariates should be handled in DiD. This step is super important because until now most methods implicitly assumed two-way CCC without stating it.</p></li><li><p>Second, they show the problem with existing estimators</p><p>Through theory and Monte Carlo simulations, they show that when CCC fails, widely used estimators (TWFE<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>, CS-DiD<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>, imputation<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>, FLEX<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>) can produce biased ATT estimates. This bias appears even if you have correctly specified the rest of your model and your data meets other standard DiD assumptions. </p></li><li><p>Third, they propose the DiD-INT</p><p>DiD-INT is designed to adjust for covariates in a way that allows CCC to be violated, recover parallel trends hidden by misspecified covariate adjustments, work with staggered adoption and heterogeneous treatment effects, and avoid &#8220;forbidden comparisons&#8221; and negative weighting issues<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>. DiD-INT runs in five steps starting with a model selection algorithm that visually checks pre-trends under different ways of interacting covariates with group and time. This identifies the correct functional form *before* estimating the ATT.</p></li><li><p>Fourth,  they develop a model selection algorithm</p><p>The authors provide a structured sequence to decide how to model covariates: 1) start without covariates and plot pre-trends; 2) if they fail, try covariates under the two-way CCC assumption; 3) if they still fail, interact covariates with group (state-varying DiD-INT) or time (time-varying DiD-INT); 4) if that fails, interact with both (two-way DiD-INT); 5) if pre-trends still look implausible, no DiD method is recommended. This approach is meant to replace the common &#8220;give up when pre-trends fail&#8221; practice with a systematic check for hidden parallel trends.</p></li></ol><p><em>Why is it important?</em></p><p>Many of us include time-varying covariates in our DiD models without thinking too much about whether the relationship between those covariates and the outcome is actually &#8220;stable&#8221;. If that stability (aka the CCC assumption) fails, the resulting ATT estimates can be biased even if everything else about the design looks good<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a>.</p><p>The problem is that CCC violations are not rare. In real datasets, covariate effects often change over time (due to shifting macroeconomic conditions, policy changes or industry composition) or vary between groups (because of structural differences in demographics, institutions or markets), and in some cases both happen.</p><p>This paper&#8217;s message is that CCC should be treated as a first-order identification issue rather than a minor technical detail. It also offers a way forward when CCC fails: DiD-INT broadens the set of situations where PT can plausibly hold, which then will probably reduce the number of projects that get abandoned just because pre-trends &#8220;don&#8217;t look parallel&#8221; under a misspecified covariate adjustment.</p><p><em>Who should care?</em></p><p>Lots of people (pretty much anyone running DiD models with covariates should care). Applied researchers running DiD models with covariates in repeated cross-sections or administrative datasets (this includes anyone working with survey microdata, labour force stats, education records or health data where both outcomes and covariates change over time); policy evaluators working with interventions that roll out across locations or institutions at different times, where standard estimators might be biased by shifting covariate effects; data scientists in government agencies and research institutes who produce official impact evaluations and must navigate both methodological correctness and practical constraints on variable selection.</p><p><em>Do we have code?</em></p><p>Yes, the authors say: &#8220;to ease in the implementation of the DID-INT estimator we have a package available in <a href="https://github.com/ebjamieson97/DiDInt.jl">Julia</a>. We also have a wrapper for Stata which calls the Julia program to perform the calculations, using the approach in Roodman (<a href="https://journals.sagepub.com/doi/pdf/10.1177/1536867X251341105">2025</a>). The Stata program is available [<a href="https://github.com/ebjamieson97/ didintjl">here</a>]. A wrapper in R is forthcoming. The software package allows for cluster robust inference using both a cluster-jackknife and randomization inference. The details of these routines, and their finite sample performance is discussed in the companion paper Karim et al. (2025).&#8221;</p><p>In summary, this paper reframes something many of us treat as a minor &#8220;specification choice&#8221; (adding covariates) into a core identification concern for DiD. By making the CCC assumption explicit, it shows that the stability of covariate effects is not a given and that violations can bias even modern estimators. This bias can appear whether or not PT hold unconditionally, meaning that covariates can sometimes turn a valid design into an invalid one. So rather than abandoning a project when pre-trends look implausible, the authors show how re-specifying the covariate adjustment can uncover &#8220;hidden&#8221; PT. Their DID-INT estimator (combined with a model selection algorithm) provides a structured way to test for and adjust to CCC violations while avoiding forbidden comparisons in staggered adoption settings<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a>. For applied work, the contribution is twofold: a warning that covariates can do harm if misused and a practical method to recover unbiased estimates when a key stability assumption fails<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a>.</p><h3>Bayesian Sensitivity Analyses for Policy Evaluation with Difference-in-Differences under Violations of Parallel Trends</h3><p><em>(Moving away from our frequentist framework&#8230;)</em></p><h5>TL;DR: this paper presents a Bayesian sensitivity analysis for DiD when the PTA is likely violated. The approach models the size and persistence of violations with an AR(1) process and allows these parameters to be fixed, fully estimated within a Bayesian model or calibrated from pre-treatment data. The authors demonstrated how treatment effect estimates change under different assumptions about violations by using the example of beverage sales data from Philadelphia and Baltimore.</h5><p><em>What is this paper about?</em></p><p>In this paper the authors examines how to use DiD when the PTA is likely violated. They propose a Bayesian sensitivity analysis that introduces a formal parameter for the size and persistence of these violations, modeled with an AR(1)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a> process to capture temporal correlation. They present three ways to set the priors for this process: fixed values, fully Bayesian estimation and empirical Bayes calibrated from pre-treatment data. The approach is illustrated by estimating the effect of Philadelphia&#8217;s sweetened beverage tax using Baltimore as a control city, showing how treatment effect estimates change under different assumptions about the violation.</p><p><em>What do the authors do?</em></p><p>The authors start by extending the standard DiD setup to include a term that measures how much the treated group&#8217;s trend could diverge from the control group after treatment. This deviation term is modeled with an AR(1) process, which captures both the average size of the violation and how persistent it is over time. They then specify three strategies for setting the priors on the AR(1) parameters: 1) fixed values, where the level of persistence is set in advance and only the size of the deviations varies; 2) fully Bayesian estimation, where the model learns both persistence and deviation size from the data within chosen prior ranges; and 3) empirical Bayes, where these parameters are first estimated from pre-treatment data and then used as priors in the post-treatment model. They then apply these models to monthly beverage sales data from Philadelphia (treated) and Baltimore (control) before and after the 2017 sweetened beverage tax. The authors tested how large the violation of PT would need to be to make the beverage tax&#8217;s effect statistically insignificant. For most of their models the required violation was *implausibly* large (e.g. for supermarkets, one model required a 980-fold increase in counterfactual sales), which strengthens their conclusion that the tax did have an effect.</p><p><em>Why is this important?</em></p><p>In the two previous paper we went on and on about violation of the PTA. It is super important, but less discussed in Bayesian settings<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a>. When the PTA does not hold, the estimated treatment effect can reflect underlying differences between treated and control groups rather than the impact of the policy. Many studies check for PT by looking at pre-treatment data but these tests are often underpowered and can&#8217;t guarantee that trends would have remained aligned after treatment. The Bayesian framework in this paper gives us a more structured way to relax the assumption, quantify the size and persistence of violations and see how the results change under different plausible scenarios. This makes it possible to present policy conclusions alongside a transparent assessment of how much they depend on the PTA.</p><p><em>Who should care?</em></p><p>I feel like I am repeating myself here, but this paper will interest applied researchers who use DiD and worry that the PTA may not hold particularly in policy evaluations where treated and control groups have different pre-treatment dynamics. It is also relevant for statisticians and econometricians working on sensitivity analysis methods and for policy analysts who need to present results with a clear statement of how robust they are to key assumptions. Anyone who works with short panels, noisy outcomes or limited pre-treatment data will find the discussion of prior choices and empirical Bayes calibration really useful (I did).</p><p><em>Do we have code?</em></p><p>The paper does not share replication files, but it states that the models were implemented in PyMC (version 4.0) using the No-U-Turn Sampler, a form of Hamiltonian Monte Carlo. Since the main addition to a standard Bayesian DiD is the AR(1) process for violations, the method could be reproduced by combining a DiD setup with an AR(1) prior on the deviation term. PyMC has <a href="https://www.pymc.io/projects/examples/en/latest/time_series/AR.html">public</a> examples of AR(1) time-series modeling that can serve as a starting point for implementing this framework.</p><p>In summary, this paper adapts DiD to situations where the PTA is unlikely to hold by modeling violations directly in a Bayesian framework. The AR(1) process for the deviation term allows both the size and persistence of violations to be estimated or calibrated from pre-treatment data. The authors show that different prior choices (fixed, fully Bayesian and empirical Bayes) can lead to different conclusions about the policy effect, making sensitivity analysis an essential part of interpretation. The application to Philadelphia&#8217;s sweetened beverage tax demonstrates how this approach can make the robustness of DiD results explicit, rather than leaving it as an untested assumption.</p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>CBPS is an alternative way of estimating the propensity score (the probability that a unit receives treatment given its observed characteristics) that directly enforces covariate balance between treated and control groups at the estimation stage, which is different to the standard approach which estimates the propensity score first (often using a logit or probit model) and then checks balance afterward. CBPS produces weights that better align the pre-treatment characteristics of treated and untreated units by making balance a &#8220;target&#8221; rather than an &#8220;afterthought&#8221;, which in turn improves the plausibility of the conditional PTA when covariates are important.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>For example, when treated units differ systematically from controls in ways that affect how outcomes evolve over time, such as having younger populations, higher baseline income, or different industry structures, but may be more plausible once these differences are accounted for through covariates. In these cases, simply comparing pre and post-treatment changes between treated and untreated groups would risk conflating treatment effects with the influence of those underlying characteristics, whereas conditioning on covariates can &#8220;help&#8221; isolate the policy&#8217;s impact. Let&#8217;s think of a concrete example related to my area (Econ of Education): consider a policy that expands access to after-school tutoring but is rolled out first in schools serving lower-income communities (let&#8217;s not get into what the policymaker might have in mind such as their specific goals with the policy, whether there is even a real problem this policy would solve, or whether they have a detailed plan for measuring its effectiveness). If we just compared changes in test scores at these schools to changes in more affluent schools without the program, the trends might differ even without the policy because achievement growth often follows different trajectories across socioeconomic groups. But if we adjust for covariates (such as baseline test scores, parents&#8217; education or school resources) the assumption that treated and control schools would have evolved in parallel becomes more plausible, but not completely. Remember that misspecification is a huge issue that can arise from omitted variable bias, wrong functional form, mismeasured variables, incorrect distributional assumptions or inappropriate fixed effects structure. If any other important factors remain unobserved (or if they can&#8217;t be measured and theory offers no guidance - aka the unobserved confounders), even conditioning on a &#8220;good&#8221; set of variables won&#8217;t guarantee that the PTA holds. The PTA is an assumption about the true data generating process. Misspecification means your model doesn&#8217;t correctly represent that process, and if it leaves out factors that actually drive differences in trends, the *conditional* PTA will still be violated even if you think you&#8217;ve adjusted for &#8220;enough&#8221; covariates. This is such a good topic, I can recommend some reading: &#8220;What&#8217;s Trending in Difference-in-Differences? A Synthesis of the Recent Econometrics Literature&#8221; <a href="https://arxiv.org/pdf/2201.01194">(Roth, Sant&#8217;Anna, Bilinski, Poe, 2023)</a>; &#8220;Difference-in-differences when parallel trends holds conditional on covariates&#8221; (<a href="https://arxiv.org/pdf/2406.15288">Caetano, Callaway, 2024</a>); &#8220;Nothing to see here? Non-inferiority approaches to parallel trends and other model assumptions&#8221; (<a href="https://arxiv.org/pdf/1805.03273">Bilinski, Hatfield, 2019</a>).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>In simpler terms, the semiparametric efficiency bound is the &#8220;smallest possible variance an estimator can achieve in a given setting without making strong, unrealistic assumptions&#8221;, (e.g. knowing the exact functional form of the model, assuming perfectly normally distributed errors or ruling out heteroskedasticity). Hitting that bound means you&#8217;re extracting the maximum precision possible from your data under credible assumptions. <a href="https://arxiv.org/pdf/1906.10221">The type of modeling used is based on how much information are available about the form of the relationship between response variable and explanatory variables, and the random error distribution</a>. Remember the difference: in a parametric model you fully specify the functional form AND distribution of the errors (e.g. linear regression with normally distributed, homoskedastic errors), and all parameters are finite-dimensional; in a nonparametric model you make NO functional form assumptions, you let the data determine the shape of the relationship (often at the cost of efficiency since you need large samples to get precise estimates); and in a semiparametric model you specify a finite-dimensional parameter of interest (like a treatment effect) BUT leave part of the model (it can be the error distribution or the functional form of covariates) unspecified. Going back to our education examples, let&#8217;s consider estimating the effect of reducing class size on test scores (you can think of a couple of mechanisms in place). In a parametric model, you might assume a linear (up or down) relationship between class size and scores, control for other factors in a fixed way and assume the errors are normally distributed and homoskedastic. In a nonparametric model, you would make no assumption about the shape of the relationship, letting the data reveal it (by having lots of fun plotting the data) but you would need a much larger sample to get a precise answer (often applied micro people we don&#8217;t have that). In a semiparametric model, you could focus on estimating the ATE of small classes while leaving the relationship between other covariates (like parents&#8217; income or teacher experience) and test scores unspecified, which give you a bit more flexibility while retaining more precision than a fully nonparametric approach.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;In a missing data model [where the missingness mechanism can be MCAR, MAR or MNAR], an estimator is doubly robust (DR) or doubly protected if it remains consistent when either a model for the missingness mechanism or a model for the distribution of the complete data is correctly specified. In a causal inference model, an estimator is DR if it remains consistent when either a model for the treatment assignment mechanism or a model for counterfactual data is correctly specified. Because of the frequency and near inevitability of model misspecification, double robustness is a highly desirable property&#8221;. (<a href="https://academic.oup.com/biometrics/article/61/4/962/7296220">Bang, Robins, 2005</a>)</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>Check &#8220;LaLonde (1986) after Nearly Four Decades: Lessons Learned&#8221; (<a href="https://yiqingxu.org/papers/2024_lalonde/imbens_xu24.pdf">Imbens, Xu , 2024</a>) for a look at the LaLonde dataset. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>Even with better balance and double robustness, CBPS-DiD (like any DiD estimator) can&#8217;t account for unobserved confounders that affect trends differently between treated and control groups. The conditional PTA must still hold for the set of observed covariates you include.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>This is an assumption which says that the effect of each covariate on the outcome is constant either across time, across groups or both. There are three forms: 1) two-way CCC, which is constant across time and groups; 2) state-invariant CCC, which is constant across groups but can vary over time; and 3) time-invariant CCC, which is - as you can guess by - constant over time but can vary between groups. Violation means that the &#8220;adjustment&#8221; provided by covariates is not the same for all observations, which leads to bias. Think of it as a recipe for a cake that works the same way and produces the same result whether you bake it in California or New York (it&#8217;s state-invariant), and whether you bake it today or in five years (it&#8217;s time-invariant). The recipe is &#8220;stable&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>TWFE &#8220;suffers&#8221; from negative weighting and forbidden comparisons in staggered designs (Goodman-Bacon, 2021).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Callaway&#8211;Sant&#8217;Anna avoids forbidden comparisons by estimating group-by-time ATTs and aggregating, and it defaults to a doubly robust DiD approach that still assumes CCC when covariates are time-varying.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>Imputation is an approach that predicts untreated counterfactual outcomes for treated units using a model fit on controls, then takes differences. Can use time-varying covariates but still assumes CCC holds.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>A flexible regression specification allowing for time-varying covariates, but like others, does not address CCC violations.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>In staggered adoption some treated units serve as controls for others which introduces bias when treatment effects are heterogeneous. Negative weights mean some group-time effects get subtracted in aggregation.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>Even in designs where PT unconditionally hold, adding covariates that violate CCC can create bias that wouldn&#8217;t otherwise exist. Covariates are not always &#8220;safe&#8221; to include and bad covariates can turn a valid design into an invalid one.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>The paper makes it clear that while DiD-INT is unbiased even when CCC fails, this robustness comes at the cost of efficiency. The estimates from DiD-INT will have higher variance (i.e. wider CIs) than a standard estimator. We are familiar with this trade-off: gaining accuracy at the cost of some precision.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>Given the bias-variance trade-off, the authors suggest a practical path forward: researchers can either use the model selection algorithm to find the most efficient (parsimonious) model that satisfies PT or simply default to the two-way DiD-INT estimator, which is unbiased across all potential CCC violation scenarios - albeit at the cost of <a href="https://x.com/JohnHolbein1/status/1954250222043005225">statistical power</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>Macro friends will get this one hehe ;) AR(1) stands for autoregressive process of order 1, it&#8217;s a way to model how a value at one time point depends on its own value in the previous time period + some random noise. Here it means that deviations from PT are assumed to change gradually over time rather than jumping around randomly. Broader explanations <a href="https://econweb.rutgers.edu/ctamayo/teaching/AR(1)_process.pdf">here</a> and <a href="https://julia.quantecon.org/introduction_dynamics/ar1_processes.html">here</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>Bayesian methods are used here because they make it straightforward to incorporate prior beliefs about the size and persistence of violations into the model and to quantify uncertainty about those violations in a coherent way. In a frequentist framework, sensitivity analyses require separate runs for each assumed violation level and they *do not* produce a single posterior distribution that integrates over uncertainty about these parameters. The Bayesian approach can treat the violation parameters as random variables (which are neither variable nor random, iykyk), update their distributions with the data and directly propagate that uncertainty into the treatment effect estimates.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[DiD+: how to handle data constraints, some timing issues and other TWFE limitations]]></title><description><![CDATA[What to do when data can't be pooled, timing gets messy, and TWFE isn't enough]]></description><link>https://www.diddigest.xyz/p/did-handling-data-constraints-timing</link><guid isPermaLink="false">https://www.diddigest.xyz/p/did-handling-data-constraints-timing</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Fri, 01 Aug 2025 15:30:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Gult!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Gult!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Gult!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Gult!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Gult!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Gult!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Gult!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Gult!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!Gult!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!Gult!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!Gult!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f4a40fa-86ee-45ef-9bad-82ad33686d15_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>(Claude creates the image prompt based on the post and it gets crazier by the day. I appreciate it tho)</em></p><p>Hi! </p><p>Before we start today&#8217;s post, I&#8217;d like to recommend this one &#8220;<a href="https://causalinf.substack.com/p/omitted-variables-versus-replicating">Omitted Variables versus Replicating the RCT</a>&#8221; by prof Scott. For those of you who don&#8217;t know, prof Rubin is the Rubin behind the Rubin causal model<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. He&#8217;s one of the founding figures of modern causal inference, who introduced the potential outcomes framework that is now in all discussions of treatment effects, from RTs to DiD models. Also we got a mention on my friend Sam&#8217;s <a href="https://samenright.substack.com/p/links-for-july">post</a>, thanks! To conclude the introduction, I know this is not the purpose of this newsletter so I avoid adding links not directly related to the post (for useful ones I have an entire <a href="https://beagietner.github.io/webpage/LearningResources.html">section on my website</a>), but I found a <a href="https://www.linkedin.com/posts/vladislavvmorozov_datascience-statistics-econometrics-activity-7356292489034588160-R0Ul">link</a> on LinkedIn by professor <a href="https://www.linkedin.com/in/vladislavvmorozov/overlay/about-this-profile/">Vladislav Morozov</a> on his Advanced Econometrics course and it is reeeeally good and worth checking out :)</p><p>The content we will be covering today is:</p><ul><li><p><a href="https://arxiv.org/abs/2403.15910">Difference-in-Differences with Unpoolable Data</a>, by Sunny Karim, Matthew D. Webb, Nichole Austin, and Erin Strumpf</p></li><li><p><a href="https://arxiv.org/html/2507.20415v1">Staggered Adoption DiD Designs with Misclassification and Anticipation</a>, by Clara Augustin, Daniel Gutknecht, and Cenchen Liu</p></li><li><p><a href="https://arxiv.org/abs/2507.19099">Interactive, Grouped and Non-separable Fixed Effects: A Practitioner's Guide to the New Panel Data Econometrics</a>, by Jan Ditzen and Yiannis Karavias</p></li></ul><p>And there are three &#8220;applied&#8221; papers I won&#8217;t go through but that might be interesting for you to read:</p><ul><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5337463">Using did multiplegt dyn in Stata to Estimate Event-Study Effects in Complex Designs: Four Examples Based on Real Datasets</a>, by Cl&#233;ment de Chaisemartin and Bingxue Li (I found this on <a href="https://www.linkedin.com/posts/clement-de-chaisemartin-b076135_using-did-multiplegt-dyn-in-stata-to-estimate-activity-7356326827105165313-CkBH">LinkedIn</a>, it&#8217;s a guide that shows you how to estimate dynamic treatment effects in complex DiD designs using the <code>did_multiplegt_dyn</code> Stata command, with real examples involving staggered, continuous, and multi-treatment settings)</p></li></ul><ul><li><p><a href="https://link.springer.com/epdf/10.1007/s11113-025-09968-w?sharing_token=aV9-kUdX2bT7sQzc7BJD5_e4RwlQNchNByi7wbcMAY5isHukhWD934n8c8jM5iu3ZUrNBVx0lwkEvxksRN8cELRkzgOOSWzSy4DVD3ntvPwrHxaUoWF9SwUixTry_rKLxfyuOsxzU-bN-7w80yIqZ6b078Y28sifWNl2uCapIQo%3D">Re-Assessing the Impact of Brexit on British Fertility Using Difference-in-Difference Estimation</a>, by Ross Macmillan and Carmel Hannan (this paper offers a forensic re-assessment of a published study, showing how a statistically significant &#8220;Brexit effect&#8221; on fertility is entirely an artifact of inappropriate control group selection. It provides a practical template for how rigorously ensuring pre-treatment PT can completely reverse a study&#8217;s conclusions. Read it if you, like me, love a good discussion on internal validity).</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5338512">Emerging Techniques for Causal Inference in Transportation: Integrating Synthetic Difference-indifferences and Double/Debiased Machine Learning to Evaluate Japan&#8217;s Shinkansen</a>, by Jingyuan Wang, Shintaro Terabe and Hideki Yaginuma (this paper has all the buzzwords I like + it&#8217;s about Japan and trains :-) What to do when your treated units are unique megacities like Tokyo or Osaka? The authors show how traditional DiD can give unstable or biased answers and instead use both Synthetic DiD and Double/Debiased Machine Learning to evaluate Japan&#8217;s bullet train network. They argue that the real power comes from seeing these two completely different, advanced methods point to the same conclusion, which is a great way to build confidence in your findings).</p></li></ul><h3>Difference-in-Differences with Unpoolable Data</h3><p><em>(This one hit close to home)</em></p><h5>TL;DR: <em>s</em>tandard DiD often requires merging datasets, but privacy rules sometimes forbid this. This paper introduces UN-DID, a method that works around this by calculating changes <em>within</em> each secure dataset first. You then combine just the summary results (not the sensitive data) to get a valid treatment effect. The authors provide code to do it.</h5><p><em>What is this paper about?</em></p><p>This paper is about how to estimate treatment effects using DiD when data from treatment and control groups cannot be combined. In many applications, we have to work with administrative data that is stored in separate secure environments such as in different provinces, agencies or countries. Legal or privacy restrictions may prevent pooling these datasets into a single file for analysis. The authors here propose a method called UN-DID that would allow us to estimate DiD models without needing to combine datasets. It works by estimating pre-post differences separately within each &#8220;silo&#8221;, then combining these differences externally to recover the ATT. The method supports covariates, multiple groups, staggered adoption and cluster-robust inference, making it suitable for a wide range of applied settings where conventional DiD is infeasible due to data restrictions.</p><p><em>What do the authors do?</em></p><p>They start with the simple 2x2 case and show that UN-DID gives the same estimate as conventional DiD when there are no covariates, then they extend the method to allow for covariates, staggered treatment timing and multiple jurisdictions.</p><p>UN-DID works by running separate regressions in each data silo to estimate within-group pre-post differences. These differences, along with their standard errors and covariances, are extracted and combined outside the silos to calculate the treatment effect. The method remains valid even when the effect of covariates differs across silos which is a setting where conventional DiD may be biased.</p><p>They support their method with formal proofs, Monte Carlo simulations and two empirical applications using real-world data. In both examples they take pooled datasets and artificially treat them as unpoolable to show that UN-DID produces results that closely match those from standard DiD even in the presence of covariates and staggered adoption.</p><p><em>Why is this important?</em></p><p>Many administrative datasets like health records, tax files, or student registers are stored in secure environments with privacy protections that prohibit exporting or combining data across jurisdictions. These rules are designed to prevent the risk of re-identifying individuals, and as a result, we can&#8217;t pool treatment and control data into a single file, which is what conventional DiD methods require.</p><p>UN-DID makes it possible to estimate DiD effects in these settings without violating data-sharing rules. UN-DID works within the boundaries of legal and ethical data use while still enabling credible policy evaluation. The method also handles situations where covariates affect outcomes differently across silos, which helps avoid bias that can affect standard DiD models. For example, the authors note that researchers have been unable to evaluate policies like Canada&#8217;s national cannabis legalization because they couldn&#8217;t combine provincial data. UN-DID is designed for precisely this kind of setting.</p><p><em>Who should care?</em></p><p>You, me, and anyone who works with administrative data/records behind a government &#8220;paywall&#8221;, which includes even health policy analysts, education researchers, and anyone using secure data that can&#8217;t be exported or merged. UN-DID also matters for applied researchers dealing with staggered adoption or treatment heterogeneity where credible counterfactuals lie in datasets they can&#8217;t directly combine. It gives them a way to recover treatment effects without weakening their design or dropping research questions entirely due to data restrictions.</p><p><em>Do we have code?</em></p><p>To everybody&#8217;s happiness, yes. The authors provide UN-DID software packages in <a href="https://cran.r-project.org/web/packages/undidR/index.html">R</a>, <a href="https://github.com/ebjamieson97/undidPyjl">Python</a>, <a href="https://github.com/ebjamieson97/undid">Stata</a>, AND <a href="https://github.com/ebjamieson97/Undid.jl">Julia</a> (!) which are described in Section 7.3 of the paper. These tools are designed to work within the constraints of siloed data environments, meaning each analyst can run their part locally, and only summary estimates (like pre-post differences and standard errors) need to be exported. A separate user guide with a full empirical example using real-world siloed data is forthcoming.</p><p>In summary, UN-DID is a very practical extension of DiD for settings where legal or privacy rules prevent combining datasets across jurisdictions. It produces valid treatment effect estimates by working within each silo and then combining results externally. The method supports covariates, staggered timing, and heterogeneous effects, and performs well in both simulations and real data. For researchers working with administrative data that can&#8217;t be pooled, it offers a way to keep using DiD without compromising identification (or breaking the law).</p><h3>Staggered Adoption DiD Designs with Misclassification and Anticipation</h3><h5>TL;DR:<em> </em>when treatment dates are misclassified or people anticipate a policy before it officially starts, standard staggered DiD estimators produce biased results. This paper offers a solution by introducing new, bias-corrected estimators that account for these issues and provides a diagnostic test to detect the presence and timing of such misspecification in the data.</h5><p><em>What is this paper about?</em></p><p>This paper studies two issues that many of us face when using staggered DiD: misclassification of treatment timing and anticipation effects. Such problems aren&#8217;t new or even overlooked, but in practice we often do the best we can with the data we have, knowing damn well that policy implementation may not be cleanly observed or that people can react to treatment announcements before the official start date. What&#8217;s been missing is a formal framework to understand how these issues threaten identification and how to adjust for them.</p><p>Staggered adoption designs are now widely used in policy evaluation. They extend the basic DiD setup to cases where groups are treated at different times. Recent papers, like <a href="https://www.sciencedirect.com/science/article/pii/S0304407620303948?casa_token=siMgZNpx3aUAAAAA:H1eUmDW-_hg_u1Y3_I2dFL7YLT1NEimYIwxyuhtVFVWHUA6szH2clMS5ukcbd2HLwmZSpekwmg">Callaway and Sant&#8217;Anna (2021)</a> and <a href="https://www.nber.org/system/files/working_papers/w25904/w25904.pdf">De Chaisemartin and D'Haultfoeuille (2020)</a> have improved how we deal with treatment effect heterogeneity and timing, but they assume that treatment dates are correct and that outcomes don&#8217;t change before treatment starts.</p><p>This paper builds on that work and shows what goes wrong when those assumptions don&#8217;t hold. It explains how misclassification and anticipation can distort the comparisons made in staggered DiD and how that leads to biased estimates. It also connects to other research on errors in treatment status, like <a href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1468-0262.2006.00756.x?casa_token=q12aufpW1ugAAAAA:4pIHbmyPU0fyII9svBZyUETNewNF0AfN52Ff-Gt87mL-RjkutEXYNtbhwSSXtI6IX4OhgCPMF_6d1LU-">Lewbel (2007)</a> and <a href="https://arxiv.org/pdf/2207.11890">Denteh and Kedagni (2022)</a> and to recent work on checking pre-trends before running DiD such as <a href="https://academic.oup.com/restud/article-pdf/90/5/2555/51356029/rdad018.pdf?casa_token=3_QLmDFhKswAAAAA:AVwP0GDEhVUocpmdeEoLFFHxzBku7wVKczLY_sQMBtygBQIjJ3DGxuqusnkbA--jwJBipYJoKXYEVGY">Rambachan and Roth (2023)</a>.</p><p>In short, the paper shows that when treatment is misclassified or anticipated even the best available DiD estimators may not be reliable<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> unless we adjust for those issues, and proposes ways to both detect and correct the resulting bias.</p><p><em>What do the authors do?</em></p><p>The authors make three main contributions. First, they formally characterize the bias introduced by misclassification and anticipation in staggered DiD. They show that TWFE and other estimators aren&#8217;t &#8220;reliable&#8221; when the timing or incidence of treatment is misspecified (even under homogeneous treatment effects). This happens because these estimators end up comparing units that shouldn&#8217;t be treated as untreated, or vice versa, thus creating what the authors call &#8220;forbidden comparisons&#8221;. </p><p>Second, they propose bias-corrected estimators that recover causal effects under these conditions. One version adjusts an existing staggered estimator (from De Chaisemartin and D&#8217;Haultfoeuille, 2020) to account for misclassified timing, while another version targets the ATE among actually treated units (even when actual treatment dates aren&#8217;t observed) by imposing a weak homogeneity assumption on misclassification probabilities, which lets them estimate how many units were really treated at each point in time and then adjust accordingly.</p><p>Third, they introduce a 2-step testing procedure to detect the presence, extent, and timing of misspecification. The first step checks whether the PTA holds using standard pre-trend comparisons. If it passes, then the second step compares outcomes across switchers and not-yet-switched groups in earlier periods to flag potential misclassification or anticipation. These tests help determine when the bias corrections are needed and whether the assumptions behind them are plausible.</p><p>They support all of this with simulation evidence and an application to the rollout of a computer-based testing system in Indonesian schools. Interestingly, their results suggest that schools were adjusting behaviour even before officially adopting the system, and that failing to account for this leads to underestimating the true treatment effect.</p><p><em>Why is this important?</em></p><p>Staggered adoption designs are popular in applied economics for estimating treatment effects when policies roll out at different times. They handle heterogeneity and exploit timing variation, but rely on hard-to-verify assumptions.</p><p>2 key challenges are misclassification and anticipation. Misclassification happens when recorded treatment timing doesn&#8217;t match actual implementation due to measurement error or delayed adoption. Anticipation occurs when units respond before treatment formally begins, perhaps due to advance announcements. Both distort the comparisons staggered DiD depends on, biasing even sophisticated estimators designed for treatment heterogeneity.</p><p>These problems appear often in practice (you might have come across them yourself). Studies show institutional behaviour shifting before formal implementation (<a href="https://assets.aeaweb.org/assets/production/files/6526.pdf">Bindler and Hjalmarsson, 2018</a>), misalignment between administrative eligibility and actual responses (<a href="https://academic.oup.com/qje/article-pdf/132/3/1165/30646702/qjx008.pdf?casa_token=fSrro-kMLtsAAAAA:3MeYtdO5CDYDWHn_e3KOV4BuZEcByMDBVFkg21B0Luib0e1cV_po8fZMw6NWo_AEeT7yM0wUnXPlCsQ">Goldschmidt and Schmieder, 2017</a>), and anticipatory adjustments before official adoption (<a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11216028/">Berkhout et al., 2024</a>). In each case, recorded dates poorly reflect actual exposure or behavioural change. So it&#8217;s welcoming to see this paper addressing how timing misspecification creates biased comparisons and further providing diagnostic tools and corrected estimators. Rather than treating timing as a secondary issue, it integrates these concerns directly into identification and estimation, and this is particularly valuable for researchers using administrative or policy data where treatment definitions are inherently noisy (aren&#8217;t them all, to some extent?).</p><p><em>Who should care?</em></p><p>Anyone using staggered DiD designs with administrative or policy data should look this one up. That includes applied researchers studying policy reforms, education interventions, labour regulations or any other setting where treatment doesn&#8217;t arrive cleanly and uniformly. If your treatment variable comes from government records, institutional reports, or eligibility rules (and not from a randomized rollout) then the risk of misclassification or anticipation is probably already in your data. This is also relevant if you&#8217;re working with event study plots and interpreting pre-trends. The paper shows that violations from misclassified timing or early behavioral responses can sneak into the pre-treatment window and look like non-parallel trends even when the actual causal structure is sound. So if you&#8217;ve ever seen a bumpy pre-trend and wondered whether to throw out your design, this gives you a reason to pause and test before you abandon the setup.</p><p><em>Do we have code?</em></p><p>Not yet. As of the first version of the arXiv post (July 2025), there&#8217;s no replication repository linked, and the paper doesn&#8217;t mention any software package or implementation details. The estimators and tests are well defined and nothing looks too difficult to implement, but it would be helpful to have code or even a worked example. If that shows up in a later version or companion repo, I&#8217;ll update this post.</p><p>In summary, this paper strengthens the foundation of staggered DiD by tackling something most of us have seen but rarely model directly: the gap between recorded treatment timing and actual exposure. It formalizes how misclassification and anticipation distort group comparisons, shows that even advanced estimators are vulnerable and offers tools to detect and correct those problems. The framework is useful, the diagnostics are intuitive and the estimators are practical. If you work with policy data where timing is a bit messy (which is most of it) this is worth your attention.</p><h3>Interactive, Grouped and Non-separable Fixed Effects: A Practitioner's Guide to the New Panel Data Econometrics</h3><p><em>(Speaking of TWFE, prof Jeffrey was on about it today <a href="https://x.com/jmwooldridge/status/1951269173692092435">her</a>e, and prof Scott&#8217;s 3 latest posts also discuss it <a href="https://causalinf.substack.com/p/which-is-easier-to-interpret-canonical">here</a> and <a href="https://causalinf.substack.com/p/a-design-argument-to-not-use-twfe">here</a> and <a href="https://causalinf.substack.com/p/dropping-treated-units-from-the-control">here</a>. You have a lot of reading to do, I don&#8217;t make the rules :-)).</em></p><h5>TL;DR: this paper is a practical guide arguing that the standard TWFE model is often too restrictive for modern panel data which can lead to biased results when unobserved characteristics have complex, time-varying impacts. It provides an accessible overview of more flexible methods like interactive, grouped, and non-separable fixed effects, and shows through empirical examples that these newer estimators often fit the data better and can significantly change your conclusions.</h5><p><em>What is this paper about?</em></p><p>This is not a &#8220;paper&#8221; per se but a practical guide to the new generation of panel data models that better capture unobserved heterogeneity. The standard TWFE approach assumes that unobserved traits are either fixed over time or shared across units in a given period. But this often misses how units might respond differently to common shocks or how the value of unobserved characteristics can evolve.</p><p>The authors walk us through 3 more flexible frameworks: interactive fixed effects (IFE), which let unobserved unit traits interact with time-varying shocks; grouped fixed effects (GFE), which cluster units into latent groups that share patterns over time; and non-separable two-way (NSTW) models, which allow for nonlinear interactions between units and periods. So rather than introducing new theory, the paper focuses on helping us understand when and how to use these models. It covers the main estimators, lays out key assumptions, compares performance and discusses how to test which model fits best. It also includes two empirical examples and a software appendix to make the methods easier to apply.</p><p><em>What do the authors do?</em></p><p>(They do a lot and I learned so many new terms by reading this paper)</p><p>The provide a very structured overview of the main estimation methods for modeling unobserved heterogeneity in panel data using interactive, grouped and non-separable FE. Again, they don&#8217;t propose a new estimator but consolidate and explain a wide set of recent contributions (many of which are *technically* demanding), and translate them into accessible terms for applied people. They cover how to estimate models with IFE, where unobserved unit characteristics interact with time-varying common shocks. For that, they walk through estimation techniques like: iterative least squares (ILS), which alternates between estimating factors and coefficients; penalized least squares (PLS), which uses adaptive LASSO to select the number of factors; nuclear norm regularization (NNR), which reframes factor estimation as a convex optimization problem; IV approaches, which include two-step IV estimators and GMM estimators that avoid estimating incidental parameters; and common correlated Effects (CCE), which &#8220;sidestep&#8221; direct factor estimation by using cross-sectional averages as proxies. They also explain how GFE can be estimated by clustering units into latent groups, reducing dimensionality and avoiding bias when the number of time periods is limited. Finally they discuss non-separable models that go beyond linear IFE structures, allowing for flexible, possibly nonlinear interactions between units and time effects.</p><p>The good thing is that throughout the paper they emphasize the trade-offs: whether you need to estimate the number of factors, how estimation bias behaves in small samples, and which estimators work best when regressors are endogenous, when N and T are large (I never had this issue hehe) or when models include lags.</p><p>They back this up with two empirical applications, one looking at the inflation-growth nexus and another reassessing the Feldstein-Horioka puzzle in international finance, to show that newer models like NSTW can give quite different (and more plausible) results than standard TWFE.</p><p><em>Why is this important?</em></p><p>This is an important guide because most applied people still rely (or just really like) TWFE even when the assumptions behind it don&#8217;t hold (not our fault!). TWFE assumes that unobserved characteristics are either constant over time (like geography or ability) or affect all units the same way in a given year (like a national shock), but in many panels (like when T is large) these assumptions are too restrictive. Unobserved factors often evolve over time and different units may respond to the same shock in different ways. If we *ignore* (sometimes we don&#8217;t have an option) this and keep using TWFE, the estimates can be biased and misleading, which means we might be drawing the wrong conclusions from the data. And in many cases, even using IFE isn&#8217;t enough: the paper shows that nonlinear and group-based models often fit the data better. This guide matters because it lowers the barrier to using these newer methods, it explains what each estimator assumes, when it &#8220;breaks down&#8221; and how to choose the right &#8220;model&#8221; (prof Jeffrey might have an issue with this nomenclature). It also shows that applying these methods doesn&#8217;t have to be intimidating! There are diagnostics, workflows and open-source tools that make it manageable.</p><p><em>Who should care?</em></p><p>Anyone working with panel data where unobserved heterogeneity might evolve over time or affect units differently, and that includes applied micro people using survey or firm-level data, macro people modeling country or regional dynamics and financial people studying markets or portfolios over time.</p><p>If you&#8217;re running panel regressions with many time periods or even just worried that your unobserved variables aren&#8217;t neatly time-invariant or uniform across units, this guide gives you the tools to do better. It&#8217;s also useful if you&#8217;re dealing with potential cross-sectional dependence, incidental parameter bias or unclear model fit and want diagnostics that go beyond residual plots. Even if you don&#8217;t plan to use these estimators like right now, understanding them helps you interpret existing empirical work more critically and helps you avoid relying on TWFE out of &#8220;habit&#8221; (or convenience).</p><p><em>Do we have code?</em></p><p>The paper includes a super helpful appendix listing open-source implementations for most of the methods discussed, available in both Stata and R, with some support in Matlab as well. For example: for CCE, you can use the popular <code>XTDCCE2</code> command in Stata or the <code>plm</code> package in R; for the iterative IFE estimator, there's <code>REGIFE</code> in Stata and <code>INTERFE</code> in R; for GFE, the authors point to code available on St&#233;phane Bonhomme's webpage; and for the newer NSTW models, you can find <code>PCLUSTER</code> in R. Some of the very latest estimators may require more hands-on coding, but the paper points to companion codebases and replication files for their applications. It&#8217;s definitely worth checking out!</p><p>In summary, this is a practical and well-organized guide to modern fixed effects models for panel data. It doesn&#8217;t introduce new estimators, but it brings together a wide set of tools that applied people should know about, more so if they&#8217;re still relying on TWFE.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For a chapter about it written by him and prof Imbens, email me or access <a href="https://link.springer.com/chapter/10.1057/9780230280816_28">here</a>. I also highly recommend this <a href="https://projecteuclid.org/journals/statistical-science/volume-29/issue-3/A-Conversation-with-Donald-B-Rubin/10.1214/14-STS489.pdf">interview</a> he did with profs Fan Li and Fabrizia Mealli. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>For years researchers used TWFE as the default &#8220;DiD&#8221; estimator, even in staggered settings. It wasn&#8217;t until Goodman-Bacon (2021), Sun and Abraham (2021) and others showed it could mix already-treated units into control groups (the negative weights!) that we realized it was doing more than we thought. Now most new work tries to avoid these comparisons but legacy papers and habits still linger.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[Heterogeneity, Small Sample Inference, Anticipation, and Local Projections]]></title><description><![CDATA[On uniform confidence bands for conditional effects, a simple approach for small-N inference, refining the "no anticipation" assumption, and the LP-DiD estimator for staggered designs.]]></description><link>https://www.diddigest.xyz/p/heterogeneity-small-sample-inference</link><guid isPermaLink="false">https://www.diddigest.xyz/p/heterogeneity-small-sample-inference</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Tue, 22 Jul 2025 14:55:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!prUs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!prUs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!prUs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!prUs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!prUs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!prUs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!prUs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!prUs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!prUs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!prUs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!prUs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1391507b-ad9a-4e59-a888-9e17d8ae6b66_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! I am back from <a href="https://www.linkedin.com/posts/beatrizgietner_some-asked-me-what-i-was-doing-in-china-activity-7351251238140764161-1Wkq?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAACTHmhoBpQRxz0hUy8fhU1GuRQuRyKVjdd0">China</a> now, and I have a few <a href="https://x.com/BeatrizGietner/status/1946134194960183653/photo/1">posts</a> lined up on ML and CI in general, but first we will go through the ones related to DiD. </p><p>A bit of housekeeping first. We got two package updates when I was away.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5ai_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5ai_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 424w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 848w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 1272w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5ai_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png" width="997" height="954" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:954,&quot;width&quot;:997,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!5ai_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 424w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 848w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 1272w, https://substackcdn.com/image/fetch/$s_!5ai_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb131fbe7-d5b3-4bca-ab53-9db351c2c83c_997x954.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Link to post <a href="https://www.linkedin.com/posts/brantly-callaway-b5297092_difference-in-differences-with-a-continuous-activity-7348331291672539139-e__K/">here</a> </figcaption></figure></div><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FH9p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FH9p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 424w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 848w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 1272w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FH9p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png" width="900" height="897" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a231842-3913-4875-bc7e-200328098cbd_900x897.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:897,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:192643,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/168625655?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FH9p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 424w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 848w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 1272w, https://substackcdn.com/image/fetch/$s_!FH9p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a231842-3913-4875-bc7e-200328098cbd_900x897.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Link to post <a href="https://x.com/pedrohcgs/status/1944797761540391026">here</a></figcaption></figure></div><p>We spoke about these papers <a href="https://diddigest.substack.com/p/coming-soon">here</a> and <a href="https://diddigest.substack.com/p/ddd-estimators-distributional-effects">here</a> :) </p><p>And today&#8217;s post is about:</p><ul><li><p><a href="https://arxiv.org/pdf/2305.02185">Doubly Robust Uniform Confidence Bands for Group-Time Conditional Average Treatment Effects in Difference-in-Differences</a>, by Shunsuke Imai, Lei Qin, and Takahide Yanagi.</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5325686">Simple Approaches to Inference with Difference-in-Differences Estimators with Small Cross-Sectional Sample Sizes</a>, by Soo Jeong Lee and Jeffrey M. Wooldridge.</p></li><li><p><a href="https://arxiv.org/abs/2507.12891">Refining the Notion of No Anticipation in Difference-in-Differences Studies</a>, by Marco Piccininni, Eric J. Tchetgen Tchetgen, and Mats J. Stensrud.</p></li><li><p><a href="https://onlinelibrary.wiley.com/doi/10.1002/jae.70000">A Local Projections Approach to Difference-in-Differences</a>, by Arindrajit Dube, Daniele Girardi, &#210;scar Jord&#224;, and Alan M. Taylor (this is the JAE published version - congrats!! - of the WP that <a href="https://x.com/arindube/status/1653046276790054914">went around</a> in 2023. Also for Stata users, there&#8217;s an update to the locproj package <a href="https://www.bbvaresearch.com/en/publicaciones/locproj-and-lpgraph-stata-commands-to-estimate-local-projections/">here</a>).</p></li></ul><h3>Doubly Robust Uniform Confidence Bands for Group-Time Conditional Average Treatment Effects in Difference-in-Differences</h3><p><em>(<a href="https://sites.google.com/view/shunsuke-imai/home">Shunsuke</a> is a second-year PhD student at Kyoto University)</em></p><h5>TL;DR: this paper shows how to study treatment effect heterogeneity in staggered DiD designs when you care about a continuous covariate, like the pre-treatment poverty rate. The authors build on Callaway and Sant'Anna (2021) to estimate group-time Conditional Average Treatment Effects (CATTs) that vary with that covariate. They construct uniform confidence bands so we can see which parts of the curve are meaningful. The method is doubly robust and works under standard identification conditions. The procedure combines parametric estimation for nuisance functions with nonparametric methods for the main parameter of interest. The authors provide an R package (<code>didhetero</code>) to help implement the methods.</h5><p><em>What is this paper about?</em></p><p>This paper is about how to study treatment effect heterogeneity in staggered DiD designs. The goal is to move beyond group-time averages and see how treatment effects vary depending on the value of a continuous pre-treatment covariate, like the poverty rate. The authors focus on estimating group-time Conditional Average Treatment effects on the Treated (CATTs). This tells us, for each group and period, how the effect changes across the distribution of a pre-treatment variable. Instead of asking whether a policy worked, we can ask for whom it worked, and how strongly .</p><p>The second part of the paper is about inference. It is one thing to estimate a curve, but we also want to say which parts of that curve are statistically meaningful. The authors construct uniform confidence bands for the CATT function. These bands help us see where effects are large, where they are noisy, and where they are indistinguishable from zero. The method adapts the doubly robust estimand from <a href="https://www.sciencedirect.com/science/article/pii/S0304407620303948">Callaway and Sant&#8217;Anna (2021)</a> to a more granular, conditional setting. It works under standard assumptions and uses a three-step procedure that combines both parametric and nonparametric estimation. The result is a tool that can handle covariate heterogeneity, treatment timing, and proper inference at the same time.</p><p><em>What do the authors do?</em></p><p>They start with the doubly robust estimand from Callaway and Sant&#8217;Anna (2021) and build on it to estimate conditional treatment effects. The main innovation is extending this framework to recover how effects vary with a continuous covariate. To do this, they propose a three-step procedure that combines parametric estimation of nuisance components (the outcome regression and generalized propensity score) in the first stage with nonparametric local polynomial regressions in the second and third stages. This setup captures nonlinearity without forcing a specific shape onto the CATT function.</p><p>For inference, they develop two ways to construct uniform confidence bands: one based on an analytical approximation and another using weighted or multiplier bootstrapping . This part is technically demanding. They show how nonparametric smoothing and the presence of estimated nuisance terms affect the distribution of their estimator, and they prove that the bands have the correct asymptotic coverage . The statistical theory extends results from recent work on conditional average treatment effects in unconfoundedness setups to the staggered DiD setting. The paper includes simulations and discusses practical aspects like how to pick bandwidths and estimate standard errors. It also defines several summary measures based on the CATTs, which can help with interpretation when there are many groups and time periods.</p><p><em>Why is this important?</em></p><p>When we employ DiD we often end up reporting some average effect for a group or a time period and move on. But averages aren&#8217;t always helpful because they smooth over interesting variation, and sometimes that variation is the whole point. This paper shows how we can move beyond average effects. Instead of asking if a policy worked on average, we can ask who it worked for. Did high-poverty counties see a benefit? This paper shows how to estimate treatment effects that change with covariates like the poverty rate and then build confidence bands around them so we know what we&#8217;re seeing is real.</p><p>It&#8217;s useful because this kind of question comes up all the time in applied work. Think about the minimum wage. The extent to which minimum wage increases reduce poverty depends on the structural relationship between wage gains and job losses. A group average can&#8217;t tell you that, but this method can help assess it. And what&#8217;s nice is that it works with staggered treatment timing, which is the norm in many real-world applications. You don&#8217;t have to pretend that everyone was treated at once or that effects are constant across space and time. You get a picture that&#8217;s both more honest and more informative. That&#8217;s the kind of thing you want when the stakes are high and the outcomes matter.</p><p><em>Who should care?</em></p><p>If you&#8217;re working with staggered treatment and you already use group-time ATT methods like Callaway and Sant&#8217;Anna (2021), you should check this paper. It adds another layer: instead of just estimating how effects vary by group and time, you can now see how they shift across something continuous, like the poverty rate. It&#8217;s helpful if you&#8217;re trying to answer questions like: does this policy work better in richer or poorer counties? Is the effect stronger where unemployment was already high? This paper gives you a way to answer those kinds of questions without needing a separate model for every subgroup. And it gives you proper confidence bands so you can see where the heterogeneity is real and where it&#8217;s just noise. The authors also mention that even though their main example is minimum wage and poverty, this method is general and you can apply it to any staggered DiD setup where treatment effects might depend on a baseline variable.</p><p><em>Do we have code?</em></p><p>We have a package: <a href="https://tkhdyanagi.github.io/didhetero/">didhetero</a> (Treatment Effect Heterogeneity in Staggered Difference-in-Differences). It &#8220;provides tools to construct doubly robust uniform confidence bands (UCB) for the group-time conditional average treatment effect (CATT) function given a pre-treatment covariate of interest and a variety of useful summary parameters in the staggered difference-in-differences setup of Callaway and Sant&#8217;Anna (2021)&#8221;.</p><p>In summary, this paper is a nice reminder that treatment effects are more than numbers, they&#8217;re functions, and once we start thinking about them that way, we need the right tools to estimate and interpret them. The authors take a standard DiD setup and give us a way to ask questions like how effects shift with poverty or unemployment, while still keeping the core assumptions intact. You don&#8217;t need to change your identification strategy or build a different model for every subgroup, this method fits into what you already do and just makes it more informative.</p><h3>Simple Approaches to Inference with Difference-in-Differences Estimators with Small Cross-Sectional Sample Sizes</h3><p><em>(<a href="https://econ.msu.edu/about/directory/Lee-Soo-Jeong-grad-2019">Soo-jeong</a> is Professor Jeffrey&#8217;s soon-to-be-former student. <a href="https://x.com/jmwooldridge/status/1940979034168742186">Here</a>&#8217;s his thread about this paper)</em></p><h5>TL;DR: standard statistical methods for DiD studies are unreliable when you have a small number of treated or control groups (e.g., a single state or a few hospitals). This paper proposes a simple solution: for each group, calculate the average outcome before the policy and the average outcome after. Then, run a basic cross-sectional regression on these averages. This simple trick provides statistically valid confidence intervals and t-tests, even with extremely small samples, such as one treated unit and two control units. It offers a straightforward and easy-to-implement alternative to more complicated methods like Synthetic Control, giving us a reliable tool for policy evaluation when data is limited.</h5><p><em>What is this paper about? </em></p><p>This paper looks at a familiar setup (tracking outcomes for treated and control units before and after some intervention) and asks what to do when you don&#8217;t have many units on either side. Usually DiD works fine if you have lots of treated and control units since you can lean on large&#8209;sample theory and cluster&#8209;robust errors. Here the authors show a simple trick: collapse each unit&#8217;s time series into two numbers (the pre&#8209;intervention average and the post&#8209;intervention average) then run an ordinary cross&#8209;sectional regression of that time&#8209;collapsed outcome on a treatment indicator. Under the usual linear&#8208;model assumptions (normal errors, constant variance) you get exact inference even if you have just one treated unit and two controls. They then extend that to remove unit&#8209;specific trends (fitting a little time trend in the pre&#8209;period, subtracting it off, then averaging), and show we can still do exact t&#8209;tests on that transformed data. Finally, they compare this to synthetic control and synthetic DiD approaches, run simulations and apply the method to California&#8217;s 1989 smoking restrictions (one treated state) and to staggered &#8220;castle law&#8221; rollouts. The upshot is a very easy&#8208;to&#8208;implement alternative when sample sizes in the cross&#8209;section are small.</p><p><em>What do the authors do? </em></p><p>The authors take a hands&#8209;on look at how to get valid confidence intervals when you have very few treated or control units. They start by showing that you can collapse each unit&#8217;s entire time path into just two numbers (its average outcome before the policy and its average outcome after). You then run an ordinary cross&#8209;sectional regression of those two&#8209;period averages on a simple indicator for whether a unit was treated. Under the familiar assumptions of a linear model with normal, homoskedastic errors, that single regression gives you exact t&#8209;tests and confidence intervals even if you have as few as one treated unit alongside two controls. Next they make the method more flexible by carving out linear trends at the unit level. They fit a straight&#8209;line trend to each unit&#8217;s pre&#8209;treatment history, subtract it from the full series, and then average the residuals before and after treatment. You can then plug those de&#8209;trended averages into the same cross&#8209;sectional regression and still get exact inference. After laying out these core ideas they run through simulations that confirm the small&#8209;sample accuracy of their approach, and then they demonstrate it on two real cases: California&#8217;s 1989 smoking ban and a staggered rollout of castle&#8209;law enactments. That empirical work shows the method is fast to implement and gives results that line up well with more elaborate techniques</p><p><em>Why is this important? </em></p><p>A lot of times we applied researchers don&#8217;t have the benefit of working with large samples due to the lack of proper data. For example if you&#8217;re studying a law change in one state or a corporate rollout in a handful of markets, the usual cluster&#8209;robust approach can give wildly misleading confidence intervals when you only have a few clusters. This paper offers a way to get honest measures of uncertainty in those settings without complicated bootstrap schemes or heavy-duty synthetic control machinery. By boiling each unit&#8217;s history down to a before&#8209;and&#8209;after average (or its de&#8209;trended version), you end up with a tiny cross&#8209;sectional regression where the standard t&#8209;statistic really follows a t&#8209;distribution. That means you can report confidence intervals you can trust, even with just one treated unit. It gives us a straightforward fallback when we lack large samples or when more elaborate methods are hard to justify in small&#8209;n contexts. </p><p><em>Who should care? </em></p><p>This paper will matter to anyone running a policy evaluation where you don&#8217;t have the luxury of dozens of treated and control units. If you&#8217;re an applied economist looking at a single state&#8217;s law change or a public health researcher tracking a handful of hospitals before and after an intervention, you&#8217;ll run into the limits of cluster&#8209;robust inference with small samples. It&#8217;s also useful for consultants or corporate analysts who roll out a pilot program in just a few markets and need reliable uncertainty measures without wrestling with heavy bootstraps or complex synthetic controls. In those cases this simple before&#8209;and&#8209;after averaging trick gives you a clear way to get honest confidence intervals even when your cross&#8209;section is tiny.</p><p><em>Do we have code? </em></p><p>For this paper there isn&#8217;t a dedicated R or Python library you need to install. Once you&#8217;ve collapsed each unit&#8217;s time series into a pre&#8209; and post&#8209;treatment average (or its de&#8209;trended counterpart), you just run an OLS of that two&#8209;period outcome on your treatment indicator. In Stata you could type something like &#8220;reg post_pre_diff treated, robust&#8221; and the built&#8209;in t&#8209;statistic is exact under the paper&#8217;s normal&#8209;error assumptions. If you want an exact p&#8209;value without leaning on normality, you can use the user&#8209;written <code>ritest</code> command for randomization inference. To compare against synthetic DiD you&#8217;d load the <code>sdid</code> package in Stata&#8239;18, but you don&#8217;t need any fancy routines for the core approach. A few lines of code in R or Python (compute the averages, run <code>lm()</code> or <code>statsmodels.OLS</code>) and you&#8217;re done.</p><p>In summary, this paper serves as a reminder that sometimes the simplest solution is the most effective. In a field that often trends toward more complex estimators, the authors show how a clever data transformation combined with foundational econometric principles can solve a difficult and common inferential problem. Here we can see a &#8220;DiD problem&#8221; through a different lens. Soo-jeong and Prof Jeffrey provide a method that is transparent, easy to implement and statistically sound even in small samples. It empowers us to make credible claims in the exact settings where evidence is often most needed but hardest to analyze.</p><h3>Refining the Notion of No Anticipation in Difference-in-Differences Studies</h3><p><em>(I really enjoyed reading this paper - it did have more words than the others)</em></p><h5>TL;DR: this paper resolves a common point of confusion about the &#8220;no anticipation&#8221; assumption in studies that employ DiD. We often worry that people might change their behaviour in anticipation of a policy, but the formal assumption simply states that future treatments can&#8217;t affect past outcomes, which is a trivial point about the arrow of time. The authors argue this mismatch stems from conflating the policy&#8217;s implementation with the decision to implement it. By introducing a separate variable for the &#8220;plan&#8221; or &#8220;decision&#8221; (P), they provide a clearer framework. This clarifies what &#8220;no anticipation&#8221; really means, shows how the standard DiD estimator can be biased if people react to the plan (as expected) and helps us specify whether we are estimating the effect of the plan itself or its ultimate implementation.</h5><p><em>What is this paper about?</em></p><p>This paper tackles a subtle but important ambiguity at the heart of DiD methodology. Many DiD guides and papers state that a no anticipation assumption is required for the method to be valid. Formally, this assumption is often written in a way that says an intervention at a future time point does not affect an outcome in the past. The problem is that in any standard causal model the future can&#8217;t affect the past &#8220;by definition&#8221;. This makes the no anticipation assumption seem either trivially true or completely unnecessary which may be why some foundational DiD papers don&#8217;t even mention it. This has led to widespread confusion: is this an important identifying assumption we need to worry about or a redundant statement about time travel? </p><p>The authors argue that this confusion arises because the formal assumption fails to capture what researchers are really concerned about. When we talk about anticipation, they don&#8217;t mean that the policy itself reached back in time. They mean that knowledge of a future policy (the plan, the announcement, the waiver application) caused people or firms to change their behaviour before the policy was officially implemented. For example, in a study of Medicaid expansion, insurers might change their premiums as soon as a state announces its plan, well before the expansion actually begins. This paper&#8217;s goal is to resolve this ambiguity by providing a new, expanded causal model that formally separates the policy&#8217;s implementation from the prior decision to implement it. </p><p><em>What do the authors do?</em></p><p>Ok, let&#8217;s go in parts. The authors&#8217; main contribution is to clarify the causal model underlying DiD. They do this by introducing a new variable,P, which represents the plan or decision to implement a policy. This decision P occurs at a specific point in time and is distinct from the policy&#8217;s actual implementation, A_2&#8203;, which it causes. The key is that the decision P can happen before the pre-treatment outcome Y_1&#8203; is measured. With this expanded model, they propose a new, more meaningful no anticipation assumption (Assumption 5): the plan P has no average effect on the pre-treatment outcome Y_1&#8203; for the group that plans to adopt the policy. Unlike the standard assumption, this is a non-trivial statement about behaviour that could be violated in the real world. This framework has a few consequences. </p><p>They show that if a researcher (perhaps implicitly?) assumes parallel trends with respect to the plan (P=0) but there are anticipation effects (their new assumption is violated), then the standard DiD estimator is biased. Proposition 1 demonstrates that the DiD estimator identifies the true Average Treatment Effect on the Treated (ATT_A_2&#8203;&#8203;) minus a bias term, &#968;, which is exactly the effect of the plan on the pre-treatment outcome. They also argue that in some cases, the researcher might be more interested in the effect of the decision itself (ATT_P&#8203;) rather than the effect of the implementation (ATT_A_2&#8203;&#8203;). For example, an announcement of a car recall may cause people to stop driving the faulty cars, meaning the decision had a huge effect even if the subsequent act of seizing the cars had none. They show that under their new no-anticipation assumption, the standard DiD functional identifies ATT_P&#8203;.</p><p>Finally, they extend this logic from the simple two-period case to the more complex staggered adoption setting, showing how their framework can clarify the limited anticipation assumptions used in modern DiD estimators like Callaway and Sant&#8217;Anna (2021).</p><p><em>Why is this important?</em></p><p>This paper provides essential clarity on a foundational assumption that has been a source of confusion for decades. It closes the gap between the informal, intuitive meaning of anticipation and its flawed mathematical formulation. This matters for applied work for two main reasons. </p><p>First, it makes the potential for bias explicit. If you are studying a policy that was announced well in advance, you must now seriously consider whether that announcement itself changed behaviour. If it did, this paper shows that your standard DiD estimate of the implementation&#8217;s effect is likely biased, and it formalizes the source and structure of that bias.</p><p>Second, it forces us to be more precise about what causal question we are asking. Is the goal to measure the effect of a law actually being implemented (ATT_A_2&#8203;&#8203;)? Or is it to measure the total effect of the policy process, starting from its public announcement (ATT_P&#8203;)? These are different questions with different answers and the choice has real consequences for interpretation. This paper provides the formal language needed to distinguish between them and to state the assumptions required to identify each one.</p><p><em>Who should care?</em></p><p>This paper is a must-read for any applied researcher who uses or teaches DiD. This includes economists, political scientists, epidemiologists, and other social scientists who evaluate policies or interventions. It is especially important for those studying policies that are announced publicly before they take effect, which covers the vast majority of new laws and regulations. If your research involves a scenario where people could plausibly react to the knowledge of a future treatment (whether it&#8217;s a minimum wage increase, a new environmental rule, or a coming tax change) this paper provides the conceptual tools to handle it correctly. It fundamentally clarifies the assumptions discussed in popular practitioner guides and recent econometric surveys, making it essential reading for both students and experts in the field.</p><p><em>Do we have code?</em></p><p>No, this is a conceptual and theoretical paper, not a new estimator. It does not come with a new R or Stata package because its purpose is to refine the thinking that precedes the coding. The paper clarifies the causal model, the estimand of interest, and the identifying assumptions. The tools used to calculate the DiD functional (e.g., packages like <code>did</code> and <code>fixest</code> in R, or <code>csdid</code> and <code>reghdfe</code> in Stata) are unchanged. The contribution of this paper is to help you decide which estimand you are targeting (ATT_A_2&#8203;&#8203; vs ATT_P&#8203;) and to be explicit about the no anticipation assumption that your identification strategy relies on.</p><p>In summary, this paper solves a long-running methodological puzzle: why does the no anticipation assumption in DiD appear to be a self-evident statement about the impossibility of time travel? The authors show that researchers have simply been using the wrong formal language for the right intuition. The real concern is not about the policy itself affecting the past, but about the plan for the policy affecting behaviour in the present. By formally separating the policy&#8217;s implementation from the decision to implement it, the paper provides a clear and coherent framework for thinking about anticipation effects. It&#8217;s a nice reminder that before we rush to estimate, we must first be precise about what it is we are trying to estimate and under what assumptions. The paper&#8217;s contribution is not a new command to type, but a new clarity of thought.</p><h3>A Local Projections Approach to Difference-in-Differences</h3><h5>TL;DR: this paper offers a solution to the &#8220;negative weighting&#8221; problem that biases standard DiD estimates in settings with staggered treatment adoption. The authors propose an approach based on local projections (LP), a method common in macroeconomics. The &#8220;LP-DiD&#8221; estimator runs a separate, simple regression for each post-treatment period. The innovation is to restrict the sample in each regression to only include newly treated units and &#8220;clean controls&#8221; (units that have not yet been treated). This transparently avoids the problematic comparisons that cause bias, is computationally fast, highly flexible, and can replicate the results of more complex modern DiD estimators.</h5><p><em>What is this paper about?</em></p><p>This paper enters the ongoing conversation about how to properly conduct DiD analysis when treatment is &#8220;staggered&#8221;. A recent wave of econometric research has shown that the traditional two-way fixed-effects (TWFE) regression fails in this setting. The core issue is often called the &#8220;negative weighting&#8221; problem. In a staggered design the TWFE estimator implicitly uses already-treated units as controls for more recently treated units. For example a state that adopted a policy in 2005 might be used as part of the control group for a state that adopts the same policy in 2010. This is a &#8220;forbidden comparison&#8221; because the 2005 state may still be experiencing its own dynamic treatment effects, making it a contaminated or &#8220;unclean&#8221; control. This can lead to the TWFE estimate being a weighted average of the true effects where some weights are negative, sometimes producing an average effect that is nonsensical or even has the wrong sign. In response, this paper proposes an intuitive framework called LP-DiD. It leverages the Local Projections (LP) method to estimate dynamic effects and combines it with a straightforward &#8220;clean control&#8221; condition to sidestep the negative weighting problem entirely. The result is a regression-based tool that is easy to implement and understand, yet powerful enough to stand alongside other recently developed DiD estimators.</p><p><em>What do the authors do?</em></p><p>The authors&#8217; approach breaks the problem down by estimating the treatment effect for each post-treatment period, or horizon, separately. For each horizon, say, two years after treatment, they run a distinct regression. The outcome variable in this regression is the change in the outcome from the period just before treatment up to that two-year mark. The key variable of interest is an indicator that flags units at the exact moment they become treated.</p><p>The real innovation lies in how they construct the sample for each of these regressions. To get a clean estimate they only include two groups of units: the newly treated units (at the moment they switch into treatment) and the clean controls. A clean control is a unit that remains completely untreated all the way through the specific horizon being estimated. For the two-year effect, for example, the controls are units that are still untreated two years later. This simple rule is powerful because it guarantees that already-treated units are never used as controls, which is the source of the negative weighting bias in older methods.</p><p>The authors show that this basic procedure estimates a variance-weighted average of the effects across different treatment cohorts. They then demonstrate how to easily recover the more standard, equally-weighted average effect using familiar techniques like weighted least squares or a two-step regression adjustment. This flexible framework is also shown to easily accommodate covariates, situations where treatment isn&#8217;t permanent and different ways of defining the pre-treatment baseline.</p><p><em>Why is this important?</em></p><p>The clean control condition is highly intuitive. It makes it easy to understand and explain exactly which comparisons are being made to identify the treatment effect, in contrast to more &#8220;black-box&#8221; estimators. At its core, the method is just a series of OLS regressions on different subsamples of the data. This makes it computationally fast, which is an advantage when working with very large panel datasets where more complex estimators can be slow. The LP-DiD approach helps demystify the new DiD literature. The authors show that their estimator, with different weighting schemes, is numerically equivalent or very similar to other leading methods, such as those proposed by Callaway and Sant&#8217;Anna (2021) and Borusyak, Jaravel, and Spiess (2024). This reveals the common logic underlying these different approaches. The framework is easily adapted to different empirical contexts. Researchers can modify the clean control definition for non-absorbing treatments, include covariates, or pool estimates over various horizons, all within a straightforward regression setup.</p><p><em>Who should care?</em></p><p>Any applied researcher using DiD with panel data, especially in settings with staggered treatment adoption, should be aware of this paper. It will be particularly useful for:</p><ol><li><p>Economists, political scientists, and public health researchers looking for a robust, easy-to-implement alternative to traditional TWFE.</p></li><li><p>Researchers who value transparency and want to be able to clearly articulate the identifying assumptions and comparisons underlying their estimates.</p></li><li><p>Analysts working with large datasets (e.g., using administrative or worker-level panel data) who would benefit from a computationally efficient estimation method.</p></li><li><p>Instructors of econometrics courses, as LP-DiD provides a very clear and teachable example of how to solve the negative weighting problem in modern DiD.</p></li></ol><p><em>Do we have code?</em></p><p>Yes. The authors have released a Stata command, <code>lpdid</code>, which implements the estimators discussed in the paper. The paper also provides STATA example files on GitHub. While the method is simple enough to be implemented manually with basic regression commands, the dedicated package streamlines the process of estimation, reweighting, and inference. The paper also provides a clear example of how to use the built-in <code>teffects ra</code> command in Stata to obtain the reweighted estimates.</p><p>In summary, this paper provides an elegant and intuitive bridge from the world of traditional DiD to the modern literature on robust estimation with staggered treatment. By framing the problem through the lens of local projections, a familiar tool from macroeconomics, the authors show that the notorious &#8220;negative weighting&#8221; bias that plagues TWFE can be solved with a simple and intuitive sample selection rule. The LP-DiD approach estimates dynamic effects by running a series of simple regressions, each focused on a specific post-treatment horizon and using only clean comparisons between newly treated units and those not yet treated. This approach demystifies what many newer DiD methods are doing under the hood, offering us a tool that is not only robust but also exceptionally flexible, transparent, and computationally fast.  </p>]]></content:encoded></item><item><title><![CDATA[Efficient Estimation, Sequential Synthetic Control DiD, and the TWFE Debate]]></title><description><![CDATA[Some practical insights for a more precise and robust estimation.]]></description><link>https://www.diddigest.xyz/p/efficient-estimation-sequential-synthetic</link><guid isPermaLink="false">https://www.diddigest.xyz/p/efficient-estimation-sequential-synthetic</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Fri, 27 Jun 2025 14:02:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SwxZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SwxZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SwxZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SwxZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!SwxZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!SwxZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc75444b-3a1b-459c-a63f-61eb225e8aee_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! We are back to talking about DiD in specific :) Here are the latest papers we will discuss:</p><ul><li><p><a href="https://arxiv.org/pdf/2506.17729">Efficient Difference-in-Differences and Event Study Estimators</a>, by Xiaohong Chen, Pedro H. C. Sant&#8217;Anna, and Haitian Xie</p></li><li><p><a href="https://arxiv.org/abs/2404.00164">Sequential Synthetic Difference in Differences</a>, by Dmitry Arkhangelsky and Aleksei Samkov</p></li><li><p><a href="https://arxiv.org/abs/2402.09928">When Can We Use Two-Way Fixed-Effects (TWFE): A Comparison of TWFE and Novel Dynamic Difference-in-Differences Estimators</a>, by Tobias R&#252;ttenauer and Ozan Aksoy</p></li></ul><h3>Efficient Difference-in-Differences and Event Study Estimators</h3><p><em>(After all, who never heard someone saying &#8220;this DiD should&#8217;ve been an event study&#8221; at a seminar&#8230;)</em></p><h5>TL;DR: modern DiD estimators handle heterogeneity but often ignore efficiency. This paper builds estimators that hit the semiparametric efficiency bound under standard assumptions, with no functional form restrictions, no extra modelling, just better use of the data. Confidence intervals then shrink, power goes up, and all the pieces are there for anyone working with short panels and staggered treatment.</h5><p><em>What is this paper about?</em></p><p>This paper is about improving the precision of DiD and Event Study estimators. Modern methods account for treatment effect heterogeneity and staggered adoption, but they tend to handle pre-treatment periods and control groups in ad hoc ways. Many estimators either drop baseline periods or assign equal weights to them, even though these choices are rarely supported by theory or data.</p><p>The authors offer a formal solution<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>: they build a framework for efficient estimation of DiD and ES parameters that works with short panels, staggered treatment timing, and standard identification assumptions like PT and no anticipation. They show how to characterise these assumptions as conditional moment restrictions and use that structure to derive semiparametric efficiency bounds.</p><p>The main insight is that some pre-treatment periods and control groups carry more information than others. By using optimal weights based on the conditional covariance of outcome changes, the proposed estimators deliver tighter confidence intervals and smaller root mean squared error with no added assumptions and no loss in consistency, which then gives us a principled way to use all available data more effectively.</p><p><em>What do the authors do?</em></p><p>This paper is dense, so we&#8217;ll go section by section (each answers one question). </p><p>After the Introduction, section 2 (Framework, causal parameters, and estimands) answers &#8220;what are we estimating?&#8221;. The authors formalise the setup: short panels, staggered treatment, absorbing states, and potential outcomes indexed by treatment timing. They define the causal parameters of interest (group-time ATTs and event-study summaries) and lay out standard identification assumptions such as random sampling, overlap, no anticipation, and PT. They allow for both post-treatment-only and full-period parallel trends (PT), with or without covariates. </p><p>In section 3 (Semiparametric Efficiency Bound for DiD and ES), they answer &#8220;what&#8217;s the best we can hope for?&#8221;. This is the theoretical core of the paper. They derive semiparametric efficiency bounds for both ATT and ES estimators under the two types of PT. These bounds represent the lowest possible asymptotic variance for any estimator that satisfies the DiD identification conditions. The result is that most DiD estimators are inefficient because they either ignore or misweight pre-treatment and comparison groups. They also provide closed-form expressions for the efficient influence function (EIF). These formulas make clear how to combine periods and groups using weights based on the conditional covariance structure of outcome changes. </p><p>In section 4 (Semiparametric efficient estimation and inference), they translate the theory into practice by proposing two-step estimators that use plug-in methods for nuisance components (like outcome regressions and group probabilities) and then apply the EIF weights to get point estimates. They show that this approach can be implemented using either flexible nonparametric methods (like forests or splines) or simple parametric working models. They also explain how to stabilise the procedure by estimating ratios of group probabilities directly. </p><p>In section 5 (Monte Carlo simulations), they calibrate simulations using CPS and Compustat data and compare their efficient estimators to popular alternatives (TWFE, Callaway-Sant&#8217;Anna, Sun-Abraham, Borusyak et al.). Across designs, the efficient estimators reduce root-mean-squared-error and confidence interval width (often by more than 40%) without increasing bias. </p><p>In section 6 (Empirical Illustration) they reanalyse data from <a href="https://www.aeaweb.org/articles?id=10.1257%2Faer.20161038&amp;utm_source=TrendMD&amp;utm_medium=cpc&amp;utm_campaign=American_Economic_Review_TrendMD_1">Dobkin et al. (2018)</a> on the effect of hospitalisation on out-of-pocket medical spending. Using the publicly available Health and Retirement Study (HRS), they estimate treatment effects based on variation in the timing of hospitalisation. Their efficient estimators produce tighter confidence intervals than existing methods. To match the same level of precision with standard DiD approaches, researchers would need about 30% more data. Their analysis also showcases a fascinating result of the method: the ability to use a post-treatment period as a baseline by &#8220;bridging&#8221; through the never-treated group. They also use a form of specification curve analysis to visually assess the stability of the estimates across all valid modeling choices, thus providing evidence for the plausibility of the PTA. </p><p>They conclude by providing practical advice and future directions. They suggest that we should test whether some pre-treatment periods or control groups are less informative or potentially invalid. To support this, they introduce Hausman-type overidentification tests and offer visual tools to assess sensitivity to different weighting schemes. They also point to extensions: nonlinear DiD models, switching treatments, unbalanced panels, and setups without PT. The framework is designed to be flexible, but it&#8217;s built for the standard large&#8209;n, fixed&#8209;T world. When that changes, so should the estimator.</p><p><em>Why is this important?</em></p><p>I love everything about efficiency, but sometimes that&#8217;s too much to ask. Most DiD papers focus on identification and stay there. Estimators are consistent, robust to heterogeneity, and safe, but they leave a lot of precision on the table. In a refreshing way, this paper argues that we can have both: credible identification and minimal variance. By working directly with the structure implied by the PTA, the authors show how to build estimators that are semiparametrically efficient. No extra assumptions, no black boxes, no functional form restrictions, and no need for large samples, just better use of the data we already have. One of the key insights of this paper is that DiD setups are nonparametrically overidentified (which means there&#8217;s more information available than most estimators use). That&#8217;s what makes these efficiency gains possible. I also have to note that while negative weights in some DiD estimators can be problematic, here they are not a concern: they arise naturally from the covariance structure and are a feature of optimal estimation, not a bug. Efficiency matters because in this case because it gives us tighter confidence intervals, more stable estimates, and more power by using the design more intelligently. The fact that the estimators are Neyman orthogonal adds to their credibility as it makes them less sensitive to misspecification of the preliminary estimation steps.</p><p><em>Who should care?</em></p><p>Anyone working with DiD or ES designs using short panels, staggered treatment or small samples. If you&#8217;ve ever looked at your standard errors and thought &#8220;this feels bigger than it should be&#8221;, you should read this one carefully. It&#8217;s also useful for methodologists and software developers who want to understand what makes an estimator efficient, and how to design tools that make full use of available data without relying on strong assumptions. Even if you don&#8217;t implement the estimators right away, this paper gives you a benchmark: what is the best you could be doing, given your design?</p><p><em>Do we have code?</em></p><p>Professor Pedro <a href="https://x.com/pedrohcgs/status/1937509086968578419">said</a> &#8220;we will work on packaging everything so you can adopt all this seamlessly&#8221;. The estimators are presented in closed form and the implementation relies on standard tools (regression adjustments, propensity scores and conditional covariances), so while there&#8217;s no package, the ingredients are all there for anyone comfortable with two-step semiparametric estimation (not sure how many of us to be honest). The simulation and empirical application sections hint at how to put the pieces together and the estimators can be implemented in R or Python with off-the-shelf tools. </p><p>In summary, this paper takes the designs we already trust and shows how to push them further. The authors prove that standard DiD setups contain more information than we usually use, and they show how to recover it with estimators that are efficient, transparent, and grounded in the assumptions we already make.</p><h3>Sequential Synthetic Difference-in-Differences</h3><h5>TL;DR: this paper introduces the Sequential Synthetic Difference-in-Differences (Sequential SDiD) estimator, a new method for event studies with staggered treatment adoption (particularly useful when you doubt the PTA). It works by applying SC principles sequentially to cohort-aggregated data; it estimates effects for early-adopting groups and uses those results to impute counterfactuals for later-adopting groups. The authors&#8217; key theoretical result is that this estimator is asymptotically equivalent to an ideal but infeasible &#8220;oracle&#8221; OLS estimator that has access to the unobserved factors driving the violation of PT. </h5><p><em>What is this paper about?</em></p><p>The rise of synthetic control (SC) methods, starting with Abadie and Gardeazabal (2003) and Abadie, Diamond, and Hainmueller (2010)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> offered applied researchers a new way to construct counterfactuals when PT are unlikely to hold. In 2021 Arkhangelsky, Athey, Hirshberg, Imbens, and Wager gave us the combination of DiD and SC (adding up the structure of SC with the intuition of DiD) in their seminal <a href="https://www.aeaweb.org/articles?id=10.1257/aer.20190159">paper</a>, and now in this one, Arkhangelsky and Samkov extend that framework to settings with staggered treatment adoption (the setting behind many modern policy evaluations).</p><p>They propose a method, Sequential Synthetic DiD, that builds counterfactuals *cohort by cohort*, using the information from early adopters to inform the estimation for later ones. It operates on cohort-aggregated data, estimating treatment effects one horizon at a time, and updating the data after each step to reflect imputed counterfactuals. This structure helps correct for violations of PT caused by unobserved time-varying confounders while preserving the transparency we value in DiD.</p><p>A key theoretical result is that this estimator behaves like an &#8220;oracle&#8221; (a benchmark) OLS regression that &#8220;knows&#8221; the unobserved interactive fixed effects. This connection provides both valid inference and formal efficiency guarantees. The method handles staggered timing, treatment effect heterogeneity, and violations of PT and it includes standard DiD and recent imputation estimators (like <a href="https://academic.oup.com/restud/article-abstract/91/6/3253/7601390">Borusyak et al. 2024</a>) as special cases, depending on how the weights are chosen.</p><p><em>What do the authors do?</em></p><p>They introduce a new estimator, Sequential Synthetic Difference-in-Differences (Sequential SDiD), designed for event studies with staggered treatment adoption, where PT may fail due to unobserved time-varying confounders. The method builds on Synthetic DiD (Arkhangelsky et al., 2021)<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>, but adapts it for aggregated cohort-level data and implements it sequentially. They formalize causality by interpreting the observed outcomes using potential outcomes (Neyman, 1990; Rubin, 1974; Imbens and Rubin, 2015).<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a></p><p>For each adoption cohort, they estimate treatment effects at each horizon and update the observed outcomes with imputed counterfactuals. These updated values are then used in the estimation for later cohorts. This recursive structure is the paper&#8217;s core innovation: it reduces the risk of bias accumulation by preventing treated observations from contaminating the control pool in subsequent steps. </p><p>Also, to make the method directly useful for applied work (you might have asked yourself by no &#8220;but how do I handle control variables?&#8221;), the authors provide significant practical guidance on how to incorporate covariates (in Section 4<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>). They focus on time-invariant characteristics and outline three different strategies: full stratification, a recommended hybrid model that allows for direct application of their algorithm, and a specific approach for cases with group-level treatment assignment.</p><p>They formalise this estimator by connecting it to a theoretical benchmark: an &#8220;oracle&#8221; OLS regression that has access to the true unobserved interactive fixed effects. This oracle represents an idealised model of what applied researchers try to do when including unit-specific trends or unobserved factor structures. The authors then prove that their estimator is asymptotically equivalent to this oracle under &#8220;mild conditions&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>. Their results provide a strong theoretical backing, which includes valid inference, asymptotic normality, and the first formal efficiency guarantees for an SC-type method.</p><p>A key intermediate result of the paper, presented in Proposition 3.1, is that the infeasible oracle OLS estimator can be represented and computed by a sequential algorithm (Algorithm 2) that mirrors the main Sequential SDiD estimator. The authors note this is a &#8220;result of independent interest&#8221; because it reveals the underlying mechanics of OLS in staggered settings and provides &#8220;new insight into the mechanics of modern imputation estimators&#8221; like Borusyak et al.</p><p>Inference is conducted using the Bayesian bootstrap, and the authors also propose a placebo-style validation check by artificially shifting adoption dates backward (a way to assess model fit in the absence of PT).</p><p>They demonstrate the method&#8217;s performance through both an empirical application (on Community Health Centers, using <a href="https://www.aeaweb.org/articles?id=10.1257/aer.20120070">Bailey and Goodman-Bacon 2015</a>) and two simulation exercises. In settings where standard DiD performs well, Sequential SDiD yields similar results, but in scenarios with unobserved confounding, it remains stable while DiD becomes severely biased.</p><p>At the end they outline two key limitations. The first being that the method requires reasonably large adoption cohorts to ensure that aggregation averages out idiosyncratic noise. Second, it assumes that idiosyncratic errors are independent across units so that residuals concentrate around zero when averaged. They say that while these assumptions are common in modern DiD applications, they may be restrictive in contexts where individual-level shocks have strong aggregate effects.</p><p><em>Why is this important?</em></p><p>Most modern DiD estimators still depend on some version of the PTA, even when considering those designed for staggered adoption and heterogenous treatment effect. PTA means that untreated units can serve as a valid counterfactual for treated ones, which often doesn&#8217;t hold up when treatment is timed by factors we cannot observe or measure. In practice we would try to deal with this by adding unit-specific trends or other proxies for unobservables, but these adjustments only go so far. </p><p>Sequential SDiD offers an alternative for us because it addresses violations of PT by working with aggregated cohort data and estimating effects sequentially, using SC-style weighting to build credible counterfactuals one step at a time. This setup is way more robust when unobserved time-varying confounders drive both outcomes and treatment timing, which is exactly the kind of threat that undermines &#8220;traditional&#8221; DiD.</p><p>Their method is also grounded in familiar econometric foundations. The authors work in an asymptotic regime with many units and a fixed number of time periods, connecting their framework to the classic moment-based panel data literature (like <a href="https://www.sciencedirect.com/science/article/pii/S1573441284020146">Chamberlain, 1984</a>). They operate in a low-noise environment, similar to the conditions studied in the SC literature (by averaging within large cohorts), which makes the method both practically feasible and theoretically sound. </p><p>Sequential SDiD also connects back to simpler methods. If the regularisation is set very high, the weights collapse to uniform ones, and the estimator reduces to a version of imputation-based DiD, closely related to Borusyak et al. (2024). This makes it easy to interpret and benchmark against more familiar designs.</p><p>Finally, the paper fills a gap in the SC literature by establishing formal efficiency guarantees<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a>. The authors prove that Sequential SDiD is consistent, unbiased, and achieves first-order efficiency by showing it mimics an oracle OLS estimator that knows the unobserved factors. This is the first time an SC-type method has been shown to reach this level of statistical performance. It offers us a principled and robust estimator for settings with staggered adoption and unobserved confounding.</p><p><em>Who should care?</em></p><p>Anyone working with event studies, staggered adoption or policy evaluations where PT are unreliable should pay attention to this paper. If your treatment timing is potentially related to unobserved factors or if you&#8217;re worried that units are changing at different underlying rates, Sequential SDiD gives you a way to build more credible counterfactuals. This is useful in applications like health, labour, and education policy where adoption often happens gradually and for reasons we can&#8217;t fully observe.</p><p><em>Do we have code?</em></p><p>No, but the paper includes full algorithmic descriptions and detailed pseudocode for implementation. Algorithm 1 describes the Sequential SDiD procedure and Algorithm 2 outlines the oracle OLS estimator used for benchmarking. You can use the structure outlined in your preferred programming language. </p><p>In summary, this paper extends Synthetic DiD to settings with staggered treatment, unobserved time-varying confounding and cohort-level aggregation. The proposed Sequential SDiD estimator builds counterfactuals one step at a time, drawing on the structure of synthetic control and the logic of DiD. The authors prove that it behaves like an oracle OLS estimator, offering consistency, valid inference, and formal efficiency guarantees. </p><h3>When Can We Use Two-Way Fixed-Effects (TWFE): A Comparison of TWFE and Novel Dynamic Difference-in-Differences Estimators</h3><p><em>(Here we are again, discussing heterogeneous treatment effects&#8230;)</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!swTf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!swTf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 424w, https://substackcdn.com/image/fetch/$s_!swTf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 848w, https://substackcdn.com/image/fetch/$s_!swTf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!swTf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!swTf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg" width="500" height="625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:625,&quot;width&quot;:500,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:83571,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/166324473?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!swTf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 424w, https://substackcdn.com/image/fetch/$s_!swTf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 848w, https://substackcdn.com/image/fetch/$s_!swTf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!swTf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c0ee404-cc71-4edd-bc3e-879ff3f7316c_500x625.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">By prof <a href="https://x.com/KhoaVuUmn/status/1939632186799362466/photo/1">Khoa Vu</a> here :)</figcaption></figure></div><h5>TL;DR: this paper is about the ongoing debate around using TWFE in staggered treatment settings, where units receive treatment at different times. The authors explain how TWFE can become biased when treatment effects vary across time or groups. They compare TWFE to five newer DiD estimators using Monte Carlo simulations, testing each one under a range of scenarios, including violations of common assumptions.</h5><p><em>What is this paper about? </em></p><p>This paper walks us through the growing debate over using TWFE estimators in staggered treatment settings (cases where some units are treated earlier than others). The authors explain how TWFE can be biased if treatment effects vary over time or across groups. They then compare TWFE to five alternative DiD estimators using Monte Carlo simulations and show how each performs under different scenarios, including violations of key assumptions. </p><p><em>What do the authors do? </em></p><p>They say they have three main goals in this research. First, they want to make the recent staggered DiD literature more accessible to social scientists. They start by explaining the traditional TWFE estimator and the problems that recent research has identified with it<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>. Then they walk through the new alternative estimators in a way that&#8217;s easier to understand<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>. Second, they test how these different estimators perform using real panel data<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a>. They run Monte Carlo simulations with a large sample size and a staggered design where the treatment is spread out over the entire time period, which differs from macro studies where treatment only happens in a few specific periods. In their simulations, they gradually add &#8220;realistic&#8221; complications that deviate from the perfect conditions these estimators were designed for. This lets them see how well each estimator holds up when things aren&#8217;t ideal, which gives practical insights for researchers. At the end, they provide recommendations for best practices when analyzing this type of data.</p><p><em>Why is this important? </em></p><p>TWFE is common practice in applied work. In simple 2x2 settings, it works as it was designed for, but once treatment is staggered across units, it becomes unclear what TWFE is estimating. Instead of a single contrast, it averages over many group comparisons and some of those comparisons involve already-treated units acting as controls, which introduces bias. This matters because many real-world treatment effects are not static. In individual-level panel data, for example, effects often build gradually and then fade out. That inverted-U shape is common in studies of life events, shocks and policy reforms. If you only include a single treatment dummy, you are misrepresenting the data-generating process, and TWFE becomes biased as a result.</p><p>This is one of the reasons why we have so many DiD new methods, but the authors point out that the backlash against TWFE has gone a bit too far. There is nothing wrong with using it as long as you model treatment effect heterogeneity in the right way. Event-time indicators are one simple fix, and while the new estimators are designed to handle dynamic effects, they still rely on assumptions<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a>. When that assumption fails, all estimators perform poorly (and sometimes worse than TWFE).</p><p><em>Who should care? </em></p><p>Anyone doing DiD with more than two time periods. If your treatment is staggered, your effects are dynamic, or your units might anticipate treatment, this paper is super relevant. The audience includes applied economists, sociologists and political scientists working with panel data, particularly those deciding between TWFE and one of the newer DiD estimators.</p><p><em>Do we have code? </em></p><p>The authors say at the end that &#8220;a replication package with the simulation and analysis code is available on the author&#8217;s Github repository [the link will be added]&#8221;. I will update this post when it&#8217;s released. </p><p>In summary, this paper helps bring clarity to a debate that has confused a lot of applied researchers. TWFE is not inherently flawed but it needs to be specified correctly. If treatment effects vary over time, a single treatment dummy won&#8217;t cut it. Switch to an event-time specification or use one of the newer estimators. But remember: all of them rely on assumptions like PT and no anticipation. Each has strengths and weaknesses, and the right choice depends on what you&#8217;re most worried about in your data.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>&#8220;We provide the first semiparametric efficiency bounds for DiD and ES estimators in settings with multiple periods and varying forms of the parallel trends assumption, including covariate-conditional and staggered adoption designs&#8221;. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>Professor Abadie is <a href="https://www.nber.org/system/files/working_papers/w22791/w22791.pdf">generally</a> credited as the primary architect of the SC approach. The method creates a &#8220;synthetic&#8221; version of the treated unit by taking a weighted combination of control units that best matches the pre-treatment characteristics of the treated unit, which makes it possible to estimate causal effects in comparative case studies where traditional experimental methods aren&#8217;t feasible. If I&#8217;m not mistaken and didn&#8217;t miss anything, I think the order is: Abadie and Gardeazabal&#8217;s 2003 &#8220;The Economic Costs of Conflict: A Case Study of the Basque Country&#8221;; Abadie, Diamond and Hainmueller&#8217;s 2010 &#8220;Synthetic Control Methods for Comparative Case Studies: Estimating the Effect of California&#8217;s Tobacco Control Program&#8221; and 2015 &#8220;Comparative Politics and the Synthetic Control Method&#8221;; and finally Abadie&#8217;s 2021 &#8220;Using Synthetic Controls: Feasibility, Data Requirements, and Methodological Aspects&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>A key difference from the original SDiD estimator is that the weights for the synthetic controls are not constrained to be non-negative.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8sWk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8sWk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 424w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 848w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 1272w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8sWk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png" width="1456" height="420" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:420,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:92463,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/166324473?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!8sWk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 424w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 848w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 1272w, https://substackcdn.com/image/fetch/$s_!8sWk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e992092-eae8-4f6c-886b-b54938287df7_1806x521.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>I highly recommend <a href="https://academic.oup.com/ectj/article-abstract/27/3/C1/7701402">this</a> paper.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>The authors say that time-varying covariates can present theoretical challenges by acting as additional treatment variables and are thus beyond the scope of the analysis of their paper. They propose three distinct strategies for applied researchers. The first, full stratification, involves running the analysis separately for each covariate stratum, though the authors advise that this is often impractical as it may result in subsamples that are too small. Their recommended strategy is a practical hybrid model where additive fixed effects can depend on the covariate but the multiplicative factors are assumed to be common across all units. This allows us to aggregate the data in a way that makes it directly compatible with the main Sequential SDiD algorithm. The final strategy is designed for the specific case where treatment adoption is a deterministic function of the covariates (e.g., a policy assigned at the state level). In this setting, the authors suggest it is more natural to average the data within the groups defined by the covariate rather than by adoption cohort.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>This means that the theoretical results are shown to hold even in a &#8220;weak factor&#8221; setting, where the identifying variation from the interactive fixed effects can be small and vanish asymptotically, only as long as it does so slower than the statistical noise. The authors point to the fact that this is especially appealing for empiricists.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>The paper frames its results as an answer to a &#8220;long-standing critique of SC methods&#8221;. The critique questions why a researcher should rely on SC weights instead of directly estimating the underlying factor model that motivates the method. The paper&#8217;s theoretical equivalence result (Theorem 3.1) fixes this tension by formally showing that in large samples, their SC-based method is the same as an oracle that does use the unobserved factors directly. It clarifies that &#8220;the choice is not between balancing and direct estimation, but rather how to feasibly approximate the same ideal oracle benchmark&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>When treatment is staggered over time (some units get treated earlier than others), TWFE doesn&#8217;t estimate a clean average treatment effect, it&#8217;s &#8220;just&#8221; a weighted mix of many 2&#215;2 DiD comparisons. Some of these comparisons are valid (like early-treated vs never-treated), but others are &#8220;forbidden&#8221; (late-treated units compared to already-treated ones), and these can bias results even if treatment effects are constant. The authors argue that the issue is not that TWFE is inherently broken per se, but that using a single dummy (0 = untreated, 1 = treated) assumes effects are the same everywhere and constant over time. We all know that&#8217;s rarely the case because real-world treatment effects fade in, fade out, or vary across groups. When we ignore this, TWFE gets it wrong. </p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p><a href="https://www.sciencedirect.com/science/article/pii/S0304407620303948">Callaway and Sant&#8217;Anna (2021)</a> breaks treatment effects into clean group-by-time contrasts and avoids problematic comparisons; <a href="https://www.sciencedirect.com/science/article/pii/S030440762030378X">Sun and Abraham (2021)</a> is a regression-based version of that same logic; <a href="https://academic.oup.com/restud/article-abstract/91/6/3253/7601390">Borusyak et al. (2024)</a> uses untreated and not-yet-treated units to impute what the treated group would have looked like without treatment; <a href="https://www.tandfonline.com/doi/abs/10.1080/01621459.2021.1891924">Matrix Completion (Athey et al. 2021)</a> is a ML-based imputation method that fills in the untreated matrix using patterns over units and time; and <a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3906345">Wooldridge&#8217;s Extended TWFE (2021)</a> is a more flexible version of TWFE that interacts treatment with time and group indicators. On a funny note, I started my PhD in 2020 and all of this was published in the subsequent years. I still haven&#8217;t finished my PhD, so who knows what will happen in a year from now.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>They simulate 6 realistic scenarios that vary along 4 dimensions: how treatment effects evolve (step, trend-break, inverted-U), whether effects differ across groups, whether units anticipate treatment, and whether treated and untreated units were on PT. For each setup, they look at two things: bias in the overall ATT estimate and bias in time-specific effects. TWFE works fine when effects are static and homogeneous, but once you introduce heterogeneity it breaks (unless you use event-time indicators). Even then, late-period bias can come in. The main finding is quite an intuitive one: no estimator is perfect, some are better with anticipation (Borusyak, Matrix Completion, ETWFE), others handle non-PT better (Callaway and Sant&#8217;Anna, Sun and Abraham). The right choice? As always, &#8220;it depends&#8221; (more specifically, it depends on what you&#8217;re most worried about).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>&#8220;Violations of the parallel trends assumption appear to be more consequential than issues of treatment effect heterogeneity&#8221;. </p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[All About Heterogenous Treatment Effects]]></title><description><![CDATA[The power of disaggregated causal effects.]]></description><link>https://www.diddigest.xyz/p/all-about-heterogenous-treatment</link><guid isPermaLink="false">https://www.diddigest.xyz/p/all-about-heterogenous-treatment</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Fri, 20 Jun 2025 13:31:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!m3yS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m3yS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m3yS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 424w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 848w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 1272w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m3yS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png" width="963" height="1387" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1387,&quot;width&quot;:963,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2541070,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/166140258?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m3yS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 424w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 848w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 1272w, https://substackcdn.com/image/fetch/$s_!m3yS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ec191b7-871d-490a-8d1b-91b4e2d0ecf0_963x1387.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Hello! Today&#8217;s post (somewhat unintentionally) ended up being about four papers that all revolve around the same goal: to push DiD methods further so we can say more than just &#8220;it worked&#8221; or &#8220;it didn&#8217;t&#8221;. You will read about how they help us answer &#8220;for whom did it work? How did it work? And are we even measuring the right thing?&#8221;</p><p>Here are they:</p><ul><li><p><a href="https://arxiv.org/pdf/2502.19620">Triple Difference Designs with Heterogeneous Treatment Effects</a>, by Laura Caron</p></li><li><p><a href="https://arxiv.org/pdf/2505.09706">Forests for Differences: Robust Causal Inference Beyond Parametric DiD</a>, by Hugo Gobato Souto and Francisco Louzada Neto</p></li><li><p><a href="https://arxiv.org/pdf/2506.12207">Estimating Treatment Effects With a Unified Semi-Parametric Difference-in-Differences Approach</a>, by Julia C. Thome, Andrew J. Spieker, Peter F. Rebeiro, Chun Li, Tong Li, and Bryan E. Shepherd</p></li><li><p><a href="https://www.rfberlin.com/wp-content/uploads/2025/06/25019.pdf">Child Penalty Estimation and Mothers&#8217; Age at First Birth</a>, by Valentina Melentyeva and Lukas Riedel</p></li></ul><h3>Triple Difference Designs with Heterogeneous Treatment Effects</h3><p><em>(<a href="https://laurakcaron.github.io/">Laura</a> is a PhD candidate at Columbia)</em></p><h5>TL;DR: triple-difference (3DiD) designs are everywhere, but the way we usually interpret the estimates requires careful consideration. In this paper Laura shows that the usual approach (comparing subgroups assuming one is unaffected) can mislead us and the reader, especially when people in different subgroups respond differently to the treatment (heterogenous treatment effects). She proposes a new way to think about it: instead of comparing average effects (DATT), she focuses on causal differences (causal difference in average treatment effects on the treated, or CDATT), which isolates how much of the difference is actually because of subgroup status, not just because the groups were made up of people who were already likely to respond differently. Laura lays out what needs to be assumed for this to be valid, proposes estimators that still work when models are misspecified, and shows through simulations from real data that this actually changes what we take away from some widely cited studies.</h5><p><em>What is this paper about? </em></p><p>This paper is about how we interpret triple-difference (3DiD) estimates and how easily we can get them wrong if we&#8217;re not careful about subgroup comparisons. </p><p>The ATT (Average Treatment Effect on the Treated) compares treated and untreated groups over time. The key assumption here is that the untreated group shows us what would&#8217;ve happened to the treated group in the absence of treatment. That&#8217;s what makes it &#8220;causal&#8221;: you&#8217;re treating the untreated group as a valid counterfactual. The DATT (Difference in ATT) takes that one step further and compares the magnitude of treatment effects across two subgroups. And most importantly, it doesn&#8217;t assume either subgroup is unaffected. It just compares how strongly each group responded to the treatment. </p><p>3DiD designs use these kinds of comparisons across three dimensions, which usually encompasses time (before vs. after the introduction of a policy), treatment (treated vs. untreated units), and subgroup (which can be a demographic or structural characteristic, e.g. men vs. women, South vs. North). If we were to take a policy mandating paid maternity leave, for example, time would be considered pre vs. post-policy, treatment could be states that implemented it vs. those that didn&#8217;t, and subgroups would be women of childbearing age vs. older women. </p><p>What we usually do is that we assume one subgroup (e.g., older women) is unaffected by the policy. That group then serves as a kind of within-treatment counterfactual. If their outcomes don&#8217;t change post-policy, any observed differences in the other group (e.g., younger women) get attributed to the policy. This gives you an ATT-style interpretation: &#8220;this group was affected by the policy, relative to a group that wasn&#8217;t&#8221;. But that&#8217;s a strong assumption. And more often than not, it&#8217;s incomplete. What if the &#8220;unaffected&#8221; subgroup was actually affected in some way, like indirectly, or just to a lesser degree? In that case, the 3DiD estimate no longer tells you what you think it does. Laura shows that once you allow for heterogeneity in treatment effects, i.e., once you admit that different subgroups might be differently sensitive to the same policy, the DATT no longer captures the causal effect of being in one group versus the other. It just tells you that outcomes moved differently. But that difference might be driven by other things like occupation, baseline risk, or other characteristics correlated with subgroup status.</p><p>To get around this, she defines a new parameter: the causal DATT (CDATT). Instead of just comparing observed differences, it asks: what would the difference in treatment effects have been if both groups had the same underlying sensitivity to treatment, and the only thing that varied was subgroup status itself? That&#8217;s a much &#8220;cleaner&#8221; question, and also a much harder one to answer. The rest of the paper is about how to answer it properly.</p><p><em>What does the author do? </em></p><p>Laura starts with a common practice: using a 3DiD design to compare how different subgroups respond to a policy, assuming one subgroup can stand in as a kind of internal control, which is the standard 3DiD logic used in studies like <a href="https://www.jstor.org/stable/2118071">Gruber (1994)</a>, <a href="https://www.sciencedirect.com/science/article/pii/S092753710300037X">Baum (2003)</a>, and more recently <a href="https://academic.oup.com/qje/article-abstract/136/1/169/5905427">Derenoncourt and Montialoux (2021)</a>. But she shows that this logic breaks down fast when subgroups differ in how sensitive they are to the treatment.</p><p>Even if identification assumptions are satisfied and your DATT is correctly estimated, the result doesn&#8217;t have a clean causal interpretation. Why? Because if people in one subgroup would have responded more strongly to the policy even if they had been in the other subgroup, the difference in outcomes reflects more than just subgroup status: it reflects unobserved differences in treatment sensitivity. That&#8217;s not a subtle footnote because it completely changes what the parameter means.</p><p>To fix this, she defines a new estimand: the causal DATT (CDATT). It asks: what&#8217;s the treatment effect because of belonging to one subgroup versus another, ceteris paribus, including how reactive people are to the policy? That&#8217;s a much more policy-relevant question, but it comes at a cost: stronger assumptions, more careful identification, and better estimation tools.</p><p>The DATT is still identifiable under the usual assumptions (no anticipation and parallel trends across subgroups). So if you&#8217;re just interested in whether two groups responded differently to a policy, the DATT will give you that. But if what you really want to know is why they responded differently (like whether being in subgroup A versus subgroup B caused that difference) then you need to go further. That&#8217;s what the CDATT is for. To identify it, you need one more assumption: that treatment effect heterogeneity isn&#8217;t itself correlated with subgroup membership. Without that, you&#8217;re just comparing aggregates.</p><p>Laura&#8217;s paper includes both simulations and an empirical re-analysis of Gruber&#8217;s maternity leave data. In the simulations, she shows how the usual DATT and her proposed CDATT can lead to very different conclusions depending on the data-generating process. In the empirical application, she revisits a classic 3DiD design and shows how accounting for treatment effect heterogeneity shifts the interpretation, sometimes quite substantially.</p><p>The technical framework is grounded in the <a href="https://books.google.com/books?hl=en&amp;lr=&amp;id=Bf1tBwAAQBAJ&amp;oi=fnd&amp;pg=PR17&amp;dq=Causal+inference+in+statistics,+social,+and+biomedical+sciences&amp;ots=jf_D9c-TCB&amp;sig=1Z0seY73BE3Rk6D3huatX2raqLk">Imbens and Rubin (2015)</a> potential outcomes model, so if that&#8217;s familiar territory, the paper is easier to follow. But even if it&#8217;s not, the intuition behind her argument is very clear: comparing subgroup outcomes doesn&#8217;t tell you anything causal unless you&#8217;re explicit about what you&#8217;re holding constant.</p><p><em>Why is this important? </em></p><p>3DiD designs are <a href="https://academic.oup.com/ectj/article-abstract/25/3/531/6545797">everywhere</a> in applied work. When the usual DiD assumptions do not quite hold (say, when you cannot confidently claim parallel trends between treated and untreated groups) 3DiD offers a way out. It adds a third dimension (typically a subgroup split) and tries to recover the treatment effect by leveraging variation across time, treatment, and some structural or demographic distinction. But that third layer of complexity opens the door to a whole new set of problems.</p><p>While the DiD literature has already wrestled with treatment effect heterogeneity (between treated and control groups, or over time) there was less attention paid to what happens when subgroups themselves differ in how they respond to treatment. And yet, that is exactly what most 3DiD designs rely on: that one subgroup can stand in as a control for another.</p><p>Laura&#8217;s paper fills that gap. It shows that if we are not careful, we can end up interpreting subgroup differences as causal when they are really just driven by selection, sorting, or other underlying differences in treatment sensitivity. And if the whole point of 3DiD is to recover cleaner effects in messy empirical settings, then failing to address this undermines the design at its core.</p><p>Laura offers a formal framework for thinking about causal subgroup comparisons in 3DiD, shows what needs to be assumed, and provides robust estimators that work in the presence of heterogeneity. It is a nice correction to how subgroup analyses are often handled in 3DiD settings, and it is really useful now that staggered treatments and increasingly rich data structures are becoming more common.</p><p><em>Who should care? </em></p><p>Anyone using 3DiD to compare subgroups. If your paper assumes one group is not affected by the policy, or if you say &#8220;this group was more affected than that one&#8221;, then you should check Laura&#8217;s paper out. It matters most if treatment happens at different times (staggered), there&#8217;s the possibility of spillovers, and you&#8217;re comparing groups like men vs. women, or states in the North vs. South. Also if your story is about why the policy worked more for one group than another, this paper shows you what you need to assume to make that claim.</p><p><em>Do we have code? </em></p><p>No public code, but the paper walks through both a simulation and an empirical example using Gruber (1994). The simulation shows how DATT and CDATT can tell different stories depending on how treatment effects vary. The empirical example shows how ignoring that variation can lead you to the wrong conclusion.</p><p>In summary, 3DiD designs are everywhere, but the way we interpret subgroup comparisons often leans too hard on assumptions we do not check. This paper shows how easily treatment effect heterogeneity can throw things off and offers a better way to frame and estimate what we really care about. If you&#8217;re using subgroup comparisons to tell a causal story, this is the paper that tells you what that story actually needs.</p><h3>Forests for Differences: Robust Causal Inference Beyond Parametric DiD</h3><h5>TL;DR: this paper introduces DiD-BCF, a new Bayesian ML method that makes it easier to estimate treatment effects in DiD settings, particularly when we &#8220;have&#8221; heterogenous treatment effects and staggered adoption. It builds on Bayesian Causal Forests, but reworks the model to fit panel data and common policy setups. The key innovation is a way to make the estimation task simpler by using the parallel trends assumption more effectively. The result is more accurate and flexible estimates of average, group-specific, and conditional treatment effects.</h5><p><em>What is this paper about?</em></p><p>The paper proposes a new non-parametric method for causal inference in DiD settings. It is called DiD-BCF and is designed to deal with two major challenges in modern DiD applications: staggered treatment adoption and treatment effect heterogeneity. To do this, the authors extend Bayesian Causal Forests to panel data settings. Their method estimates: average treatment effects (ATT), group-level effects (GATT), and conditional effects (CATT), all within a single flexible framework. A key feature of the method is a new way of using the parallel trends assumption to simplify the estimation task. This helps improve the accuracy and stability of the results, especially when standard parametric DiD methods struggle.</p><p><em>What do the authors do?</em></p><p>They develop a new method (DiD-BCF) that combines the credibility of DiD with the flexibility of Bayesian Causal Forests. Instead of relying on fixed-effects regressions or linear models, they propose a fully non-parametric approach that can handle staggered treatment adoption and treatment effect heterogeneity in a single unified framework.</p><p>They start by generalising the standard DiD model, allowing for complex, nonlinear relationships between outcomes, time, and covariates. Then they extend Bayesian Causal Forests to panel data, so the model can recover dynamic and covariate-specific treatment effects over time. A key part of their strategy is to reparameterise the model using the parallel trends assumption, so instead of forcing the model to learn that treatment effects should be zero before treatment begins, they build that into the structure from the start. This makes estimation more stable and less prone to error.</p><p>To make the model more practical for applied researchers, they also include a procedure that speeds up convergence using a more efficient tree-fitting algorithm before switching to full Bayesian estimation. Finally, they test their method across a range of simulated settings that mirror real-world challenges, such as nonlinear trends, selection into treatment, and heterogeneity in effects, and compare it to leading alternatives like TWFE, DiD-DR, DiD2s, SDID, and DoubleML. Across the board, DiD-BCF performs well, especially in the kinds of cases where standard methods tend to break down. </p><p><em>Why is this important?</em></p><p>In observational settings, we rarely get (to do) randomised experiments. Treated and control units usually differ in ways that matter, whether in background characteristics, exposure, or timing, which makes simple comparisons misleading. That is why quasi-experimental methods like DiD have become central to applied work. DiD gives us a way to estimate causal effects from policy changes or discrete events, as long as the identifying assumptions can be justified.</p><p>Over the past few years, there has been a wave of new methods that improve DiD by dealing with common problems like staggered treatment timing or variation in treatment intensity. But most of these still focus on estimating average effects, either the overall ATT or group-time averages, which is useful, but often not enough.</p><p>In many real-world applications, the question goes beyond by how much did the policy work on average, it also matters for whom did it work. We need to understand the heterogeneity in treatment effects if we want to say something about mechanisms, or about which groups benefit more or less from an intervention. This is where CATT come in: they tell us how effects vary across observable characteristics.</p><p>To estimate CATTs, researchers have started turning to ML tools like Causal Forests. These models are flexible enough to recover treatment effect variation without needing to specify it upfront, and recent work has adapted them for use in DiD settings. If done carefully, this kind of approach lets us combine the identification logic of DiD with the flexibility of non-parametric estimation, enabling us to detect whether something worked, for whom it worked, and when.</p><p>This paper fits right into that agenda. It pushes the literature forward by offering a way to estimate CATTs in panel data under staggered adoption, all while maintaining interpretability and robustness. That makes it especially useful for policy evaluations where average effects might hide meaningful variation across time, space, or population subgroups.</p><p><em>Who should care?</em></p><p>This paper will be especially useful to applied researchers working with panel data, where treatment doesn&#8217;t happen all at once and where you suspect that not everyone is affected in the same way. If you work in labour, education, health, or policy evaluation more generally, and you&#8217;re kinda frustrated with the limitations of TWFE or concerned about heterogeneous effects being averaged away, this model is worth looking into.</p><p>It&#8217;s also useful for people who want unit-level or group-specific effects in addition to a single, aggregate ATT. And for researchers who want to use ML tools but still work within a familiar causal inference framework like DiD &#8594; this is a bridge between the two.</p><p><em>Do we have code?</em></p><p>The authors rely  on existing packages like <code>did2s</code>, <code>did</code>, and <code>DoubleML</code>, and while their proposed model is based on well-established tools like BCF and XBART, some of the benchmark methods (like CFFE and MLDID) couldn&#8217;t be included in the simulations due to GitHub installation issues. Still, the core components for reproducing DiD-BCF and the comparisons are all available or well-documented. If you&#8217;re familiar with R and Bayesian modelling, you should be able to adapt their setup pretty easily.</p><p>In summary, DiD-BCF is an interesting new tool that brings flexibility to DiD. It replaces rigid parametric assumptions with a fully Bayesian, nonparametric model that can handle staggered treatment timing, heterogeneity in effects, and selection on observables, all in one go. By leaning on Bayesian Causal Forests and smartly reparameterising the treatment effect, the method improves estimation accuracy without giving up the core logic of DiD. The result is a model that performs well even when traditional approaches break down, and that opens the door to much richer treatment effect analysis in applied work.</p><h3>Estimating Treatment Effects With a Unified Semi-Parametric Difference-in-Differences Approach</h3><h5>TL;DR: most DiD methods focus on average treatment effects (ATT) and assume parallel trends in means. But when the outcome is skewed, ordinal, or censored, estimating means can be misleading or hard to interpret. In this paper the authors introduces a new semi-parametric DiD estimator that allows researchers to estimate four treatment effects: average, quantile, probability, and Mann-Whitney, using a single model and a unified assumption. The method performs well in simulations and is applied to evaluate the impact of Medicaid expansion on CD4 counts among people living with HIV. </h5><p><em>What is this paper about?</em></p><p>In this paper the authors present a new DiD estimator that can recover a quite comprehensive range of causal effects, not just averages but also quantiles, probabilities, and rank-based measures. And they do using a single model and a single identification assumption. Their key idea is to replace traditional mean-based comparisons with a <a href="https://www.tandfonline.com/doi/full/10.1080/01621459.2024.2315667">semi-parametric cumulative probability model (CPM)</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. Instead of assuming a specific transformation (like logs) to deal with skewed outcomes, the CPM treats the outcome as a monotonic transformation of a latent variable. </p><p>The authors focus on four causal estimands, all defined in terms of the marginal distributions of potential outcomes for the treated group: average effect on the treated (ATT) the effect on a specific quantile of the treated outcome distribution (QTT), the change in the probability that the outcome is below a given threshold (PTT), and a rank-based measure that captures the probability a treated person would outperform a comparable untreated person (MTT). Because these estimands depend only on marginal, not joint, potential outcome distributions, the method requires fewer assumptions than alternative approaches. All four effects are then identified using a single conditional parallel trends assumption, made on the latent scale (rather than on the original outcome scale or after transformation). This assumption states that, conditional on covariates, the untreated change over time for the treated group would have followed the same trend as the control group not in observed outcomes, but in the underlying latent variable used by the model.</p><p><em>What do the authors do?</em></p><p>The authors&#8217; proposed DiD estimator can recover the four aforementioned types of causal effects (ATT, QTT, PTT, and MTT) using a single model and a single identification strategy, which is remarkable because it contrasts most existing approaches where estimating each of these effects would require a separate set of assumptions and often a different estimation method. </p><p>To operationalise their approach, the authors define a semi-parametric DiD estimator based on a CPM, allowing them to handle skewed, ordinal, or otherwise non-linear outcome distributions without requiring a pre-specified transformation. Then they move on to formally define the four estimands using only marginal potential outcome distributions for the treated group. This keeps things simpler by not trying to model how each treated person&#8217;s outcome would have compared to their own untreated potential outcome, which usually requires stronger, less realistic assumptions. </p><p>As mentioned, they rely on a single identification assumption, which is a conditional parallel trends assumption on the latent outcome scale. Unlike our parallel trends in means (or transformed means), their assumption is made on the underlying latent variable.  They then evaluate performance through simulation, varying sample sizes (n = 200 to 2,000) and generating a large pseudo-population to benchmark the true values of each estimand. The simulations show good performance even in small samples and under data skewness.</p><p>In the last part of the paper, they apply the method to real data by estimating the impact of Medicaid expansion in the U.S. on CD4 cell count at HIV care entry (a classic example of a skewed, clinically relevant outcome). They find that the policy had a broad positive impact, affecting all four estimands.</p><p><em>Who should care?</em></p><p>Have you log-transformed your variables before? Does your data look not-normal? If so, you should read this paper. Also pretty much anyone else using wanting to model DiD with &#8220;non-standard outcomes&#8221; (skewed, ordinal, censored, etc) will find this paper useful. I&#8217;d say it&#8217;s particularly useful for applied researchers in health and labour economics, where we often find rank-based or distributional effects to matter more than mean shifts.</p><p><em>Do we have code?</em></p><p>Not yet, but hopefully in the future.</p><p>In summary, this paper introduces a flexible, semi-parametric DiD method that lets you estimate a wide range of treatment effects using a single model and a single identification assumption. It&#8217;s a compelling alternative to the usual mean-based approaches, especially when dealing with skewed or messy outcomes. And by including estimands like the Mann-Whitney<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> treatment effect, it opens the door to more interpretable, rank-based causal questions that standard DiD methods often lack.</p><h3>Child Penalty Estimation and Mothers&#8217; Age at First Birth</h3><h5>TL;DR:  ever heard economists saying &#8220;the gender pay gap is mostly due to motherhood penalty&#8221;? This paper takes this question and shows we might be underestimating the motherhood penalty (in wage terms) by about 30% because we&#8217;re averaging across mothers who are nothing alike (or, in other words, standard methods bias the estimated penalty by ignoring staggered timing and treatment heterogeneity). The authors then propose a cleaner DiD approach that estimates the penalty separately by age at first birth, and as a result they find big differences both in the penalty&#8217;s size and in what it really means for younger vs older mothers.</h5><p><em>What is this paper about?</em></p><p>&#8220;Motherhood is still costly for the careers of women&#8221;. In this paper, the authors take something we all say (that the gender pay gap is largely driven by the career costs of motherhood) and ask a simple, yet a bit uncomfortable, question: what if we&#8217;ve been measuring those costs incorrectly? The standard approach so far has been a big event study centred on the first birth, tracking earnings before and after, and showing that women&#8217;s earnings fall sharply while men&#8217;s barely move a centimeter. The problem, the authors argue, is that this kind of model implicitly assumes that all mothers are the same. But of course they&#8217;re not. The age at which someone has their first child is correlated with education, career stage, earnings, occupation, and parental background (and lots of other unobserved decisions). So when we pool all mothers together, we&#8217;re collapsing both very different types of women and very different types of penalties into one average, and that average not only misleads but also obscures meaningful variation policymakers should care about.</p><p>They find lots of interesting bits. First, the penalty is actually larger than previously estimated. In their preferred specification, there&#8217;s a cumulative earnings loss of nearly &#8364;30,000 by year four (about &#8364;10,000 more than what a standard event study would suggest), which isn&#8217;t a small correction. It reflects the fact that conventional event studies systematically understate what women&#8217;s earnings would have looked like in the absence of children, largely because their control groups include women who have already had children themselves. That violates the parallel trends assumption and biases the counterfactual downwards, making the penalty look smaller than it actually is.</p><p>Second, the penalty grows with age in absolute terms, but shrinks in relative terms. Older mothers lose more in euros because they were earning more to begin with. But younger mothers lose a larger share of their pre-birth income because they&#8217;re cut off just as their wage growth would have accelerated.</p><p>Third (and this is what I found most interesting), the nature of the penalty differs. For older mothers, the penalty is mostly about reducing hours, exiting temporarily, or giving up seniority. For younger mothers, it&#8217;s about missing the steepest part of the wage trajectory entirely. It&#8217;s not the same shock. One is a level shift. The other is a slope change. This is super important from a policy perspective: helping someone who missed a promotion track requires a different tool than helping someone who stepped back from senior management, for example.</p><p><em>What do the authors do?</em></p><p>The paper starts from a now-familiar problem in applied work: when treatment timing is staggered and treatment effects vary across units, conventional event study models can break down. In this case, those models do two things we definitely want to avoid. First, they make forbidden comparisons by including already-treated women in the control group. Second, they suffer from contamination, where estimates for one time period &#8220;bleed&#8221; into another, especially when pre- and post-birth windows aren&#8217;t cleanly separated.</p><p>The root of the problem is that standard event study models pool together younger and older first-time mothers, implicitly assuming they&#8217;re comparable and that the effects of childbirth are uniform across them. But that assumption, as the authors put it, is &#8220;unlikely to hold.&#8221; Mothers at different ages differ in earnings, career stage, education, and trajectories. Pooling across them &#8220;masks&#8221; both the size and the shape of the penalty.</p><p>To fix this, the authors propose a stacked DiD design (closely following <a href="https://www.nber.org/papers/w32054">Wing, Freedman, and Hollingsworth 2024</a>, with attention to overlapping cohorts), combined with a rolling window of control groups by age at first birth. Instead of estimating one big average effect, they estimate separate effects for each age-at-birth group, and they&#8217;re very careful about who gets used as a counterfactual. For each group of treated mothers, the control group is made up of not-yet-treated women who will give birth at slightly older ages, observed before they become mothers. So if you&#8217;re looking at women who give birth at 25, the control group might be 26- or 27-year-olds who are about to give birth but haven&#8217;t yet. That way, everyone in the comparison is close in age, close in life stage, and still on the pre-birth trajectory.</p><p>As they describe it, this &#8220;combination of a stacked DiD with a rolling window of control groups enables [them] to eliminate the issues present in conventional event studies and estimate the age-at-birth-specific effects of childbirth on post-birth labor market outcomes.&#8221; No already-treated mothers in the control group. No artificial smoothing across life stages. Just clean, age-specific estimates of what motherhood does to earnings.</p><p>Once the age-specific penalties are estimated, they&#8217;re aggregated using sample shares as weights. The resulting average is substantially larger than the pooled estimate, and, most importantly, it makes clear that instead of being one number, the penalty is a set of distinct experiences depending on when motherhood begins.</p><p><em>Why is this important?</em></p><p>The motherhood penalty is one of the largest contributors to gender inequality in the labour market, but if we measure it incorrectly, any policy built on that estimate risks being too late, too generic, or simply misdirected. What this paper shows is that how we estimate the penalty matters. Conventional event studies often use control groups that include women who have already had children, violating parallel trends and biasing the counterfactual downwards. That alone is enough to underestimate the true cost of motherhood, especially because what is being missed isn&#8217;t just earnings levels but the growth that would have happened absent childbirth. The timing of motherhood matters because the nature of the penalty changes with age: reducing hours or exiting work hits harder when your earnings are high, but missing wage growth early on can permanently derail a trajectory. This is a methodological paper, but it goes well beyond that. By analysing the effects of motherhood by age at first birth, it opens the door to more targeted and better-informed policy, recognising that women who have children at different stages in their life and career also respond differently to support and constraints.</p><p><em>Who should care?</em></p><p>If you work on gender gaps and use event studies, this paper is a must-read. Same goes if you&#8217;ve ever used the phrase &#8220;motherhood penalty&#8221; without checking whether your estimate holds across women with different life paths. But it also matters more broadly for anyone interested in how we measure inequality, because it reframes the timing of a treatment as a source of information, and not merely a nuisance to be adjusted for. If we keep pooling heterogeneous groups, we risk erasing the very dynamics we claim to study. So even if you do not study gender, if you work with staggered treatments, event-time estimators, or stacked DiD designs, this is a paper worth reading and assigning to your students.</p><p><em>Do we have code?</em></p><p>It&#8217;s an application paper, so no.</p><p>In summary, this paper is a methodological critique with real-world stakes. It shows that when we average across all mothers to estimate the cost of motherhood, we get a number that is not just wrong but misleading. The timing of motherhood shapes the size, shape, and meaning of the penalty, and by pooling across it, we risk flattening nuance into noise. The stacked DiD approach they propose fix the bias and brings to the surface an economically and policy-relevant heterogeneity that would otherwise be lost. If we care about inequality, then we also need to care about how we measure it.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>A cumulative probability model (CPM) models the probability that the outcome is less than or equal to a given value (P(Y &#8804; y)), based on covariates. Instead of modelling Y directly, it assumes there&#8217;s a latent variable Y* that follows a linear model, and the observed outcome Y is a monotonic transformation of Y*. The function linking Y* and Y (denoted H) is unspecified and estimated from the data, making the model semi-parametric. This allows the CPM to flexibly handle skewed, ordinal, or censored outcomes without needing to pre-specify a transformation like log(Y).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>This paper is a joint work between authors from Biostats, Public Heath and Economics departments, so the Mann-Whitney parameter might not be familiar to us. The Mann-Whitney parameter comes from the Mann-Whitney U test (also called Wilcoxon rank-sum test), and it&#8217;s often interpreted as the probability that a randomly selected observation from group A is larger than one from group B. So if group A is treated and group B is untreated, and the probabilistic index is 0.65, it means there&#8217;s a 65% chance a treated person scores higher than an untreated one. In the paper they define a Mann-Whitney treatment effect among the treated (MTT), which is a probabilistic measure of treatment impact. The Mann-Whitney parameter is a descriptive, non-causal rank-based comparison between two groups, while the MTT turns that idea into a causal estimand in DiD by comparing treated individuals&#8217; actual outcomes to their estimated counterfactuals under no treatment, using the rank-based structure of the Mann-Whitney statistic. Instead of asking how much the outcome increased (like the ATT does), it asks a different question: what is the probability that a treated person has a better outcome than a comparable untreated person? As an example, if the MTT is 0.70, it means there&#8217;s a 70% chance that a randomly selected treated individual will have a higher outcome than a randomly selected untreated individual. If the MTT is 0.50, it means treatment had no effect on rank &#8594; treated and untreated units perform equally, on average. The authors show how to estimate the MTT in a DiD setup, and according to them, this is the first time anyone has done this.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[In defense of Machine Learning in Economics]]></title><description><![CDATA[Why the common criticisms are outdated, and how new methods are making causal inference stronger.]]></description><link>https://www.diddigest.xyz/p/in-defense-of-machine-learning-in</link><guid isPermaLink="false">https://www.diddigest.xyz/p/in-defense-of-machine-learning-in</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Tue, 10 Jun 2025 11:19:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Pc8A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Pc8A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Pc8A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Pc8A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:69729,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/165538392?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Pc8A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Pc8A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b78580a-83aa-4338-be7b-8fb557cde4b7_1024x1024.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Edit:</em> hello! It&#8217;s August 2026 now and we need some updates. Nothing here needs retracting because - unfortunately - some criticisms I listed are still the ones you&#8217;ll hear in seminars. The replies should still work, so that should be ok. What has changed is defining where the frontier of the argument sits. In 2025 I was defending LASSO and causal forests, and in 2026 the same debate is replaying itself, almost beat for beat, with LLMs. So consider this<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> an 8th entry for the list. A few housekeeping notes on the original post: the Chernozhukov et al. practical introduction I linked in footnote 8 has since turned into a full open textbook, <a href="https://causalml-book.org">Applied Causal Inference Powered by ML and AI</a>, which is now the standard entry point to the DML literature. And on the &#8220;it&#8217;s just a fad&#8221; point: the 2026 Northwestern <a href="https://www.law.northwestern.edu/research-faculty/events/conferences/causalinference/advanced/">advanced causal inference workshop</a> lists ML methods for causal inference as one of its principal topics, alongside the DiD material we usually cover. When the methodological establishment puts your &#8220;fad&#8221; on the syllabus next to well-established causal inference methods, the argument is pretty much settled.  The rest of the post continues as originally written, below.</p><div><hr></div><p>Hi there. This post is a &#8220;response&#8221; to a <a href="https://arxiv.org/abs/2409.14202">pre-print</a> I tweeted about last Friday. I read it and thought it was great, and I&#8217;ll write about it in the next newsletter post. But that paper per se is not the topic today, and I apologize if the subject is not of your interest. Please feel free to ignore it if you don&#8217;t need convincing :)</p><p>This post is about me making the case FOR the use of ML in Econ. It feels kind of dated to have to still argue for this, and if &#8220;<a href="https://www.annualreviews.org/content/journals/10.1146/annurev-economics-080217-053433">our goal as a field is to use data to solve problems</a>&#8221; (which I know some will say it isn&#8217;t), then the criticisms warrant even less consideration. I believe the reluctance to adopt ML techniques in Econ stems from two sources<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>: orthodoxy (the tendency to stick to familiar models and econometric traditions) and purity (the idea that if a method isn&#8217;t explicitly tied to a structural model or doesn&#8217;t yield a &#8220;clean&#8221; causal estimate, it&#8217;s not real economics).</p><p>Part of this reluctance also comes from a misperception: that ML is <em>only</em> useful for prediction and therefore has little to contribute to causal inference or theory-driven analysis. But that&#8217;s increasingly out of step with how ML is actually used in empirical research. Consider causal forests which use tree-based methods to identify heterogeneous treatment effects while maintaining the rigor economists demand. Or double machine learning, which combines ML&#8217;s predictive power with traditional causal identification strategies to reduce bias in treatment effect estimation. Regularization techniques (e.g., TMLE) help us focus on the parameters we care about most while still leveraging ML&#8217;s ability to handle high-dimensional data. Rather than talking about replacing economic reasoning, we can reason from the point of view of enhancing it in a way that helps us extract more nuanced insights from complex data while maintaining causal rigor. </p><p>I like to think of these tools as complements, not substitutes, for economic theory. Just like we once borrowed techniques from Stats and Maths, we&#8217;re &#8220;now&#8221; borrowing from CS and engineering. That&#8217;s how fields evolve and strengthen. Professor Athey <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">writes</a>: &#8220;I believe that regularization and systematic model selection have many advantages over traditional approaches, and for this reason will become a standard part of empirical practice in economics&#8221;. </p><p>Another point that makes it harder to bridge these two fields, though, is that the culture around publication and validation in CS is so different: pre-prints are fast, feedback is open, and iteration happens in public. In Econ, we move at glacial speed, and nothing feels &#8220;legit&#8221; until it survives multiple rounds of refereeing. Another point to consider is that criticism often comes from academia. It barely ever comes from Economists at Amazon or Alphabet. Many of the advances in causal ML (like uplift modeling, heterogeneous treatment effects, online experimentation) were incubated in tech companies because they had the scale, the data, and the incentives to care. </p><p>Now, I&#8217;m obviously not suggesting we should abandon academic rigor for industry speed, and I believe both approaches have their strengths. Industry settings often provide the scale and urgency that drive innovation, while academic settings provide the time and incentives for thorough validation. But we can learn from both. </p><p>Just last year, 2/3 of the Nobel in Chemistry was awarded to researchers at Google DeepMind for the creation of AlphaFold<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. This kind of breakthrough didn&#8217;t come out of a &#8220;traditional&#8221; academic lab. It came out of industry research, using ML tools to push the frontier of a completely different field. If Chemistry is giving out Nobels for this work, maybe it&#8217;s time for us to stop treating ML like it&#8217;s not serious science<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>.</p><p>And beyond the science, there&#8217;s a very practical angle too: the academic job market is brutal. There are more brilliant PhDs than tenure-track positions, and even fewer in policy-facing or methodological roles that reward creativity. If we don&#8217;t expose students to modern tools (*especially* ones that are widely used in industry) we&#8217;re not preparing them (and I include myself here) for the full range of careers they may need or want to pursue. The industry isn&#8217;t waiting for us to catch up. The least we can do is make sure our students aren&#8217;t left behind.</p><p>I&#8217;ve had to defend these positions a couple of times already, and I think writing about the experience (in order to prepare others for what they&#8217;ll likely face if they try to do the same) is worth doing. To make this case more concrete, let me center you in the current state of affairs. Understanding the context helps clarify why the current resistance feels both familiar and ultimately misguided.</p><p>But before we move to the critical points, it&#8217;s necessary to acknowledge the giants pushing the fields to intersect: Professors Athey at Stanford and Chernozhukov at MIT. They&#8217;re leading figures at two of the most prestigious Econ departments in the world, representing a geographic and intellectual complementarity that&#8217;s been central to the field&#8217;s development<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>. Professor Athey, who has one foot in Silicon Valley and the other in academia (she was Chief Economist at Microsoft Research), has done a lot to make ML legible to economists (especially by focusing on causal questions and by making the theory accessible to <a href="https://d2cml-ai.github.io/mgtecon634_r/intro.html">undergrads and postgrads</a>). Her work on causal trees and causal forests is all about taking the structure economists care about and building in the flexibility that real-world data demands. Professor Chernozhukov, meanwhile, has taken the inference side (which makes sense since he&#8217;s an econometrician). His work on high-dimensional econometrics (aka the double/debiased ML framework with Belloni and others) showed how you can combine ML&#8217;s predictive power with solid causal identification. It&#8217;s basically the answer to the &#8220;but what about inference?&#8221; pushback you still hear in seminars. Together with Professor Athey&#8217;s frequent collaborator Professor Imbens and early adopters like Professor Mullainathan, they&#8217;ve helped shift the conversation from &#8220;prediction versus causation&#8221; to &#8220;prediction in service of causation&#8221;.</p><p>To put their contributions in perspective, it&#8217;s worth stepping back and tracing how ML and economics first started to intersect. It&#8217;s no easy task to figure out where these two &#8220;fields&#8221; first intersected, or to pinpoint the single <em>first</em> economics paper that used what we now broadly categorize as ML techniques, partly because the definition of "machine learning" has evolved, and partly because some foundational methods were adopted gradually. If we look at when specific techniques came into being, for example, we can start with Ridge Regression (a form of L2 regularization), introduced by Hoerl and Kennard in 1970, and LASSO (Least Absolute Shrinkage and Selection Operator, a form of L1 regularization), introduced by Robert Tibshirani in 1996. </p><p>These regularization methods weren&#8217;t created with economists in mind, but they quickly became useful for us as datasets got bigger and models more complex. Still, for a long time, they mostly lived on the margins of applied work. You might see LASSO show up in a robustness check or in a variable selection appendix, but it wasn&#8217;t central to the analysis. That started to change when the tools stopped being just about prediction and started helping us answer causal questions more transparently.</p><p>The real shift came when economists began to recognize that ML could be used reduce overfitting and improve out-of-sample performance, and also to solve problems we already cared about like estimating treatment effects in high-dimensional settings or uncovering heterogeneity we suspected was there but couldn&#8217;t capture with linear models. Suddenly, ML wasn&#8217;t a detour away from economic reasoning, it was more like a way to get closer to the truth in messy, real-world data.</p><p>This is where the crossover with causal inference really took off. Methods like causal forests, double ML, and targeted regularization didn&#8217;t just borrow from CS, they adapted those tools to meet our standards for identification and interpretation. That&#8217;s why they stuck. And that&#8217;s why people like Professors Athey, Chernozhukov, and others were so effective at building the bridge: they spoke the language of economists <em>and</em> they understood the technical machinery involved. So, when people say &#8220;ML is just prediction&#8221;, you can reply that this is an outdated perspective that overlooks the substantive evolution happening in the empirical literature. The goal was never to abandon causality. The goal was (and is) to do better causal inference <em>with better tools</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a>. That&#8217;s a framing I think is worth emphasizing.</p><p>Now that you more or less know where we are at the moment and why we need to talk about this, I have a list of the most common criticisms I&#8217;ve heard and the way I think we can address them is by pointing at the literature while we try to make the most our of it without losing sight of reality. </p><p><strong>1. &#8220;It&#8217;s just curve-fitting/overfitting&#8221;</strong></p><p>You&#8217;ve probably heard some version of this: ML models are too flexible and will fit noise rather than signal, they don&#8217;t generalize well out-of-sample, and economic relationships should be &#8220;parsimonious&#8221;, not complex. This criticism fundamentally misunderstands how modern ML works. The claim that ML models &#8220;overfit&#8221; ignores that the entire ML workflow is designed around out-of-sample validation<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>, which is something traditional econometrics often skips entirely. <a href="https://pubs.aeaweb.org/doi/pdf/10.1257/jep.31.2.87">Mullainathan and Spiess (2017)</a> make this point directly: ML&#8217;s obsession with predictive performance makes models more reliable, not less, because it forces researchers to test whether their results generalize beyond the training data. When Belloni et al. (<a href="https://onlinelibrary.wiley.com/doi/pdf/10.3982/ecta9626">2012</a>, <a href="https://pubs.aeaweb.org/doi/pdf/10.1257%2Fjep.28.2.29">2014</a>) applied LASSO regularization to causal inference problems, they improved treatment effect estimation by selecting relevant controls while discarding noise, and thus avoiding overfitting. The &#8220;parsimony&#8221; argument is equally misguided. As <a href="https://pubs.aeaweb.org/doi/pdf/10.1257%2Fjep.28.2.3">Varian (2014)</a> points out, when economic relationships are genuinely complex, forcing them into overly simple models creates more bias than allowing appropriate complexity. The solution is to use principled methods like regularization that balance complexity with generalizability<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> and not to pretend the world is simple. We see this in practice: <a href="https://academic.oup.com/qje/article-pdf/133/1/237/30636517/qjx032.pdf">Kleinberg et al. (2018)</a> showed that ML models predicting judicial decisions outperformed simple heuristics precisely because they could handle complexity without overfitting. The irony is that economists criticize ML for &#8220;curve-fitting&#8221; while often using specification searches and robustness checks that are far more prone to overfitting than properly cross-validated ML approaches. </p><p><strong>2. &#8220;Black box problem&#8221;</strong></p><p>This one comes always comes up: you can&#8217;t interpret the results, there&#8217;s no economic intuition behind the relationships, and policymakers need to understand <em>why</em> something works, not just <em>that</em> it works. But this criticism conflates older ML methods with modern approaches that are explicitly designed for interpretability. <a href="https://muse.jhu.edu/pub/56/article/793356">Athey and Wager&#8217;s (2019)</a> causal forests give you treatment effects and tell you exactly which covariates drive heterogeneity and how. You can&#8217;t call it a black box; that&#8217;s more interpretable than most traditional econometric models that assume homogeneous effects. Tools like SHAP values (<a href="https://proceedings.neurips.cc/paper/2017/file/8a20a8621978632d76c43dfd28b67767-Paper.pdf">Lundberg and Lee, 2017</a>) and LIME (<a href="https://arxiv.org/pdf/1606.05386">Ribeiro et al., 2016</a>) can decompose any model&#8217;s predictions into interpretable components, showing exactly how each variable contributes to the outcome. When the aforementioned Kleinberg et al. (2018) built ML models to predict judge decisions, they extracted insights about which case characteristics actually matter for judicial outcomes, information that was invisible in traditional analyses. We can&#8217;t simply consider every single variable alone and for its relevance. The deeper issue is that &#8220;interpretability&#8221; often means &#8220;fits my priors&#8221; rather than &#8220;reveals true mechanisms&#8221;. A linear regression with 20 control variables isn&#8217;t inherently more interpretable than a well-designed tree-based model that shows you which interactions matter the most. As <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">Athey (2018)</a> argues, ML tools can enhance economic intuition by revealing patterns and relationships that traditional methods miss. The question is whether we&#8217;re willing to update our understanding when the data suggests more complex relationships than our simple models assumed.</p><p><strong>3. &#8220;Prediction &#8800; causation&#8221;</strong></p><p>This is probably the most fundamental objection: ML is only good for forecasting, not understanding causal relationships; economics is about identifying causal effects, not just correlations; you can&#8217;t do policy analysis without causal identification. After all, we did have an entire revolution concerning this worry. But this criticism is based on a false dichotomy that ignores how modern causal ML works. <a href="https://academic.oup.com/ectj/article-pdf/21/1/C1/27684918/ectj00c1.pdf">Chernozhukov et al. (2018)</a> showed exactly how to combine ML&#8217;s predictive power with rigorous causal identification in their double/debiased ML framework<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a>: you use ML to estimate nuisance parameters (like propensity scores or outcome models) while maintaining all the causal identification you care about. <a href="https://muse.jhu.edu/pub/56/article/793356">Athey and Wager (2019)</a> took this further with causal forests by demonstrating how tree-based methods can identify heterogeneous treatment effects with full causal rigor. These are causal inference methods that happen to use ML tools rather than &#8220;prediction&#8221; methods. A deeper insight comes from <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC4869349/pdf/nihms776714.pdf">Kleinberg et al. (2015)</a>, who pointed out that many policy problems ARE prediction problems: predicting who will benefit most from a treatment, predicting which interventions will work in which contexts, predicting the consequences of policy changes. When you frame it this way, the prediction vs. causation distinction starts to look fake. <a href="https://academic.oup.com/restud/article/81/2/608/1523757">Belloni et al. (2014)</a> solved the lingering inference concerns by showing how to do valid statistical inference after using ML for variable selection. The result? &#8220;Better causal inference through better prediction&#8221; rather than &#8220;prediction instead of causation&#8221;. When <a href="https://www.jstor.org/stable/pdf/44250458.pdf">Davis and Heller (2017)</a> used causal forests to understand treatment heterogeneity in job training programs, they were making causal inference more flexible and informative than traditional approaches ever could, not abandoning it.</p><p><strong>4. &#8220;No economic theory&#8221;</strong></p><p>Ok, we are halfway through these. Here&#8217;s another one you hear all the time: ML methods are atheoretical: they let the data speak without economic reasoning, economics should be driven by theory not algorithmic pattern recognition (whatever it means), and results lack the structural interpretation needed for counterfactuals. But this gets the relationship between theory and ML backwards. <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC4869349/pdf/nihms776714.pdf">Kleinberg et al. (2015)</a> showed that prediction problems are often at the heart of economic theory: optimal taxation requires predicting behavioral responses, welfare analysis requires predicting who benefits from policies, market design requires predicting how agents will behave under different rules. ML makes this reasoning more rigorous by forcing us to be explicit about our assumptions and test whether our theories predict well. <a href="https://pubs.aeaweb.org/doi/pdf/10.1257/jep.31.2.87">Mullainathan and Spiess (2017)</a> put it perfectly: ML improves traditional econometric practice by providing better tools for the things we were already trying to do. When <a href="https://www.nber.org/system/files/working_papers/w23276/w23276.pdf">Gentzkow et al. (2019)</a> used text analysis to study media bias, they were testing theories about how competition affects content in ways that were impossible before ML tools existed, rather than discarding abandoning economic theory. The &#8220;structural interpretation&#8221; concern is equally misguided. <a href="http://proceedings.mlr.press/v70/hartford17a/hartford17a.pdf">Hartford et al. (2017)</a> showed how to combine deep learning with IVs for structural estimation (akin to what the paper that started this whole conversation was trying to do). <a href="https://www.jstor.org/stable/pdf/43821932.pdf">Bajari et al. (2015)</a> demonstrated how ML can improve demand estimation while maintaining full economic interpretation. ML doesn&#8217;t lack theory (although practice tends to precede it), and it forces us to confront whether our theories really work when tested against complex, real-world data. Sometimes they do, sometimes they don&#8217;t, and sometimes the data reveals relationships our theories missed entirely. That&#8217;s not atheoretical, that's how science progresses.</p><p><strong>5. &#8220;External validity concerns&#8221;</strong></p><p>This criticism also has it backwards. The argument is that models trained on one dataset don&#8217;t work elsewhere, that economic relationships vary across time and place, and that we need models that capture stable fundamentals. While the concern is valid, Professor Athey <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">notes</a> that an algorithm might find unstable relationships, like the fact that for a time &#8220;the presence of a piano in a video may thus predict cats&#8221;, the idea that this makes ML <em>worse</em> than traditional methods is wrong. In fact, traditional approaches that assume homogeneous treatment effects and stable parameters across contexts have a much bigger external validity problem; they just assume it away. ML tools do the opposite: they explicitly model heterogeneity, telling you exactly <em>why</em> and <em>where</em> effects vary. This is the core of external validity. When Athey and Wager developed causal forests they were providing a systematic way to &#8220;discover forms of heterogeneity&#8221;. <a href="https://academic.oup.com/ectj/article-pdf/21/1/C1/27684918/ectj00c1.pdf">Chernozhukov et al. (2018)</a> took this further by developing formal inference procedures that help us understand when results will generalize. And we see this in practice: <a href="https://academic.oup.com/qje/article-pdf/133/1/237/30636517/qjx032.pdf">Kleinberg et al. (2018)</a> built judge prediction models that worked across different courts and time periods precisely because they could identify which factors were stable and which were context-specific. The deeper point is that external validity is about understanding *systematic* variation. When <a href="https://www.jstor.org/stable/pdf/44250458.pdf">Davis and Heller (2017)</a> used causal forests to study job training programs, they identified which participant characteristics predicted where the program would work and where it wouldn&#8217;t. We should not see this as a threat to external validity, but as external validity done right. The irony is that we sometimes worry about ML models not generalizing while using methods that assume away the very heterogeneity that determines whether results will generalize in the first place. Similarly, <a href="https://academic.oup.com/rfs/article-pdf/33/5/2223/33209812/hhaa009.pdf">Gu et al. (2020)</a> demonstrated that ML methods in asset pricing consistently outperform traditional models across different time periods and market regimes. <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC5540263/pdf/nihms847251.pdf">Mullainathan and Obermeyer (2017)</a> found the same pattern in healthcare, where ML models proved more robust across different patient populations than traditional risk-adjustment methods that rely on stable, homogeneous relationships.</p><p><strong>6. &#8220;Statistical inference problems&#8221;</strong></p><p>Ok, I will give them that - to some extent. This criticism gets at something real but misunderstands the solution. The concerns are legitimate: traditional standard errors &#8220;break down&#8221; when you use ML for variable selection, multiple testing becomes a serious issue when algorithms explore thousands of potential relationships, and post-selection inference creates well-known statistical problems. As Professor Athey <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">puts it</a>, a central theme of the new literature is that ML algorithms &#8220;have to be modified to provide valid confidence intervals for estimated effects when the data is used to select the model&#8221;. What I would say here is that we shouldn&#8217;t avoid ML. We should use the statistical innovations that have specifically solved these problems.  The modern literature has done exactly that, often using techniques like &#8220;sample splitting&#8221; and &#8220;orthogonalization&#8221; to ensure valid inference. For example, <a href="https://www.sciencedirect.com/science/article/pii/S0304407614000918">Hansen and Kozbur (2014)</a> demonstrated how to do valid inference in high-dimensional panel models. <a href="https://arxiv.org/pdf/1501.03430">Chernozhukov, Hansen, and Spindler (2015)</a> provide a comprehensive framework for post-selection inference, while <a href="https://www.cambridge.org/core/services/aop-cambridge-core/content/view/A20409C7B5A9413C16716B389FED3402/S0266466618000245a.pdf/the-factor-lasso-and-k-step-bootstrap-approach-for-inference-in-high-dimensional-economic-applications.pdf">Hansen and Liao (2019) </a>developed bootstrap-based approaches for these complex settings. The double/debiased ML framework from Chernozhukov et al. solves the problem by separating the prediction task (where ML excels) from the inference task (where econometric rigor is maintained), giving you the best of both worlds. The multiple testing critique is particularly ironic because traditional econometrics is often far more guilty of this. We frequently check &#8220;dozens or even hundreds of alternative specifications behind the scenes&#8221; without correcting for it, a practice that invalidates reported p-values (the incentives are there<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a>). In contrast, ML-based methods are systematic and transparent. The causal forests developed by Athey and Wager (2019) comes with an asymptotic theory that provides valid inference even after the algorithm explores complex interactions. The deeper point is that these are not uniquely &#8220;ML problems&#8221;; they are statistical problems that exist whenever a researcher engages in model selection. The difference is that the ML and modern econometrics literature, including work like selective inference from <a href="https://www.tandfonline.com/doi/pdf/10.1080/01621459.2015.1108848">Tibshirani et al. (2016)</a>, has developed principled and transparent solutions where traditional practice often ignored the problem entirely.</p><p><strong>7. &#8220;It&#8217;s just a fad&#8221;</strong></p><p>I will finish with this one (also because the e-mail length has been nearly achieved), and I hope I was able to convince you somewhat of the benefits of using ML techniques in Econ.</p><p>My problem with people saying &#8220;it&#8217;s just a fad&#8221; is that this criticism reveals a fundamental misunderstanding of how scientific progress works. The argument goes that Econ has seen many methodological fashions come and go, that core economic insights don&#8217;t require fancy new tools, and that traditional econometric methods have worked fine for decades. But this perspective confuses *stability* with *stagnation*. Every major methodological advance in Econ was once dismissed as a fad. Regression analysis itself was once controversial. IVs are still viewed with suspicion. Even RCTs faced resistance before becoming the &#8220;gold standard&#8221; for causal inference (you can include me on that). The pattern is always the same: initial skepticism, gradual adoption by leading researchers, then widespread acceptance as the new normal. ML is following exactly this trajectory, but with one special difference: the scale and speed of adoption. </p><p>Let&#8217;s consider the institutional evidence: top economics departments are hiring ML-trained economists, leading journals are starting to publish ML-based research, and the most prestigious conferences feature ML sessions. The Nobel Committee awarded prizes to researchers using computational methods that were unimaginable decades ago. We can&#8217;t call this a fad; it&#8217;s a permanent expansion of the economist&#8217;s toolkit. The &#8220;core insights&#8221; argument is particularly misguided. </p><p><a href="https://pubs.aeaweb.org/doi/pdf/10.1257%2Fjep.28.2.3">Varian (2014)</a> argued that traditional econometric tools, while effective in many contexts, face serious limitations when confronted with the scale and complexity of modern datasets. How do you study personalized pricing algorithms? How do you analyse text from millions of social media posts or model outcomes using thousands of predictors? Do not label these as &#8220;fringe cases&#8221; when they are, in reality, a focus point to understanding today&#8217;s economic activity. ML offers practical solutions where conventional approaches often fall short. The idea that &#8220;old methods worked fine for decades&#8221; ignores the empirical reality that those methods frequently break down in the face of high-dimensional or unstructured data. <a href="https://www.aeaweb.org/articles?id=10.1257%2Fjep.31.2.87&amp;ref=ds-econ">Mullainathan and Spiess (2017)</a> demonstrated that many standard econometric practices perform poorly in out-of-sample tests, and that what &#8220;worked fine&#8221; may have been an illusion created by never testing predictive performance. The deeper issue is that dismissing ML as a fad reflects a conservative bias that mistakes familiarity for validity. As Professor Athey <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">argues</a>, the question isn&#8217;t whether traditional methods are good enough, but whether we can do better. The evidence increasingly suggests we can.\</p><div><hr></div><p>Some recommended readings, in no specific order:</p><p><a href="https://docs.iza.org/dp17014.pdf">A Hands-on Machine Learning Primer for Social Scientists: Math, Algorithms and Code</a>, by Nikos Askitas, 2024 <em>(if you are a social scientists that got to this point and want get your hands dirty, I would recommend reading this guide as a starter).</em></p><p><a href="https://wol.iza.org/articles/machine-learning-for-causal-inference-in-economics">Machine Learning For Causal Inference In Economics</a>, by Anthony Strittmatter, 2025.</p><p><a href="https://www.annualreviews.org/content/journals/10.1146/annurev-economics-080217-053433">Machine Learning Methods That Economists Should Know About</a>, by Susan Athey and Guido W. Imbens, 2019</p><p><a href="https://www.nber.org/papers/w31502">Financial Machine Learning</a>, by Bryan T. Kelly and Dacheng Xiu, 2023</p><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3702412">Lawrence R. Klein&#8217;s Principles in Modeling and Contributions in Nowcasting, Real-Time Forecasting, and Machine Learning</a>, by Roberto S. Mariano and Suleyman Ozmucur, 2020</p><p>Statistical Modeling: The Two Cultures (<a href="https://projecteuclid.org/journals/statistical-science/volume-16/issue-3/Statistical-Modeling--The-Two-Cultures-with-comments-and-a/10.1214/ss/1009213726.full">paper </a>/ <a href="https://ledaliang.github.io/journalclub/_static/breiman2001/Presentation.pdf">slides</a>), by Leo Breiman, 2001 (thanks <a href="https://x.com/CampedelliGian/status/1932871863065026950">Gian</a> for this tip!)<br></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><ol start="8"><li><p>&#8220;LLMs hallucinate, so you can&#8217;t use their outputs as data&#8221;: this one I&#8217;ll concede a bit because LLMs do hallucinate, but the problem is disappearing in newer models. <a href="https://deploymentsafety.openai.com/gpt-5-6">OpenAI&#8217;s GPT-5.6 system card (2026)</a>, for example, reports fewer factual errors than GPT-5.5 on conversations that users had flagged, although errors remain and the sample doesn&#8217;t represent normal usege. For empirical work, though, hallucination is part of a wider measurement problem. An LLM may invent a fact, misclassify a document, change its answer when the prompt changes or reproduce information from its training data. Each of these failures affects the analysis in a different way, so grouping them under &#8220;hallucination&#8221; doesn&#8217;t really tell us what has gone wrong or how the error will affect the estimate. This is where <a href="https://www.nber.org/papers/w33344">Ludwig, Mullainathan and Rambachan (2025</a>) is fitting because they distinguish between using an LLM for prediction and using it to construct a measure. When the task is prediction, the researcher needs to show that the evaluation sample didn&#8217;t appear in the model&#8217;s training data; otherwise, a conventional test set no longer represents unseen cases. When the model is used to classify text or create a variable for a regression, the researcher needs an independent validation sample in which the same concept has been measured through another method. From there, the analysis can establish how large the errors are, whether they differ across groups and whether the measure changes across prompts or model versions. Without this evidence, asking a chatbot for classifications and treating the answers as observed data is difficult to defend. The same point applies when LLM predictions are used inside a causal procedure. <a href="https://academic.oup.com/ectj/article/21/1/C1/5056401">Chernozhukov et al. (2018)</a> show how cross-fitting and orthogonal scores can limit the effect of prediction errors in DML, while <a href="https://arxiv.org/abs/2510.09684">Engh and Aronow (2025)</a> use LLM guesses to estimate conditional expectations and <a href="https://arxiv.org/abs/2509.25709">Gui and Kim (2025)</a> use LLMs to form strata before treatment assignment. These methods may improve precision, but the identifying assumptions continue to come from the research design. An LLM can&#8217;t create a valid comparison group, remove training leakage or establish that a generated variable measures the concept the paper claims to study, but that&#8217;s not on them. If we can measure these failures and account for them in the analysis, its outputs <em>can</em> still be used as data.</p></li></ol></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>One can even argue about a third driver: institutional incentives (e.g., tenure metrics tied to causal identification papers). Some of the mechanisms in play we can think of are: top journals are more likely to accept papers with clean causal designs, tenure committees count those &#8220;causal hits&#8221; and discount purely predictive studies, PhD advisors channel students toward DiD rather than neural networks, grant reviewers ask &#8220;where&#8217;s the treatment?&#8221; and mark down ML-only proposals, core graduate courses leave ML out so the entry cost stays high, and authors often graft an instrument onto their models just to satisfy referees. My friend Prashant and Professor Fetzer have an excellent <a href="https://arxiv.org/abs/2501.06873">paper</a> on the &#8220;rise in the share of causal claims (papers) - from roughly 4% in 1990 to nearly 28% in 2020 - reflecting the growing influence of the &#8220;credibility revolution&#8221;.&#8221; They find that &#8220;causal narrative complexity (e.g., the depth of causal chains) strongly predicts both publication in top-5 journals and higher citation counts, whereas non-causal complexity tends to be uncorrelated or negatively associated with these outcomes&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>AlphaFold utilizes ML (particularly deep learning) to predict the 3D structure of proteins from their amino acid sequences. It leverages large datasets of known protein structures and sequences to train a neural network that can identify patterns and relationships in these structures.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;A prediction I have is that there will be an active and important literature combining ML and causal inference to create new methods, methods that harness the strengths of ML algorithms to solve causal inference problems. (&#8230;) This new literature takes many of the strengths and innovations of ML methods, but applies them to causal inference. Doing this requires changing the objective function, since the ground truth of the causal parameter is not observed in any test set&#8221;. <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">Athey, 2018</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>I&#8217;d like to thank <a href="https://sites.google.com/view/maricavalente/about">Professor Marica Valente</a> for this insight. I attended her course in Machine Learning last summer at ISEG and had a great time, she&#8217;s awesome.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>&#8220;Researchers often check dozens or even hundreds of alternative specifications behind the scenes, but rarely report this practice because it would invalidate the confidence intervals reported (due to concerns about multiple testing and searching for specifications with the desired results). There are many disadvantages to the traditional approach, including but not limited to the fact that researchers would find it difficult to be systematic or  comprehensive in checking alternative specifications, and further because researchers were not honest about the practice, given that they did not have a way to correct for the specification search process. I believe that regularization and systematic model selection have many advantages over traditional approaches, and for this reason will become a standard part of empirical practice in economics&#8221;. <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">Athey, 2018</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>&#8220;The ML literature uses a variety of techniques to balance expressiveness against over-fitting. The most common approach is cross-validation whereby the analyst repeatedly estimates a model on part of the data (a &#8220;training fold&#8221;) and then evaluates it on the complement (the &#8220;test fold&#8221;). The complexity of the model is selected to minimize the average of the mean-squared error of the prediction (the squared difference between the model prediction and the actual outcome) on the test folds. Other approaches used to control over-fitting include averaging many different models, sometimes estimating each model on a subsample of the data (one can interpret the random forest in this way)&#8221;. <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">Athey, 2018</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>&#8220;There are discussions of what interpretability means, and whether simpler models have advantages. Of course, economists have long understood that simple models can also be misleading. In social sciences data, it is typical that many attributes of individuals or locations are positively correlated&#8211;parents&#8217; education, parents&#8217; income, child&#8217;s education, and so on. (&#8230;) So, simpler models can sometimes be misleading; they may seem easy to understand, but the understanding gained from them may be incomplete or wrong&#8221;. <a href="https://www.gsb.stanford.edu/sites/default/files/publication-pdf/atheyimpactmlecon.pdf">Athey, 2018</a>.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>Chernozhukov et al. have a new <a href="https://arxiv.org/pdf/2504.08324?">pre-print</a> where they provide a practical introduction to Double/Debiased Machine Learning (DML).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>&#8220;We examine how the evaluation of research studies in economics depends on whether a study yielded a null result. Studies with null results are perceived to be less publishable, of lower quality, less important and less precisely estimated than studies with large and statistically significant results, even when holding constant all other study features, including the sample size and the precision of the estimates. The null result penalty is of similar magnitude among PhD students and journal editors. The penalty is larger when experts predict a large effect and when statistical uncertainty is communicated with <em>p</em>-values rather than standard errors. Our findings highlight the value of a pre-result review&#8221;. <a href="https://academic.oup.com/ej/article/134/657/193/7238466">Chopra et al., 2024</a>.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Machine Learning for Continuous Treatments, Bayesian Methods for Staggered Timing, and Using Experiments to Fix Observational Data ]]></title><description><![CDATA[How new methods handle treatment intensity, small samples with staggered adoption, and selection bias in observational studies]]></description><link>https://www.diddigest.xyz/p/machine-learning-for-continuous-treatments</link><guid isPermaLink="false">https://www.diddigest.xyz/p/machine-learning-for-continuous-treatments</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Fri, 30 May 2025 16:20:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gU7e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gU7e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gU7e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 424w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 848w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 1272w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gU7e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png" width="1198" height="621" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:621,&quot;width&quot;:1198,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1038923,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/164226519?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gU7e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 424w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 848w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 1272w, https://substackcdn.com/image/fetch/$s_!gU7e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab53304-8eed-4c4a-860a-3560ccf6061a_1198x621.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! Today we have two new DiD papers, one DiD guide, and one paper by 3 heavy-weights in the field on how to leverage experiments to correct for observational bias:</p><ol><li><p><a href="https://arxiv.org/pdf/2408.10509">Continuous Difference-In-Differences With Double/Debiased Machine Learning</a>, by Lucas Zhang </p></li><li><p><a href="https://www.arxiv.org/abs/2505.18391">Model-based Estimation of Difference-in-Differences with Staggered Treatments</a>, by Siddhartha Chib and Kenichi Shimizu</p></li><li><p><a href="https://www.nber.org/papers/w33817">The Experimental Selection Correction Estimator: Using Experiments to Remove Biases in Observational Estimates</a>, by Susan Athey, Raj Chetty and Guido Imbens (this paper is *not* about DiD but it&#8217;s about an estimator :) also it&#8217;s a great reading so I decided to add it to the post)</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5187998">A Concise Guide to Difference-in-Differences Methods for Economists with Applications to Taxation, Regulation, and Environment</a>, by Bruno Bosco and Paolo Maranzano (I added this one here because it&#8217;s so good! If you&#8217;re just starting to get familiar with DiD, this guide is a great place to start)</p><p></p></li></ol><h3>Continuous Difference-In-Differences With Double/Debiased Machine Learning</h3><p><em>(<a href="https://lucaszz-econ.github.io/">Lucas</a>&#8217; <a href="https://lucaszz-econ.github.io/notes/JMP.pdf">JMP</a> is also on DiD. He&#8217;s a Ph.D. candidate at UCLA)</em> </p><h5>TL;DR: in this paper, Lucas extends DiD in settings with continuous treatments (where treatment intensity varies rather than being binary). To identify and estimate the average treatment effect on the treated (ATT) at each intensity level, he proposes a new estimator built on double/debiased machine learning (DML). The method he proposes uses orthogonal scores, kernel smoothing, and cross-fitting to reduce bias (particularly in high-dimensional settings). The result is a flexible and theoretically grounded approach to estimating how treatment effects vary with dose, complete with valid confidence bands and a real-world policy application.</h5><p><em>What is this paper about?</em></p><p>There is growing interest in extending DiD to continuous treatments, where one can further investigate how outcomes vary across different treatment intensities within the treated group. Think of hospitals exposed to different levels of policy reform, or individuals experiencing varying degrees of exposure to an intervention. This paper tackles this growing interest. The core object of interest is the average treatment effect on the treated (ATT) at any given intensity level, d, defined as </p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;\\text{ATT}(d) := \\mathbb{E}[Y_t(d) - Y_t(0) \\mid D = d]\n&quot;,&quot;id&quot;:&quot;BTUMTRHJKI&quot;}" data-component-name="LatexBlockToDOM"></div><p>Where ATT at level d is the expected difference between the treated and untreated potential outcomes at time t, for units that actually received treatment level d. To identify this effect, the paper adopts a conditional parallel trends assumption (analogous to the binary case in the &#8220;sister&#8221; papers <a href="https://academic.oup.com/restud/article-abstract/72/1/1/1581053">Abadie, 2005</a> and <a href="https://www.sciencedirect.com/science/article/pii/S0304407620301901?casa_token=DJ6v85VxpFcAAAAA:r2ZQFSS_SEVEoaoVYjmaAVUyr3xIFjShilAtrbpG-3U1qbuhq7mlyuZrN8xiOYABnlKSL455">Sant&#8217;Anna and Zhao, 2020</a>) where untreated potential outcomes evolve similarly across treatment levels, *conditional on covariates*. This allows the ATT to be expressed in terms of observable quantities even when treatment is continuous. But as Lucas notes, we have some issues: &#8220;estimating the ATT in this framework requires first estimating infinite-dimensional nuisance parameters,&#8221; making the task more complex than in binary DiD (think of nuisance parameters are parts of the model that aren&#8217;t the main focus of your estimation - like the ATT - but that must be estimated to recover it, e.g. conditional means or treatment probabilities). He shows how to handle this challenge in a rigorous way.</p><p><em>What does the author do?</em></p><p>To estimate ATT(d) in practice, Lucas adopts the double/debiased machine learning (DML) framework developed by <a href="https://academic.oup.com/ectj/article-abstract/21/1/C1/5056401">Chernozhukov et al. (2018)</a>, which uses orthogonalisation (where we rewrite the estimator so it&#8217;s robust to errors in nuisance parameter estimates) and cross-fitting (where we split the sample so that nuisance parameters are estimated on one part and used on another) to reduce bias in causal inference. This allows him to flexibly estimate complex nuisance functions (e.g., conditional densities and probabilities) without introducing first-order bias, even in high-dimensional settings. The estimator is then constructed in two steps:</p><ol><li><p>Orthogonal (or "locally insensitive") scores: ATT(d) is re-expressed using a score function that remains valid even if the nuisance parameters are estimated with error (a property known as <a href="https://maxkasy.github.io/home/files/teaching/ML_Oxford_2022/debiased_ml_slides.pdf">Neyman orthogonality</a> - like taking a derivative that&#8217;s zero at the true value, so small estimation mistakes have no first-order impact).</p></li><li><p>Cross-fitting: the sample is split so that nuisance functions are estimated on one part and used on another, which helps mitigate overfitting and improves finite-sample performance (think of this like using one set of data to &#8220;learn&#8221; the adjustment, and a different set to &#8220;apply&#8221; it).</p></li></ol><p>A central challenge here is that ATT(d) depends on the conditional density of the treatment, which is difficult to estimate directly. To address this, Lucas approximates ATT(d) with a smoothed version using kernel weights (think of it like a weighted average centred around d, where the weights smoothly drop off as you move away  &#8594; this helps us estimate quantities like ATT(d) even when we don&#8217;t have many observations with exactly that treatment level), and shows that this approximation converges to the true ATT as the bandwidth shrinks. He then derives the estimator&#8217;s asymptotic properties, showing it is asymptotically normal, and provides valid confidence bands using a multiplier bootstrap.</p><p><em>Why is this important?</em></p><p>Most real-world policies do not treat units in a binary way, and interventions often vary in intensity or exposure, whether it's funding amounts, pollution levels, or programme uptake. Yet more traditional DiD methods are built around binary treatment comparisons, limiting what we can learn. This paper offers a clean way to generalise DiD to continuous treatments, while keeping identification credible and inference valid even in high-dimensional settings where machine learning tools are used. It bridges the gap between theory and application by adapting recent advances (like DML) to a DiD context, which is still relatively underdeveloped for continuous treatments. This isn&#8217;t a theoretical exercise: the paper applies the method to the 1983 Medicare PPS reform, showing how the treatment effect of the reform varied with hospitals' share of Medicare patients, something standard DiD would miss entirely.</p><p><em>Who should care?</em></p><p>Anyone working with policy variation in intensity, not just presence or absence, and that includes:</p><ul><li><p>Applied economists studying heterogeneous policy effects (e.g. health, education, environment)</p></li><li><p>Labour and public economists evaluating programme rollouts where take-up or exposure varies</p></li><li><p>Causal inference researchers interested in extending DiD tools to more realistic settings</p></li><li><p>Data scientists and ML users who want valid inference when using flexible, nonparametric models</p></li></ul><p>If you&#8217;ve ever asked &#8220;does the effect depend on how much treatment someone got?&#8221;, then Lucas&#8217; paper gives you a way to answer it.</p><p><em>Do we have code?</em></p><p>No replication code yet, but the methods are compatible with existing double machine learning toolkits in both R and Python. For example, the nuisance functions in the paper are estimated using random forests in scikit-learn, and the estimator uses standard components like kernel regression, cross-fitting, orthogonal scores, and multiplier bootstrap for inference. So if you&#8217;re already using DML packages like <code>econml</code>, <code>DoubleML</code>, or <code>grf</code>, this setup should feel familiar, just adapted to a continuous treatment setting.</p><p>In summary, this paper pushes DiD into more realistic territory where treatment isn&#8217;t all-or-nothing, but continuous. By combining conditional parallel trends with double machine learning, it gives us a principled way to estimate how effects vary with treatment intensity. If you're studying policies with uneven exposure or variable implementation, this method lets you recover meaningful causal effects without relying on oversimplified binary comparisons.</p><h3>Model-based Estimation of Difference-in-Differences with Staggered Treatments</h3><h5>TL;DR: in this paper the authors introduces the first fully model-based approach to estimating DiD effects in staggered treatment settings using Bayesian methods. Instead of relying on TWFE or modern nonparametric estimators, professors Chib and Shimizu specify a hierarchical state-space model where treatment effects can evolve over time and vary across treated units. The model captures latent trends, unit-specific dynamics, and treatment heterogeneity all at once. Bayesian estimation via MCMC then yields full posterior distributions for group-time average treatment effects, along with credible intervals and dynamic treatment profiles (even in small samples, where frequentist methods may be unreliable). The method is then applied to a job training program dataset and to crime outcomes following a state-level policy change. The result is a coherent and flexible Bayesian alternative to recent DiD estimators, which is especially valuable when treatment effects are expected to be dynamic or heterogeneous, or when we want to do probabilistic modelling. </h5><p><em>What is this paper about?</em></p><p>When treatments are rolled out at different times across units (what we call a staggered adoption design) standard DiD methods like TWFE can produce misleading results. They assume constant treatment effects and parallel trends, and break down when effects vary across time or groups. Alternatives like <a href="https://www.sciencedirect.com/science/article/abs/pii/S0304407620303948">Callaway and Sant&#8217;Anna, 2021</a> and <a href="https://www.sciencedirect.com/science/article/abs/pii/S030440762030378X">Sun and Abraham, 2021</a> correct for some of these issues, but they still rely on large-sample, nonparametric estimation and don&#8217;t always perform well when sample sizes are small or data are noisy. This paper takes a different route. Profs Chib and Shimizu propose a Bayesian model-based approach that directly models how treatment effects evolve over time and vary across units. They set up a state-space model where the average treatment effect for each group and time period ATT(g, t) is treated as a latent variable, to be inferred from the data alongside other unknowns. This allows the model to smoothly track treatment effects over time, account for latent trends and unit-specific dynamics, while also handling staggered treatment in a natural, joint estimation setup. Rather than adjusting after the fact, the model builds in treatment timing and variation from the start. Bayesian inference (via MCMC) then delivers posterior distributions for the effects, offering not just point estimates but full uncertainty quantification. And because it borrows information across units and time, it can be especially helpful in small-sample settings, where many modern DiD estimators struggle.</p><p><em>What do the authors do?</em></p><p>Profs Chib and Shimizu build a hierarchical Bayesian model where treatment effects ATT(g,&#8239;t) are treated as latent variables within a state-space structure<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>. This means they explicitly model how effects evolve over time and differ across treated groups, instead of estimating them indirectly like most DiD methods do. The key idea here is that observed outcomes are assumed to follow a data-generating process with latent (unobserved) components (unit fixed effects, time effects latent ATT(g,&#8239;t) values). These latent effects are assumed to evolve smoothly over time via a stochastic process (e.g. random walk or autoregressive process). The model is estimated using Bayesian MCMC methods, which generate posterior distributions for each ATT(g,&#8239;t), allowing for dynamic patterns, full uncertainty quantification, and inference even in small samples. They then apply the method to a subset of the Callaway and Sant&#8217;Anna (2021) <a href="https://bcallaway11.github.io/did/">minimum wage dataset</a>, looking at log teen employment across 500 counties between 2003&#8211;2007 (the treatment is defined by the year a state first increased its minimum wage), and to a simulation study designed to assess how well the model recovers ATT(g,&#8239;t) under various data-generating processes. In both examples they demonstrate how their method produces smoother, more stable, and better-identified effect paths than TWFE or recent nonparametric DiD alternatives.</p><p><em>Why is this important?</em></p><p>Most methods for staggered DiD focus on correcting biases in TWFE, but they rely on large samples and assume you can nonparametrically estimate everything you need. That&#8217;s often not true, especially in small datasets or when effects change over time. This paper shows that you don&#8217;t have to give up on credible DiD just because your data are sparse or noisy. By directly modelling how treatment effects evolve, the method can pick up patterns that other estimators miss and provide honest uncertainty around those estimates. It also avoids the common headaches with event-study plots, like wild fluctuations or imprecise comparisons across groups and periods. Instead, it delivers smooth, interpretable trajectories of treatment effects along with full posterior intervals. The big takeaway: if you're in a setting where standard DiD feels &#8220;fragile&#8221;, a Bayesian approach can give you clearer answers with fewer assumptions about what the data can do.</p><p><em>Who should care?</em></p><p>This paper is for you if:</p><ul><li><p>You work with panel data where treatment happens at different times, and effects might change over time</p></li><li><p>Your sample size is small or your data are noisy, and standard DiD methods feel unstable</p></li><li><p>You&#8217;re interested in Bayesian methods but want something grounded in causal questions</p></li><li><p>You&#8217;ve used tools like Callaway and Sant&#8217;Anna or Sun and Abraham, but want smoother, more precise estimates (especially for group-time effects)</p></li></ul><p>It&#8217;s especially useful if your event-study plots are bumpy, hard to interpret, or if you&#8217;re trying to understand treatment dynamics but the data are too thin for conventional methods.</p><p><em>Do we have code?</em></p><p>Prof Shimizu <a href="https://www.linkedin.com/posts/kenichi-shimizu_model-based-estimation-of-difference-in-differences-activity-7333565041767497729-kWok?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAACTHmhoBpQRxz0hUy8fhU1GuRQuRyKVjdd0">said</a> that a user-friendly package 'bdid' will be available soon in R, Matlab, and Stata.</p><p>In summary, most DiD methods try to fix issues with staggered timing by adjusting the estimator. This paper takes a different route: model the treatment effects directly using Bayesian methods. Profs Chib and Shimizu show how you can estimate group-time treatment effects as latent variables, letting them evolve smoothly over time while accounting for uncertainty. The result is especially useful when you&#8217;re working with small samples, messy data, or want more than just point estimates. If your usual tools are giving you noisy ATT plots or unstable dynamics, this is a a good alternative, especially if you&#8217;re open to bringing more structure (and full posteriors) into your analysis.</p><h3>The Experimental Selection Correction Estimator: Using Experiments to Remove Biases in Observational Estimates</h3><p><em>(I don&#8217;t think the authors need an introduction)</em></p><h5>TL;DR: as more researchers rely on large observational datasets, a recurring issue emerges: these datasets capture rich outcomes and wide populations but lack randomisation (it turns out that even with larger Ns you still *do* need randomisation). At the same time, smaller experiments offer clean causal identification but only for a narrow set of outcomes or people. In this paper the authors introduce the Experimental Selection Correction (ESC) estimator, a way to use the experimental data to estimate and correct the bias in observational estimates. Think of it as training a correction function on the experiment, then applying it to the observational data. It&#8217;s a modern rethinking of Heckman-style selection correction, but with fewer assumptions and better performance when combined with ML tools.</h5><p><em>What is this paper about?</em></p><p>There is a common, real-world, problem many of us applied researchers face: there is have an experiment, but it&#8217;s small or limited in scope; there&#8217;s also a big observational dataset, but it&#8217;s potentially biased. How do we combine them to get accurate causal estimates? Profs Athey, Chetty, and Imbens propose the Experimental Selection Correction (ESC) estimator, which uses the experimental data to learn the selection bias in the observational estimates, and then corrects for that bias across the broader sample. Their set-up is: in the observational data, one observes treatment and a wide set of outcomes, but treatment is not random, whereas in the experimental data, treatment is randomly assigned, but outcomes are more limited or less representative. Their key idea is: if one can measure the difference between the experimental and observational estimates for units observed in both, one can estimate a bias function. Then one can use that function to correct the observational estimates in the larger dataset. Mathematically, it&#8217;s similar in spirit to <a href="https://rlhick.people.wm.edu/stories/econ_407_notes_heckman.html">Heckman&#8217;s selection model</a>, but the correction term is learned from experimental data rather than assumed from a structural model. </p><p><em>What do the authors do?</em></p><p>Profs Athey, Chetty, and Imbens introduce the Experimental Selection Correction (ESC) estimator, which adjusts observational treatment effect estimates using experimental data. We even have a DAG to exemplify the idea:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3PLd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3PLd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 424w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 848w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 1272w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3PLd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png" width="1093" height="1042" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1042,&quot;width&quot;:1093,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:132246,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/164226519?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3PLd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 424w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 848w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 1272w, https://substackcdn.com/image/fetch/$s_!3PLd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f779246-93c1-4a25-b4f9-a5298f0611af_1093x1042.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Panels A and B represent traditional surrogate models using experimental and observational data. Panels C and D show how ESC brings in the direct effect of treatment and how unobserved confounding affects the observational estimates. ESC uses the clean experimental setting (Panel C) to estimate the bias present in Panel D and remove it.</figcaption></figure></div><p>As one can see from the illustration, they build on the surrogate model literature, where researchers estimate the effect of a treatment on a surrogate outcome (like test scores), and then use the observed relationship between the surrogate and a long-term outcome (like graduation rates) to back out the treatment effect. This approach has been influential (<a href="https://www.nber.org/system/files/working_papers/w26463/w26463.pdf">Athey et al., 2019</a>), but it assumes no unobserved confounding, which is an assumption often violated in observational settings. The ESC estimator extends the surrogate approach by explicitly modelling and correcting the bias in the observational estimates of long-term outcomes. It works in a few steps: you use the experimental sample (where treatment is randomly assigned) to estimate the bias in observational treatment effect estimates (i.e. the difference between experimental and observational effects, conditional on covariates); then you estimate a bias function (possibly using ML), that maps observed covariates to the size of the bias; and finaly you apply this bias correction to the observational sample to get a debiased treatment effect estimate. Then we can use a the large, rich observational dataset (but with bias removed) effectively combining the precision and scale of big data with the credibility of randomised experiments.</p><p><em>Why is this important?</em><br>There's this <a href="https://x.com/rabois/status/1242642669840429056">tweet</a> that went viral once which reflects a common misconception that -especially in the era of Big Data - if your sample is large enough, you don&#8217;t need randomisation. But this paper flips that idea on its head. The authors show that selection bias doesn&#8217;t vanish with more data, it just becomes more precisely wrong. You can collect all the observational data in the world, but if treatment assignment is biased, the resulting estimate will be too. That&#8217;s why this paper matters. It offers a way to quantify and remove that bias using a well-designed experiment, even if the experiment is small or doesn&#8217;t cover every outcome. The idea is to enhance observational data for causal inference by using experimental insights to correct its flaws, thereby avoiding the need for complete replacement. It&#8217;s a flexible, scalable solution that keeps the credibility of experiments while unlocking the scope of big data.</p><p><em>Who should care?</em></p><p>This is for researchers who:</p><ul><li><p>Work with large observational datasets (e.g. admin data, education records, firm-level panels)</p></li><li><p>Have access to some experimental variation, but not at the scale or granularity they need</p></li><li><p>Worry about selection bias but don&#8217;t want to rely on strong structural assumptions</p></li><li><p>Use ML for causal inference, but want a method that won&#8217;t be too difficult to justify when someone asks in a seminar</p></li></ul><p><em>Do we have code?</em><br>We have an entire <a href="https://github.com/OpportunityInsights/Experimental-Selection-Correction-Replication-Code">repo</a> (in R) dedicated to it.</p><p>In summary, this paper delivers a powerful message: randomisation is still useful, even in the age of Big Data. Profs Athey, Chetty, and Imbens show how you can use experimental evidence to identify effects, as well as to repair the observational ones. Instead of throwing out observational data or blindly trusting it, their estimator lets you calibrate it. You borrow just enough credibility from the experiment to subtract out the bias, and then scale up your estimates with confidence. It&#8217;s a modern take on selection correction, grounded in theory but ready for real-world data.</p><p></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Latent variables within a state-space structure means that the treatment effects aren&#8217;t directly observed or estimated from a single regression equation. Instead, they&#8217;re treated as unobserved time series that evolve over time according to a probabilistic process. The model uses the observed data to infer these hidden values step by step, borrowing information across time and units.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[DDD Estimators, Distributional Effects, Causal Diagrams, and Cross-Field Counterfactuals]]></title><description><![CDATA[New methods and perspectives in causal inference research]]></description><link>https://www.diddigest.xyz/p/ddd-estimators-distributional-effects</link><guid isPermaLink="false">https://www.diddigest.xyz/p/ddd-estimators-distributional-effects</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Fri, 23 May 2025 12:37:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7GB5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7GB5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7GB5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7GB5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1244938,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/164225524?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7GB5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7GB5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85ae8ec-2b5d-40d8-917e-299875c8edb9_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! Hope you&#8217;re all doing well :) </p><p>Today we have all sorts of studies, and one &#8220;treat&#8221; at the end :)</p><ol><li><p><a href="https://arxiv.org/pdf/2505.09942">Better Understanding Triple Differences Estimators</a>, by Marcelo Ortiz-Villavicencio and Pedro H. C. Sant&#8217;Anna</p></li><li><p><a href="https://arxiv.org/pdf/2408.01208">Distributional Difference-in-Differences with Multiple Time Periods</a>, by Andrea Ciaccio</p></li><li><p><a href="https://arxiv.org/pdf/2505.03526">Using Causal Diagrams To Assess Parallel Trends In Difference-In-Differences Studies</a>, by Audrey Renson, Oliver Dukes and Zach Shahn</p></li><li><p><a href="https://arxiv.org/pdf/2505.13324">From What Ifs to Insights: Counterfactuals in Causal Inference vs. Explainable AI</a>, by Galit Shmueli, David Martens, Jaewon Yoo, and Travis Greene</p></li></ol><p><em>(I apologize if there are any mistakes here - let me know in the comments! I confess I haven&#8217;t properly slept in the past 5 days, and I&#8217;ve been writing between airports lol terrible idea)</em></p><h3>Better Understanding Triple Differences Estimators</h3><p><em>(I am so happy to cover this paper! <a href="https://github.com/marcelortizv">Marcelo</a> is my friend and prof Pedro&#8217;s - who needs no introduction - PhD student. <a href="https://x.com/pedrohcgs/status/1923447096444694545">Here</a> you can check a thread about it in prof Pedro&#8217;s words :))</em></p><h5>TL;DR: DDD estimators are commonly used to recover treatment effects when standard DiD PT assumptions are too strict. DDD allows for certain violations (e.g., group-specific or eligibility-specific trends) while still producing &#8220;valid&#8221; estimates if the identifying assumptions are correctly specified. But here&#8217;s the problem: in many realistic settings, those assumptions only hold after conditioning on covariates. Failing to account for this - by erroneously proceeding as if DDD was just the difference between 2 DiD estimators - leads to bias. Prof Pedro and Marcelo propose new regression-adjusted, inverse probability weighted, and doubly robust estimators that remain valid in these more realistic scenarios. They also show that in staggered adoption settings, pooling not-yet-treated units (as is practice in DiD) invalidates DDD. The result is a practical, modern framework for credible DDD estimation, particularly in applications with covariates, staggered timing, or event studies.</h5><p><em>What is this paper about?</em></p><p>DDD (Difference-in-Difference-in-Differences) is a common strategy when treatment depends on two conditions, like living in a treated state and being eligible for a policy. It&#8217;s attractive because it lets you relax the standard DiD assumption that treated and control groups follow the same trends in the absence of treatment. Instead, DDD allows for different trends across groups, as long as the difference in trends between treatment-eligible and ineligible units is the same in treated and control areas (and vice versa). But the problem is that, in practice, most applications assume *too much*. When covariates matter, or when treatment is staggered, the typical DiD logic fails. This paper shows why and what to do instead.</p><p><em>What do the authors do?</em></p><p>They show that the usual approaches (subtracting two DiDs or running a three-way fixed effects regression) don&#8217;t work when covariates are important. Worse, if treatment is staggered, pooling all not-yet-treated units (as in modern DiD) can introduce bias due to changing group composition. To fix this, they introduce three estimators that stay valid under the right assumptions: regression-adjusted, inverse probability weighted (IPW), and doubly robust (DR). The DR version is especially flexible because it can be used with machine learning and allows combining multiple comparison groups using GMM to boost precision.</p><p><em>Why is this important?</em></p><p>DDD is meant to be more &#8220;flexible&#8221; than a 2x2 DiD, but if you implement it like a 2x2 DiD (by ignoring covariates or picking up controls like crazy), it won&#8217;t give you the correct estimates. Marcelo and Prof Pedro show why these choices lead to bias, and offer both a diagnosis and a solution: tools that let you actually relax the PT assumption without breaking identification. They also make a key point that&#8217;s often overlooked: DDD designs with staggered adoption cannot be treated as simple extensions of DiD methods. They emphasize that the interaction between timing, eligibility, and changing group composition introduces complexities that demand more careful attention to identification and control selection.</p><p><em>Who should care?</em></p><p>Anyone using DDD in applied work, especially if:</p><ul><li><p>Your treatment depends on both where and who (e.g., geography and eligibility),</p></li><li><p>Covariates are essential for identification,</p></li><li><p>Your design involves staggered treatment,</p></li><li><p>You&#8217;re running event studies or subgroup analyses.</p></li></ul><p>Empirical economists, policy evaluators, health and education researchers, and methodologists in general will find this useful. And honestly, so will we grad students running triple differences.</p><p><em>Do we have code?</em></p><p>Marcelo and prof Pedro say that they are &#8220;finishing an R package that will automate all these DDD estimators, hopefully making it easier to adopt&#8221;. So we wait :)</p><p>In summary,  DDD isn&#8217;t just &#8220;two DiDs&#8221;, and treating it that way can seriously bias your results. This paper shows exactly why that happens, and more importantly, how to fix it. Marcelo and Prof Pedro give us toolkit for doing DDD right: estimators that are flexible, robust, and grounded in solid identification logic. If your setting involves staggered timing, covariates, or layered eligibility rules, this is the framework you should check out. And when the R package drops, I&#8217;ll let you know but keep an eye out for it.</p><h3>Distributional Difference-in-Differences with Multiple Time Periods</h3><p><em>(<a href="https://sites.google.com/view/andreaciaccio">Andrea</a> is a PhD candidate at Ca' Foscari University of Venice)</em></p><h5>TL;DR: most DiD estimators focus on average treatment effects. But what if you care about how a policy affects different parts of the outcome distribution, like the bottom 25% vs the top 25%? This paper extends the Quantile Treatment Effect on the Treated (QTT) framework to multiple time periods and staggered adoption settings. Andrea generalises and builds on <a href="https://onlinelibrary.wiley.com/doi/full/10.3982/QE935">Callaway and Li (2019)</a> to recover the full counterfactual distribution (not just means or quantiles!) without needing rank invariance. The result? A flexible, nonparametric method to estimate distributional effects using either panel data or repeated cross-sections.</h5><p><em>What is this paper about?</em></p><p>This paper is about moving beyond averages. In many applied settings, we care <strong>both</strong> about whether a policy had an effect <strong>and</strong> where in the outcome distribution those effects occurred. Imagine a minimum wage hike. It might raise *average wages*, but does it help low-income workers more, or mostly benefit those in the middle of the distribution? Average treatment effects won't tell you. Andrea proposes a method to estimate distributional treatment effects in settings with multiple periods, staggered adoption, and non-experimental data. The goal is to recover the full counterfactual distribution of untreated outcomes for the treated group. To do this, he builds on the QTT framework and generalises it to work in more realistic settings. The nicest part is that his approach does not rely on rank invariance, which is often violated in practice. Instead, he also introduces tools like stochastic dominance tests to compare distributions more credibly.</p><p><em>What does the author do?</em></p><p>He starts by generalising the QTT estimator from Callaway and Li (2019) to a more realistic setting with multiple time periods, staggered treatment, either panel or repeated cross-section data. Andrea uses a combination of distributional parallel trends and a copula invariance assumption to identify the full counterfactual distribution of untreated outcomes for the treated group. He then proposes new estimands that are robust to violations of rank invariance, such as comparisons using stochastic dominance rankings and inequality measures like the Gini index. He also provides an actual estimator. The only other paper addressing QTT identification in staggered DiD settings -<a href="https://www.sciencedirect.com/science/article/pii/S0165176524002763"> Li and Lin (2024)</a>, which I mentioned in another post - stops at theory and does not propose an estimator, which limits its applicability. Andrea&#8217;s contribution fills that gap with a method researchers can use.</p><p><em>Why is this important?</em></p><p>This paper is important because heterogeneous effects matter. Policies rarely impact everyone the same way, and when we focus only on averages, we risk missing who actually benefits (or loses). Andrea&#8217;s method lets us go beyond the ATT and estimate how the entire distribution shifts, which is especially relevant for policies that aim to reduce inequality or target disadvantaged groups. It also matters methodologically: while earlier papers offered theoretical identification of distributional effects under staggered DiD, Andrea is the first to provide a fully worked-out estimator. The fact that his approach works with repeated cross-sections and avoids rank invariance makes it much more usable in real-world applications.</p><p><em>Who should care?</em></p><ul><li><p>Researchers working with staggered DiD who want to go beyond average treatment effects</p></li><li><p>Applied microeconomists studying income, education, labour markets, or health, where heterogeneous impacts are expected</p></li><li><p>Anyone evaluating policies with distributional goals such as minimum wages, subsidies, school tracking, etc </p></li><li><p>People using repeated cross-sections instead of panel data</p></li><li><p>And of course: grad students who keep reading &#8220;QTT on staggered DiD&#8221; papers and wonder, okay, but how do I actually estimate that?</p></li></ul><p><em>Do we have code?</em></p><p>Andrea says that &#8220;all the MC simulations were run in STATA. The ado files of the command used for implementing the methodology presented in the paper, qtt, are available upon request at the time of writing&#8221;. </p><p>In summary, Andrea&#8217;s paper gives applied researchers a way to estimate how the entire outcome distribution would have evolved in the absence of treatment, even in complex staggered DiD settings. It works with repeated cross-sections, allows for conditioning on covariates, and avoids the strong rank invariance assumption that most distributional methods rely on. Once you have the full counterfactual distribution, you&#8217;re not limited to quantiles, you can compute Gini coefficients, Lorenz curves, or test for stochastic dominance. This makes the method flexible and genuinely useful for policy evaluation, especially when heterogeneity is central.</p><h3>Using Causal Diagrams to Assess Parallel Trends in DiD</h3><p><em>(This paper has brought flashbacks from Econometrics 2 hehe hi prof Ben if you&#8217;re reading! Also, it&#8217;s quite &#8220;visual&#8221;, which is nice)</em></p><h5>TL;DR: DiD relies on the PT assumption, but we usually just hope it holds, or eyeball pre-trends. This paper proposes a more &#8220;principled&#8221; approach: the use of causal diagrams to assess whether parallel trends are plausible before running the analysis. The authors derive three graphical conditions under which parallel trends likely fails, and show how to apply them using partially directed SWIGs (a tool used in causal inference to represent potential within a graphical model) and a linear faithfulness assumption.</h5><p><em>What is this paper about?</em></p><p>Most of us assess parallel trends by plotting outcomes before treatment and hoping the lines run in parallel. But this only works if we have multiple pre-treatment periods, and even then it might not be sufficient. This paper proposes an alternative: use causal diagrams to assess whether PT can plausibly hold, based on what we know about the data-generating model (DGM). The authors focus on the standard 2&#215;2 DiD setup (two groups, two periods) and ask a simple but deep question: under what structural conditions can we actually expect PT to hold? To answer it, they use nonparametric structural equation models and tools from causal graph theory. One key idea is linear faithfulness, which is a principle that says if two variables are connected in the DAG (i.e., d-connected), then they should also be statistically correlated in the data. This lets the authors derive graphical conditions under which the PT assumption is likely violated, just by inspecting the structure of the DAG.</p><p><em>What do the authors do?</em></p><p>The authors show how to use causal diagrams (more specifically something called a Single World Intervention Graph, or SWIG) to evaluate whether the PT assumption is even plausible, based on what we know about the DGM. A SWIG is a type of causal diagram that represents potential outcomes under specific interventions within a single possible world, encoding counterfactual relationships in graphical form while accounting for how unmeasured confounding might affect outcomes. It&#8217;s a way to visualise whether treated and control units would have followed the same trend, had neither been treated. They identify three structural features in a SWIG that signal PT is likely violated:</p><ol><li><p>If pre-treatment outcomes influence who gets treated</p></li><li><p>If different unmeasured confounders affect pre- and post-treatment outcomes</p></li><li><p>If pre-treatment outcomes directly influence post-treatment outcomes</p></li></ol><p>If any of these are in your graph, then standard DiD assumptions likely fail even if your pre-trends look fine. The authors also describe the largest possible causal structure that still permits PT, and provide R code (using the <code>dagitty</code><a href="https://cran.r-project.org/web/packages/dagitty/index.html"> package</a>) so you can test these conditions using your own diagrams. (While I was doing research for this paper, I found this <a href="https://dagitty.net/dags.html">website</a> which is super fun to play with).</p><p><em>Why is this important?</em></p><p>What this paper shows is that we can do better than just eyeball pre-trends. Instead of relying on visual checks or intuition, we can use causal diagrams to *reason formally* about whether parallel trends is *even* plausible to begin with, based on our assumptions about the DGM. This is especially useful in applied settings where we already use DAGs to justify identification strategies like unconfoundedness or IV. Now, we can apply the same logic to DiD, and potentially avoid relying on a violated assumption that undermines our entire design. It&#8217;s also a step toward a more unified causal inference framework, one where DiD assumptions are made transparent and testable before estimation begins.</p><p><em>Who should care?</em></p><ul><li><p>Researchers using DiD with limited pre-treatment periods, especially in observational data</p></li><li><p>Applied economists who already use causal diagrams for IVs or unconfoundedness and want to bring the same rigour to DiD</p></li><li><p>Methodologists interested in making identification assumptions explicit</p></li><li><p>Reviewers and instructors who want a clearer way to teach or check the PT assumption</p></li><li><p>Anyone who has ever said &#8220;the pre-trends look parallel, so we&#8217;re probably fine&#8221; </p></li></ul><p><em>Do we have code?</em></p><p>I mean, you can kind of draw it on paper, but it does help to be able to do it on the computer. The authors provide R code in the Appendix (they use the dagitty package) to help you check whether your DAG satisfies their conditions. It&#8217;s a helpful way to make the theory actionable. </p><p>In summary, if you&#8217;re using DiD and assuming PT, this paper gives you a smarter way to ask: is that actually plausible? Instead of relying on visual pre-trend checks or hoping for the best, you can use a causal diagram to make your assumptions explicit and formally test whether PT is likely to hold. It&#8217;s clean, insightful, and genuinely useful. One of those papers that quietly upgrades how we think about something we thought we already understood, I enjoyed reading it :)</p><h3><strong>From What Ifs to Insights: Counterfactuals in Causal Inference vs. Explainable AI</strong></h3><h5>TL;DR: &#8220;Counterfactual&#8221; means different things depending on who you ask. In causal inference (CI), it&#8217;s about estimating what would have happened under a different treatment. In explainable AI (XAI), it&#8217;s about tweaking inputs to change a model&#8217;s prediction. This paper lays out a unified framework to understand both, compares how counterfactuals are used and evaluated across fields, and points to ways CI and XAI can learn from each other.</h5><p><em>What is this paper about?</em></p><p>Both CI and XAI are built on &#8220;what if&#8221; logic, but they use counterfactuals in very different ways and often talk &#8220;past each other&#8221;. This paper provides a much-needed comparison between the two, showing how counterfactuals are defined, used, evaluated, and interpreted in each field. CI focuses on the THEN part - what happens under an alternative treatment. XAI focuses on the IF part - what input values would lead to a different predicted outcome. This shift in emphasis leads to big differences: in the quantity of interest, the assumptions made, the level of aggregation, and what is being modified (the model vs. the data). The authors propose a general definition of counterfactuals that works for both domains and lay out a roadmap for how ideas from each can inform the other.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eozZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eozZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 424w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 848w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 1272w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eozZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png" width="1439" height="442" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:442,&quot;width&quot;:1439,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:197609,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/164225524?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eozZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 424w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 848w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 1272w, https://substackcdn.com/image/fetch/$s_!eozZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7fb6c33-50bf-4ce5-ab10-7b4de863ce40_1439x442.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>What do the authors do?</em></p><p>They break down the counterfactual into its core parts (inputs and outcomes) and then walk through how each field uses it:</p><ul><li><p>Purpose: CI estimates causal effects; XAI explains a prediction</p></li><li><p>Assumptions: "CI assumptions are about the data-generating process; XAI typically treats the model as given for the purpose of explanation</p></li><li><p>Quantity of interest: CI compares outcomes; XAI changes inputs</p></li><li><p>Aggregation: CI typically works at the group level (e.g. ATT, ATE); XAI focuses on individual-level predictions</p></li><li><p>Modification target: CI modifies the model or treatment variable; XAI modifies the input data</p></li><li><p>Evaluation: CI cares about confidence intervals and robustness; XAI cares about feasibility, proximity, and sparsity</p></li></ul><p><em>Why is this important?</em></p><p>The word &#8220;counterfactual&#8221; is everywhere, but depending on your background, you may be using the term without realising how much baggage it carries from other fields. This paper clears that up. More importantly, it opens the door to &#8220;cross-fertilisation&#8221;: CI researchers can borrow ideas from XAI about individual-level interpretability, XAI researchers can adopt tools from CI to anchor their explanations in causal logic, And both sides can contribute to a broader understanding of actionable, responsible decision-making in high-stakes settings like healthcare, education, or finance. </p><p><em>Who should care?</em></p><ul><li><p>Empirical researchers using counterfactuals in any form (policy, medicine, business, etc.)</p></li><li><p>Anyone working at the intersection of prediction and explanation (ML folks!)</p></li><li><p>XAI researchers who want to build more meaningful and robust explanations</p></li><li><p>CI researchers curious about individual-level guidance and model-based insights</p></li><li><p>Anyone writing about algorithmic fairness, recourse, or model transparency</p></li></ul><p><em>Do we have code?</em></p><p>No, it is a conceptual paper, but it *does* give you a framework you can apply to your own work. The examples (hiring decisions, loan rejections, ad targeting) are accessible and very usable for teaching or presentations.</p><p>In summary, this is a paper about translation. If you've ever used the word counterfactual in a paper, a model, or a policy brief, this helps clarify what you're really doing. CI and XAI both rely on &#8220;what ifs,&#8221; but this paper shows that they mean different things, serve different goals, and run on different assumptions. I think that understanding this is the first step toward doing better science, in both fields.</p>]]></content:encoded></item><item><title><![CDATA[More DiD Drops: Continuous Treatments, Sample Selection, and Inference Tests]]></title><description><![CDATA[What to do when your treatment isn&#8217;t binary, your sample leaks, and your SEs collapse]]></description><link>https://www.diddigest.xyz/p/more-did-drops-continuous-treatments</link><guid isPermaLink="false">https://www.diddigest.xyz/p/more-did-drops-continuous-treatments</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Thu, 15 May 2025 14:12:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gxjZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gxjZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gxjZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gxjZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2635694,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/162680543?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gxjZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gxjZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08312196-d813-4408-9249-2b0feeab52fe_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there!<br>We got a few updates a few days after I published the other post. <br>Here they are:</p><ul><li><p><a href="https://arxiv.org/abs/2201.06898">Difference-in-Differences for Continuous Treatments and Instruments with Stayers</a>, by de Chaisemartin, D&#8217;Haultf&#339;uille, Pasquier, Sow, and Vazquez-Bare (This paper was originally circulated 3 years ago (- Sow). de Chaisemartin, D&#8217;Haultf&#339;uille and Vazquez-Bare have one published in the AEA P&amp;P named <a href="https://www.aeaweb.org/articles?id=10.1257/pandp.20241049">Difference-in-Difference Estimators with Continuous Treatments and No Stayers</a>).</p></li><li><p><a href="https://arxiv.org/abs/2502.08614">Difference-in-Differences and Changes-in-Changes with Sample Selection</a>, by Javier Viviens</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5221387">Inference with Modern Difference-in-Differences Methods</a>, by Yuji Mizushima and David Powel</p></li><li><p><a href="https://www.youtube.com/playlist?list=PLXs0Z86ZOIlHc8ApDD1v_wFi1GNChN284">Prof. Cl&#233;ment de Chaisemartin&#8217;s in-depth training on DiD</a>, a series of videos organized by Prof. Andrea Albanese and hosted by LISER &amp; DSEFM, University of Luxembourg.</p></li></ul><p>Let&#8217;s go :)</p><h3>Difference-in-Differences for Continuous Treatments and Instruments with Stayers</h3><p><em>(The 2024 AEA P&amp;P paper laid foundational work for DiD estimators in settings with continuous treatments and no stayers (i.e., cases where all units change treatment intensity across time). In contrast, this 2025 paper generalises and complements that work by allowing for the presence of stayers (i.e., units whose treatment remains unchanged across periods), while also providing better identification and robust estimation. These stayers serve as the comparison group in a more &#8220;classical&#8221; DiD ~spirit~.)</em></p><h5><strong>TL;DR: this paper proposes new DiD estimators for settings where treatments are continuous (like prices or taxes) and some units don&#8217;t change their treatment over time (i.e. "stayers"). It defines two interpretable slope-based estimands (one for local effects (AS) and one for realised-policy evaluations (WAS)) and shows how to identify them using switchers vs. stayers with the same baseline treatment. The estimators are doubly robust, nonparametric, and come with pre-trends tests and an IV extension. If the 2024 AEA P&amp;P paper gave us DiD for continuous treatments with no stayers, this paper gives us the rest of the picture.</strong></h5><p><em>What is this paper about?</em></p><p>In many situations where we are interested in applying DiD methods, we assume binary treatments and rely on having a control group that never receives treatment. But many real-world policies (e.g. tax rates, prices, or even pollution levels) are continuously distributed. And sometimes, everyone is &#8220;treated,&#8221; just to different degrees. This paper tackles this problem we have. It proposes DiD estimators for continuous treatments in settings where some units change their treatment (&#8220;switchers&#8221;) and others stay at the same level (&#8220;stayers&#8221;). You can think of it as when a state that raises its gasoline tax while others keep theirs unchanged. The stayers provide the comparison group, just like in a &#8220;classic&#8221; DiD. Instead of estimating treatment effects in levels, the paper focuses on slopes: how outcomes change as treatment intensity changes. </p><p>The authors define two estimands:</p><ul><li><p>AS (Average of Slopes): the average effect of treatment changes among switchers; it gives equal weight to all switchers regardless of treatment change magnitude.</p></li><li><p>WAS (Weighted Average of Slopes): a slope summary that gives more weight to larger treatment changes, which is helpful for cost-benefit analysis. WAS weights by the absolute value of treatment changes.</p></li></ul><p>They also extend the method to IV settings (e.g. using tax changes to estimate price elasticity), allow for covariates, dynamic treatment effects, and multi-period panels, and provide doubly-robust, nonparametric estimators with valid inference and pre-trends tests.</p><p><em>What do the authors do?</em></p><p>They propose two estimators: the Average of Slopes (AS), which captures how outcomes change with treatment for switchers on average, and the Weighted Average of Slopes (WAS), which weights those changes by how much treatment actually changed. Both estimators compare switchers to stayers with the same baseline treatment and rely on a new parallel trends assumption that allows for flexible heterogeneity. They also show how to estimate these slopes using doubly-robust, nonparametric estimators with valid inference, even in small samples. The paper extends the framework to handle instrumental variables, multiple time periods, covariates, and even dynamic treatment effects. They wrap it all up with an application on gasoline taxes and a Stata package for implementation.</p><p><em>Why is this important?</em></p><p>This is really important because most real-world treatments are not binary. Prices, taxes, subsidies, exposure level, they all often vary in intensity, not just presence. Traditional DiD estimators break down in these cases, especially when there&#8217;s no clear control group. This paper offers a practical fix. By comparing switchers and stayers at the same starting point, it builds credible counterfactuals without needing everyone to be untreated. It also delivers interpretable slope-based effects that can answer both &#8220;what-if&#8221; and &#8220;was-it-worth-it&#8221; questions. And unlike many continuous-treatment approaches, these estimators are nonparametric, doubly-robust, and testable using pre-trends. That makes them much more usable in applied work.</p><p><em>Who should care?</em></p><p>Basically anyone studying policies where treatment varies by degree and not just presence. That includes researchers working on:</p><ul><li><p>Taxes and subsidies</p></li><li><p>Prices and price regulation</p></li><li><p>Pollution exposure or health risk levels</p></li><li><p>Education or welfare policies with varying intensity</p></li></ul><p>It&#8217;s especially useful for applied micro folks frustrated by DiD methods that fall apart when there&#8217;s no clean control group. If you&#8217;ve ever had to drop observations just because &#8220;everyone got some treatment,&#8221; this paper is for you.</p><p><em>Do we have code?</em></p><p>Yes, partially. The authors have released a Stata package called <code>did_multiplegt_stat</code> (currently available via SSC) that implements some of their estimators. It&#8217;s still under development, but enough is there to apply the method to standard two-period and multi-period settings. No R version yet.</p><p>In summary, this paper fills a big gap in the DiD toolbox: how to estimate treatment effects when everyone&#8217;s treated, but to different degrees. It introduces slope-based estimands that are interpretable, testable, and robust, and opens the door to more credible applied work with continuous policies. If the earlier AEA P&amp;P paper gave us DiD without stayers, this one gives us the stayers and the estimators we need. Just keep in mind some things that the authors themselves noted: while this approach offers clear advances over prior DiD work, it&#8217;s not without constraints. The AS estimator can&#8217;t handle "quasi-stayers" (units with tiny treatment changes), which prevents consistent estimation. In applications with many time periods, the method assumes no dynamic effects, meaning outcomes are only influenced by current treatment, not past treatment. Though the authors suggest fixes for this, these workarounds shrink the sample size. And as with all DiD methods, the underlying parallel trends assumption remains fundamentally untestable for the actual treatment period, even if placebo tests look promising.</p><h3>Difference-in-Differences and Changes-in-Changes with Sample Selection</h3><p><em>(<a href="https://www.javierviviens.com/">Javier</a> is a 3rd-year PhD student at the European University Institute in Italy. This is his WP. I had fun reading it, Javier&#8217;s writing is really pleasant and easy to follow)</em></p><h5>TL;DR: when treatment affects who stays in your sample, DiD can break down. This paper shows that DiD estimates can become non-causal (or even undefined) under sample selection, especially when outcomes aren't observed or even well-defined for some units. It adapts Lee bounds to both DiD and Changes-in-Changes (CiC) settings and proposes a method to estimate causal effects for the &#8220;Always-Observed&#8221; stratum (e.g. the people for whom outcomes exist under both treatment and control). The paper also relaxes monotonicity assumptions and offers a partial identification strategy that holds up even when treatment is confounded.</h5><p><em>What is this paper about?</em></p><p>This working paper tackles a classic problem DiD designs often can&#8217;t account for: what if treatment affects who&#8217;s in your sample? Think dropout, death, non-response, or employment status. If treatment changes who shows up in the data, standard DiD estimates can become biased or even meaningless. The paper focuses on units for whom the outcome is always defined (what Javier calls the Always-Observed stratum). For this group, he develops a way to estimate both average and quantile treatment effects, even when the treatment is confounded and sample selection is not ignorable. He also adapts Lee bounds to both DiD and Changes-in-Changes (CiC) frameworks, providing partial identification results under weaker assumptions than existing methods.</p><p><em>What does the author do?</em></p><p>Javier reworks DiD for cases where sample selection isn&#8217;t ignorable, aka endogenous (e.g. when people drop out of a study because of the treatment, or when the outcome just doesn&#8217;t exist, like wages for the unemployed). He focuses on the subset of people for whom the outcome is well-defined regardless of treatment status and builds identification and inference strategies around that group. He generalises Lee bounds to the DiD and CiC settings, allowing for partial identification of:</p><ul><li><p>The average treatment effect for Always-Observed units (ATTAO), and</p></li><li><p>The quantile treatment effect (QTTAO), to explore distributional impacts.</p></li></ul><p>In the paper he also extends the method to relax monotonicity assumptions by using multiple sources of sample selection (like attrition and dropout), which lets him point-identify stratum proportions in more realistic settings. Finally, he puts everything to work in a job training program evaluation in Colombia, showing how na&#239;ve DiD would have overstated the treatment effect. It&#8217;s the kind of paper that quietly fills a gap many applied researchers have likely tiptoed around without realising how deep it goes.</p><p><em>Why is this important?</em></p><p>In real-world data, who is observed often depends on the treatment. Job training affects employment; education policies affect dropout. And when outcomes don&#8217;t exist for the unobserved (like wages for the unemployed), even the causal estimand can fall apart. Most DiD applications quietly assume this isn&#8217;t a problem. This paper says otherwise, and gives you tools to deal with it. By focusing on Always-Observed units and adapting Lee bounds, it lets you estimate effects that are well-defined, interpretable, and policy-relevant, even when selection is messy and treatment is confounded. If you&#8217;ve ever run DiD on a shrinking sample and just hoped for the best, this might be your way out of it.</p><p><em>Who should care?</em></p><p>Anyone doing DiD with panel or survey data where attrition, dropout, or censoring might be related to treatment. That includes:</p><ul><li><p>Labour economists studying training, unemployment, or job loss</p></li><li><p>Education researchers tracking dropout or test participation</p></li><li><p>Health economists dealing with mortality or survey non-response</p></li></ul><p>It&#8217;s also relevant for RCTs with imperfect compliance or missing outcomes, especially when researchers try to squeeze DiD or CiC frameworks onto real data that leaks. If your outcomes only exist for a non-random slice of your sample, this is for you.</p><p><em>Do we have code?</em></p><p>An R package is in development (we should give Javier a RA), though not yet public. The paper is very clear on the estimation steps, so if you&#8217;re comfortable with trimming and quantile bounds, you could replicate the method manually. </p><p>In summary, when treatment affects who&#8217;s in your sample, your DiD estimates might not mean what you think they do (or anything at all). This WP offers a sharp correction: estimate effects only for units with well-defined outcomes under both treatment arms. It adapts Lee bounds to DiD and CiC, covers both mean and quantile effects, and relaxes classic assumptions like monotonicity. A clean, careful contribution that quietly plugs a major gap in applied DiD work. Javier also identifies many promising extensions to build upon this work like: extending the method to settings with multiple pre- and post-treatment periods, which would help test identification assumptions more thoroughly; adapting the approach for staggered adoption designs through pairwise comparisons or integration with doubly robust estimators; incorporating covariates to tighten bounds and increase credibility; and even adopting a Bayesian perspective when bounds are too wide to be informative. Something to look forward to.</p><h3>Inference with Modern Difference-in-Differences Methods</h3><p><em>(Yuji is a PhD student at RAND Graduate School and an assistant policy analyst at RAND, while David is a a senior economist at RAND and a member of the RAND School of Public Policy faculty)</em></p><h5>TL;DR: most new DiD estimators fix bias from staggered adoption and heterogeneity, but what about inference? This paper runs large-scale simulations to test how well these estimators behave when the number of treated units is small. The authors find that most over-reject the null, especially with fewer than 10 treated units. The one approach that consistently stays close to the nominal 5% rejection rate? An imputation-based estimator paired with a wild bootstrap. If you&#8217;ve been assuming the standard errors are fine, you might want to read this.</h5><p><em>What is this paper about?</em></p><p>In this paper the authors evaluate how well modern DiD estimators perform when it comes to inference, especially in small-sample settings. Over the last few years, a wave of new DiD methods have corrected for issues like staggered adoption and treatment heterogeneity, but most of the focus has been on point estimates. Yuji and Powell ask: do these methods produce reliable standard errors and valid p-values when the number of treated units is small? To answer that, they simulate thousands of placebo policies (in CPS data and synthetic panels) and compare how often each estimator falsely rejects the null when it shouldn&#8217;t. The paper covers seven popular estimators (including Callaway and Sant&#8217;Anna, Sun and Abraham, Borusyak et al., Gardner&#8217;s 2SDID, and DCDH) and evaluates different inference procedures: cluster-robust SEs, wild bootstrap, randomization inference, and loads more.</p><p><em>What do the authors do?</em></p><p>They run a battery of simulations (both in real CPS data and fully synthetic panels) to see how different DiD estimators perform under the null. The key question is: do these methods reject the null too often when the number of treated units is small? Most do. They test seven leading DiD estimators, including: Callaway and Sant&#8217;Anna (2021), Sun and Abraham (2021), Borusyak et al. (2024), Gardner&#8217;s 2SDID, Wooldridge (2024), de Chaisemartin &amp; d&#8217;Haultfoeuille (DCDH), and stacked DiD. Each is evaluated with its default inference method (usually asymptotic or cluster-robust), and sometimes also with wild bootstrap or randomization inference. They vary sample size, number of treated units, and treatment heterogeneity. Their best-performing combo? Imputation-based estimators paired with wild bootstrap, which is stable even with just 5 treated units. They also call attention to two competing biases at work: small samples lead to over-rejection, while treatment heterogeneity inflates standard errors and reduces rejection rates. By comparing simulations with and without heterogeneity, they show these biases don't really cancel each other out. Surprisingly, they also find that for some estimators, when we increase the number of untreated units, it actually worsens rejection rates rather than improving them, which goes against the usual intuition that more untreated units = better inference.</p><p><em>Why is this paper important?</em></p><p>It&#8217;s important for a couple of reasons: most of us don&#8217;t have the time (or the human resources) to run 1,000 simulations every time we worry about small-sample inference. This paper does that work for us. It shows that many modern DiD estimators (despite fixing bias from staggered timing and treatment heterogeneity) still rely on inference procedures that fall apart when the number of treated units is small. Cluster-robust SEs and asymptotic formulas often lead to massive over-rejection. That means our p-values might be &#8220;lying&#8221; to us, and that published findings could be reporting spurious significance just because inference was off. This paper also points to safer alternatives, like pairing imputation methods with wild bootstrap, that behave much better in these settings.</p><p><em>Who should care?</em></p><p>Anyone doing DiD with a small number of treated units, which includes cases where researchers are:</p><ul><li><p>Evaluating state- or district-level policies with a few adopters</p></li><li><p>Analysing corporate or regional rollouts</p></li><li><p>Working with staggered adoption in small samples</p></li></ul><p>It&#8217;s especially relevant if you&#8217;ve moved beyond TWFE but still rely on default SEs, or assume that new methods automatically fix inference too. This paper is a reminder that point estimates and p-values live separate lives, and we need to validate both.</p><p><em>Do we have code?</em></p><p>This is less a &#8220;new code&#8221; paper and more a &#8220;here&#8217;s how to use the tools you already have <em>properly</em>&#8221; paper. The authors use existing Stata packages for all the estimators (like <code>did2s</code>, <code>csdid</code>, <code>did_imputation</code>, etc.) and pair them with inference strategies like wild bootstrap and randomization tests. No new package is released, but if you&#8217;re already using these estimators, you can replicate their setup pretty easily. </p><p>In summary, Yuji and Powell take the DiD estimators many of us already use and ask a simple question: do they behave when treated units are few? In most cases, the answer is no. Inference breaks down, p-values overstate significance, and standard errors underperform. But not all hope is lost: pairing imputation-based estimators with a wild bootstrap comes closest to nominal rejection rates, even with just 5 treated units. If you&#8217;re working with small samples or rare treatment, this paper is a practical guide to doing inference that actually holds up.<br></p>]]></content:encoded></item><item><title><![CDATA[Four new DiD studies + a textbook update]]></title><description><![CDATA[In celebration of 500 subs :) happy to have you all here !]]></description><link>https://www.diddigest.xyz/p/four-new-did-studies-a-textbook-update</link><guid isPermaLink="false">https://www.diddigest.xyz/p/four-new-did-studies-a-textbook-update</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Mon, 28 Apr 2025 15:32:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9OJm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9OJm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9OJm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9OJm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:134227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/162030655?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9OJm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!9OJm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6121055a-4633-43f4-985d-3c4c7f16f204_1024x1024.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi there! It&#8217;s been a while :) It turns out some interesting new developments hit the web in the meantime, and they&#8217;re the ones we will be talking about today.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.diddigest.xyz/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">DiD Digest is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I highlight 5, in no specific order: </p><ol><li><p><a href="https://arxiv.org/pdf/2502.16126">Conditional Triple Difference-in-Differences</a>, by Dor Leventer </p></li><li><p><a href="https://arxiv.org/pdf/2501.14405">Triple Instrumented Difference-in-Differences</a> by Sho Miyaji</p></li><li><p><a href="https://arxiv.org/pdf/2504.13057">Covariate Balancing Estimation and Model Selection For Difference-in-Differences Approach</a>, by Takamichi Baba and Yoshiyuki Ninomiya</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4736857">Spatial Synthetic Difference-in-Differences</a>, by Renan Serenini and Frantisek Masek (it was posted last year but some of you might not be aware)</p></li><li><p><a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4487202">Credible Answers to Hard Questions: Differences-in-Differences for Natural Experiments</a>, by Cl&#233;ment de Chaisemartin and Xavier D'Haultf&#339;uille (this is a textbook originally published two years ago, but it has just been updated.)</p><p></p></li></ol><h3>Conditional Triple Difference-in-Differences</h3><p><em>(<a href="https://sites.google.com/mail.tau.ac.il/dor-leventer/home">Dor</a> is a PhD candidate - advised by prof Saporta-Eksten - and apparently a very prolific one!)  </em></p><h5><strong>TL;DR:</strong> your trusty triple-difference (TDID) regression stuffed with dozens of controls can still feed you biased estimates whenever the treatment and comparison groups carry different mixes of those controls. This pre-print pins down the bias mathematically, then fixes it with a <strong>re-weighted, doubly-robust TDID estimator</strong>: one set of weights forces covariate balance, another piece mops up any remaining outcome misspecification. Dor gives us a R package, <code>tdid</code>, so you can swap the new estimator into your code and get triple differences that really reflect the causal effect you care about.</h5><p><em>What is this paper about?</em></p><p>Researchers often extend DiD to a triple difference (TDID) when an extra contrast is needed (e.g., time &#215; treated &#215; group). In practice, many studies add covariate controls and then difference the two DiD estimates, assuming the residual bias cancels. Dor shows this fails whenever the groups have different covariate distributions (because the group-specific bias in conditional parallel trends generally does not wash out). Merging the unconditional TDID framework of <a href="https://academic.oup.com/ectj/article-pdf/25/3/531/45842047/utac010.pdf">Olden &amp; M&#248;en (2022)</a> with the conditional-DiD framework of <a href="https://www.sciencedirect.com/science/article/pii/S0304407620303948">Callaway-Sant&#8217;Anna (2021)</a>, he formalises a conditional TDID setting, proves the conventional estimator is biased, and derives an alternative estimand that is recoverable: ATT for group A minus a covariate-reweighted conditional ATT for group B (using group A&#8217;s X-distribution). A double-robust/weighted-double-robust (DR/WDR) estimator with influence-function theory delivers consistent inference, and the accompanying R package <code>tdid</code> makes implementation and Monte-Carlo replication easy.</p><p><em>What does the author do?</em></p><p>Dor develops a conditional TDID framework that unifies two literatures: the unconditional triple&#8208;difference set-up of Olden &amp; M&#248;en (2022) and the covariate-adjusted DiD framework of Callaway-Sant&#8217;Anna (2021). Within this framework he:</p><ol><li><p>Diagnoses the problem:<br>Theorem 1 shows that the familiar recipe (e.g., run DiD + controls separately in each group and difference the two estimates) does not deliver the desired causal contrast whenever the groups&#8217; covariate distributions differ. The bias stems from group-specific deviations in conditional parallel trends that fail to cancel.</p></li><li><p>Defines an estimand that is identifiable:<br>He shifts attention to</p><p></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;ATT_A&#8197;&#8202;&#8722;&#8197;&#8202;E&#8201;&#8291;[CATT_B(X)&#8739;G=A]&quot;,&quot;id&quot;:&quot;ERCKHNPEAU&quot;}" data-component-name="LatexBlockToDOM"></div><p>i.e., the treatment effect for group A minus a covariate-reweighted conditional ATT for group B, where the weights replicate group A&#8217;s covariate mix. Theorem 2 proves this quantity is identified under the same conditional-parallel-trends assumption.</p></li><li><p>He pairs the usual DR estimator for group A with a weighted DR (WDR) estimator for group B, then differences the two. Because either the outcome model or the generalized propensity score can be misspecified (but not both), the estimator remains consistent and asymptotically normal; an influence-function derivation yields valid standard errors.</p></li><li><p>A toy analytical example and a Monte-Carlo exercise show that the conventional estimator is biased (direction and magnitude match the theory), while the DR/WDR estimator is centered on the true parameter.</p></li></ol><p><em>Why is this important?</em></p><p>Triple&#8208;difference designs are everywhere in applied work. Dor reviews 66 highly cited TDID papers in top journals and finds that 73 % add controls and 70 % rely on a fully interacted three-way specification. Yet the paper proves that this widespread practice is generically biased whenever the covariate mix differs across groups, which is exactly the scenario that motivates adding controls in the first place. The contribution therefore closes a serious identification gap: it shows when and why the usual estimator fails and replaces it with a re-weighted, double-robust alternative whose consistency can survive misspecification in either the outcome model or the generalized propensity score. That safeguard is valuable for empirical researchers who increasingly pair DiD methods with high-dimensional or machine-learning covariates. Beyond theory, the paper is practice-ready. The accompanying R package <code>tdid</code> implements the weighted DR estimator and reproduces the simulation evidence, lowering the cost of adoption for graduate students, replication teams, and journal reviewers alike. </p><p><em>Who should care?</em></p><ul><li><p>Applied micro-economists in labour, public, development, health, education, or environmental economics who already lean on triple differences to sharpen identification.</p></li><li><p>Empirical researchers and journal referees who worry that adding controls may quietly re-introduce bias they thought the third &#8220;D&#8221; removed.</p></li><li><p>Methodologists extending the DiD toolkit to high-dimensional controls, ML adjustments, or heterogeneous treatment effects.</p></li><li><p>Graduate students and replication teams tasked with validating published TDID results that rely on the conventional &#8220;DiD-with-controls, then difference&#8221; workflow. &#8203;</p></li></ul><p><em>Do we have code?</em></p><p>We have a package, <code>tdid</code><a href="https://github.com/dorleventer/tdid"> </a>on GitHub. With it we can implement the double-robust (DR) and weighted DR estimators laid out in the paper. It provides a vignette that walks through the toy example and Monte-Carlo simulation, and exposes helper functions for influence-function standard errors and covariate-reweighting diagnostics.</p><p>Installation is one line:</p><pre><code>devtools::install_github("dorleventer/tdid")</code></pre><p>After that, estimating a bias-corrected TDID is essentially:</p><pre><code><code>fit &lt;- tdid(y, time, treat, group, x_covariates, data = df) summary(fit)</code></code></pre><p>The README and vignette reproduce every figure in the paper, so users can trace each step from identification to inference.</p><p>When your empirical design leans on a triple difference and covariate adjustment, those estimates are only as good as the identification behind them. This paper (and the accompanying <code>tdid</code> package) supply a theoretically sound and implementation-ready fix, ensuring the third &#8220;D&#8221; does what you intended: isolate causality, not compound bias.</p><h3>Triple Instrumented Difference-in-Differences</h3><p><em>(<a href="https://sites.google.com/view/sho-miyaji/home?authuser=0">Sho</a> will be starting their PhD at Yale this fall, which is really impressive! Keep an eye out for them :) Being familiar with Econometrics notation will help you understand their work better)</em></p><h5>TL;DR: standard DID-IV compares two groups over two periods; Sho shows how to add a third dimension (e.g., season &#215; state &#215; time) and still recover a local ATE. The key is a triple Wald-DID ratio (DDD in outcomes divided by DDD in treatment) that is valid under monotonicity plus &#8220;common acceleration&#8221; (a triple-difference analogue to parallel trends). The framework extends to staggered roll-outs and ordered (non-binary) treatments, and comes with clear guidance on estimation and asymptotic inference.</h5><p><em>What is this paper about?</em></p><p>Triple-difference designs sometimes face endogeneity in the treatment itself; researchers therefore instrument and difference, but the identification conditions had not been formally spelled out. In this paper Sho defines a triple DID-IV set-up, pins down the target estimand (LATET for the subgroup affected by the instrument in the third dimension), and states the assumptions (monotonicity and common acceleration in both treatment and outcome) that make the triple Wald-DID ratio valid. The analysis covers both two-period and staggered-adoption panels. </p><p><em>What does the author do?</em></p><ol><li><p>Sho introduces notation for a two-period, three-group setting (exposed/unexposed x instrumented/non-instrumented &#215; demographic group) and defines the triple Wald-DID estimand.</p></li><li><p>They prove that the ratio of DDDs equals the LATET under monotonicity and common-acceleration assumptions (Theorem 1), and generalise to ordered treatments (Theorem 2).</p></li><li><p>Then they partition units into cohorts by instrument-start date, define cohort-specific LATETs (CLATTs), and show how to estimate them with either a never-exposed or last-exposed control cohort (Theorems 3&#8211;5).</p></li><li><p>Finally, they provide influence-function formulas and show that an IV regression with triple-interaction dummies in both first stage and reduced form delivers the estimator and its standard error.</p></li></ol><p><em>Why is this important?</em></p><p>Instrumented DiD is already popular when no clean control group exists; empirical work is increasingly layering a third difference (season, demographic subgroup, geography) on top of an instrument without a clear blueprint for identification. Sho supplies that blueprint, shows the precise conditions under which the familiar &#8220;DDD-IV&#8221; regression isolates causal effects, and offers a robust alternative to two-way-fixed-effects IV estimators that can be badly biased under staggered timing.</p><p><em>Who should care?</em></p><ul><li><p>Applied researchers using instruments that switch on only for a subset within a treated group (e.g., environmental regulations applied only in summer months in participating states).</p></li><li><p>Econometricians extending DiD/IV methods to heterogeneous treatments, staggered timing, or multi-dimensional policy variation.</p></li><li><p>Reviewers assessing DDD-IV studies who need a clear checklist of assumptions and estimators.</p></li></ul><p><em>Do we have code?</em></p><p>No public package yet. Estimation boils down to a standard IV regression with fully interacted group &#215; time &#215; instrument dummies; the paper&#8217;s appendix gives ready-to-copy equations for both two-period and staggered panels. (If a replication package appears, it will likely be linked from the Sho&#8217;s homepage or the arXiv record.)</p><h3>Covariate Balancing Estimation and Model Selection for Difference-in-Differences</h3><h5>TL;DR: when you run a DiD you usually weight observations by a propensity score, and if that score is even a little wrong your estimate can drift. Baba &amp; Ninomiya show how to choose weights that force the weighted treated and control samples to match on the chosen covariate moments (even quadratic ones), so bias disappears even when the propensity model is misspecified <strong>as long as the before-minus-after outcome change is linear in those covariates</strong>. They also derive a smarter information criterion that tells you which covariates are really worth including.</h5><p><em>What is this paper about?</em></p><p>Abadie&#8217;s semiparametric difference-in-differences (<a href="https://www.jstor.org/stable/3700681">SDID</a>) estimator is unbiased only when the propensity-score model is correctly specified; misspecification introduces bias. Baba and Ninomiya propose an alternative SDID procedure, covariate balancing for DID (CBD), that chooses propensity-score weights by solving moment conditions that force the weighted treated and control samples to match on selected covariate moments, including second-order terms such as xx&#7488;. The resulting average-treatment-effect-on-the-treated (ATT) estimator is doubly robust: it stays consistent if either (i) the propensity-score model is correct or (ii) the before-minus-after change in outcomes is linear in the covariates. Baba and Ninomiya derive its large-sample distribution and show analytically and via simulations that CBD removes the bias seen with conventional maximum-likelihood weights when the propensity model is misspecified. Because the weights themselves are estimated, standard information criteria (AIC, GIC) are inappropriate for covariate selection. The authors develop an asymptotically unbiased risk-based information criterion whose penalty depends on the estimated weighting matrix and the heteroskedasticity of the weighted outcomes; in practice this penalty is often much larger than the familiar 2 &#215; (number of parameters). Simulations confirm that the new criterion selects sparser, lower-risk models than an AIC-style adaptation (QICw). A re-analysis of the LaLonde job-training data illustrates the practical gains from both CBD weighting and the new model-selection rule. </p><p><em>What do the authors do?</em></p><p>Instead of guessing a propensity-score formula and hoping for the best, they choose weights that make the weighted treated and control samples exactly match on selected covariate moments&#8212;including quadratic (covariate &#215; covariate) terms. Think of it as forcing the scales to balance before comparing outcomes. With those balanced weights in place, they run the usual before/after comparison, but with a twist: the ATT estimate stays consistent if either the weight recipe is correct or the before-minus-after change in outcomes is linear in those covariates. One correct piece is enough. Adding covariates can help or just add noise. Classic AIC-style rules under-penalise that noise once weights are estimated, so the authors derive a larger, data-dependent penalty that tells you when an extra variable earns its keep. In simulations, their method wipes out the bias that appears when the usual propensity model is misspecified, and the new information criterion selects lean, low-risk models. Re-analysing the famous LaLonde job-training data, the CBD criterion selects a noticeably different (much sparser) set of covariates than an AIC-style rule, illustrating the practical impact of both the balancing weights and the new model-selection penalty.</p><p><em>Why is this important?</em></p><p>Applied DiD studies often spend pages justifying that treatment and control &#8220;look similar.&#8221; CBD makes them similar by construction, so that diagnostic &#8220;headache&#8221; disappears. If your usual propensity-score equation is even slightly misspecified, traditional weights can push your estimate off target. CBD cushions that risk; one correct ingredient (either the weights or a simple outcome trend) is enough to keep the estimate on track. Adding controls can just as easily inflate variance as reduce bias. The new information criterion tells you quantitatively when a covariate earns its keep, instead of relying on ad-hoc *p-value fishing* or AIC rules that are too forgiving once weights are estimated. In simulations and the LaLonde re-analysis, CBD removed bias, kept models lean, and produced a noticeably larger and more precise treatment effect than the standard approach. In other words: fewer assumptions, tighter answers. &#8203;</p><p><em>Who should care?</em></p><ul><li><p>The method is most directly useful in two-period DiD designs (pre-/post- only). Multi-period settings will need future extensions, which the authors mention in their discussion.</p></li></ul><ul><li><p>Researchers running DiD in economics, public health, education, or policy who rely on propensity-score weights and worry about whether those weights truly balance their samples.</p></li><li><p>Any analyst who already uses &#8220;covariate-balancing&#8221; tools (CBPS, kernel balancing, etc.) is also a natural user, because CBD brings that logic into the DiD world.</p></li><li><p>Replication teams and journal reviewers who want a transparent check that treated and control groups are comparable rather than trusting a black-box propensity model.</p></li><li><p>Methodologists developing robust causal estimators. CBD adds a practical, doubly-robust tool to the DiD arsenal. &#8203;</p></li></ul><p><em>Do we have code?</em></p><p>The article gives formulas, theorems, and simulation/empirical results, but it does not include any replication code, software appendix, or GitHub link. The only software-related remark is that the LaLonde data come from the R package <em>Matching</em>; everything else is presented algebraically and in tables. (The authors do mention that the weighting matrix can be computed with GMM or GEL, but they do not supply code for doing so.)</p><h3>Spatial Synthetic Difference-in-Differences</h3><p><em>(Thanks prof Renan for sending me this!)</em></p><h5>TL;DR: this paper extends Synthetic Difference-in-Differences (SyDiD) to settings where policies spill across space, violating the no-interference (SUTVA) assumption. By embedding a spatial-weights matrix in SyDiD&#8217;s weighted two-way-fixed-effects regression, the authors create <strong>Spatial SyDiD (SpSyDiD)</strong>, which simultaneously estimates <strong>direct effects on treated units</strong> and <strong>indirect (spillover) effects</strong> on their neighbours. Monte-Carlo experiments show SpSyDiD outperforms both &#8220;vanilla&#8221; SyDiD (which ignores spillovers) and Spatial DiD (which uses uniform weights) in bias and precision. A re-analysis of Arizona&#8217;s 2007 employer-sanctions law illustrates the method. </h5><p><em>What is this paper about?</em></p><p>This paper gives the classic DiD toolkit a badly needed &#8220;GPS upgrade&#8221; for the real world, where policies rarely stay inside state lines. Think minimum-wage hikes that nudge workers across borders, congestion charges that reroute traffic into the suburbs, smoking ban in one city that push smokers next door, or a jobs program in one region that can poach workers from its neighbours. In standard DiD we kind of assume these cross-border ripples don&#8217;t exist, which can bend our estimates.  To tackle this problem, Serenini and Masek fuse two well-known ideas:</p><ul><li><p>Who&#8217;s my neighbour?&#8194;A spatial-weights matrix lists, for every region, which other regions lie close enough to feel its policy shock and how strongly they are connected.</p></li><li><p>Who&#8217;s my best match?&#8194;Synthetic DiD already re-weights control regions and pre-treatment periods so the treated region&#8217;s trend is mimicked as closely as possible.</p></li></ul><p>Merge the two and you get Spatial Synthetic DiD (SpSyDiD), a simple weighted regression that delivers two headline numbers instead of one:</p><ol><li><p>Direct effect (&#964;): what happens inside the region that actually adopts the policy.</p></li><li><p>Spillover effect (&#964;&#8347;): how much of that impact seeps into its neighbours.</p></li></ol><p>The authors show, with simulations covering &#8220;many treated counties&#8221; and the classic &#8220;one treated state&#8221; setup, that omitting &#964;&#8347; can induce noticeable bias and extra variance; in their county experiments the relative error reaches a few percentage points and grows with stronger spillovers. SpSyDiD, by contrast, keeps both direct and indirect estimates on target without giving up the nice robustness features of Synthetic DiD. A case study makes it concrete: Arizona&#8217;s 2007 employer-sanctions law cut the share of working-age non-citizen Hispanics in the state by &#8211;2.7 percentage points, while neighbouring Nevada, Colorado and New Mexico saw a +0.9 pp uptick, which is evidence that the law displaced people rather than making them disappear from the data.</p><p><em>What do the authors do?</em></p><p>Serenini and Masek take Synthetic DiD (the re-weighting framework that already softens the parallel-trends assumption) and graft onto it the core machinery of spatial econometrics. They let the treatment indicator bleed through a row-standardised spatial-weights matrix W, so the regression now contains two coefficients: &#964;, the change experienced by the directly treated region, and &#964;&#8347;, the change that reaches its neighbours. They preserve SyDiD&#8217;s unit weights (&#969;) and time weights (&#955;), which means the new estimator, Spatial SyDiD, can be implemented with the same weighted least-squares routine once those weights and W are in hand. The paper spells this out as a six-step recipe that needs nothing fancier than ordinary matrix algebra. &#8203;After building the estimator, they show why it matters. A short algebraic detour reveals that if a researcher ignores spillovers, the average treatment effect is biased by the factor (1 + &#961; W&#772;), where &#961; measures how strongly shocks diffuse and W&#772; is the average share of treated neighbours. In other words, the more porous the border, the further your na&#239;ve DiD slides from the truth. &#8203;The authors then stress-test their method in two simulation worlds. In a county-level setting with many treated units, Spatial SyDiD recovers both the Average Treatment on the Treated and the Average Indirect Treatment Effect with errors under three percent, while plain SyDiD (which has no spillover channel) can only speak to the direct effect and spatial DiD (which uses uniform weights) is less precise. Repeating the exercise at the state level with a single treated unit produces the same ranking: Spatial SyDiD is markedly closer to the data-generating truth and roughly halves the root-mean-squared error relative to its competitors. &#8203; To show the estimator at work in real data, they revisit Arizona&#8217;s 2007 Legal Arizona Workers Act. Using Current Population Survey microdata, Spatial SyDiD attributes a 2.7-percentage-point fall in working-age non-citizen Hispanics to the law inside Arizona, alongside a 0.9-point rise in neighbouring Nevada, Colorado and New Mexico&#8212;evidence that the legislation displaced people rather than erasing them from the labour force. &#8203; Finally, the paper adapts Synthetic DiD&#8217;s placebo-based inference to the spatial setting. By repeatedly re-assigning &#8220;treated&#8221; and &#8220;neighbour&#8221; labels among control units, the authors build a finite-sample variance for both &#964; and &#964;&#8347; without leaning on additional distributional assumptions, giving practitioners an off-the-shelf way to attach standard errors to each component of the effect. &#8203;</p><p><em>Why is this important?</em></p><p>Most policy studies assume each region is an island; when that fails, estimates can be way off. Spatial SyDiD gives researchers a plug-and-play fix: it keeps SyDiD&#8217;s relaxed parallel-trend strengths while explicitly measuring how much impact leaks across borders. That means: (i) cleaner causal claims when spillovers exist; (ii) a transparent split between &#8220;what happened here&#8221; and &#8220;what we pushed onto the neighbours,&#8221; which is exactly what policymakers worry about; and (iii) no need for exotic software, you just add a spatial weight matrix to workflows people already use. In short, it turns an unchecked threat to validity into something you can quantify, interpret, and report. &#8203;</p><p><em>Who should care?</em></p><ul><li><p>Applied micro-economists evaluating policies that plausibly shift jobs, prices, or people across county or state lines.</p></li><li><p>Urban and regional planners measuring knock-on effects of zoning changes, congestion charges, or transit investments.</p></li><li><p>Public-health and environmental scientists tracking how smoking bans, pollution controls, or epidemics spill into neighbouring areas.</p></li><li><p>Political scientists studying policy diffusion and electoral spillovers.</p></li><li><p>Impact-evaluation teams at governments and NGOs that rely on DiD/SCM workflows but face obvious &#8220;border leakage.&#8221;</p></li></ul><p><em>Do we have code?</em></p><p>Yes! We have a <a href="https://github.com/serenini/spatial_SDID">repo</a> on GitHub that you can clone to reproduce the findings of the paper (and maybe adapt the code to your own research). </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.diddigest.xyz/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">DiD Digest is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[CME + RDD + IV because DiD has been quiet ]]></title><description><![CDATA[Meanwhile, in the "rest" of causal inference...]]></description><link>https://www.diddigest.xyz/p/cme-rdd-iv-because-did-has-been-quiet</link><guid isPermaLink="false">https://www.diddigest.xyz/p/cme-rdd-iv-because-did-has-been-quiet</guid><dc:creator><![CDATA[Beatriz Gietner]]></dc:creator><pubDate>Thu, 10 Apr 2025 13:05:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!g2jt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!g2jt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!g2jt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 424w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 848w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!g2jt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg" width="500" height="625" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:625,&quot;width&quot;:500,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74648,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://diddigest.substack.com/i/161003654?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!g2jt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 424w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 848w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!g2jt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6bd4328-4160-4207-ba72-5468fbb60f30_500x625.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>                                        (Image by prof <a href="https://x.com/chelseaparlett/status/1458461737431146500">Chelsea Parlett-Pelleriti</a>)</p><p>Hi there!</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.diddigest.xyz/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading DiD Digest! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I found 2 papers and 1 guide in the past two weeks that I thought some of you would appreciate. The IV one is related to Health Economics, the RDD is actually a job market paper (full of colorful plots, we love it), and the guide was actually born out of Political Economy (co-written by 2 PhD students!), so I guess there is something for everyone here :)</p><p>Here they are:</p><ol><li><p><a href="https://acrobat.adobe.com/id/urn:aaid:sc:EU:6ff45cfa-8e3e-484e-ad9c-5a002835f67c">IV in Randomized Trials</a>, by Joshua D. Angrist, Carol Gao, Peter Hull, and Robert W. Yeh</p></li><li><p><a href="https://arxiv.org/pdf/2504.03992">RDD with Distribution-Valued Outcomes</a>, by David Van Dijcke (<a href="https://x.com/packlesshepherd/status/1909793301760205288">thread</a>)</p></li><li><p><a href="https://arxiv.org/pdf/2504.01355">A Practical Guide to Estimating Conditional Marginal Effects: Modern Approaches</a>, by Jiehan Liu, Ziyi Liu, and Yiqing Xu (<a href="https://x.com/xuyiqing/status/1907664132175966605">thread</a>)</p></li></ol><p>We will go from shortest to longest today.</p><p></p><h3>IV in Randomized Trials</h3><p><em>(You do not need to be an econometrician to get what is going on here, though it helps to have one nearby - it can be on X/BSky).</em></p><h5>TL;DR: clinical trials are sometimes messy. Patients do not always do what they are assigned to do: some in the treatment group never take the treatment, while some in the control group sneak off and take it anyway. This messes with both intention-to-treat (ITT) and per-protocol analyses (the intended treatment is randomly assigned but treatment received is not): the first underestimates the true effect, and the second suffers from selection bias. IV methods offer a middle way, and this paper shows how to apply them in real-world trials. The authors explain IV theory using the ISCHEMIA trial as a motivating example and walk through how IV estimates recover local average treatment effects (LATE) for compliers (the participants whose behaviour was actually influenced by their random assignment). The paper argues that IV methods should be standard in pragmatic, strategy, and nudge trials where adherence is imperfect.</h5><p></p><p><em>What is this paper about?</em></p><p>This paper revisits a well-known challenge in clinical trials: what do we do when patients do not stick to their assigned treatments? Traditional approaches like intention-to-treat (ITT) analyses estimate the effect of being assigned to treatment, regardless of whether patients actually receive it. This preserves the benefits of randomization but often underestimates the treatment&#8217;s true effect when adherence is imperfect. On the other hand, per-protocol or as-treated analyses attempt to estimate the effect of receiving treatment, but at the cost of introducing selection bias, since treatment uptake is no longer randomized. To navigate this trade-off, the authors advocate for IVs as a middle-ground solution. They treat random assignment as an instrument for actual treatment receipt and use it to estimate the Local Average Treatment Effect (LATE, the causal effect of treatment for the subgroup of participants who comply with their assigned treatment). The paper explains the theory clearly, demonstrates its application using the ISCHEMIA trial data, and argues that IV should be a standard tool in modern clinical trials, especially in settings where nonadherence or crossover are common. <em>(Best of all? Health researchers get to use IV without having to write three paragraphs defending the exclusion restriction. Must be nice, I would not know).</em></p><p><em>What do the authors do?</em></p><p>They lay out a simple yet compelling case for using IVs to estimate causal effects in clinical trials with imperfect adherence. They begin by explaining the core intuition behind IV: when random assignment affects treatment uptake, it can be used to isolate the causal effect of treatment, specifically for the subset of participants who comply because of their assignment. They walk us through the basic mechanics: divide the ITT effect (difference in outcomes by assignment) by the compliance rate (difference in treatment uptake by assignment). The result is a per-protocol effect for compliers (those whose treatment behaviour was influenced by randomization). To make this concrete, they apply IV to the ISCHEMIA trial, a large cardiovascular study where about 20% of patients assigned to invasive treatment did not actually get it, and about 12% of patients in the conservative group did. Using IV, they show that the estimated treatment effect on health-related quality of life (SAQ score) is substantially larger than the ITT estimate (and clinically meaningful). They also provide adjusted estimates using regression-based IV methods, reinforcing the idea that this is more than a back-of-the-envelope trick. </p><p><em>Why is this important?</em></p><p>This is important because most clinical trials do not go exactly as planned and pretending otherwise does not help. Nonadherence, crossover, and treatment contamination are common, especially in pragmatic or strategy trials where patient choice and real-world logistics come into play. The standard ITT approach is clean but conservative, often underestimating the true effect of treatment. Per-protocol analyses try to adjust for this, but they sacrifice the core benefit of randomization &#8594; unbiased comparison. IVs offer a &#8220;principled&#8221; way out of this dilemma. By using assignment as an instrument for treatment receipt, researchers can recover valid causal estimates for the subset of participants whose treatment was influenced by randomization &#8594; the compliers.  And while economists have long treated IV as part of the standard causal toolkit, this paper brings that logic to a clinical audience, with an accessible explanation, a real trial application, and a strong case for making IV a routine part of analysis when nonadherence is more than a footnote.</p><p><em>Who should care?</em></p><p>This paper is for clinical researchers and trialists who know their sample size calculations by heart but start sweating when compliance drops below 80%. It is for anyone designing or analyzing trials where real-world behaviour gets in the way of ideal randomization (which, let us face it, is most of them). More specifically, it is relevant to:</p><ul><li><p>Clinical trialists working on pragmatic, strategy, or nudge trials, where adherence is not enforced and treatment effects depend on behaviour.</p></li><li><p>Health economists and outcomes researchers who want causal estimates that are both internally valid and clinically meaningful.</p></li><li><p>Biostatisticians looking for alternatives to biased per-protocol analyses that do not involve violating the randomization.</p></li><li><p>Regulatory scientists and policy analysts trying to understand not just whether a treatment works, but how much it works for patients who actually receive it.</p></li></ul><p>If your trial has a CONSORT flowchart with more arrows than a subway map, IV might be exactly what you need.</p><p>In summary, if you are working with nonadherence, crossover, or anything resembling the real world, you do not need to choose between an ITT that is too soft and a per-protocol that is too biased. You can run an IV. You probably already know how. Now is the time to do it.</p><p></p><h3>RDD with Distribution-Valued Outcomes</h3><p>(<a href="https://www.davidvandijcke.com/">David</a> is on this year&#8217;s market! This is his JMP)</p><h5>TL;DR: what if your outcome is a whole distribution rather than a single number? This paper extends RDDs to handle distribution-valued outcomes (think income brackets, price spreads, or test score distributions) where you want to know how an intervention shifts the shape of the outcome, going beyond just the average. David introduces a new causal estimand (LAQTE), proposes two estimators (one local polynomial, one in Wasserstein space), and shows how to construct valid confidence bands for inference. The method is applied to U.S. gubernatorial elections, where Democratic wins reduce income inequality by compressing the top end of the distribution. The theory is sharp, the plots are gorgeous, and the contribution is genuinely useful if you care about distributional effects.</h5><p><em>What is this paper about?</em></p><p>This paper introduces a very cool (and very visual) extension of the classic RDD for situations where the outcome of interest is a distribution. Think test score distributions in schools, wage distributions within firms, or price distributions across products, cases where you care about the whole shape. Traditional RDDs are not usually equipped to handle these settings, because they assume scalar outcomes and ignore the structure (and randomness) inside each unit (there is a two-level randomness: one from how units are assigned around the cutoff, and one from sampling within each unit&#8217;s distribution). This paper provides the tools to handle both. David proposes a new framework, which he calls R3D (Regression Discontinuity with Distribution-Valued Outcomes), that treats these within-unit distributions as functional data. He defines a new causal estimand, the Local Average Quantile Treatment Effect (LAQTE), which captures how the average quantile function shifts around the cutoff. The goal is to estimate how an intervention affects the shape of the distribution, beyond just a point on it. Along the way, he introduces two estimators (one based on local polynomial regression, the other on local Fr&#233;chet regression in Wasserstein space &#8594; Maths people will have a ball with this one), proves asymptotic normality, derives uniform confidence bands, and shows us why standard quantile RDD estimators fail in these cases. The paper closes with an application to gubernatorial elections and income inequality in the U.S., showing that Democratic wins reduce top-end incomes (a classic equality&#8211;efficiency tradeoff). </p><p><em>What does the author do?</em></p><p>David develops a new framework called R3D , designed for settings where each unit (e.g., a state, school, hospital) is associated with a full empirical distribution. In these cases, traditional RDDs fall short, they cannot account for the internal variation or the structure of the distribution itself. He then defines a new causal estimand (LAQTE), which captures how the average quantile function changes at the cutoff; then he shows that standard quantile RDD estimators (e.g., estimating the 10th, 50th, 90th percentiles one by one) do not work in this setting, because they fail to aggregate information across units with different internal structures. The author proposes two estimators: a na&#239;ve estimator that estimates quantiles across units and applies standard local polynomial RDD logic, and a Fr&#233;chet/Wasserstein-based estimator, which treats the quantile function as a curve and applies local <a href="https://arxiv.org/pdf/1608.03012">Fr&#233;chet regression</a> in <a href="https://library.oapen.org/bitstream/id/b27ca94b-41a7-486c-863f-8de6b3a8f914/2020_Book_AnInvitationToStatisticsInWass.pdf">Wasserstein space</a> (yes, the space of probability distributions). It sounds fancy (maybe because Ren&#233; Maurice Fr&#233;chet was French?), but it works beautifully and gives valid inference. He also derives uniform confidence bands for the full quantile function, letting us ask questions like: where in the distribution does the treatment hit? Finally, he applies the method to a gubernatorial RDD in the U.S., showing that when Democrats win, the top end of the income distribution shrinks, even though the average income stays flat &#8594; evidence of distributional shifts without mean effects: exactly what this method is built for.</p><p><em>Why does this matter?</em></p><p>Not all treatment effects show up in the mean. In many real-world settings (from education to health to inequality research) the most meaningful changes occur in the spread, shape, or tails of a distribution. An intervention might compress inequality without raising average income, lift the bottom of the score distribution without changing the top, or reduce risk exposure at the extreme end of a health outcome. Traditional RDDs miss these effects entirely. What David shows is that we do not need to flatten rich outcome structures into single numbers. His proposed method lets us treat distributions as what they are: distributions, with estimators that are both theoretically grounded and visually intuitive. It gives us a way to ask: where is the effect happening? and How does the shape change across the cutoff? This is especially valuable for fields that care about inequality, heterogeneity, or tail risk (not just the average). And the fact that the method &#8220;nests&#8221; easily within familiar RDD logic makes it much more usable than it first sounds.</p><p><em>Who should care?</em></p><p>Basically anyone working with outcomes that are more than just averages, but more specifically:</p><ul><li><p>Education researchers analyzing test score distributions across schools or students (that would be me and a lot of you).</p></li><li><p>Health economists looking at risk distributions, not just mean outcomes.</p></li><li><p>Public finance folks studying income inequality or top-end wealth effects.</p></li><li><p>Applied microeconomists using RDDs who are somewhat frustrated by their &#8220;scalar-only&#8221; worldview.</p></li></ul><p><em>Do we have code?</em></p><p>We have a package! I do not know how he did this without RAs because it seems like a lot of work, so kudos to him! <em>(By the way do you have a moment to talk about out Lord and Saviour of <a href="https://r-pkgs.org/">R packages</a>, <a href="https://hadley.nz/">Hadley Wickham</a>?)</em></p><p>The <a href="https://www.davidvandijcke.com/R3D/">R package </a><code>r3d</code> implements both the na&#239;ve and Fr&#233;chet/Wasserstein estimators, along with tools for plotting and inference. With it you can estimate LAQTEs using either local polynomial or Fr&#233;chet regression, generate quantile function plots that show where the treatment effect lies across the distribution, and construct uniform confidence bands for inference.</p><p>In sum, this paper is a reminder that not all outcomes fit in a spreadsheet cell. When your data come as distributions (not just point estimates) you need tools that treat them that way. David proposed R3D framework gives us exactly that: a method that respects the structure of complex outcomes, builds on familiar RDD logic, and opens the door to asking richer questions about where effects happen, not just if they do. If you are working with quantiles, histograms, or any outcome with internal shape, you should check this paper out. And it comes with code, plots, and clean theory. <em>What is there not to like?</em></p><p></p><h3>A Practical Guide to Estimating Conditional Marginal Effects: Modern Approaches</h3><p><em>(I particularly enjoyed reading this one. It is a &#8220;self-explanatory&#8221; guide and the authors do their best in elucidating concepts step by step, but if you are not familiar with stats notation I suggest you try to grasp the idea of what they are trying to do first. Familiarity with R is essential, but the code chunks are very easy to follow. I also learned that &#8220;kernel&#8221; means &#8220;core&#8221;, &#8220;center&#8221;, &#8220;basis&#8221;. The more you know!)</em></p><h5><strong>TL;DR:</strong> This guide introduces robust methods for estimating how treatment effects vary with a moderating variable (&#8220;conditional marginal effects,&#8221; or CME). Traditional approaches like linear interaction models often rely on unrealistic assumptions, suffer from lack of overlap, and impose overly rigid functional forms. The authors then present a progression of methods, from parametric to more flexible ML approaches, that address these limitations while preserving statistical validity. They clearly define the CME estimand, enhance semi-parametric kernel estimators (a method that uses smooth curves to capture patterns without assuming a strict relationship between variables), and introduce modern, robust techniques like <strong>AIPW-LASSO</strong> (which combines weighting and model selection for added robustness) and <strong>Double Machine Learning</strong> (which uses ML to estimate treatment and outcome models separately, then combines them to reduce bias). The paper also offers practical recommendations based on simulations and real-world examples, and all methods are implemented in the <code>interflex</code> R package.</h5><p><em>(Before we move forward, I just wanted to say that there is one formal definition of CME in the causal inference literature - as defined in this guide - but the term "marginal effect" or "conditional marginal effect" is sometimes used differently in statistical modeling contexts, especially in GLMs or GLMMs. So the label might be the same, but the meaning depends on the framework. Here, CME is defined as:</em></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;&#952;(x)=E[Y_i(1)&#8722;Y_i(0)&#8739;X_i=x]&quot;,&quot;id&quot;:&quot;CYCCMIQCIY&quot;}" data-component-name="LatexBlockToDOM"></div><p><em>for binary treatment, or</em></p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;&#952;(x)=\\mathbb{E}\\left[\\frac{\\partial Y_i(d)}{\\partial d} \\middle| X_i = x\\right]&quot;,&quot;id&quot;:&quot;CTTBKLPZQP&quot;}" data-component-name="LatexBlockToDOM"></div><p></p><p><em>for continuous treatment. It answers "what is the average treatment effect at a specific value of the moderator X?" which is a causal estimand grounded in the potential outcomes framework.)</em></p><p></p><p><em>What is this paper about?</em></p><p>In this paper the authors tackle a fundamental and increasingly relevant challenge in applied research: how to estimate treatment effects that vary across subgroups or contexts, formally known as CME. In many fields, from political science to economics and public policy, we want to understand not just whether a policy or intervention works, but for whom and under what conditions it works best. Unfortunately, the tools available to answer these questions are often limited, outdated, or misused. </p><p>Nowadays most applied researchers turn to linear interaction models (which assume that the effect of a treatment changes in a linear way with a moderating variable &#8594; this variable tells us &#8220;for whom&#8221; or &#8220;under what conditions&#8221; the treatment has a stronger or weaker effect, it is one that we condition on to examine how treatment effects vary). These models are nice because they are simple to estimate and interpret. But in practice, they come with serious limitations: they often rely on strong, mostly unrealistic assumptions about the shape of relationships, while also being highly sensitive to model misspecification (e.g., omitted interactions and nonlinearity) &#8594; leading to misleading conclusions. They also sometimes fail to assess whether there is enough overlap in the data (i.e., whether treated and control groups are comparable across the range of the moderator &#8594; lack of common support), and they typically lack a clear connection to the causal estimands we care about. Who has not struggled to answer seminar questions about nonlinearity and high-dimensional covariates? Just add a polynomial to your already kitchen sink regression, it should work (not).</p><p>These limitations are especially problematic in observational studies, where researchers do not have the &#8220;luxury&#8221; of randomized assignment and must rely on modeling assumptions to recover causal effects.</p><p>To address these challenges, the authors build on prior work by Hainmueller, Mummolo, and Xu (the author of this one himself) <a href="https://www.cambridge.org/core/services/aop-cambridge-core/content/view/D8CAACB473F9B1EE256F43B38E458706/S1047198718000463a.pdf/how_much_should_we_trust_estimates_from_multiplicative_interaction_models_simple_tools_to_improve_empirical_practice.pdf">(2019)</a>, who introduced semiparametric kernel estimators (SKE from here onwards &#8594; they are statistical methods that relax functional form restrictions by using local weighted averaging to estimate relationships between variables, they are a middle ground between fully parametric models - with strict assumptions - and nonparametric approaches - with no functional form assumptions) to allow for more &#8220;flexible&#8221; relationships between treatment (D), moderators (X), and outcomes (Y). But even those methods left key gaps in terms of clarity, robustness, and ease of use. In this guide they clearly define the CME estimand, aligning what researchers want to estimate with how they estimate it.</p><p><em>What do the authors do?</em></p><p>The authors develop a comprehensive and unified framework for estimating CMEs,  that is, how the effect of a treatment varies with a moderating variable X, holding other covariates constant. A key contribution of the paper is to clarify what researchers are trying to estimate when they talk about treatment effect heterogeneity. As I wrote in the introduction, the authors define the CME as a special case of the conditional average treatment effect (CATE) for both binary and continuous treatments. For both cases they rely on the assumptions of Unconfoundedness and Strict Overlap (given Random Sampling and SUVTA assumptions). </p><p>Next they propose a progression of estimation strategies, from simpler ones to more flexible ones, aka a tiered approach, which allows us to choose an estimator that matches our data structure, research question, and level of complexity. </p><p>In Chapter 2 they go through classical approaches to estimating and visualizing the CME (with examples from applied papers), while &#8220;emphasizing the importance of clearly defining the estimand and explicitly stating identifying and modeling assumptions&#8221;. They address the issue of multiple comparison by introducing uniform CIs constructed via bootstrapping. They propose diagnostic tools to detect lack of common support and model misspecification and discuss strategies to minimize them. Finally, they introduce a SKE to relax the functional form assumptions (I love to say &#8220;functional form&#8221;) of linear interaction models. They also summarize the challenges and how to solve them.</p><p>In Chapter 3 the authors aim to solve Chapter 2 limitations (the fact that SKEs still rely solely on outcome modeling and may struggle with nonlinearities or complex interactions in covariates) by introducing Augmented Inverse Propensity Weighting (AIPW ) and its extensions (especially the AIPW-LASSO approach) to improve the robustness, efficiency, and flexibility of CME estimation. AIPW, introduced by <a href="https://www.jstor.org/stable/pdf/2290910">Robins, Rotnitzky, and Zhao (1994)</a>, improves on IPW by combining the outcome model and the treatment assignment model, making it doubly robust, that is, it yields consistent estimates as long as either model is correctly specified. (We have been through this before in the previous post, so you know how valuable this is, especially in observational settings where modeling assumptions are hard to verify.) When both models are correctly specified, AIPW achieves the lowest asymptotic variance within the class of IPW estimators. However, as <a href="https://academic.oup.com/aje/article-pdf/188/1/250/27238730/kwy201.pdf">Li, Thomas, and Li (2019)</a> show, in finite samples, its performance depends heavily on the degree of overlap. If overlap is poor and propensity scores are extreme, AIPW may actually underperform compared to simpler outcome-based estimators. The authors then introduce the idea of &#8220;signals&#8221; - transformed versions of the data that isolate the part relevant for estimating the CME. They propose a three-step estimation strategy based on these signals that is intuitive and straightforward to implement. Then they close the chapter with Section 3.4, which focuses on inference (a step that is often overlooked when we get excited about flexible estimators). Because the AIPW approach involves multi-step estimation (outcome modeling, propensity scores, signal construction, and smoothing), getting valid confidence intervals is non-trivial. The authors propose bootstrap-based procedures to construct both pointwise and uniform CIs for the CME. We know that uniform intervals are especially helpful when visualizing treatment effect heterogeneity across the range of the moderator, since they adjust for multiple comparisons and avoid misleading &#8220;significance at a glance&#8221; conclusions. This emphasis on valid inference ensures that the flexibility gained by AIPW-LASSO or kernel smoothing does not come at the cost of statistical rigor, something applied researchers will appreciate when facing skeptical referees (or seminar rooms).</p><p>In Chapter 4 (my favourite one!) they extends the AIPW framework into a more flexible and scalable setting using Double Machine Learning (DML). The key idea is to use machine learning algorithms (like random forests, neural nets, or gradient boosting) to estimate nuisance functions (the outcome and treatment models), while applying orthogonalization techniques to isolate the CME and preserve valid inference, which then allows us to model complex, high-dimensional data without overfitting or violating identification assumptions. The authors explain the partialling-out procedure, discuss cross-fitting to avoid overfitting, and go into more detail on how DML can be implemented using off-the-shelf ML tools. They also show how to plug DML into the same projection-and-smoothing framework introduced earlier, making it compatible with their general approach to CME estimation.</p><p>In Chapter 5 we get to work! After introducing the &#8220;full menu&#8221; of estimation strategies (linear models, kernel estimators, AIPW-LASSO, and DML) the authors turn to the obvious question: which method should I use? That is exactly what Chapter 5 addresses through a series of Monte Carlo simulation studies designed to compare the methods under different conditions. The authors simulate data from three different Data Generating Processes (DGPs) that vary in complexity: from linear to nonlinear covariate effects and finally to a high-dimensional, highly nonlinear scenario. Each simulation evaluates how well the different estimators recover the true CME, and how their performance changes with sample size, model complexity, and hyperparameter tuning. The first DGP is fairly simple: the CME is nonlinear, but covariates enter the outcome model linearly. The goal here is to assess how well each estimator captures a quadratic treatment effect when most of the rest of the model is well-behaved. They find that Kernel and AIPW-LASSO both perform well when the sample size is moderate (n &#8805; 1,000), that DML methods (e.g., DML-neural networks, DML-histogram gradient boosting ) need larger samples (n &#8805; 5,000) to reach similar accuracy, linear models are clearly misspecified here and exhibit high bias, and in terms of confidence intervals, kernel and AIPW-LASSO produce narrower bands with better coverage at smaller sample sizes compared to DML. In the second DGP, things get more complex. Now, other covariates (not just the moderator) enter the outcome model nonlinearly, through functions like exp(2Z + 2). This setup challenges estimators that assume or rely on linearity in covariates. They find that Kernel estimators break down in this setting, because they assume linearity in non-moderator covariates; AIPW-LASSO and DML methods handle this complexity much better, flexibly approximating nonlinearities through basis expansions (AIPW) or ML models (DML); even at n = 1,000, AIPW-LASSO produces stable and accurate estimates; DML improves substantially with n &#8805; 3,000, but struggles in small samples due to the variance introduced by flexible learning. In the final study they explore a more realistic and challenging DGP with four additional covariates, nonlinear interactions, and a highly nonlinear propensity score model. It also compares the effect of using default vs. tuned hyperparameters in DML models (neural nets, random forests, histogram gradient boosting). They find that tuning matters, especially for neural networks (NN), where cross-validation significantly improves CME estimates; histogram Gradient Boosting (HGB) performs well at larger sample sizes, but tuning does not help much beyond defaults; random forest (RF) does not improve much with tuning and generally underperforms compared to NN and HGB; and that computational cost increases substantially with tuning, raising the classic trade-off: better fit vs. longer runtime. The final take-aways are:</p><ol><li><p>When the CME is nonlinear but covariates are linear, kernel and AIPW-LASSO work well even in small to moderate samples.</p></li><li><p>When covariates are also nonlinear, kernel breaks down, and AIPW-LASSO or DML are needed.</p></li><li><p>AIPW-LASSO is the best all-rounder (yey!), especially when sample sizes are limited and you want flexibility without the computational burden of full DML.</p></li><li><p>DML is powerful, but only if you have lots of data (n &#8805; 5,000, usually not the case for most observational studies, but macro people will be happy) and can afford to tune models carefully.</p></li><li><p>Tuning DML learners (especially neural nets) improves accuracy, but not always dramatically, and it takes time.</p></li></ol><p><em>Why is this guide important?</em></p><p>Understanding how treatment effects vary across individuals or contexts is at the ~kernel~ of applied social science research. Whether you are studying the impact of a job training program, a public health intervention, or a policy reform, knowing for whom and under what conditions the intervention works is often more important than knowing whether it works on average.</p><p>This guide is a goldmine! Their proposed framework gives us the tools to avoid common analytical pitfalls, such as extrapolating treatment effects to regions without common support or imposing rigid linear assumptions when the true relationship may be more complex. By presenting a progression from linear models to semiparametric and fully machine-learning-based approaches, their guide gives us a toolbox that can be tailored to our data's structure and complexity.  It is not just a case of estimating effects, it is about trusting them (by properly constructing valid pointwise and uniform confidence intervals, handling multiple comparisons, and diagnosing violations of overlap or model misspecification). The authors show how recent advances in DML can be applied in causal settings without sacrificing interpretability or statistical rigor, something that most economists are worried about when they hear &#8220;machine learning&#8221;. This helps make cutting-edge methods usable by applied researchers, not just methodologists.</p><p><em>Who should care?</em></p><p>Pretty much everybody (even the macroeconomists)! </p><ul><li><p>Applied researchers studying heterogeneous treatment effects across social science fields like political science, economics, sociology, education, and public policy.</p></li><li><p>Methodologists working at the intersection of causal inference and ML, especially the ones interested in effect heterogeneity and flexible estimation strategies.</p></li><li><p>Policymakers and practitioners who need to understand whether, how, and for whom an intervention works, beyond just headline average effects (they might not need to know the details tho).</p></li><li><p>Data scientists working with complex or observational data who want to go beyond basic treatment-control comparisons and explore nuanced causal relationships.</p></li></ul><p><em>Do they have code? </em></p><p>Yes! They even have a package (we love packages): the <code>interflex</code><a href="https://yiqingxu.org/packages/interflex/RGuide.html"> for R</a>.</p><p>With <code>interflex</code>, you can:</p><ul><li><p>Implement various estimators (linear, kernel, AIPW-LASSO, DML) with straightforward function calls.</p></li><li><p>Generate diagnostic plots to check overlap, functional form assumptions, and model misspecification.</p></li><li><p>Visualize CMEs with appropriate CIs.</p></li><li><p>Conduct sensitivity analyses to test the robustness of findings.</p></li><li><p>And even export plots and summaries for inclusion in papers or presentations.</p></li></ul><p>In sum, I had a great time going through this guide. If you have ever sat through a seminar, got grilled on nonlinearities, or wondered whether your interaction term was doing anything meaningful, this one is for you. Whether you are team kernel, curious about LASSO, or finally ready to get into DML, this guide gives you a roadmap grounded in strong theory and practical recommendations. And best of all, it comes with code. So yes, CMEs may sound intimidating at first, but by the end of this guide, they are just a way of answering the question we care about most: who benefits, and how?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.diddigest.xyz/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading DiD Digest! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>