A Field Guide to Missing Data (Part 1)
What economists do when the data they need aren’t there.
(This is part 1 of 2 because it was getting way too long, even by my standards. I will format both articles in LaTeX and add the link to the .pdf for download in part 2).
Empirical work often starts with an object we can’t observe. Sometimes that object is a variable or an observation. Sometimes it is a counterfactual, a comparable definition, a price, a target population, an assignment mechanism, a joint distribution, or a reproducible computational record. I use missing data in this broad sense. There’s a reason I think we should stop and think a bit deeper about these “absences”: they require different responses. An economist might report a set instead of a point, stress an estimate under a departure from an assumption, reconstruct an unavailable object, combine information across sources, generate a reference distribution, change the target question, regularise a sparse model, or improve the reporting and audit trail. These are different inferential activities, and because they are different, in nature and in substance, they don’t become interchangeable when a results table places them under one “robustness” heading.
So to get more out of the imperfect situation we find ourselves in, the first question we should ask is: what exactly is unavailable? This requires a lot of training, actually, and the next ones aren’t easier either. After we’ve thought about what is unavailable, one way “out” is to ask: what side information is available? Which assumption links it to the missing object? What can the proposed exercise establish under that assumption? What remains unknown afterwards? And has the target estimand (the quantity we’re trying to estimate) changed along the way?
With all of this in mind, I tried my best to come up with a field guide organised around 22 recurring problems. Several entries below contain exercises that do different jobs, and I keep them separate whenever the choice between them changes what the analysis can claim: a bound, a reconstruction and a stress test answer to different things, even when they address the same problem.
Here’s a list of the entries we will be covering today:
P01. A relevant confounder is unobserved
P02. The variable you want is represented by a proxy
P03. Some ordinary observations or items are missing
P04. Outcomes disappear through attrition or selection
P05. A continuous variable is measured with error
P06. Either treatment or outcome status is misclassified
P07. A price or quality adjustment is unavailable
P08. The counterfactual is unavailable
P09. Inputs or computational objects are inaccessible
P10. Only aggregates are observed
Let’s start:
P01. A relevant confounder is unobserved
We shall make this easier to understand with a familiar setup: we want to know whether a job-training programme raises earnings, and people chose whether to enrol. One worry for us would be “motivation”: one would assume more motivated people are more likely to sign up and more likely to earn more anyway, and no .csv has a column called “motivation”. That’s the missing object here: the part of who-gets-treated that’s related to the outcome but that our controls don’t capture. Confounding has other responses, and a convincing instrument, a discontinuity or randomised variation can change the design itself (Angrist and Pischke 2010), but this entry is about what we can learn when the original observational comparison is kept in place.
Since we can’t observe motivation directly (and don’t have a “suitable” proxy), the best we can do is ask the “what if” question: how strong would the link between motivation, enrolment and earnings have to be before our estimate stopped meaning what we say it means? That’s what sensitivity calculations do. Robustness values and coefficient-stability exercises (Cinelli and Hazlett 2020; Oster 2019) put numbers on that question. Oster’s version asks it in a more specific way: suppose the unobservables (e.g. the “motivation”) are related to enrolment in the same way as the observables we do observe in the data, like education and work history. If controlling for the things we can see barely moves the estimate while those controls explain a good share of the outcome, it would take an implausibly strong hidden factor to overturn it; if the estimate jumps around as we add controls, a hidden factor could plausibly finish the job. Both halves of that condition are needed since stable coefficients alone can be misleading: controls that carry little information also leave the estimate unmoved, without telling us anything about the confounder (Pei, Pischke and Schwandt 2019). To run this we have to declare two inputs: δ1, which says how strong the hidden selection is allowed to be relative to the observed selection, and R-max2, which caps how much of the earnings variation a complete model could ever explain. The calculation won’t choose those values for us, and what it returns is a threshold that stresses the estimate, not an estimate of the missing confounding itself.
Rosenbaum sensitivity analysis asks a similar question from the assignment side, and it’s an easy one to over-read. Take two people who look identical on every variable we observe. Its Γ parameter3 says how much more likely one of them could be to enrol because of the hidden factor: Γ of 2 means one could face twice the odds of treatment. The procedure then reports whether our conclusion would survive that much hidden imbalance. So what it bounds is the inference under that specified hidden-bias model rather than a causal parameter in general (Imbens 2003).
A negative control borrows a habit from epidemiology: find an outcome the programme couldn’t possibly affect, and check whether we “find” an effect there anyway. Earnings in the years before the programme existed are a natural candidate. If our design shows the training “raising” earnings that were paid out before anyone was trained, we haven’t discovered time travel; we’ve diagnosed a bias pattern, most likely the motivated people pulling ahead in any period (Lipsitch, Tchetgen Tchetgen and Cohen 2010). The catch is that the diagnosis is only as good as our claim that the control outcome really is untouchable by the treatment. And going further, actually using negative controls to identify the causal effect, is a different exercise with its own requirements: proxy conditions, conditional independence, and rank or completeness conditions (Miao, Geng and Tchetgen Tchetgen 2018).
This family also has two more members worth talking about. When economic knowledge supports a sign or relative-magnitude restriction on the omitted bias, the output can be an identified set rather than a threshold (Chalak 2019; Krauth 2016). And for ratio measures, epidemiology has what they call the E-value: it asks how strongly an unmeasured confounder would have to be associated with both treatment and outcome to explain away the observed association, which is one more threshold, and not evidence that no such confounder exists (VanderWeele and Ding 2017).
Whichever of these we run, none of them observes motivation for us, and none of them proves exchangeability, meaning that enrolled and non-enrolled people are comparable once we condition on the controls. They tell us how fragile our conclusion is to the confounder we’re worried about, which is valuable, but different from removing the worry.
P02. The variable you want is represented by a proxy
This time the missing object is the variable itself. Staying with our training study: suppose earnings are self-reported, and what we actually want is what employers paid. People round, forget, and sometimes flatter themselves, so the column called “earnings” in our dataset is really a proxy for earnings. The same situation shows up everywhere once we look for it: test scores standing in for learning, self-reported health standing in for health, patents standing in for innovation, etc.
The standard way out is a validation sample: a dataset where we observe both the proxy and a better benchmark for the same people - like a subset of our survey linked to tax records. A replicate measure or a reliability study can recover parts of the noise, while a linked benchmark can also reveal systematic bias, so these tools give us different kinds of side information rather than substitutes for one another4. With one of these in hand, we can estimate how the proxy relates to the target, which is our model of the error process5, and then use that estimated relationship to correct the analysis in the main sample.
This move rests on an assumption we should say out loud: it needs both the measurement process and the relationship we estimated to transport from the validation setting to ours, meaning they hold in our sample too (Bound, Brown and Mathiowetz 2001). If people misreport differently in the validation study than in our survey (because the years, the interview and/or the people themselves differ), then the correction we imported was calibrated for a different problem, and the apparent fix has just moved the problem from one dataset to the other. This is not ideal. One practical detail follows: for a regression correction, the validation records need to hold the proxy, the benchmark, the outcome and the relevant covariates together, and if the benchmark itself is imperfect, its error has to be modelled too, with the uncertainty from the calibration step carried into the final interval (Bound, Brown and Mathiowetz 2001).
We should also be clear about what the correction returns when it does work: a regression parameter under the assumed error model. We can recover, let’s say, the average effect of training on true earnings, and still have no idea what any particular person truly earned. And none of this tells us whether the proxy measures the concept we care about in the first place: a validation sample can tell us how noisy our earnings measure is, but it can’t tell us whether earnings were the right thing to measure.
P03. Some ordinary observations or items are missing
Back to our training survey: some people answered everything except the earnings question, and a few interviews never happened at all. The missing object here is ordinary, a value that should have been sitting in our analysis sample, and the response to it is where most of us first met the phrase “missing data”.
There are three standard routes here: imputation, weighting and observed-data likelihood, and the first two work in opposite directions. Multiple imputation fills each “empty space”, several times, with plausible values drawn from a model of the data6. Inverse-probability weighting doesn’t fill anything: it makes the people we did observe count for more so each respondent also stands in for the similar people who didn’t respond. The third route, observed-data maximum likelihood, neither fills nor reweights but writes down a likelihood for the data we saw and integrates over the values we didn’t, and it inherits the same ignorability and modelling assumptions as the other two (Little and Rubin 2019). (The simplest response of all - dropping the incomplete cases - is transparent but it pays in precision and can easily change the target to the kind of person who has complete records.) All three routes can support valid inference, but only under an assumption called missing at random, or MAR7, and the fine print on MAR is where the argument lives: it’s always relative to the information we observe. Given what we see about a person, their age, education and work history, the chance that their earnings answer is missing can’t depend on the earnings themselves. Whether that holds can’t be checked from the observed data alone (Rubin 1976; Robins, Rotnitzky and Zhao 1994); adding more variables to the model can make it more plausible, but never verified.
There’s also a “belt-and-braces” version, augmented inverse-probability weighting, which combines a model of the outcome with the model of who responds and stays consistent when either of the two is correctly specified (Bang and Robins 2005). This property, double robustness8, is a genuine insurance against picking the wrong model, but the insurance operates inside the maintained MAR assumption: it gives us two routes through model specification, and it doesn’t repair data that are missing not at random (when the chance of a space depends on the value hiding in it), nor a positivity problem (when some kinds of people never respond at all, so nobody is available to stand in for them).
Imputation carries one more requirement that sounds bureaucratic and turns out to be important: the imputation model has to be compatible with the analysis we are planning to run9. If our analysis studies an interaction (e.g., training effects that differ by gender, and the imputation model doesn’t include it), we’ve quietly built the absence of that interaction into the completed data before the analysis even starts. With chained equations - which is the popular variable-by-variable default - keeping the two compatible may need a substantive-model-compatible imputation scheme rather than the out-of-the-box settings (Bartlett et al. 2015).
Diagnostics have limits here too. We can spot completed values that look implausible or weights that “explode”, which diagnoses instability or thin support, though it doesn’t tell us which model or assumption caused it, and no diagnostic computed on the observed data verifies the missingness assumption itself. The assumption stays an argument, which is why the choice of variables in the imputation or response model is a substantive decision and belongs in the paper’s text rather than in a replication file.
P04. Outcomes disappear through attrition or selection
Our training study has one more trick to play on us. At the follow-up interview, some people have left the study, and wages exist only for people who are working. This looks like P03, but it’s a different thing because now the treatment itself is “editing” the sample: training changes who finds a job, and having a job determines whose wage we get to see. If we naively compare the wages of trained workers with the wages of untrained workers, we’re mixing two things: the effect of training on wages AND the effect of training on who shows up in the wage column at all.
There are four responses here, and they answer different questions. The first borrows from P03: if dropping out is ignorable given everything we observed along the way, inverse-probability-of-censoring weights let the people who stayed stand in for the people who left, with the same MAR and positivity requirements as before, and no help against dropout that depends on the missing wages themselves (Robins, Rotnitzky and Zhao 1994). The second is to give up the point estimate and report a range. Worst-case bounds fill the spaces with the most extreme plausible values in both directions (Horowitz and Manski 1998), which I think is honest and often uncomfortably wide. Lee trimming narrows the range under one extra assumption (monotone selection), meaning training can move people into employment but never out of it. It also needs the treatment itself to be as good as random, or exogenous given the covariates we use, so it doesn’t fix the self-selection into training we worried about in P01. The narrower bound comes at a price in interpretation: it describes the effect for the always-observed group, the people who would be working with or without training10, and not for everyone originally assigned (Lee 2009). Fittingly for our example, Lee’s original application was a job-training programme, the US Job Corps.
The third response is to reconstruct the missing wages with a selection model. A Heckman-style model writes down who ends up employed and what they earn as one joint system with related errors, and then uses the system to correct the wage equation (Heckman 1979). For this to rest on more than an assumption about the shape of the error distribution, we need an exclusion restriction11: some variable that shifts whether a person works but has no direct business in their wage. Without a convincing one, the identification leans heavily on distributional form, which is a polite way of saying the estimate can move a lot when we change assumptions we can’t test.
The fourth response is to stress the missingness assumption directly. Pattern-mixture and tipping-point exercises12 say: suppose the people we lost differ from the people we kept by some amount, and let’s see how large that amount has to be before our conclusion changes (Scharfstein, Rotnitzky and Robins 1999). If it takes a difference that is implausibly large on a scale we can defend, the result is less fragile under that sensitivity model; if a small plausible difference is enough, we’ve learned our conclusion leans heavily on the missing-outcome assumption (aka, it’s held up by hope only).
So: a reweighting, a set, a reconstruction and a stress test, four different answers to “what happened to the people we can’t see?”. And we shouldn’t expect the data to referee between them because model fit can’t do it: every missing-not-at-random model has a missing-at-random counterpart that fits the observed data exactly as well (Molenberghs et al. 2008). The choice between stories is an argument about behaviour, and not a statistic.
P05. A continuous variable is measured with error
Let’s change what we’re asking of the training study. Suppose we now want to know how earnings gains relate to the hours of training a person actually attended, and hours are self-reported. You see where I’m going with this, right? Nobody remembers precisely; people round to whole weeks and guess. The missing object is the error-free value - or really - the regression relationship we would run if we had it.
Most of us were taught one fact about this situation: noise in a regressor pulls the estimated coefficient towards zero, so our estimate is at least conservative. That fact is real but much narrower than its (big) reputation13. It holds for classical error, meaning noise unrelated to the true value and to everything else, in a simple linear regression with one mismeasured variable. Add more regressors, let the errors correlate with the truth (remember from P02 that real survey errors do exactly that), make the model nonlinear, or put the error in the outcome instead, and even the direction of the bias can change (Hausman 2001). “Our estimate is biased towards zero, so the true effect is bigger” is a sentence that needs its assumptions checked before it leaves the seminar room.
The tools for repairing this all work by learning something about the error process, and each one assumes something different about it. A validation sample, as in P02, estimates the error process directly. Repeated measures use a second noisy report of the same variable, and one clean use is as an instrument: the second measure is correlated with the true hours but, if its error is independent of the first measure’s error and of the disturbance in our earnings equation, not with the parts that cause the trouble14 (Griliches and Hausman 1986). Regression calibration replaces the noisy value with our best estimate of the true one given everything observed, and SIMEX runs the regression several times with extra noise deliberately added, watches how the coefficient degrades, and projects that path backwards to the no-noise case15 (Cook and Stefanski 1994). Likelihood methods write the error process into the model and estimate everything jointly16.
What all of these can deliver (when their assumptions hold!) is a corrected population parameter: the average relationship between true hours and earnings. What they don’t deliver is each person’s true hours. The correction works at the level of the regression, and that’s usually what we need, but it’s worth saying out loud, because a corrected coefficient sitting in a clean table makes it easy to believe the underlying variable was fixed too, and it wasn’t.
P06. Either treatment or outcome status is misclassified
Sometimes the mismeasured variable isn’t a number but a yes-or-no, and that changes the problem enough to deserve its own entry. In our training study: some people say they attended when they didn’t, some attended and don’t say so, and on the outcome side “employed” versus “not employed” has its own wrong labels. False positives and false negatives behave differently from the continuous noise of P05 because a binary variable can only be wrong in two directions, and the mix of the two directions is what drives the damage.
The first question is what we’re actually trying to recover, because there are three different targets hiding here: the true prevalence (what share of people really attended training), a regression coefficient (how attendance relates to earnings), or a causal effect. Each needs its own correction, and a fix for one is not automatically a fix for the next. A corrected prevalence doesn’t by itself correct a regression coefficient, and a corrected coefficient isn’t automatically a causal effect (Hausman, Abrevaya and Scott-Morton 1998). This chain is easy to skip in practice: a paper corrects the share and quietly treats the regression as corrected too.
What we can do depends on what we know about the error rates, which readers from epidemiology will recognise as sensitivity and specificity17. If we know them, or can estimate them from a validation subsample, a correction can be built for our specific target and error model, and Aigner (1973) shows what that full knowledge buys us in his one-sided model for a binary regressor. If we only have partial information, like a plausible range for the error rates rather than their values, the defensible output is a bound (Bollinger 1996): a range of coefficients consistent with what we know instead of a corrected point. Between the two sits a third option: when the error rates are uncertain rather than known, probabilistic sensitivity analysis feeds defensible distributions for sensitivity and specificity through the correction and reports the resulting distribution of corrected results, though it still doesn’t learn those distributions from our own data (Fox, Lash and Greenland 2005).
The assumptions doing the work are about the error rates themselves: which ones we know, whether the validation setting is comparable to ours (the transport question from P02 again), and above all whether the rates vary with treatment, outcome or covariates. When they do vary, the misclassification is called differential18, and it needs its own model, because the convenient folk theorem that misclassification biases us towards zero belongs to the non-differential case, and not always even there.
And when there’s no gold standard at all (no validation sample, no error rates from anywhere), latent-class methods19 can seem to rescue us: give the model several imperfect measures of the same status and let it estimate the error rates and the truth together. The rescue is real but runs on assumptions we should lay out, usually that the measures err independently given the true status and that the classes mean the same thing across populations. The output can look impressively precise, and that precision can be fragile to exactly those assumptions.
P07. A price or quality adjustment is unavailable
Let’s forget our training study example for a second because this problem sits in a different corner of economics. Suppose we want to compare the cost of living across regions (spatial variation), or track inflation over time (time variation), and the goods people buy differ in quality (they’re heterogeneous): rice in one market is not the rice in another, this year’s laptop is not last year’s laptop. The missing object may be a comparable price, a quality adjustment, or the weights an index needs. What we observe instead are transactions for goods that are never quite the same twice.
Household surveys tempt us with a shortcut called the unit value: divide what the household spent on rice by the kilos it bought, and call that the price of rice. The trouble is that a unit value mixes three things: the price the household faced, the quality it chose, and the composition of its purchases. One could say that richer households buy better rice, so their unit values are higher even at identical prices, and if we treat unit values as prices we build the household’s own choices into the “price” we then use to explain those choices (Deaton 1988). Unit values can be validated and adjusted into something useful, but they don’t start out as prices. Deaton’s own method is a reconstruction in this article’s sense: it needs geographically clustered households, a model of quality choice and a measurement-error correction, so dividing expenditure by quantity is where the exercise starts rather than what it is.
For goods that change quality over time, hedonic methods20 price the characteristics instead of the good: regress observed prices on attributes (for a laptop, memory, speed, screen), and use the fitted relationship to ask what the new model would have cost with the old characteristics (Rosen 1974). The comparison this gives us depends on the characteristics we observed and the functional form we chose, which is why official price indices don’t stop at a regression: the statistical manuals spell out explicit rules for matching items over time, weighting them, handling a product’s replacement and linking new items in (Intersecretariat Working Group on Price Statistics 2020). The rules don’t remove the judgement, but they put it where we can inspect it. Two named index families do much of this work: matched-model indexes compare identical items over time, and multilateral methods pool comparisons across several periods, which also reduces the chain drift we will meet below, though both still depend on the matching rule, the product universe and how entry and exit are handled (De Haan and van der Grient 2011).
There’s also a difference between calculating a missing price and inferring one, and it decides how much the number can be trusted. Where a fully specified accounting identity pins the number down, total expenditure, quantities and the other components all observed and compatible, we can calculate the residual and carry the components’ uncertainty through to it. A price backed out of a demand curve, a production function or any other behavioural relationship is a different object: a model-based scenario whose value changes with the model. There is no general “residual price” correction. And the diagnostics here have the usual limit. Chain-drift checks21 can tell us a chained index has “wandered” in a way that pure price change can’t explain, but they can’t tell us whether the items being linked were economically comparable in the first place.
P08. The counterfactual is unavailable
Readers of this newsletter are more than familiar with this problem :) This is THE central absence in causal inference, the one every other previous entry has been circling around. Let’s take our example a level up: a state introduces a training subsidy and we want to know what happened to earnings because of it. The missing object is what earnings in that state would have been without the subsidy. No dataset ever contains that column so everything in this entry is about constructing a stand-in for it and then interrogating (a lot) the construction.
Design, diagnostics, falsification and sensitivity analysis constantly get blurred under the word “robustness”, and it helps to keep the four apart. Our chosen design constructs the comparison: these units, these periods, this assumption about why the comparison is fair. The diagnostics check parts of that design, while falsification tests look for effects in places where the design says there should be none, and sensitivity analysis varies an identifying restriction to report which conclusions remain. The four are related, and they are not the same task, so a paper that runs one of them has not automatically done the others. For our subsidy question, the two designs we would most likely use are synthetic control and difference-in-differences22.
With synthetic control, we build the counterfactual ourselves: the untreated states get combined into a weighted mix, and the weights are chosen so that the mix tracks the treated state’s earnings path before the subsidy (Abadie, Diamond and Hainmueller 2010). The comparison is only as good as its conditions, good pre-treatment fit and a donor pool of comparable, untreated states. The usual check is a placebo run23: pretend each donor state passed the policy and see whether our real state’s post-policy gap stands out among the pretend ones. That comparison produces a rank, and a rank is not a p-value by itself (Abadie 2021); it earns a probability interpretation only with an argument about how treatment was assigned. There’s also an in-time placebo: move the intervention to a date before the real one and check whether the same procedure finds an effect where there shouldn’t be one, and here a check that finds one counts against the design, while a clean one stays just a diagnostic (Abadie 2021).
With difference-in-differences, the counterfactual comes from assumptions rather than weights, and two of them work together. They are a) parallel trends, which in our example would be “without the subsidy, the treated state’s earnings would have moved like the comparison states’ did”, and b) no anticipation, in our example “earnings didn’t start responding before the subsidy took effect because people saw it coming”. (Anticipation and treatment timing get their own entry in upcoming P19; here we stay with parallel trends, which is where - one can say - the diagnostic action is - wink). The standard diagnostic is a pre-trend check, and it’s a bit weaker than its popularity suggests24: the test can have little power against the violations that would hurt us most, so passing it doesn’t validate the assumption (Roth 2022). The newer response is to stop testing and start stressing: the Honest DiD approach25 asks what we can conclude if post-treatment violations of parallel trends are allowed to be, say, no larger than the worst pre-treatment one, and reports an interval or a set for the effect under that restriction (Rambachan and Roth 2023). The restriction, and not the method’s reassuring name, is what does the identifying work.
P09. Inputs or computational objects are inaccessible
This entry is about a different kind of disappearance, and the example works best told forward. Ten years from now, someone questions our training-subsidy paper. The tax microdata we used were confidential and our access agreement has expired, the person who wrote the estimation code has left, and what survives is the published paper with its tables (this hits way too close to home). The missing objects are the inputs of the analysis itself: microdata, code, the fitted model, the intermediate files that turned one into the other. Every other entry in this guide worried about data on the world; this one is about the data trail of the research.
The first thing to keep in mind is what question each exercise answers here, because the vocabulary gets used loosely26. Reproduction asks whether the same data and the same instructions regenerate the published number (Dewald, Thursby and Anderson 1986; Chang and Li 2022). That is a question about the record, and it’s separate from whether the design identifies anything, from whether the finding is true about the world, and from whether a new study with new data would find it again. A number can reproduce and still be wrong, and a paper whose files were lost can still have been right. We could never know.
Some questions need more than the published record to answer. If we want to know whether one state, or a handful of unusual people, drove our earnings estimate, we need influence or leave-one-out calculations27, and those run on the objects themselves: unit-level contributions, the fitted model, score contributions, or something equivalent that survived. When those objects are gone, they are gone. The defensible report in that situation is a documented limit on reproducibility, saying what was checked, what could not be, and why; an influence analysis can’t be recreated from objects that no longer exist, and writing one up as if it could is a reconstruction wearing a “diagnostic’s clothes”.
Before settling for that report, there are moves we haven’t used yet, and they’re the same verbs from the beginning of this article. The first one is to reconstruct the inputs from their sources instead of from the paper: the tax microdata still exist at the agency even though our agreement expired, agreements can be renewed, and if we documented our extraction criteria, a new extract plus a reimplementation of the described pipeline can regenerate the analysis. The result needs plain labelling, because it’s a new extract of possibly revised data (this is P22’s problem, arriving a bit early), so a difference from the published number can come from the data as well as from the code.
The published record itself can also be checked without touching any lost object, because a table has arithmetic obligations: the coefficients have to cohere with their standard errors and t-statistics, the sample sizes have to agree across tables, and the subgroups have to add up to their totals. This is diagnosis rather than regeneration: it can reveal contradictions in the surviving report, and it doesn’t regenerate the missing analysis (psychology built automated versions of these checks; search “statcheck” and “GRIM test”).
There is a narrow path between giving up (and crying on the floor) and pretending. On that path, a scenario range or a set of bounds can stand in for the lost analysis, but only after we define which missing inputs we consider compatible with the published record and justify that set: “the result holds under any reasonable version of the lost code” is an empty sentence until “reasonable” has content. Even then, the published totals on their own rarely settle it because a table of coefficients doesn’t reveal how each unavailable component, a weight here, a sample restriction there, moved the original estimate. When none of this recovers the analysis, the remaining move is to change the question: a replication with new data stops asking whether the published number was computed correctly and asks whether the claim is true, which is usually the thing we cared about in the first place.
P10. Only aggregates are observed
Suppose the individual data behind our training question were never available to us in the first place, and all we have is a table by state: the share of workers who went through training and the average earnings. The missing object is the joint distribution, which is the individual relationship living inside those margins. The temptation we face is to correlate the two columns and say that trained workers earn more, and this temptation has a name and a famous cautionary tale28: the correlation between two aggregates can differ in size - and even in sign - from the correlation between the same two things across people.
Once we stop trusting the two-column correlation, the correct move is to ask less of the aggregates we have, and ecological bounds are the exercise built for that: they describe the set of individual relationships consistent with the margins we observe, and the set is often wide, because the margins don’t determine what happens inside them (Cho and Manski 2008; open-access chapter here). Sadly, the bounds don’t recover the micro relationship, and more aggregate margins alone won’t either; a narrower set or a point estimate comes only from additional restrictions, and then the restrictions carry the conclusion. King’s ecological-inference model is the best-known point-estimation route, and it works exactly that way: it selects an answer from what the margins allow by adding distributional and homogeneity assumptions that the margins themselves can’t test (King 1997). The available diagnostics tell us more about the aggregation than about individuals: re-estimating across plausible geographies (like counties instead of states, if possible) shows how sensitive the aggregate association is to where the (sometimes literal) lines were drawn, a dependence with its own name (the modifiable areal unit problem, or MAUP), and comparing the aggregate analysis with an individual-level one, wherever even a small individual sample exists, can expose the contextual and ecological biases in the aggregate estimate29 (Greenland 2001).
Multilevel and within-between models sound like the modern solution, and they are, but only for the data they actually need: relevant lower-level observations, or repeated hierarchical information such as many periods per state. Given one aggregate cross-section, a multilevel model has nothing to work with underneath the aggregates; it can’t manufacture individual data from a table of forty-something rows.
Shift-share and component decompositions30 can also mislead here, because they look like they answer the individual question and they don’t. A decomposition tells us how the aggregate change adds up, how much came from composition shifting across groups and how much from changes within groups, and that’s a descriptive answer to a descriptive question. It isn’t a solution to ecological identification, because both pieces of the sum are themselves aggregates.
The useful discipline is to state which margins we observe, which associations are missing, and what extra structure would be needed to connect the two. Papers that do this read differently, because the reader can see where the aggregates stop and the assumptions start.
That’s all for today. We will be back next week. Any comments/suggestions/corrections please send me an email or leave them below.
Fable-5 helped with the literature review and gpt-5.6-sol audited the text (twice).
δ is the coefficient of proportionality: a number that says how strong selection on unobservables is allowed to be relative to selection on observables. If δ = 1, the factors we can’t see (like motivation) are related to programme take-up just as strongly as the factors we included in the regression (like education and work history); δ = 2 would mean twice as strongly, and δ = 0 takes us back to assuming no hidden selection at all. The idea comes from Altonji, Elder and Taber (2005), who were studying Catholic schools and argued that if the variables we observe are effectively a random subset of all the factors that drive selection, then how much the observables shift our estimate is informative about how much the unobservables could. Oster (2019) turned this into the calculation applied papers now report, where the usual output is the δ that would drive the estimate to zero: if it would take hidden selection stronger than everything we measured combined (δ > 1), many authors read the result as safer, though that reading still depends on how good our observables were in the first place. A search tip: the useful terms are “proportional selection”, “coefficient stability”, and the command names, psacalc in Stata and robomit in R.
R-max is the R-squared of a hypothetical regression we can never run: outcome on treatment plus every relevant control, the observed ones and the unobserved ones together. It answers the question “how much of this outcome could we ever explain if we had all the right variables?”, and the sensitivity calculation needs it because the room left between our actual R-squared and R-max is exactly the room where a hidden confounder could live. It can’t be lower than the R-squared we actually got, and for most outcomes it shouldn’t be 1 either, since part of the variation in something like earnings is noise, luck and measurement error that no variable will ever explain. Oster (2019) suggests a default of 1.3 times the observed R-squared, calibrated so that the method would validate the majority of a sample of results from randomised experiments, but that default is a starting point for an argument, and the right move is to justify a value for your own outcome (or show a range) rather than copy the 1.3. When searching, write it out as “Rmax” together with “Oster”, or look for “maximum R-squared” and the “1.3 rule” for the discussions around the default.
Γ is Rosenbaum’s sensitivity parameter, and the easiest way to read it is through a matched pair. Take two people who look identical on every covariate we observe. If assignment were as good as random given those covariates, both would have the same odds of ending up treated, and Γ measures how far from that ideal we’re willing to imagine: Γ = 1 means no hidden bias at all, while Γ = 2 allows the hidden factor to double the odds of enrolment for one of them relative to the other (odds, and not probability, is the technically right word here). The procedure then asks how our p-value would look in the worst case allowed by each value of Γ, raising Γ until the worst-case p-value crosses our significance level, and that value of Γ is what gets reported: a study that is “insensitive up to Γ = 1.8” keeps its conclusion as long as no hidden factor comes close to doubling the odds of treatment. The method comes from Rosenbaum (1987) and is developed in his book Observational Studies (2002, second edition), whose “Sensitivity to Hidden Bias” chapter is the standard introduction. Two reading habits are worth keeping: the bound applies to the inference (the p-value or CI) under this specific model of hidden bias, and a large Γ doesn’t certify the design, it just says the rejection survives that much imagined imbalance. When searching, the terms that work are “Rosenbaum bounds”, and the implementations are rbounds in R and Stata plus sensitivitymv and sensitivitymw in R for matched sets.
A validation sample, a linked benchmark, a replicate measure and a reliability study sound interchangeable, but they observe different things, and it’s worth keeping them apart. A validation sample observes the proxy and a better benchmark for the same people, so it can estimate both the noise and any systematic gap between them. A linked benchmark is the administrative version of the same idea, like survey answers linked to tax or payroll records. A replicate measure is the same instrument applied twice, asking the same person their earnings in two interviews: comparing the answers tells us how noisy the instrument is, but if both answers share the same bias (everyone rounding their salary up), the replicate can’t see it, because the bias cancels in the comparison. A reliability study is what statistical agencies run when they re-interview a subsample and publish the resulting reliability ratio - which is the share of the measure’s variance that is signal rather than noise. The practical rule: replicates and reliability studies can tell us about noise, while only a benchmark can tell us about bias.
The textbook error model is classical measurement error: the gap between the proxy and the truth is random noise, unrelated to the true value and to everything else in the analysis. It’s the assumption behind the familiar claim that mismeasurement pulls estimates toward zero (more on when that claim holds in P05). The reason to be suspicious of it is that when economists actually checked, real survey errors turned out not to be classical. Bound and Krueger (1991) compared what people told the Current Population Survey with their Social Security payroll records and found errors that were mean-reverting: high earners tend to underreport and low earners to overreport, so the error is correlated with the truth, which is exactly what the classical model rules out. Bound, Brown, Duncan and Rodgers (1994) found the same pattern in the PSID Validation Study, a survey run on workers of a firm that shared its payroll records. This is why the transport assumption in the main text is doing real work: a correction built on the wrong error model can move an estimate in the wrong direction, not just by the wrong amount. When searching, the useful terms are “errors-in-variables”, “classical measurement error”, “mean-reverting measurement error”, “reliability ratio” and “PSID Validation Study”, and the survey covering all of it Bound, Brown and Mathiowetz (2001).
Why fill each space several times instead of only once? Because filling once with our best guess pretends we knew the value all along, and the analysis then reports standard errors as if there had been no space, which overstates our certainty. Rubin’s solution was to draw several plausible values from a model that reflects how uncertain we are, produce several completed datasets, run the analysis on each, and combine the results with a set of formulas known as Rubin’s rules, so the disagreement across the completed datasets flows into the final standard errors. The foundations are in Rubin (1976) and his 1987 book Multiple Imputation for Nonresponse in Surveys. The implementations most people use are mice in R and mi impute in Stata, and those are also the best search terms together with “multiple imputation” and “Rubin’s rules”.
The taxonomy comes from Rubin (1976), and the names are famously confusing, so here they are in plain words. MCAR, or missing completely at random: the spaces are unrelated to anything, as when a page of questionnaires is lost in the office. MAR, or missing at random: given the variables we observe, the chance of a space doesn’t depend on the missing value itself, so younger respondents may skip the earnings question more often, but among people with the same observed characteristics, skipping is unrelated to earnings. MNAR, or missing not at random: the space depends on the value hiding in it even after we condition on everything observed, as when high earners “prefer not to say”. The names mislead because MAR doesn’t mean “missing randomly”; it means “random enough once we condition on what we see”. Search terms that work: “missing at random”, “ignorability”, and for the MNAR case “nonignorable nonresponse”.
The definition is worth quoting, because the original authors say it more cleanly than most textbooks I’ve checked: “an estimator is doubly robust (DR) or doubly protected if it remains consistent when either a model for the missingness mechanism or a model for the distribution of the complete data is correctly specified. Because of the frequency and near inevitability of model misspecification, double robustness is a highly desirable property” (Bang and Robins 2005). The augmented weighting estimator itself goes back to Robins, Rotnitzky and Zhao (1994). Positivity, the other requirement in the main text, is the condition that every kind of person in the analysis has some chance of being observed; weighting can stretch respondents to cover similar non-respondents, but it can’t conjure information about groups with no respondents at all. Search terms: “AIPW”, “doubly robust estimation”, “positivity” or “overlap”.
The technical term for this compatibility is congeniality, coined by Meng (1994), and the concern is practical rather than pedantic. Imputation is often done by one person (or one agency) and analysis by another, and if the imputer’s model is poorer than the analyst’s (let’s say it leaves out interactions, nonlinearities or subgroup structure that the analysis cares about), the completed datasets systematically understate exactly the patterns being studied. The working rule that falls out: make the imputation model at least as rich as the analysis model, and when in doubt, richer. Search “congeniality multiple imputation” or “uncongenial imputation” for discussions.
The formal name for this group is a principal stratum, and the idea is worth thirty seconds because it explains who the estimate is really about. Sort people by how their employment would respond to the programme: some would work whether or not they were trained (the always-employed), some would work only if trained, some wouldn’t work either way. We can’t tell person by person who belongs where since we only ever see each person under one treatment, but the monotone-selection assumption rules out the awkward group that training pushes out of work, and that’s what lets the trimming produce a bound. The bound then belongs to the always-employed stratum: a policy claim like “training raised wages” silently becomes “training raised wages for the kind of person who’d be employed regardless”. The framework is from Frangakis and Rubin (2002). Search terms that work: “principal stratification”, “Lee bounds”, and the Stata command leebounds.
An exclusion restriction is a variable that moves the selection (whether we observe the outcome) without moving the outcome itself. In the classic female labour-supply applications, the number of young children was argued to shift whether a woman worked but not her offered wage; whether the exclusion is believable is always the fight, and it should be. What happens without one is specific: the Heckman model is still estimable, because the assumed joint normality of the errors gives the correction term (the inverse Mills ratio) a particular nonlinear shape, and that shape alone identifies the model. That means the estimate is resting on a distributional assumption we can’t check, and modest changes to it can move the results a lot. The original is Heckman (1979), the Stata command is heckman, and the useful search terms are “Heckman selection model”, “exclusion restriction” and “inverse Mills ratio”.
A pattern-mixture model splits the sample by response pattern, the people we kept and the people we lost, models the ones we can see, and then states openly how the lost ones are assumed to differ, usually through an offset: “dropouts earn 10% less than observationally similar completers”. That factorisation by response pattern comes from Little (1993). The tipping-point exercise repeats the analysis over a range of offsets and reports the value where the conclusion changes (here “changes” means whatever we declared as the criterion: usually that the effect is no longer statistically distinguishable from zero, sometimes that the estimate changes sign). Instead of defending one assumption, we show the whole range and discuss whether that value is realistic. Scharfstein, Rotnitzky and Robins (1999) is the standard reference for the related selection-model version, where a selection-bias parameter plays the role the offset plays here. One naming warning: this literature (especially in clinical trials where regulators routinely ask for these analyses) calls the offset a “delta adjustment”, and that delta has nothing to do with Oster's δ from P01, an unlucky collision of Greek letters. Maybe we should start using Chinese characters instead? :) Search terms: “pattern-mixture model”, “tipping point analysis”, “delta adjustment”, “nonignorable dropout”.
The mechanics are worth going over once. A regression coefficient is a covariance divided by a variance. Classical noise in the regressor leaves the covariance with the outcome alone but inflates the regressor’s variance, so the ratio shrinks: that’s attenuation, and the shrinkage factor is the reliability ratio from P02’s footnote, the share of the measure’s variance that is signal. The textbook result stops there, and the trouble starts where the textbook stops hehe. With several regressors, the bias from one mismeasured variable spreads to the coefficients of the correctly measured ones, in directions that depend on how everything correlates. If the error is mean-reverting rather than classical, the covariance is contaminated too, and the estimate can be biased away from zero. In nonlinear models there is no single shrinkage factor at all. And error in the outcome behaves differently again: if classical, it inflates standard errors without biasing the slope, but if it correlates with the truth (as in the earnings validation studies), it biases the slope as well. Search terms: “attenuation bias”, “errors-in-variables”, “reliability ratio”, and for the multi-regressor case is “measurement error contamination”.
We have two noisy reports of the same true variable, say self-reported hours and hours from the training provider’s records, which are both imperfect. Using one as an instrument for the other works because a valid instrument needs two things: it must be correlated with the variable we’re instrumenting (which may hold if the shared signal is strong enough to give a relevant first stage), and it must be unrelated to the errors that cause the bias (which holds only if the two errors are independent of each other and of the earnings equation’s disturbance). That second condition is the one to argue about: if both reports come from the same person’s memory, their errors probably share a component, and the instrument inherits the “disease” it was meant to cure. Griliches and Hausman (1986) developed this logic for panel data, where differencing can amplify measurement-error bias, and their paper is also the standard warning about that amplification. Search terms: “instrumental variables measurement error”, “repeated measures instrument”, “errors in variables panel data”.
I’d say these two have the most personality of the correction methods. Regression calibration asks: given the noisy report and everything else we observe, what’s our best estimate of the person’s true hours? It substitutes that estimate (a conditional expectation) into the regression and then adjusts the standard errors to account for the substitution, which makes it simple and popular in epidemiology. SIMEX, simulation-extrapolation, is the counterintuitive one: since we can’t remove noise, just add more of it, in several known amounts, and rerun the regression each time. The coefficient degrades along a path as noise grows, and fitting that path lets us project it back to the point of zero noise (Cook and Stefanski 1994). The projection step is the assumption: we never observe the zero-noise end, so the answer depends on the fitted shape of the path. Implementations: simex in R and Stata’s simex command; search “regression calibration” and “simulation extrapolation SIMEX”.
In the simple classical linear case, we also have two older tools: a known reliability ratio can correct the naive slope directly, and running the regression in both directions (earnings on hours, then hours on earnings, inverted) brackets a positive slope between the two estimates, though neither trick survives nonclassical error or richer models (Hausman 2001). One caution on regression calibration: it’s exact for a linear mean model under its assumptions and generally an approximation in nonlinear ones (Carroll et al. 2006).
Sensitivity is the probability that a true yes is recorded as a yes (a real trainee reports attending); specificity is the probability that a true no is recorded as a no. Together they pin down the misclassification process for a binary variable, and they make the prevalence correction almost arithmetic: the observed rate of yeses mixes true yeses caught by the test with false positives, so observed rate = sensitivity x true rate + (1 − specificity) x (1 − true rate), and knowing the two error rates lets us solve for the true rate. The same logic extends (with more algebra) to regression coefficients. The reason economists cite Aigner (1973) for the known-rates correction and Bollinger (1996) for the partial-information bounds is that they answered the two different situations: Aigner shows what full knowledge of the error rates buys, Bollinger shows what can still be said when we only know something. Search terms: “sensitivity specificity”, “misclassification correction”, “misclassified binary regressor”.
Misclassification is non-differential when the error rates are the same for everyone regardless of their treatment, outcome or characteristics, and differential when they aren’t. The distinction decides how much trouble we’re in. A classic example from epidemiology is recall bias in case-control studies: people who got sick search their memory for exposures harder than healthy controls do, so the exposure’s error rate depends on the outcome, and the bias can go in any direction. In our training example, people who benefited from the programme may be more likely to remember and report attending, which is the same disease. Even the comfortable non-differential case has fine print: the towards-zero result is for a binary exposure in simple settings, and Hausman, Abrevaya and Scott-Morton (1998) show that misclassification in a discrete outcome biases coefficients in probit and logit models even when the errors are non-differential, which is the nonlinear-model warning from P05 wearing a different coat. Search terms: “differential misclassification”, “non-differential misclassification bias”, “recall bias”.
The best-known design is due to Hui and Walter: with two imperfect tests applied in two populations whose true prevalences differ, the error rates of both tests and both prevalences become estimable at once, with no gold standard anywhere. It feels like magic, and the magic has a price list. The measures must err independently given the true status, which is exactly what stops holding when both measures come from the same interview or the same administrative pipeline, and the true classes must be the same objects across the populations, and the design adds two requirements of its own: the two populations need different true prevalences, and each test’s sensitivity and specificity must stay constant across them (Hui and Walter 1980). When any of these is wrong, the model may still converge and print tight standard errors (Albert and Dodd 2004), which is why an apparently precise answer can be fragile: the precision is conditional on assumptions the data can’t check. Search terms: “latent class analysis without gold standard”, “Hui-Walter design”, “conditional dependence diagnostic tests”.
The hedonic idea is that a differentiated good is a bundle of characteristics, and the market implicitly prices the characteristics. Practically it’s a regression of price on attributes, and statistical agencies use exactly this to quality-adjust fast-changing goods like computers and housing. Rosen (1974) is the theory paper behind the practice, and it comes with a warning that often gets lost: the regression coefficients are equilibrium implicit prices, formed where buyers’ preferences meet sellers’ costs, and not demand parameters or willingness to pay. Two things limit any hedonic adjustment: the characteristics we didn’t record (a camera improves in ways no spec sheet captures) and the functional form we imposed. Search terms: “hedonic regression”, “hedonic quality adjustment”, “implicit prices”.
A chained index links many short comparisons: it compares this month with last month, then multiplies that by the comparison of last month with the month before, and so on. Chain drift is when this multiplication goes wrong in a specific way: prices and quantities come back exactly to where they started, but the index doesn’t. One well-documented cause is supermarket sales in scanner data (De Haan and van der Grient 2011). During a sale the price is low and households buy a lot, so the price cut enters the chain with a big weight. When the sale ends the price goes back up, but by then purchases have returned to normal, so the price increase enters with a smaller weight. The cut and the increase don’t cancel out, the index ends the round trip below where it began, and over many sales the gap keeps growing. A chain-drift check tells us this is happening, and the multilateral methods built to avoid it can reduce the drift under a stated product universe, window and formula. What they can’t answer is whether a discounted, end-of-line item is still the same good as its full-price predecessor: that is a judgement about comparability, and - sadly - no statistic settles it. Search terms: “chain drift”, “scanner data price index”, “GEKS”.
As readers of this newsletter can remember, both designs have newer hybrids: augmented synthetic control brings in outcome modelling to patch residual pre-treatment imbalance (Ben-Michael, Feller and Rothstein 2021), and synthetic difference-in-differences combines unit and time weighting with a DiD adjustment (Arkhangelsky et al. 2021).
A step by step of the placebo logic: run the entire synthetic-control procedure once per donor state, each time pretending that donor passed the policy, and collect the resulting post-“policy” gaps. If our real state’s gap is the largest of twenty (or something) such runs, we can say its rank is 1 out of 20, and it’s tempting to call that p = 0.05, but we should be really careful because that number is a real randomisation p-value only if the policy was as good as randomly assigned among the states, and nothing in the procedure makes that true since states pass subsidies for n different reasons. Without that argument, the rank is still informative, our state stands out, but it’s a comparative statement instead of a significance level. Abadie (2021) is the survey treatment of this and of the feasibility conditions, and it’s the right first read before using the method. Search terms: “synthetic control method”, “placebo test synthetic control”, “in-space placebo”; implementations are synth (Stata, R) and tidysynth (R).
A pre-trend check regresses the outcome on time relative to treatment and asks whether the treated and comparison groups were already diverging before the policy. Roth (2022) documents two problems with it. The first one is low power: when the data are noisy, a violation large enough to bias our estimate badly can still pass the test, so a finding of “no significant pre-trend” often means we couldn’t detect the divergence rather than that there was none. The second problem comes from the pre-testing itself: if we only carry on with the analysis when the pre-test looks clean, then the estimates that end up published are a selected sample, and that selection introduces its own bias. None of this makes pre-trend plots useless since they still show what the two groups did before the policy; the trouble starts when we treat a passed test as a validated assumption. Search terms: “pre-trends test power”, “pretesting parallel trends”, and Prof Roth’s pretrends R package.
The Honest DiD idea is to replace “parallel trends holds” with “parallel trends can be violated, but only by this much”, and then to state the “this much” as a number. One version allows the post-treatment violation to be at most M times the largest pre-treatment violation, so M = 1 means the assumption can be broken after the policy by as much as it ever was before. Another version restricts how fast the violation can change from period to period (a smoothness restriction). For each choice of restriction, the method reports the range of effects consistent with the data, and the usual summary is the breakdown value: the M at which our conclusion changes, which in most applications means the M where the reported interval starts to include zero (Rambachan and Roth 2023). Reading a result then becomes an argument about whether violations of that size are plausible in our application, and that argument is on econ rather than on stats. The implementation is the HonestDiD package (R and Stata); search terms: “Honest DiD”, “Rambachan Roth”, “relative magnitudes restriction”, “breakdown value”.
Discussing the taxonomy is important here since the words get swapped quite frequently in seminars: reproduction (or computational reproducibility) means running the same data through the same code and getting the same number; replication means asking the question again with new data or a new design; and both are different from identification (does the design isolate the effect?) and from empirical validity (is the claim true about the world?). The reason economists cite Dewald, Thursby and Anderson (1986) is the story: the Journal of Money, Credit and Banking asked authors for their data and programs, and most published results could not be regenerated, often for silly reasons like lost files and undocumented steps. Chang and Li (2022) repeated the exercise decades later (sixty papers from thirteen journals) and their title answers the question: “often not”. The silly causes are the point of this entry: in those studies, results became irreproducible through ordinary decay, lost files, expiring agreements and departing coauthors, which doesn’t rule out misconduct elsewhere, but does show how far ordinary entropy gets on its own. Search terms: “computational reproducibility economics”, “replication in economics”, “data availability policy”.
An influence analysis asks how much each observation contributed to the final estimate, and a leave-one-out analysis answers it by brute force: drop one unit, re-estimate, repeat. Both need - at least - the per-unit material, the microdata or the fitted model’s unit-level pieces (in regression these are the score contributions, each observation’s push on the coefficient). This is why they can’t be run from a published table: a coefficient and a standard error are sums over units, and a sum doesn’t remember its parts. The practical lesson runs in both directions. Looking backward, if the objects are gone, the analysis is gone. Looking forward, archiving the fitted model and unit contributions alongside the code is what keeps this door open for our own papers. Search terms: “influence function”, “leave-one-out”, “dfbeta”, and for the recent automated version “finite-sample robustness metric”.
The tale is Robinson (1950), and it’s worth knowing in its original form. Across US states in 1930, the share of foreign-born residents correlated positively with English literacy rates: states with more immigrants had higher literacy. Across individuals, the correlation ran the other way since immigrants were on average less literate in English than the native-born. Both facts are true at once because immigrants had settled in states with high native literacy. Reading the state-level correlation as a statement about individuals became known as the ecological fallacy, and Robinson’s paper is the reason every methods course warns about it. Search terms: “ecological fallacy”, “ecological inference”, “cross-level inference”.
The aggregate association mixes up to three things, and Greenland (2001) is the careful accounting of them. There is the individual effect we wanted (does training raise your earnings?), a possible contextual effect (does living in a high-training state raise your earnings through the labour market, whatever you did yourself?), and confounding at the group level (high-training states differ in other ways). The contextual effect deserves respect rather than deletion because sometimes it’s the actual question since spillovers from a trained workforce are a real policy object; the fallacy here is in reading an aggregate coefficient as an individual one without saying which of the three ingredients it contains. Search terms: “contextual effects”, “ecological bias”, “multilevel analysis”.
A shift-share decomposition splits an aggregate change into parts using an accounting identity: the change in a state’s average earnings, for example, into the part from employment shifting between industries (the composition part) and the part from earnings changing within industries. The arithmetic is exact and the output is descriptive. One naming warning because economics reused the phrase: “shift-share” also names the Bartik instrument literature where industry composition is used as an identification strategy, and that is a different object with its own assumptions rather than an accounting identity. A reader searching should use “shift-share decomposition” for this footnote’s tool and “shift-share instrument” or “Bartik instrument” for the other one. Related decompositions worth knowing by name: “within-between decomposition” and “Oaxaca-Blinder”.


