The link to the .pdf is here, as promised.
Hello hello.
So in Part 1 we mostly covered the absences inside a single dataset: the unobserved confounder, proxies and measurement error, values that go missing or disappear through attrition, and the counterfactual itself. This time the problems get a bit more structural: trying to get to inference when the data are few or the instrument is weak, attempting to combine datasets that were never meant to be combined (at least initially), the world of official stats with its changing classifications, censoring, publication lags and revisions, and the design complications we kept promising, timing and anticipation, spillovers, and whether our variable measures a concept at all.
In today’s post we cover:
P11. Definitions or classifications change
P12. The sample is small or the event is rare
P13. There is little overlap or little identifying variation
P14. Records describing the same units can’t be linked exactly
P15. Variables sit in different datasets
P16. Values are censored, top-coded or altered for disclosure
P17. The official outcome arrives too late
P18. The study population differs from the target
P19. Treatment timing is uncertain or anticipated
P20. One unit’s treatment affects another unit
P21. The available measure doesn’t settle the concept
P22. Data are revised after first release
P11. Definitions or classifications change
Going back to our training study for this. Suppose we can now track earnings by occupation for twenty years and halfway through the occupation classification gets revised, e.g. “web developer” didn’t exist in the old system, “typist” is practically gone in the new one, and a category like “clerical worker” was split into others that now occupy different places. The missing object is comparability - over time or across systems. Nothing was mismeasured and nobody dropped out, but the ruler itself changed. Readers from other fields will recognise the situation because it’s the same one that affects trade data when product codes are revised and health data when ICD versions turn over: it’s the problem from P07, where this year’s laptop was not last year’s laptop, except now it’s the categories that are never quite the same twice (Heraclitus vibes).
We have three direct ways to connect the classifications: a crosswalk, a splice built from an overlap, and a backcast. Chain-linking is a related time-series operation: it joins adjacent price or volume comparisons, but it doesn’t by itself translate an old category into a new one. A crosswalk1 builds the bridge from concordance information, meaning documented links between the two systems (Pierce and Schott 2012). A splice builds it from an overlap period - a stretch where both systems ran side by side - so we can see directly how they line up. Backcasting applies an existing bridge backwards, rewriting the earlier data in the new system’s terms, so the clerical workers of 1985 get re-expressed as today’s categories. And chain linking replaces one long bridge with many short ones: it compares each year only with the next one, inside whichever system covers both, and multiplies those short comparisons into one long series (International Monetary Fund 2018). The shared assumption is the same for all the first three: a bridge is still meaningful only as far as the observations used to build it, and the employment shares we measured in the overlap years, applied to 1985, assume the occupation’s composition hadn’t changed by then, which is exactly the kind of thing that does change.
These exercises can produce a usable series, but no bridge can erase a substantive change in what a category means. If the new “software occupations” group covers different work than the old “computer operators” did, a perfectly smooth spliced series is describing a category whose meaning moved beyond it, and the smoothness is the disguise rather than the reassurance. What we then need to do is to disclose clearly what we did since we can’t expect the reader to audit a bridged series without knowing the many-to-many mappings, the break dates, the overlap years, and how entrants and exits were handled. What helps is to create a revision triangle2, which describes how a statistic’s estimates change across successive releases, with one row per reference period and one column per release date, so reading along a row shows how the estimate for that period changed as new information arrived. It’s a diagnostic for how the data arrive, but it has nothing to say about a classification break. Revisions are their own absence with their own exercises (which we will see in a second).
P12. The sample is small or the event is rare
Suppose our training programme was only a pilot: six sites, two hundred people, and the outcome we care most about - moving into self-employment - happened eleven times. Two separate problems come from this situation, and they need different tools to be handled. The first is that p-values and confidence intervals depend on large-sample approximations, and with six sites and eleven events, the reference distribution behind them can be a poor description of the uncertainty we face in reality. The second is that the model itself can become unstable or separated3, and with something like eleven events we are certainly in that territory.
For the first problem we have three standard solutions - randomisation inference, permutation tests, and the wild-cluster bootstrap4 - and each one builds its comparison world from a different source. Randomisation inference uses the assignment mechanism we are familiar with: if we know exactly how the six sites were chosen for treatment, we can then re-run that assignment thousands of times (on our software of choice), and under the null of no effect for anyone, we can fill in what each site would have shown so the p-value counts how unusual our real allocation looks among all the allocations that could have happened (Athey and Imbens 2017). A permutation test shuffles labels instead, and it needs an exchangeability argument - a reason to believe the shuffled versions are statistically interchangeable with the real one under the null (Chung and Romano 2013). The wild-cluster bootstrap resamples the data in a particular way, and the quality of what it produces depends on the details we have: how much leverage each cluster has, how treatment is spread across clusters, whether the null is imposed when resampling, the choice of weights, and the statistic being bootstrapped. Their shared commonality is that each of these builds its reference world from one source, a known mechanism, an invariance argument, or a resampling construction, and the result is only as good as that source. There’s also a fourth, albeit less glamorous, family that adjusts the cluster-robust variance and the degrees of freedom directly for having few clusters (Cameron and Miller 2015). These are corrections to an approximation rather than new reference worlds, and how well they do still depends on leverage, on balance, and on how many clusters actually got treated.
For the second problem the tools work on the estimate rather than on the reference world - Firth correction, rare-events corrections, and Bayesian partial pooling. Firth correction is a penalised version of maximum likelihood that reduces small-sample bias and keeps the estimate finite when the data are separated (Firth 1993)5, whereas rare-events corrections deal with the small-sample bias of logistic models when the outcome is uncommon (King and Zeng 2001). The right correction depends on why the events are rare and on how the sample was drawn, and recovering the population probabilities can require extra information like the population prevalence or the sampling fractions. Bayesian partial pooling shrinks noisy site-level estimates towards each other, trading them against a hierarchical model and a prior (Gelman 2006). All three estimate or regularise, and that solves the second problem only. The first one is still standing: before we can report a p-value or a confidence interval, we still need a reference distribution to compare the estimate against, so the tools in this paragraph and the tools in the previous one do different jobs, and a small-sample study often needs one from each.
Even the best reference world is still small here. With six sites split three against three, there are only twenty ways treatment could have been assigned, so the smallest p-value randomisation inference can produce is 1/20 = 0.05, and that’s for a one-sided test where our real allocation is the most extreme of them all. With the usual two-sided statistic the floor is generally 2/20 = 0.10 because every allocation has a mirror allocation with the opposite sign. A placebo-style rank deals with the same arithmetic: standing out among a handful of alternatives is a comparison, and it becomes a formal p-value only with an assignment or invariance argument behind it. None of the tools in this entry escapes this, and none of them repairs a weak design, poor measurement, or a target population the sample doesn’t support. What they do is make the most of six sites while being clear about how much that is - and with six sites, saying which population our estimate speaks for is often the hardest sentence in a study.
P13. There is little overlap or little identifying variation
Back to the observational version of our training study where people chose whether to enrol. Suppose the choosing was “lopsided”: almost everyone with a university degree signed up, and almost nobody over sixty did. For those two groups the data contain no comparison because there are no untreated graduates and no treated sixty-somethings to compare against, and no amount of regression adjustment would conjure that. The missing object is the support for the original comparison: treated and untreated people who resemble each other across the whole range of characteristics our question covers.
We have two standard responses - trimming and overlap weighting6 - and they share one consequence: both change who the estimate is about. Trimming is about dropping the observations where a comparison is unsupported, e.g. the graduates and the over-sixties, and then estimating the effect for the people who remain (Crump et al. 2009). Overlap weighting keeps everyone but counts people more the more “contested” their enrolment was so the estimate concentrates on the middle ground where similar people ended up on both sides (Li, Morgan and Zaslavsky 2018). Neither recovers the effect for the full original population, and the recommendation that follows is the same as in P11: explicitly write/say what the new estimand is and who was excluded since “the effect of training” and “the effect of training for people whose enrolment could have gone either way” are very different sentences.
Little identifying variation is the same absence in IV form. Suppose we use distance to the nearest training centre as an instrument for enrolment, and distance turns out to barely move anyone’s decision: the first stage (the relationship between the instrument and the treatment) is then fragile, and everything built on it inherits the fragility. The diagnostics, like the familiar first-stage F statistic, describe that fragility without fixing it (Staiger and Stock 1997). Weak-instrument-robust procedures7, like the Anderson-Rubin test, change how the uncertainty is computed so that the reported intervals stay valid however weak the first stage is, and that is all they change: they don’t strengthen the instrument, and they don’t validate it (Andrews, Stock and Sun 2019). Their guarantees also come from a specific setting, and the classical exactness results don’t carry over on their own to clustered or serially dependent data, so the version we run has to match the dependence structure we have.
The two halves of this entry are the same absence in two forms: the data hold little information where our question needs it. The defensible responses either shrink the question to where the information is, and are explicit about it, or keep the question and report uncertainty in a way that respects how little there is. Nothing here manufactures variation, and an estimate that ignores that will look precise for the wrong reason.
P14. Records describing the same units can’t be linked exactly
Our training survey has two hundred people, and the tax records that would tell us their true earnings sit in another file with no shared identifier. Linking means deciding which tax record belongs to which respondent using name, birth date and address, and all three can disagree for the same person: names get typed differently, people move houses, and some change their name at marriage. The missing object is unit identity across files - the fact of the matter about who is who.
Probabilistic linkage8 turns this decision into a calculation. Each candidate pair of records gets compared field by field, the pattern of agreements and disagreements goes into a comparison model, and the model returns a match probability9 - a pair that agrees on birth date and address but differs on one letter of the surname is probably the same person, and a pair that agrees only on being called “J. Silva” probably isn’t (Fellegi and Sunter 1969). For this to work we need three things: 1. linking fields that carry information, meaning fields that distinguish people rather than describe half the country, 2. a comparison model we can defend, and 3. a manageable set of candidate pairs, because comparing two hundred respondents against every tax record in the country is neither feasible nor necessary.
The linkage will make two kinds of error, and they do different levels of damage. A false link attaches someone else’s tax record to our respondent, so their “true earnings” are now a stranger’s. A missed match loses the record that was there to be found, and the people it loses are not random since the hardest people to link are the ones who - in our example - have moved, divorced or changed names. What this does to our estimate depends on the full structure of those errors, and it need not be a simple attenuation (Lahiri and Larsen 2005), so the defensible practice is to carry the linkage uncertainty into the analysis, either by keeping the match probabilities and letting them propagate, or by modelling the linkage and the analysis as one joint problem. And when the files can only match one-to-one (each person has at most one tax record), the candidate links stop being independent choices, so the model should enforce that structure and the joint versions can then average the analysis over several complete, mutually compatible linkages.
The diagnostics assess the process, and the process is not the population. Clerical review (humans checking a sample of borderline pairs) and quality metrics like match rates tell us how the linkage went for the records we tried to link. A high match score doesn’t establish that the linked sample represents the target population: if the people we failed to link are the movers and the name-changers, our linked sample is simply a sample of stable lives, and that is a selection problem in a data-management costume.
P15. Variables sit in different datasets
This is the sibling of P14 with one twist now: the files describe different people. Our training survey has enrolment and hours for two hundred respondents, and earnings are in a labour-force survey of other people entirely, so there is no pair of records to link, no matter how clever we get. The missing object is the cross-file association: nobody, anywhere, observed training hours and earnings together, and that association is exactly what our question needs.
One way out is to give every respondent a statistical twin. We look in the other file for someone with the same age, education and from the same region, we copy that person’s earnings over, and we end up with a completed dataset where everyone has both variables. This exercise is called statistical matching10, and the donated association comes from an assumption, usually conditional independence: given the shared variables, hours and earnings are assumed unrelated, so the shared variables are assumed to carry all the connection there is (Rässler 2004). The completed file can’t test this because the association it contains is the one the assumption put there. A completed synthetic microdata file is evidence that an algorithm worked, but we can’t say that the missing association was recovered.
The other way out skips completed files: instead of building records that have both variables, we take summary pieces from each file (like the relationship between distance and enrolment from one survey and between distance and earnings from the other) and combine the pieces into the estimate we want. The exercises in this family are moment combination, two-sample IV, and two-sample 2SLS11, and they are separate exercises: each uses different pieces from each file, and each needs its own conditions (Ridder and Moffitt 2007; Inoue and Solon 2010). The shared requirements are a) compatible variable definitions, b) the same population relationships holding in both samples, and c) variance conditions that differ by estimator - and these sit on top of the ordinary IV requirements12. When those conditions ask more than we can defend, we can report a set instead - a range of answers consistent with what each file separately shows, instead of a fused point estimate carrying an unverified association inside it.
Whatever we choose, the useful discipline is to name - again - the bridge: say whether the information is being combined across files, across instruments, or across populations because each link has its own weak point. For a statistical match it is the conditional-independence assumption; for a two-sample IV it is two samples that don’t come from the same population; and neither problem is revealed automatically in the output, which looks complete either way.
P16. Values are censored, top-coded or altered for disclosure
One more visit to our training survey :) This time the data arrive incomplete on purpose. Earnings above 150,000 show up as “150,000+”, some respondents answered with a bracket (“between 30,000 and 40,000”) - not their fault! - and in the public state tables, cells with too few people are left blank. The missing object depends on the release mechanism (the rule the publisher applied before letting the data out), and the mechanisms are a family with true differences13: a top-code replaces values above a limit with the limit itself, a bracket gives a range instead of a number, censoring cuts off at a known point, swapping exchanges values between similar records, rounding coarsens them, noise addition perturbs them, and suppression removes cells outright. Lots of problems here because each of these actions throws away different information, therefore… yeah, you guessed, each asks for a different exercise.
For the two oldest mechanisms, we can name the exercises. When the data give us ranges, we can fit a regression that uses the known endpoints of each range instead of a single value, and that is known as interval regression (Stewart 1983). When values pile up at a known limit, we can write a model for the outcome we would have seen without the limit and estimate it from the pile-up and the rest, and that is the Tobit model (Tobin 1958)14. Both assume the mechanism is exactly what the model says, a known endpoint or a known limit, and neither is a general answer to confidentiality treatment: a swapped or noise-injected value is not a censored one, and a model built for limits has nothing to say about it.
The first exercise is… actually reading (yourself or your agents - helpful if it’s in a language you don’t speak), and it comes before any model: the agency documents its release rules, and the documentation says which mechanism affected which variable in which year (U.S. Census Bureau 2023). For top-coded earnings, the workable repair is a tail model (a distribution fitted to the upper part of the data and used to fill in above the limit) and it earns its keep only when its family, its threshold and its fit are defended for the application at hand; multiple imputation above the top-code, for example, produces answers that are conditional on the tail model chosen, and different defensible tails give different inequality estimates (Jenkins et al. 2011). You will meet more general recipes in the wild (e.g. fit a Pareto tail above any top-code, or recover suppressed cells from the published margins with optimisation) and this is the part where I should stop short of recommending them: the case-by-case defence is the method, and the generic version I can give would simply skip it.
In this entry, the mechanism decides. Treating a top-code as if it were censoring, or noise-injected values as if they were true ones, mixes two release rules into one analysis, and the output comes without a warning. The exercise that protects us costs nothing but time: match what we do to what the documentation says was done.
P17. The official outcome arrives too late
Our training subsidy is up for renewal and the state has to decide this quarter whether it worked, but the official earnings series for this quarter will only be published next year. The missing object is the current value at the decision date: the number exists in the world, in payrolls already paid, and the statistical system hasn’t finished counting it. Meanwhile the timely data pile up (e.g. unemployment claims, job postings, payroll-processor records) and the exercise of estimating the official number from them before it’s released is called nowcasting (Giannone, Reichlin and Small 2008).
The standard tools to solve this problem are bridge equations, factor models and mixed-frequency models15, and they have one thing in common in order to make them work: the information set must be dated. A bridge equation links the quarterly target to monthly indicators through a regression, a factor model compresses many indicators into a few common movements and reads the target off them, and mixed-frequency models let monthly and quarterly data live in one equation without forcing them to the same calendar (Bańbura et al. 2013). The shared “discipline” concerns what the model is allowed to see: for any date we nowcast, only the data releases that existed on that date can enter. Later revisions and indicators published afterwards were not available to the decision we’re trying to inform.
We need to be careful because that discipline is also where we can get the evaluation wrong. If we test our nowcasting model against history using today’s data, we commit a crime in two types of hindsight: revised values of the indicators - which are cleaner than what was on the screen at the time - and series that hadn’t been released yet on the dates we’re pretending to nowcast. A fair historical test recreates what was known at each date by using real-time vintages, meaning the data exactly as they stood on each past day16 (Croushore and Stark 2001). A model that looks sharp against final data and ordinary against real-time data isn’t being treated unfairly by the second test; the second test is actually the true one.
Nowcasting also has a neighbour it gets confused with. A nowcast estimates a current value nobody has measured yet, while backcasting (see previous entries) builds a historically comparable series for periods already measured under other rules; a model can do one well and the other badly because the two jobs need different things from it. And forecast accuracy is a statement about prediction only: an indicator can predict earnings movements while measuring nothing we would interpret as the effect of training, so a model that nowcasts well tells us what the number will be, and says nothing on why.
P18. The study population differs from the target
Our pilot went well and the state now wants to roll the programme out to everyone. The missing object is the target-population effect: the effect in the population the decision is about, which is not the population we studied. The two can differ because effects differ across people (e.g. training may help younger workers more than older ones) and our six pilot sites had their own mix of ages, industries and cities, while the state has another. A variable like age here is called an effect moderator, meaning a characteristic that changes the size of the effect, and the pilot’s mix of moderators is baked into its headline number.
We have two17 ways to move from the study to the target - standardisation and weighting - and they have four requirements in common. Standardisation computes the effect within subgroups (e.g. for younger and older workers separately) and then re-averages those effects using the target population’s subgroup shares instead of the study’s (Cole and Stuart 2010). Weighting goes the other way around: it reweights the study so it resembles the target, giving more weight to the kinds of people the study under-represents, with weights built from the odds of being in the study rather than the target; these are inverse odds of sampling weights18 (Westreich et al. 2017). Both need the study effect itself to be right (internal validity), the variables to mean the same thing in both datasets, the effect to carry across given the shared variables (conditional effect exchangeability), and every kind of person in the target to have some counterpart in the study (positivity, the same requirement we met with the missing respondents). The last one is the hard limit: if the pilot enrolled nobody over sixty, no weighting scheme produces the over-sixty effect because there is nothing to reweight.
The moderators we didn’t measure get a different exercise. If we suspect the pilot sites differed from the state in something unrecorded (how motivated the local caseworkers were, for example) we can vary that unmeasured moderator in a sensitivity analysis and report how far the target effect could move. The calculation of Andrews and Oster (2019)19 belongs here: it uses how the observed characteristics relate to participation to calibrate how much the unobserved ones could move the effect, under its own selection model and a small-selection approximation, and it is a stress test under those terms, not an automatic bound, correction, or transport estimator.
A study that sits inside its target - our pilot counties are part of the state - and a target that is a different population altogether - another state asking whether our results apply there - are different problems with different names: generalisability for the first, transport for the second, and a study should say which one it is doing. And the weights themselves invite one very specific question. Two weighting schemes can produce similar-looking columns of numbers while resting on different assumptions, one about how the study was sampled and one about how effects carry across populations, so the thing to ask of any weight is which assumption it encodes because the numbers look the same either way.
P19. Treatment timing is uncertain or anticipated
When we set up the subsidy example, we assumed nobody responded before the policy took effect. This entry is for when that assumption is the problem. The subsidy legally starts in January, but the bill passed in October, the local news covered it, and training providers began advertising in November, so by the time January arrives, some of the response has already happened. The missing object is the behavioural onset of treatment: the date behaviour started responding which is not required to equal the legal date, and which no dataset column announces.
When the institutional record leaves several defensible onset dates (the bill’s passage, the news coverage, the legal start, etc), the exercise is to run the analysis under each coding and report the set. Each coding changes real things: who counts as treated from when (the cohorts), which periods serve as untreated comparison, and the estimand itself, since “the effect of the subsidy from January” and “the effect of the subsidy from October” are different questions. This exercise is a stress test tailored to our design, and there is no general correction behind it: not knowing when treatment began is not a statistical problem with a standard fix, but rather it is a piece of missing institutional knowledge, and the codings are how we show what it costs us.
Does all of this ring a bell? Anticipation has its own tools when the early response is real and we are not talking about a dating error. People and firms respond before formal adoption whenever they can see the policy coming and acting early pays, and a period we labelled “untreated” then already contains part of the effect20 (Malani and Reif 2015). The group-time estimators most of us already use can account for this to some extent: we declare an “anticipation horizon” (a stated number of periods before adoption in which effects are allowed to exist), and the estimator moves its baseline back accordingly21 (Callaway and Sant’Anna 2021). We “declare” the horizon and we try our best to defend it with institutional facts, such as when the bill passed or when coverage began, rather than estimated from the data since the data can’t distinguish early effects from differential trends on their own. We should do that.
There’s one more common thing we can do: we drop the periods right around adoption and compare “cleanly before” with “cleanly after”, which is called a donut. A donut can be informative, and it has two costs we should account for: it discards the transition data, and it can change the target contrast because an effect measured between “well before” and “well after” does not say much about the adoption period itself. A donut is a disclosed design choice, and should not be considered a universal repair for ambiguous timing. Our duty, present throughout this entry, is the same one: date the treatment with institutional evidence, show the codings, and treat the calendar as part of the design rather than as a setting nobody chose.
P20. One unit’s treatment affects another unit
Our subsidy trains thousands of workers, and trained workers compete for jobs with untrained ones, so an untrained worker in a treated state is not untouched by the programme: some of what happens to their wages happens because their competitors got trained. Across the border the story continues. Firms in the comparison state recruit workers from ours, so the “untreated” comparison may contain a “diluted dose” of the treatment. In every entry so far we assumed that each person’s outcome depends only on their own treatment. In this entry, such assumption holds no more. The missing object is the exposure mapping, which is the rule that says who can affect whom, over what distance, and through which channel. The data don’t record that rule, so the mapping has to come from an argument about how the market works.
The designs that handle interference all work by restricting it to somewhere we can see22. One restriction allows effects to travel within groups but not between them (within a local labour market but not across markets). This is called partial interference23 (Hudgens and Halloran 2008). A more general version states outright which other units count for my outcome and how (such as workers in my commuting zone, weighted by how much we compete), and this is an exposure mapping in the formal sense. Once we declare one, contrasts such as “treated versus untreated, holding neighbours’ exposure fixed” become estimable under the design (Aronow and Samii 2017). This mapping actually does a bit more than just enable estimation. It defines the estimand: the effect of the subsidy now means “the effect given who counts as exposed to whom, under this specific mapping”.
A common check re-runs the analysis with different reaches (spillovers allowed within 10 km, then 20, then 50). This kind of radius sensitivity is often called a ring analysis. What we get from it is sensitivity, meaning how much the estimates move as the assumed reach changes. What we don’t get is the true spillover radius. Each radius is a different mapping, each mapping may define a different estimand, so the rings compare answers to related questions instead of closing in on one answer (Sävje 2024). And when the mapping is misspecified, we end up with two errors at once24, each with its own name: a definitional error, because we estimated a contrast other than the one we intended, and an identification error, because even that contrast may be estimated with bias.
When the network or the geography is too weakly known to defend one mapping, the better report is the set of estimates across the defensible mappings, so the reader can see what the missing knowledge costs us. A single exposure model hides that cost. And the assumption we started this entry with should be written down in our studies like any other. No interference is an assumption like parallel trends or missing at random, it can be wrong in ways that change the answer, and it spent nineteen entries being assumed without anyone saying so - which is exactly how an assumption ends up treated as a harmless default.
P21. The available measure doesn’t settle the concept
Suppose the state asks whether our training programme improved workers’ economic security, and we answer them with earnings. Earnings is a real variable, well measured for once, but it is still not the concept: economic security has something to do with stability, benefits, and how easily a job is lost, and a higher wage on a precarious contract may not be more of it. The missing object here is a defensible operationalisation, meaning a way of turning the concept we care about into a variable we can measure, plus the argument that the two line up. Back in P02 the variable existed and was poorly measured. In this entry the measurement can be perfect and the question still remains because what’s in doubt is whether we (well, others) actually measured the right thing.
The main exercise in this entry is to check the measure against what theory expects from the concept. If we have prior evidence that earnings really tracks economic security, it should move with the things security moves with (savings, reported financial stress, job tenure) and it should not move one-for-one with things that are clearly not security (like a temporary overtime spike). This is called construct validation (Cronbach and Meehl 1955). A stronger version of it measures several concepts with several methods at once and checks the full pattern: different measures of the same concept should agree, and measures of different concepts should disagree enough to count as different. This is the multitrait-multimethod comparison25 (Campbell and Fiske 1959). Both exercises can show us where a measure behaves as the concept should. Neither can prove the indicator is the concept, because behaving as expected is evidence, and identity is not something evidence of this kind can establish.
Economics runs on examples of this gap. Patent counts are an indicator related to innovation, publication measures indicate research output or visibility, employment is one observable labour-market outcome, and nominal expenditure is a monetary input measure. These are illustrations of common practice, and none of them is a validated proxy relationship. Whether any of them measures the intended concept depends on the question and on domain evidence, which is an argument each study has to make for its own case rather than inherit from the literature.
When the validation exercise does not support the measure, we have two options, and the less obvious one is often better. We can a) hunt for a better measure of the same concept, or b) we can revise the concept, and ask a question the evidence can answer (not “did training improve economic security” but “did training raise earnings”, stated as exactly that26) (Adcock and Collier 2001; Bollen and Lennox 1991). A narrower question answered with a variable we understand beats a grand question answered with a variable we can only dream about.
P22. Data are revised after first release
This is the shortest entry in the guide because we’ve already met its tools. The revision triangle appeared with the classification breaks, and the vintages and the pseudo-real-time evaluation appeared with the nowcasts. What’s left is the rule that ties them all together. The missing object is the vintage appropriate to our question, and “appropriate” changes with the question. A real-time policy analysis should use the latest data available at the decision date while a historical measurement exercise can use first releases, later releases, or one fixed vintage, as long as it says which. The final vintage is not automatic ground truth for a real-time question because the decision-maker we’re studying never actually saw it (Croushore and Stark 2001).
The diagnostics here are the exercises we already know. Vintage panels and revision triangles show how the estimates changed as information arrived, and a pseudo-real-time evaluation reruns our whole procedure with each date’s information set (International Monetary Fund 2018). What these expose is revision risk, meaning how much our answer depends on which release we used. What they can’t do is decide which vintage is the right one for every estimand, because that choice belongs to the question, and the question is ours to state.
The eight verbs
After 22 problems, you might have noticed a pattern. The exercises we walked through are variations on a small set of responses, and the easiest way to see this is to write each response as a verb. We end up with eight of them. A single study usually does several at once, so treat the verbs as a way of reading applied work rather than boxes to sort papers into, and when a procedure combines them, we pull it apart and name the verb for each piece.
Bound it. If the data cannot give us one answer, report the range of answers they can support. The width of that range tells us something too: the wider it is, the more we would have to assume to get to a precise answer (Tamer 2010).
Stress it. Change an assumption and see when the conclusion changes. The useful result is not that the estimate moved, but how far we had to push the assumption before it did (Lu and White 2014).
Reconstruct it. Sometimes we can use other information to fill in what we cannot observe directly. The important question is then whether the link we used to fill the gap is credible. A completed dataset does not make the missing information observed.
Combine, or transport it. Sometimes the information exists, just somewhere else. We can bring together datasets or populations, but doing so requires a reason to believe that what connects them is comparable. Successfully merging two files is not evidence that this assumption holds.
Generate reference worlds. Sometimes what is missing is the comparison itself: what would our statistic look like under some alternative world? We can generate that comparison in different ways, but what we learn from it depends on how that alternative world was constructed.
Change the question. Sometimes the original question asks more than the evidence can answer. In that case, the best option may be to ask a narrower question and say clearly that the target has changed.
Estimate, or regularise it. Sometimes the problem is that there is too little information to estimate the model reliably. Regularisation can make the estimate more stable, but it does so by adding structure. That structure is part of the answer.
Report, or audit it. Sometimes the missing object is the analytical record itself. Better reporting, shared code and computational checks can tell us what was done and whether we can reproduce it (Christensen and Miguel 2018; Olken 2015). They cannot tell us whether the original research design was valid.
Writing things this way forces us to say what we actually learned. Different exercises produce different kinds of evidence, and they should not all end up in a table labelled “robustness”. They also come with a trade-off: if the data cannot answer the original question on their own, we either make an additional assumption, accept a less precise answer, or ask a different question.
This is where robustness checks can become ritualistic. A useful check starts with a specific problem and tells us something about it. A ritual check starts with a familiar method and works backwards. Some of this is probably driven by publication incentives and convention, but we should be careful about blaming referees for all of it.
There is good evidence that reasonable analytical choices can produce different results and that selective reporting affects what gets published (Steegen et al. 2016; Simonsohn, Simmons and Nelson 2020). We also know something about what editors and referees value and about selection during publication (Berk, Harvey and Hirshleifer 2017; Brodeur et al. 2023). What we do not know is how much of the robustness machinery in published papers is there because referees asked for it. That story is plausible, but the evidence does not get us that far.
The same point applies to reproducibility. Being able to rerun an analysis is quite useful, but it does not settle whether the design was good in the first place or whether the finding would survive new data (Chang and Li 2022). Reproducibility can close one gap while leaving the important one untouched.
So when something we need is missing, start there. Say what is missing and work out what the available evidence can still tell us. Sometimes the answer will be weaker or narrower than the one we wanted. That is fine. The standard robustness table can wait.
In the end, it is quite a simple rule: say what the exercise does, say what you had to assume to do it, and say what is still missing. A robustness check is useful when it changes what we know, not when it merely gives us another column for the table.
A concordance is the documented mapping between two classification systems, and the issue here is that it’s rarely one-to-one in the sense of one old category splitting into several new ones, and several old ones merging into one new one. Allocating a many-to-many mapping needs weights (something like employment, output or trade shares from the overlap period), and those weights are where the assumptions enter, since a share measured in 1997 is being asked to describe 1985. Pierce and Schott (2012) built the standard concordance between ten-digit US trade codes and industry classifications and their paper kind of lists everything that can go wrong. The same structure exists in other fields under other names: medicine crosswalks ICD-9 to ICD-10 with official equivalence mappings, and education does it between ISCED versions. Search terms: “concordance”, “crosswalk”, “classification change”, and for the trade version search “Pierce Schott concordance”.
A revision triangle is a table with one row per reference period (e.g., GDP in 2019Q4) and one column per release date, so reading along a row tells you how the estimate for that quarter changed as later releases arrived. It’s called a triangle because recent periods have had fewer releases so the table has a staircase edge. Agencies and real-time-data archives publish these, and they answer questions like “how large is a typical revision?” and “do first releases systematically under- or overstate growth?”. What a triangle can’t do is repair a classification break because a break is a change in what was being counted, and it shows up in the triangle as a discontinuity that no amount of averaging across releases eliminates. Search terms: “revision triangle”, “real-time data”, “data vintages”.
Separation is the situation where our predictors split the two outcome groups perfectly: you could draw a line with every self-employment case on one side and every non-case on the other, e.g. all eleven cases sitting in trained sites and none anywhere else, so the variable “trained site” sorts the yes’s from the no’s on its own. Ordinary maximum likelihood then has no finite best answer because making the coefficient larger always fits a little better, so the estimate runs off towards infinity, the standard errors explode, and some software stops with a warning while some reports enormous numbers. “Unstable” means that the estimate stays finite but swings considerably when one or two observations change. The fixes are in the main text (Firth correction, rare-events adjustments); the implementations are logistf in R, firthlogit in Stata, and relogit for the King–Zeng version. Search terms: “separation logistic regression”, “complete separation”, “quasi-separation”.
Cameron, Gelbach and Miller (2008) found substantial improvements in simulations with as few as five to ten clusters, though that’s no general guarantee because with treatment concentrated in one or two clusters, or badly unbalanced cluster sizes, the bootstrap can still perform poorly. The practitioner’s guide to all of it is Cameron and Miller (2015). One configuration is tough even for the bootstrap: few treated clusters, e.g. six sites of which only one or two got the programme, where no resampling scheme has much to work with and the randomisation-inference route tends to be the more defensible one. The implementation everyone uses is boottest in Stata (also available in R); search terms: “wild cluster bootstrap”, “few treated clusters”, “cluster robust inference”.
To be fully correct, Firth’s bias-reduction adjustment modifies the likelihood score (Firth 1993), and its use to obtain finite logistic estimates under separation was established by Heinze and Schemper (2002).
Both tools run on the propensity score, which is the estimated probability of enrolling given a person's observed characteristics. A score near 0 or 1 flags a person whose treatment status was close to decided in advance, and this is exactly where comparisons are unsupported. The practical trimming rule from Crump et al. (2009) is to keep observations with scores between roughly 0.1 and 0.9, and it’s a rule of thumb rather than a law. Overlap weights, from Li, Morgan and Zaslavsky (2018), weight each person by the probability of the treatment they didn’t get, so the weight peaks for people with scores near 0.5 and fades toward the extremes, and the resulting estimand even has its own name, the average treatment effect in the overlap population (ATO). Search terms: “common support”, “propensity score trimming”, “overlap weights”, “limited overlap”.
The famous F > 10 rule of thumb for the first stage comes from Staiger and Stock (1997), and it was a rough guide from the start; the modern survey treatment of what to do instead is Andrews, Stock and Sun (2019). The robust procedures work by inverting tests rather than trusting the usual standard errors: the Anderson-Rubin test asks, for each candidate effect size, whether the data reject it, and collects the survivors into a confidence set. One feature of these sets is worth knowing before you meet one: when the instrument is weak enough, the set can be wide or even unbounded, and that is the method being truthful about the data rather than the method misbehaving. Anderson-Rubin also has company: conditional likelihood ratio tests can say more in the settings they cover, and the first-stage F we all learned to check has to match the covariance structure of the data, so the usual threshold is not a certificate once heteroskedasticity, clustering or serial dependence enter the picture (Andrews, Stock and Sun 2019). The implementations are weakiv and ivreg2 in Stata and ivmodel in R. Search terms: “weak instruments”, “Anderson-Rubin test”, “conditional likelihood ratio test”, “first-stage F statistic”.
The Fellegi-Sunter framework runs on two probabilities per linking field: how often the field agrees when two records truly are the same person (high, but not 1, because of typos), and how often it agrees by coincidence when they aren’t (low for birth dates, higher for common names). Each field’s agreement or disagreement contributes evidence, the contributions add up to a score, and two thresholds turn scores into decisions: clear matches above, clear non-matches below, and a middle zone sent to clerical review. The “manageable candidate pairs” requirement has its own name: blocking. Rather than comparing everyone with everyone, we compare only within blocks that share something cheap and reliable, like a birth year, and accept that a wrong birth year now means a missed match. Implementations: fastLink in R, dtalink in Stata, and Splink in Python. Search terms: “probabilistic record linkage”, “Fellegi-Sunter”, “blocking record linkage”, “linkage error bias”.
Some later models turn that evidence into an actual match probability, but only by adding assumptions about the linkage process (Sadinle 2017).
The mechanics are usually a donor pool and a distance rule: the other file’s records are the donors, we measure similarity on the shared variables, and the closest donor’s value gets copied over (or several donors get averaged, or a model predicts the value from the shared variables - the flavours differ, the logic doesn’t). The name for what’s being assumed is the conditional independence assumption, and we can see how it works in practice in our example: if motivated people train more hours and earn more in ways age, education and region don’t capture, the matched file will understate the hours-earnings connection, and nothing in the file will show it. Without conditional independence, the data from the two files only bound the association rather than pin it down, which is why the fusion literature increasingly reports uncertainty about the match itself. Search terms: “statistical matching”, “data fusion”, “conditional independence assumption”, and for the bounds version the famous “Fréchet bounds”.
Two-sample IV runs an IV design with its two halves in different datasets: one file has the instrument and the treatment, so it can estimate the first stage, and the other has the instrument and the outcome, so it can estimate the reduced form, and the ratio of the two estimates the effect - under the usual IV assumptions - even though no single person has all three variables. In our example, one survey shows how distance to a training centre moves enrolment, the other shows how distance relates to earnings. The requirements are that both samples come from the same population and that the variables mean the same thing in both, which sounds ok and reasonable until two surveys define “earnings” differently. Drawing both samples from the same population is the cleanest case; there are methods that let the samples differ (Zhao et al. 2019), but they work by stating which relationships stay stable across the two, not by waving the difference away. Two-sample IV and two-sample 2SLS are not the same estimator: Inoue and Solon (2010) show they differ in finite samples and in efficiency, so you have to choose between them. Search terms: “two-sample IV”, “two-sample 2SLS”, “data combination econometrics”.
Splitting the moments across two files doesn’t excuse the instrument from its “day job”: it still has to move treatment, affect the outcome only through treatment, and be unrelated to the unobserved determinants of earnings.
The confidentiality mechanisms deserve their own definitions because the words look interchangeable and they aren’t, actually. Swapping exchanges values, like earnings, between records that look similar on other variables, so each record is real but some are wearing a neighbour’s number. Noise addition perturbs values with random error whose distribution the agency knows and we usually don’t. Suppression comes in two rounds: primary suppression blanks the sensitive cell, like the average earnings of the three trainees in a small county, and secondary suppression blanks other cells too since a blank cell inside published row and column totals could otherwise be recovered by subtraction. The umbrella term for all of it is statistical disclosure control, and agencies publish their general approach while keeping some parameters secret on purpose, which is a constraint our analysis has to live with rather than an oversight to complain about - we can still complain tho. Search terms: “statistical disclosure control”, “cell suppression”, “data swapping”, “top-coding”.
Interval regression and Tobit get confused because both involve limits, and the difference decides the assumptions we get. On one hand, interval regression knows each observation’s endpoints, so it needs a distributional assumption only to spread probability inside the brackets (kind of a mild request). On the other hand, the Tobit model assumes a full latent outcome (usually ~N) and its estimates lean on such normality harder than we would expect, so a heavy-tailed or skewed truth can move the results. If the question is about conditional quantiles rather than a latent mean, censored quantile regression is the standard option (Powell 1986), and it still needs the censoring rule and the quantile restriction to be right, while heavy censoring can leave some quantiles unidentified no matter what we do. There’s also a third “animal”: tail models for top-codes. Jenkins et al. (2011) fit a flexible distribution (the generalised beta of the second kind, GB2) to the upper range and impute above the limit, and their multiple-imputation machinery brings the tail uncertainty into the standard errors, which is what makes the exercise conditional-on-model in a way we can inspect. Implementations: intreg and tobit in Stata, survreg in R; search terms: “interval regression”, “Tobit model”, “top-coded earnings”, “GB2 distribution”.
The simplest one is a bridge equation: you aggregate the monthly indicators to quarterly frequency and regress the target on them, so the “bridge” crosses from what’s published monthly to what’s published quarterly. A dynamic factor model assumes many indicators move together because a few underlying forces drive them, extracts those forces, and nowcasts from the forces rather than the indicators, which is why it digests dozens of series without falling apart; the version in Giannone, Reichlin and Small (2008) also updates its nowcast release by release, so you can read how much each morning’s data changed the estimate, a decomposition the literature calls the “news”. Mixed-frequency regressions (search “MIDAS”) skip the aggregation step and let a quarterly variable sit on the left with monthly regressors on the right. Many of these models are written in state-space form and updated with a Kalman filter, which is handy when releases arrive on different calendars and the dataset has a ragged edge (Bańbura et al. 2013). Search terms: “nowcasting”, “bridge equation”, “dynamic factor model”, “nowcast news decomposition”.
A vintage is a snapshot of a data series as it stood on a given date, and vintage archives exist because researchers kept needing yesterday’s version of the truth: the Philadelphia Fed’s Real-Time Data Set, built from the work in Croushore and Stark (2001), and the St. Louis Fed’s ALFRED archive are the standard sources for US series. An evaluation that uses them is called a pseudo-real-time evaluation: pseudo because we run it now, real-time because every number the model sees is dated to what existed then. The evaluation has to close two separate gaps, and closing only one still overstates the model’s accuracy: the indicator values must be the unrevised ones that were on record at the time, and the list of series itself must match what had been published at all on that date. Search terms: “real-time vintage data”, “pseudo real-time evaluation”.
If we are being really honest, there’s a third that combines them: the doubly robust procedures. These stay consistent when either the outcome model or the participation model is right, though only inside the transport assumptions we already signed up for (Degtiar and Rose 2023).
Let’s think in numbers. Suppose training raises earnings by 800 for under-forties and by 200 for over-forties. Our pilot was 70% under-forty so its headline effect is 0.7 × 800 + 0.3 × 200 = 620. The state is 40% under-forty so the standardised target effect is 0.4 × 800 + 0.6 × 200 = 440. Same study, same subgroup effects, and a headline that drops by a third when the mix changes. This arithmetic we are doing is the entire point of the exercise, and it only needed the subgroup effects and the target’s shares. Inverse odds weighting reaches the same destination by another road: model each study person’s probability of being in the study rather than the target given their characteristics, weight by the inverse odds, and the reweighted study behaves like the target on everything the model included. How to choose? Go for standardisation when the moderators are few and the subgroup effects are estimable, and weighting when the characteristics are many. Search terms: “standardization”, “transportability”, “inverse odds of sampling weights”, “generalizability of trial results”.
The Andrews-Oster logic is: if the people and places that joined the study differ from the target in ways we can see, and those visible differences relate to the effect, then the visible selection gives us a yardstick for how much the invisible selection could move the result, under the assumption that the invisible part behaves like the visible part (the selection model) and that selection into the study is not too strong (the small-selection approximation). The answer moves with both assumptions, which is why the output reads as a stress test: it answers “how bad could unmeasured selection be if it resembles the measured kind”, and it doesn’t answer “what is the corrected target effect”. The distinction between generalisability and transport also decides the machinery: when the study is nested in the target, the target’s covariate distribution is known from the start; when the target is a separate population, its data enter as a second dataset with all of P15’s compatibility questions attached. Search terms: “Andrews Oster external validity”, “selection on observables and unobservables”, “transportability”.
Malani and Reif (2015) studied tort reform, where physicians responded while reforms were still working through legislatures, and their point cuts two ways. A “pre-trend” can be anticipation rather than a violation of comparability: the groups were comparable, and the treated one started responding early. And anticipation contaminates the baseline. If the pre-period already contains part of the effect, the before-after comparison subtracts that early response from the later one. When the anticipation runs in the same direction as the effect, the estimate usually comes out too small, but with other dynamics the bias can go either way (Malani and Reif 2015). The two readings ask for different responses, which is why the pre-trend discussion from P08’s diagnostics and this entry’s dating problem are neighbours rather than the same issue. Search terms: “anticipation effects”, “pre-trends as anticipation”, “announcement effects”.
In the Callaway and Sant'Anna (2021) framework, effects are estimated per cohort and period, and the comparison baseline is the cohort’s last untreated period. Declaring an anticipation horizon of one period tells the estimator that the last untreated period is one earlier than the adoption date, so the baseline moves back and the adoption-adjacent period gets treated as possibly affected. The declaration costs us data since the baseline now sits further from the outcomes it anchors, and that is the price of admitting the response started early. Implementations: the did package in R and csdid in Stata, both with an anticipation argument. Search terms: “group-time average treatment effects”, “anticipation horizon”, “Callaway Sant’Anna”.
In experiments we can build this variation on purpose with a two-stage design, which first randomises the treatment intensity across groups and then randomises treatment within each group (Hudgens and Halloran 2008), and that assignment structure is what gives us separate direct and spillover comparisons.
The partial-interference setting comes from vaccines, where the spillover is entirely the point: your vaccination protects me. Hudgens and Halloran (2008) showed that with interference confined to groups, four effects become definable and estimable. The direct effect compares treated with untreated people in the same group. The indirect effect compares untreated people in a heavily treated group with untreated people in an untreated group, and it is the formal version of herd immunity. The total effect combines the two, and the overall effect compares whole groups under different treatment intensities. The same four questions translate to our subsidy study: the indirect effect is what happens to untrained workers’ wages when many colleagues get trained, which for a labour economist is the displacement question. Search terms: “partial interference”, “direct and indirect effects”, “spillover effects”, and the no-interference assumption’s formal name, “SUTVA”.
The two-error decomposition is the useful lesson of Sävje (2024): when the declared mapping is wrong, part of the damage is that our estimand changed, and part is ordinary bias, and the two need different repairs - a better definition for the first, a better design for the second. This is also why ring analyses can disagree without any of them being broken, since each ring answers its own question. When we can’t defend a single mapping, the practical move is to report estimates under several and to say which institutional knowledge would narrow the set, e.g. commuting-flow data that shows how far workers travel. Search terms: “exposure mapping misspecification”, “ring analysis spillovers”, “interference causal inference”.
Our friends at the Psychology department met this problem first. Cronbach and Meehl (1955) had to validate tests for things like anxiety and intelligence, where no direct benchmark exists, so they proposed checking the measure against the relationships theory predicts (if the concept should rise with X and fall with Y, the measure should too). Campbell and Fiske (1959) turned this into a table. The idea is, we measure several concepts with several methods, then ask for two things: different methods measuring the same concept should agree (convergent validity), and the same method measuring different concepts should not agree too much (discriminant validity). The second check catches a failure we in econ rarely test for. Two supposedly different concepts can move together just because the same method produced both, e.g. everything self-reported moving together because the same person answered everything on the same afternoon. Search terms: “construct validity”, “multitrait-multimethod matrix”, “convergent and discriminant validity”, “common method bias”.
Adcock and Collier (2001) split operationalisation into steps: background concept, systematised concept, indicators, scores. The split is what tells us where the trouble is. Sometimes the indicator is fine and the concept was never pinned down, and no new variable repairs an undefined concept. Bollen and Lennox (1991) add a distinction that changes what validation means. Reflective indicators are effects of the concept (test answers reflect ability), while formative indicators build it (income, education and occupation together form socioeconomic status). The agreement logic that works for reflective measures is the wrong test for formative ones, whose components can disagree because they measure different parts on purpose. Search terms: “systematized concept Adcock Collier”, “reflective and formative indicators”, “measurement model”.


