5 Comments
User's avatar
scott cunningham's avatar

Regarding the third one. Am I understanding correctly that the authors built placebos based on a single treated unit? Why put CS and the others against that? The placebo exercise for staggered is interesting, but we already know the limitations will be smaller as the size of the never treated grows.

Beatriz Gietner's avatar

hi prof Scott!! yep, correct! they say the staggered estimators all reduce to TWFE in that case, so they include them as a check that none does worse in finite samples (and they all perform identically). they also say that a staggered version could show bigger differences :)

scott cunningham's avatar

I need to read the paper. I don’t understand the point of the exercise. Why use CS on single treated situations? It just collapses to TWFE then since cs calculates 2x2s which are numerically what TWFE will do in that situation too. Four averages and three subtractions.

Beatriz Gietner's avatar

yes, you’re right, and they say it themselves (p. 9): with one treated unit the TWFE-like estimators “all converge to the canonical TWFE estimator”. they still include them because the methods are “slightly different in their implementation”, to check “whether they perform any worse than TWFE under finite-sample randomization inference”. they also say (footnote 1) that a staggered version “may result in more discernable differences in performance for the TWFE-like estimators”. i think your observation is pertinent (re it being mechanical), but maybe their reasoning was about implementation? they normalise all of them to the last pre-treatment period (pp. 10-11), which is what makes them match TWFE. they don’t say which packages they used, and theres not much on the Fed’s page, so i can’t check to be sure

scott cunningham's avatar

What do we learn if you use four averages and three subtractions with TWFE and four r averages and three subtractions with CS, and the same with SA? They’re literally the same thing.