Regarding the third one. Am I understanding correctly that the authors built placebos based on a single treated unit? Why put CS and the others against that? The placebo exercise for staggered is interesting, but we already know the limitations will be smaller as the size of the never treated grows.
hi prof Scott!! yep, correct! they say the staggered estimators all reduce to TWFE in that case, so they include them as a check that none does worse in finite samples (and they all perform identically). they also say that a staggered version could show bigger differences :)
I need to read the paper. I don’t understand the point of the exercise. Why use CS on single treated situations? It just collapses to TWFE then since cs calculates 2x2s which are numerically what TWFE will do in that situation too. Four averages and three subtractions.
yes, you’re right, and they say it themselves (p. 9): with one treated unit the TWFE-like estimators “all converge to the canonical TWFE estimator”. they still include them because the methods are “slightly different in their implementation”, to check “whether they perform any worse than TWFE under finite-sample randomization inference”. they also say (footnote 1) that a staggered version “may result in more discernable differences in performance for the TWFE-like estimators”. i think your observation is pertinent (re it being mechanical), but maybe their reasoning was about implementation? they normalise all of them to the last pre-treatment period (pp. 10-11), which is what makes them match TWFE. they don’t say which packages they used, and theres not much on the Fed’s page, so i can’t check to be sure
What do we learn if you use four averages and three subtractions with TWFE and four r averages and three subtractions with CS, and the same with SA? They’re literally the same thing.
Regarding the third one. Am I understanding correctly that the authors built placebos based on a single treated unit? Why put CS and the others against that? The placebo exercise for staggered is interesting, but we already know the limitations will be smaller as the size of the never treated grows.
hi prof Scott!! yep, correct! they say the staggered estimators all reduce to TWFE in that case, so they include them as a check that none does worse in finite samples (and they all perform identically). they also say that a staggered version could show bigger differences :)
I need to read the paper. I don’t understand the point of the exercise. Why use CS on single treated situations? It just collapses to TWFE then since cs calculates 2x2s which are numerically what TWFE will do in that situation too. Four averages and three subtractions.
yes, you’re right, and they say it themselves (p. 9): with one treated unit the TWFE-like estimators “all converge to the canonical TWFE estimator”. they still include them because the methods are “slightly different in their implementation”, to check “whether they perform any worse than TWFE under finite-sample randomization inference”. they also say (footnote 1) that a staggered version “may result in more discernable differences in performance for the TWFE-like estimators”. i think your observation is pertinent (re it being mechanical), but maybe their reasoning was about implementation? they normalise all of them to the last pre-treatment period (pp. 10-11), which is what makes them match TWFE. they don’t say which packages they used, and theres not much on the Fed’s page, so i can’t check to be sure
What do we learn if you use four averages and three subtractions with TWFE and four r averages and three subtractions with CS, and the same with SA? They’re literally the same thing.