Double Post-Lasso High-Dimensional Covariate Selection
RCTClinical Trial
Including a large set of baseline covariates in experimental regression models creates a risk of specification searching and parameter over-fitting when evaluating treatment effects.
Picture this
Think of a detective sifting through hundreds of potential clues at a crime scene. Instead of manually picking favorite clues to build a case, the detective uses an automated scanner that systematically selects only the few specific clues that mathematically improve the investigation's accuracy without bias.
What the evidence says
The automated procedure selected only 4 baseline variables out of 43 (prior vacancies, prior permanent hires, prior fixed-term hires, and sector indicators), yielding identical treatment effect point estimates (+0.047 for permanent hires) and standard errors (0.022).
- Who was studied
- N = 7,438 firms in France evaluated across 43 candidate baseline control variables.
- How
- Double post-Lasso covariate selection protocol (Belloni et al., 2014) applied to linear treatment effect estimation.
What to do
Apply double post-Lasso automated variable selection to control for baseline high-dimensional firm attributes without introducing specification search bias.
From the source
"To select covariates, we implement the double post lasso developed in Belloni et al. (2014) in order to avoid the risk of specification search... There is almost no impact of adding the additional covariates either on the estimated coefficients or on their standard errors."
Are_Active_Labor_Market_Policies_Directed_at_Firms_Effective_Evidence.pdf