aikyam school

Pre-Experimental Out-of-Sample Predictive Profiling

RCTClinical Trial

In randomized evaluations, dividing experimental subjects using endogenously estimated control group models introduces sample overfitting and systematic treatment effect estimation biases.

Picture this

Imagine testing a new teaching method on current students. Instead of grouping students using current test scores—which might be altered by the study itself or small sample noise—a formula built on last year's student records is used to categorize current students before measuring results.

What the evidence says

The historical model successfully stratified the experimental sample into 1,688 high-employability ($\le 6$ months) and 2,475 low-employability ($> 6$ months) individuals, correctly identifying 87% of long-duration cases in baseline data without endogenous stratification bias.

Who was studied
N = 55,545 male UI entrants from 2011 (historical pre-experimental cohort) used to classify N = 4,163 male job seekers aged 25–64 in Germany.
How
Out-of-sample Weibull Proportional Hazard duration model estimated on historical 2011 administrative register data to predict individual median unemployment duration ($m(T|x)$) without using experimental sample outcomes.

What to do

Estimate baseline duration and profiling models on independent pre-experimental historical cohorts rather than in-sample experimental controls when evaluating subgroup treatment heterogeneity.

From the source

"We do not use in-sample observations to quantify the prediction model in order to avoid overfitting and the related risk of biased treatment effects (Abadie et al. 2018)."

Mandatory_integration_agreements_for_unemployed_job_seekers_a_randomized.pdf

Tags

  • Profiling
  • Econometrics
  • Predictive Modeling
  • Out-of-Sample Estimation
  • Weibull Model