aikyam school

Honest Split-Sample ML Heterogeneity Inference

RCTClinical Trial

Standard subgroup analysis in policy interventions risks data-mining and false positives when searching across multi-dimensional participant attributes. Detecting true treatment effect variation requires agnostic, out-of-sample machine learning predictions that prevent an individual's own outcome from biasing their predicted treatment effect.

Picture this

Imagine trying to guess if a new coaching style helps specific athletes. Instead of looking at an athlete's own score to decide if the coaching worked for them, you use data from half the team to build a predictive rule, and then test that rule on the other half. Because an athlete's predicted benefit is calculated strictly from everyone else's results, you guarantee that your prediction is unbiased and honest.

What the evidence says

Identified statistically significant treatment effect divergence between the top and bottom predicted individual treatment effect (ITE) quintiles for social stigma (p = 0.009 in Exp 1, p = 0.022 in Exp 2) and professional stigma (p = 0.015 in Exp 1), uncovering hidden heterogeneity where average treatment effects were near zero.

Who was studied
N = 1,460 jobseekers (Experiment 1) and N = 768 jobseekers (Experiment 2) in Cairo, Egypt.
How
Randomized Field Experiment applying Chernozhukov et al. (2022) generic machine learning framework with 100 repeated random split-sample cross-validations using Lasso, Elastic Net, Boosted Trees, and Random Forest algorithms to estimate Sorted Group Average Treatment Effects (GATES).

What to do

Apply split-sample machine learning algorithms to baseline data in randomized trials to evaluate subgroup variation instead of performing ad-hoc post-hoc subgroup tests.

From the source

"These methods use machine learning algorithms that predict the individual treatment effect for study participants using baseline data on others in the sample. This ensures that the estimated individual treatment effects are not merely a form of data-mining, but are in fact 'honest'."

Stigma and Take-Up of Labor Market Assistance: Evidence from Two Field Experiments

Tags