aikyam school

Agnostic Machine Learning Heterogeneity Detection in Field Trials

Standard average treatment effect (ATE) estimates frequently fail to detect meaningful behavioral responses because opposing positive and negative reactions across sub-populations cancel each other out in aggregate data.

Picture this

Evaluating a program using only average results is like measuring the average temperature of a room with one frozen wall and one burning wall and concluding the room is comfortable. Machine learning acts like a thermal camera, splitting the room into sections to reveal the extreme cold and extreme heat hiding behind the average.

What the evidence says

While overall average treatment effects were small and statistically insignificant (-1.4 to -2.7 percentage points, p > 0.10), generic ML inference uncovered significant treatment effect gaps between the top and bottom quintiles of 20.9 percentage points for social stigma (p = 0.009) and 12.6 percentage points for salient stigma outreach (p = 0.022).

Who
N = 1,460 to 1,470 (Experiment 1) and N = 768 (Experiment 2) unemployed/underemployed youth aged 18–35 in Cairo, Egypt [4, 23].
How
Chernozhukov et al. (2022) generic machine learning framework employing 100 iterations of split-sample cross-validation across Elastic Net, Boosted Trees, Random Forest, and Neural Networks to calculate Sorted Group Average Treatment Effects (GATES) [21, 24-26].

What to do

Apply agnostic machine learning split-sample validation to estimate Sorted Group Average Treatment Effects before concluding that an economic or social intervention lacks impact.

From the source

"Many studies have failed to find much evidence of stigma depressing take-up on average (see Currie (2006) for a survey), but as we show, this may be due to underlying heterogeneity that went undetected."

Stigma and Take-Up of Labor Market Assistance- Evidence from Three Experiments

Tagged

Nearby findings