Empirical Bayes Shrinkage Estimation for Volatility Adjustment
RCTClinical Trial
Single-year performance estimates contain random sampling noise and measurement error, causing unadjusted scores to overstate the true underlying variance in individual effectiveness.
Picture this
Imagine evaluating an archer who hits an exceptional bullseye during a single gusty day; assuming that single score defines their true skill will lead to disappointment on calm days. Empirical Bayes shrinkage scales extreme single-year scores back toward the group average based on how much performance naturally fluctuates across different years.
What the evidence says
Using raw value-added scores without shrinkage generated actual achievement differences following random assignment that were only 43% as large as predicted (coefficient = 0.430, SE = 0.057). Single-year math value-added required multiplication by shrinkage factors of 0.396 in elementary and 0.512 in middle school to predict future performance.
- Who was studied
- 782 elementary (grades 4–5) and 559 middle school (grades 6–8) math and English language arts (ELA) teachers across 6 urban school districts.
- How
- Linear regression of prior-year value-added estimates on subsequent-year value-added (Empirical Bayes shrinkage estimation) to derive volatility scaling parameters.
What to do
Multiply single-year performance metrics by cross-year stability coefficients before deploying estimates for predictive modeling or high-stakes decision-making.
From the source
"When there is a lot of error, we should be less willing to 'go out on a limb' with large positive or large negative predictions. We should rein in the estimates, shrinking both the large positive impacts and the large negative impacts back toward zero."
Have We Identified Effective Teachers? Validating Measures of Effective Teaching Using Random Assignment