Multi-Measure Composite Teacher Evaluation
RCTClinical Trial
Single-instrument teacher evaluation systems—such as relying solely on state test scores or administrator observations—suffer from high year-to-year volatility, narrow domain coverage, and susceptibility to student sorting bias.
Picture this
Evaluating a driver based only on their top speed on a clear highway gives an incomplete picture of driving ability. Combining a track timer, a passenger satisfaction survey, and an instructor's driving technique checklist creates a balanced scorecard that predicts overall driving performance far more reliably.
What the evidence says
The multi-measure composite predicted actual student achievement gains following random assignment with an IV slope coefficient of 0.955 (p < 0.01, SE = 0.123), demonstrating that composite ratings accurately match causal student achievement impacts.
- Who was studied
- N = 1,181 teachers and 27,265 randomized students in grades 4–8 across 6 urban public school districts (Charlotte-Mecklenburg, Dallas, Denver, Hillsborough, Memphis, and New York).
- How
- Randomized Controlled Trial (RCT) using Limited Information Maximum Likelihood (LIML) Instrumental Variables (IV) to test a weighted composite of prior value-added, Framework for Teaching classroom observations, and Tripod student perception surveys against post-randomization student test scores.
What to do
Construct a weighted evaluation composite combining prior achievement value-added, structured classroom observation scores, and student perception surveys to forecast instructor effectiveness.
From the source
"First, we used the data collected during 2009-10 to build a composite measure of teaching effectiveness, combining all three measures to predict a teacher's impact on another group of students."
Have We Identified Effective Teachers? Validating Measures of Effective Teaching Using Random Assignment