Worst-Case Variance Bounded Sampling
Expert TheoryReview
Adaptive multi-armed bandit algorithms dynamically reduce assignment to underperforming treatment arms, which risks severe loss of statistical power and invalidates ex-ante frequentist standard error guarantees.
Picture this
Think of testing four different types of tires on race cars over a season. If one tire performs poorly early on, you want to stop putting it on drivers' cars to avoid accidents; however, if you stop using it completely, you will not gather enough data to mathematically prove its performance. Setting a strict safety floor guarantees that every tire gets tested on a minimum percentage of cars so statistical proofs remain valid.
What the evidence says
Setting lambda = 0.00125 guarantees a worst-case standard error of 0.05 for treatment contrasts and a statistical power >= 0.80 for detecting effect sizes of 0.124 across N = 4,000 participants, enforcing a minimum assignment probability floor of 0.05 per treatment arm.
- Who was studied
- Parameterized for N = 4,000 subjects across 16 demographic strata in Jordan.
- How
- Mathematical derivation establishing an upper bound lambda on estimator variance (lambda = 0.00125) to derive the algorithm tempering parameter gamma.
What to do
Select the upper variance parameter lambda based on required statistical power and sample size N to establish the minimum treatment assignment share gamma/k.
From the source
"Our experiment has the objective of maximizing participant outcomes... subject to the constraint of allowing for sufficiently precise inference on the effect of different treatments... We require frequentist guarantees on the precision of estimates."
An_Adaptive_Targeted_Field_Experiment_Job_Search_Assistance_for.pdf