aikyam school

Methods & evidence

68findings0stories

The findings

Active Tracking Attrition MitigationAttrition was maintained at 3.4% to 7.7% across most study arms (except Vadodara Year 1 at 17-18% due to communal riots), with attritor pretest scores showing no significant difference between treatment and control groups.RCTCommercial Ad Platform Algorithmic Randomization ImbalanceA joint statistical test of baseline covariates rejected balance across treatment groups (p = 0.00), showing significant baseline differences for gender (-0.03 female diff, p < 0.01) and age (+0.03 old diff, p < 0.01), requiring normalized difference robustness re-estimation.RCTMulti-Source Administrative-Survey Data TriangulationMethodological framework integrates daily factory administrative records with individual psychometric and climate surveys; empirical findings forthcoming.RCTAdversarial Non-Stationarity Bounds in Multi-Armed BanditsProves that algorithms like the Tempered Thompson Algorithm achieve worst-case average regret on the order of $O(1/\sqrt{T})$ (up to logarithmic terms) relative to any fixed targeted assignment policy, guaranteeing robust participant welfare even under non-stationary outcome distributions.RCTAttanasio Selection-Bias Decomposition FrameworkVirtual-Within interaction yields a statistically significant intensive quality increase among motivated Milestone 0 completers ($\beta = 0.138, p < 0.05$ in Large-Country Sample; $\beta = 0.427, p < 0.05$ in Uganda Sample). Under the BAS assumption, this observed conditional effect serves as a strict lower bound for the true unconditional quality impact.RCTConservative Attrition Bounding and ImputationLong-run wage gains remained statistically significant and economically meaningful (>10% of control mean) across all differential attrition scenarios, retaining a positive point estimate even under the most conservative assumption of imputing a full 0.5 SD penalty to missing treated units.RCTIndividualized Posterior Taste Weight Recovery via Bayes' RuleIndividual posterior weights ($\hat{\beta}_i^S$) ranged from near 0 to above 3.0 across applicants. Because measurement error ($\beta_i^S - \hat{\beta}_i^S$) is orthogonal to the posterior expectation by construction, using $\hat{\beta}_i^S$ as an instrumental variable interaction avoids classical measurement error attenuation bias.Observational StudyOrder Independence in Field BDM ValuationMean willingness-to-pay bids were virtually identical across orderings (Real policy: Rs. 69.0 in Ordering 1 vs. Rs. 67.9 in Ordering 2; Basis risk policy: Rs. 38.9 vs. Rs. 39.1), confirming complete order independence.RCTTangible Goods Simulation for Field BDM AuctionsPractice rounds enabled successful implementation of 4-contract BDM elicitations across 1,978 rural subjects, resulting in zero instances of participants refusing to complete winning transactions.RCTBirthdate-Based Randomization DesignProduced balanced baseline demographic distributions across treatment and control groups while enabling evaluation of policy regime shifts.RCTCaseworker-Level Randomization in Policy EvaluationCaseworker-level randomization successfully eliminated cross-client informational spillovers while maintaining balanced treatment and control groups across observed jobseeker demographic characteristics (passing balance tests in all regions except Geneva).RCTCaste-Stratified Market Penetration RandomizationGenerated exogenous cross-village coverage variation ranging from 0% to 53% for cultivators and 0% to 100% for landless laborers, with zero statistically significant correlation to village size, caste concentration, or number of castes after conditioning on eligibility shares.RCTCatchment Cohort Size InstrumentThe first-stage F-statistic was 23.54 without controls and 10.32 with controls, yielding an IV elasticity estimate of household expenditure to rule-based grants of -0.946 to -1.124 (p < 0.01).Observational StudyConstruct Sensitivity Divergence in Non-Cognitive EvaluationIntention-to-treat estimates showed a statistically significant improvement of 0.165 to 0.175 standard deviations on targeted PLANEA scores (p < 0.05), compared to a statistically insignificant 0.017 to 0.027 standard deviation effect on the broad BarOn inventory.RCTCross-Sample Intervention ReplicabilityKey treatment effects replicated across both cohorts with striking consistency: persistent choice of challenging tasks increased by 8.5 percentage points in Sample A (p < 0.01) and 12.9 percentage points in Sample B (p < 0.01); week-ahead commitment increased by 14.3 percentage points in Sample A (p < 0.01) and 19.1 percentage points in Sample B (p < 0.01); and second-visit success increased by 8.7 percentage points in Sample A (p < 0.01) and 10.7 percentage points in Sample B (p < 0.05).RCTDual-Channel Branch and Intercept RecruitmentRetention rates remained highly balanced across recruitment channels (56% retained inside MTO branches vs. 52% retained in public areas), confirming observable baseline balance across gender (74% female), age (mean 35 years), and monthly income (55% above median).RCTEvaluation Time Horizon BiasThe estimated intervention effect was 0.085 SD (p < 0.05) after 1 year (before household reoptimization), but fell to a statistically insignificant 0.053 SD (p = 0.611) after 2 years once households adjusted private spending.RCTExogenous Failure AssignmentAmong the subset forced into the difficult task who failed in Round 1 (N = 558 in Sample A, N = 522 in Sample B), treated students re-attempted the difficult task in Round 2 at rates 17.6 percentage points higher in Sample A (p = 0.002) and 13.9 percentage points higher in Sample B (p = 0.064) than control group baseline rates of 36% and 51%.RCTExogenous Boundary Redistricting for Sorting ControlEstimated choice parameters for driving distance, home-school preference, and test scores in the reassigned subsample were quantitatively and qualitatively similar to the full sample (middle school distance parameters -0.255 vs -0.234 for white non-lunch students), confirming that residential sorting does not drive preference estimates.Observational StudyExperimental Self-Selection Verification and External ValidityProject firms showed no statistically significant differences from nonproject firms in pre-intervention assets ($12.8M vs $13.9M, p = 0.841), employees (204 vs 221, p = 0.552), total borrowings ($4.9M vs $5.5M, p = 0.756), or baseline BVR management scores (2.52 vs 2.55, p = 0.859).Observational StudyFacebook Social Proof for Field Survey ParticipationSocial proof validation helped overcome initial refusal challenges, enabling successful enrollment of 8,248 participants and sustained engagement across 51,395 weekly survey responses over 30 weeks.RCTFinancial Outcome WinsorizationWinsorization bounded weekly remittance amounts (control mean 2,338.50 PhP), stabilizing regression estimates across full sample (p = 0.81) and low-baseline subsample (+174.94 PhP, p < 0.05) specifications.RCTFloating Recall Period Data CaptureCaptured 137,927 weekly observations while eliminating date overlaps and memory gaps across 30 follow-up weeks.RCTFirst-Order Stochastic Dominance (FOSD) Communication PolarizationBusiness content and positive sentiment exhibit strong negative Pearson correlation ($\rho = -0.478$ at message level, $\rho = -0.480$ at entrepreneur level, $p < 0.01$). Kolmogorov-Smirnov FOSD test yields $p = 1.0$ for low sentiment dominating high sentiment in business content and $p = 0.0$ for the reverse, confirming strict inverse polarization.RCTHawthorne Effect Mitigation in Phone SurveysStandardized fortnightly phone surveys collected reliable longitudinal search data without generating detectable Hawthorne or priming effects, preserving true treatment-control contrast across survey rounds.Observational StudyHigh-Frequency Survey Recall ValidationDaily visits averaged 8.11 GH¢ cash in and 3.32 GH¢ cash out. Comparing monthly recall to aggregated daily data showed no statistically significant difference in revenue (p-value > 0.10, difference of -4.08 GH¢) or expenses (difference of 8.71 GH¢), proving that monthly recall surveys do not introduce strategic reporting bias in consulting evaluations, though daily data exhibited higher variance for record-keepers (squared difference 22,895 GH¢, p < 0.05).RCTHonest Split-Sample ML Heterogeneity InferenceIdentified statistically significant treatment effect divergence between the top and bottom predicted individual treatment effect (ITE) quintiles for social stigma (p = 0.009 in Exp 1, p = 0.022 in Exp 2) and professional stigma (p = 0.015 in Exp 1), uncovering hidden heterogeneity where average treatment effects were near zero.RCTRecipient Household Welcome Call Rapport StrategyThe preliminary welcome call strategy enabled completion of 2,075 detailed endline household surveys despite high phone survey attrition challenges (with only ~10% explicit refusal rate).RCTHybrid Administrative Record Linkage OptimizationThe hybrid four-tiered matching protocol achieved a 1.97% false positive error rate and a 2.46% false negative error rate overall (compared to exact matching alone which yielded a 0.8% false positive rate but an 8.0% false negative error rate), minimizing overall attenuation bias in binary outcome estimation.Observational StudyIncentivized Lab Game Behavioral ForecastingEvery 100 € allocated to direct payment in the game increased real-world EduPay take-up by 2.07 percentage points (p < 0.01; an 8.7 percentage point increase per 1 SD, representing a 32% increase relative to the 27.1% mean take-up rate).RCTIncentivized High-Frequency Micro-SurveysThe micro-survey protocol achieved a retention rate of 54% over 30 weeks while capturing 137,927 weekly survey observations verified through photo receipt submissions.RCTIncentivized Real-Effort Grit ElicitationExperimental choice of the difficult task in all 5 rounds strongly predicted baseline math test scores (coef = 0.346, p < 0.01) and Turkish scores (coef = 0.152, p < 0.05) over and above Raven cognitive matrix scores (coef = 0.437, p < 0.01).RCTIndirect Inference Structural EstimationThe structural parameters matched the reduced-form empirical moments (treatment exit rate, control exit rate, regional comparison exit rate, and vacancy response curves), validating structural policy simulation counterfactuals.Expert TheoryIndividual Fixed Effects Precision Enhancement in Panel RCTsIncluding individual fixed effects controlled for baseline individual variation, achieving an adjusted R-squared of 0.010 for weekly remittance probability and 0.012 for the low-baseline remittance subsample, ensuring identification precision across 137,927 panel observations.RCTIntention-to-Treat (ITT) InstrumentationIn Year 2 in Mumbai, 33% of schools assigned to receive a balsakhi tutor did not receive one due to administrative difficulties; 2SLS instrumentation successfully eliminated non-compliance bias to yield unbiased direct treatment estimates.RCTInstrumental Variable Estimation under One-Sided Non-ComplianceIV estimates proved that downloading predictions had zero statistically significant impact on short-term caseworker compliance (estimates ranged from -0.05 to +0.03 across specifications, all statistically insignificant).RCTKling Summary Index AggregationThe Kling labeling index demonstrated a statistically significant treatment take-up effect (mean index = 1.233, t = 44.445, p < 0.001) and a statistically significant overall remittance increase in the low-baseline subsample (+0.027 index units, p < 0.01).Observational StudyKling Summary Index Aggregation in Multi-Outcome RCTsThe overall labeling treatment effect on the aggregated Kling remittance index across the full sample was +0.013 standard deviations (p > 0.10), while for the low baseline remittance subsample it reached +0.027 standard deviations (p < 0.01).RCTLottery Prize Allocation for Preference ElicitationFood expenses (26%), utility bills (25%), and education expenses (20%) were the most frequently designated expenditure priorities, followed by medical expenses (10%), rent payment (6%), mortgage payment (5%), and business expenses (5%).RCTLottery Prize Preference ElicitationThe lottery allocation task elicited baseline spending priorities across food (26%), utility bills (25%), education (20%), and medical expenses (10%) while identifying target beneficiary households in the Philippines.RCTMarket Tightness InvarianceSupply-side windfall ($\omega = 0.259$) and demand-side substitution ($\psi = 0.227$) were nearly equal ($\omega - \psi = 0.032$), and program scale relative to the broader labor market was under 10% (9.9% mean ratio), resulting in near-zero market tightness adjustment and validating control group spillover independence.RCTModular Unbundling of Financial Education ComponentsHighlights that traditional aggregate training evaluations obscure underlying mechanisms, recommending cross-randomization of sub-modules to isolate distinct cognitive and behavioral channels.Expert TheoryMulti-Register Administrative Data Linkage for Counterfactual EstimationLinked multi-register database captured 11-year pre-unemployment labor histories and 12-to-24 month post-unemployment follow-up outcomes with zero panel attrition across the entire national population sample.Observational StudyMulti-Trait Intervention DecouplingPatience training alone produced no significant effect on gritty task choices or perseverance after failure, whereas standalone grit training in Sample B replicated the full magnitude of treatment effects observed in Sample A's combined intervention (+19.1 vs +14.3 percentage point increase in week-ahead task commitment).RCTMultiple Comparison with the Best (MCB) for Statistical Treatment SelectionMCB routines dynamically categorized interventions based on individual covariate precision; the set of significantly best programmes varied from a single clear optimal choice for some jobseekers to all available programmes for jobseekers with high variance in predicted potential outcomes.RCTMultiple-Ranked Choice Identification in Exploded Logit SystemsBetween 27.4% and 50.9% of parents who listed their guaranteed home-school first or second submitted subsequent choices beyond the default, providing within-person substitution patterns that isolated preference parameters from random unobserved utility shocks.Observational StudyNon-Response Bounding and Weighting RobustnessPost-intervention wage returns remain positive and statistically significant under weighted adjustments (IV estimate = $4.20/day, p < 0.05; IHS daily wage = 0.605, p < 0.05) and conservative bounding scenarios, confirming that survey non-response does not account for the observed labor market returns.RCTObjective Behavioral Metric via Voucher RedemptionTreatment group students demonstrated a 28% redemption rate compared to 18% in the control group, representing a 55% relative increase in condom demand (p = 0.07).RCTOmnibus Super-Family Index InferenceConstructing standardized family indices confirmed that secondary domains (e.g., total financial expenditure $p = 0.797$, life satisfaction $p = 0.901$) showed no spurious treatment effects, maintaining statistical rigor across multi-domain evaluations.RCTPhased Treatment Enablement ProtocolThe phased enablement protocol effectively isolated non-complying recruits without introducing observable baseline trait imbalances (p > 0.10 across balance tests on gender, age, income, and remittance history).RCTReceipt-Verified Point Incentive SystemReaching 100 points awarded a 25 AED ($6.81 USD) gift certificate; the verification mechanism successfully generated high-quality, verifiable weekly panel data across 137,927 observations.RCTSecond-Order Peer Instrumentation (G²X Identification)In Large-Country Virtual-Within networks, peer outcomes increase proposal submission rates ($\beta = 0.533, p < 0.01$). In Small-Country Virtual-Across networks, peer outcomes increase intensive business proposal quality ($\beta = 0.476, p < 0.01$).RCTSeemingly Unrelated Regressions (SURE) for Joint Covariate BalanceSURE joint hypothesis testing yielded p-values of 0.163 across the initial randomized sample, 0.282 for the state test sample, and 0.405 for subsequent peer characteristics, statistically confirming successful experimental balance across all baseline features simultaneously.RCTShift-Isolated Educational Unit PartitioningValidated shift-level experimental independence across 13 multi-shift facilities without cross-shift contamination or social interaction overlap.RCTSingle-Industry Sectoral Standardisation in Enterprise RCTsConcentrating on a single trade allowed four business consultants to deliver 1,200 hours of sector-specific coaching (averaging 10 hours per tailor), eliminating cross-sector noise and permitting precise tracking of 35 tailoring-specific operational practices over eight survey rounds.RCTSmartphone-Based Remittance Survey IncentivizationThe micro-incentive tracking framework achieved a 54% retention rate over 30 weeks, generating a take-up index of 1.233 (t = 44.45, p < 0.001) with treatment migrants actively sending 0.054 labeled remittances per week (t = 41.69, p < 0.001).RCTSocial Transparency in Peer ReportingUnder no-stakes conditions, public reporting doubled survey response accuracy, but under high-stakes grant competition, public visibility had limited impact on reducing strategic misreporting.RCTStability Selection via Bootstrapped Regularized Logistic RegressionHigh business content is reliably flagged by terms like product, project, business, market, idea, need, customer, service, plan, and cost, whereas low business content is flagged by hello, hi, token, leaderboard, status, and thanks.RCTStaggered Treatment Activation ProtocolEstablished a uniform 5-week pre-treatment baseline window per respondent, enabling precise subsample categorization between low and high pre-treatment remitters.RCTStock vs. Flow Sampling in Program EvaluationShort-term compliance rates remained identical across both subgroups (11% stock vs 14% flow for Definition 1; 29% stock vs 30% flow for Definition 2), confirming that timing of prediction availability did not alter caseworker compliance behavior.RCTStratified Cluster Randomization DesignStratification achieved complete baseline balance across treatment and comparison groups, keeping all initial pretest score differences below 0.10 standard deviations.RCTTarget Beneficiary Identification via Lottery ProtocolSuccessfully identified key target household respondents—48% parents, 20% siblings, 10% spouses, and 12% children—establishing paired migrant-household linkages for 1,377 households at endline.RCTTwo-Wave Strategic Baseline SuspensionWave 2 branch co-location resolved baseline refusal challenges, allowing enrollment to reach 8,248 total migrants and successfully retaining 4,458 participants.Observational StudyUnannounced Paper-and-Pencil Survey Measurement ProtocolAttrition was limited to 13% between baseline and first follow-up, and 10% between baseline and second follow-up, with no differential attrition across treatment and control arms.RCTWithin-Firm Spillover Dilution ControlRemoving multi-establishment corporate branches prevented internal information spillovers, maintaining orthogonal treatment variation across 7,438 independent establishments.RCTWithin-Participant Nested Choice ArchitectureIsolating features revealed that soft labeling accounted for a 15% increase (93.66 €) over the basic remittance baseline of 614.6 €, while adding direct payment added only an incremental 2.2% (13.74 €), demonstrating diminishing marginal returns for hard commitment features.RCTWorkplace Batching and Peer Isolation ProtocolControl group workers showed zero output decline on Day 9 after witnessing Wave A treatment workers receive early cash payouts, confirming the absence of relative fairness spillover.RCTWorst-Case Variance Bounded SamplingSetting lambda = 0.00125 guarantees a worst-case standard error of 0.05 for treatment contrasts and a statistical power >= 0.80 for detecting effect sizes of 0.124 across N = 4,000 participants, enforcing a minimum assignment probability floor of 0.05 per treatment arm.Expert Theory