education
tutoring the wrong percentile
When a district scaled tutoring to 6,800 Students, effects concentrated in the middle of the achievement distribution and vanished for the students already getting other help — and no one has a way to allocate scarce tutor slots accordingly
Problem statement
High-dosage tutoring is the best-evidenced academic-recovery strategy from small and medium randomized trials, and U.S. districts spent heavily on it after 2020. But when Metro Nashville Public Schools built and scaled its own program — over 125,000 hours to more than 6,800 students across five semesters, mostly delivered by the district's own teachers — the rigorous evaluation found small-to-medium reading effects (0.04–0.09 SD) and no average effect on math or grades. Two mechanisms the authors identify are the operational core of this brief: first, roughly 55% of tutored students were pulled from a "personalized learning time" block in which non-tutored students were already receiving computer-adaptive practice or Tier II/III small-group instruction, so the counterfactual "approximates individualized instruction … to a large degree" — the district was replacing one form of individualized help with another; second, the standards-based, universal-curriculum tutoring served the full performance range while "the most sizable effects in reading are concentrated in the 40th to the 60th percentiles" and math effects in the 50th–70th, yet "only 30% of students tutored in reading and 33% of students tutored in math were in the percentile ranges where effects were most concentrated." Districts have no practical tool for deciding which students should get scarce tutor slots given what those students would otherwise receive, and small-scale "best practice" design features (1:1, during the day, certified teachers) proved infeasible or non-decisive at scale.
Why this matters
Tutoring at MNPS cost about $750 per student per semester (about $1,500 per year), with stipends for staff tutors "accounting for 80% of total costs," and was funded almost entirely by time-limited external money (ESSER, foundations, a state corps program). Under a permanent budget, a district cannot tutor everyone; it must choose. If it chooses by a low-score cutoff — the intuitive rule — it concentrates tutoring on students who are already receiving specialized supports (weakest treatment-control contrast) and, at least under a standards-based model, outside the percentile band where the program moved scores. Kraft et al. also document a broader "clear pattern of declining effect sizes when comparing the pooled effects of smaller versus larger tutoring programs," and NSSA's research agenda lists as open questions "which students receive tutoring," "how many students can a tutor handle," and displacement — "the counterfactual: what the student would have experienced without the tutoring." Getting allocation wrong wastes the one lever districts are most willing to fund.
What’s been tried and why it hasn’t worked
MNPS iterated its design across three years in ways typical of scale-up: it pivoted from virtual volunteers to compensated district teachers (85% of students tutored by an MNPS teacher by spring 2023), moved from 100% one-to-one to a modal 3:1 ratio, and shifted 53% of sessions to before/after school because teacher planning periods rarely aligned with students' intervention block and because "there are few such spaces in most schools" for daytime small groups. Targeting used a percentile band (students scoring between the 15th and 60th percentiles nationally on diagnostics, excluded when tutoring conflicted with Tier II/III, special-education, or English-language services during the intervention block; students below the 25th/10th percentiles were eligible for Tier II/III supports) — a sensible triage rule that nonetheless produced the counterfactual problem above. The authors tested whether effects varied by tutor type, timing, modality, or ratio and found "no compelling evidence that tutoring effects varied systematically across these flexible design features," which undercuts the field's design-feature checklists as allocation guidance. The NSSA agenda notes that because "the vast majority of research on tutoring has evaluated the effect of individual tutoring programs, we cannot definitively say whether or not a specific program characteristic leads to increases in student learning" — the evidence base was never built to answer allocation questions.
What would unlock progress
Treating tutor slots as a constrained allocation problem with heterogeneous treatment effects and heterogeneous counterfactuals: districts need (a) an estimate, by student, of expected gain from tutoring relative to the support that student would otherwise receive in that time slot, and (b) an assignment mechanism that respects tutor supply, space, and schedule constraints. Adjacent fields solved analogous problems — precision-medicine and policy-learning methods (uplift modeling, optimal policy trees) choose whom to treat given heterogeneous effects, and operations research handles scheduling under room and staff constraints — but neither has been fitted to a school master schedule with a moving counterfactual. Also needed: diagnostic-driven tutoring content, since MNPS's universal standards-based curriculum "likely designed for the average student" may explain the middle-of-distribution effect concentration.
Entry points for student teams
A data-science/operations team could build a simulated district (or use one partner district's de-identified data) with a master schedule, tutor supply, and per-student baseline scores plus existing-support flags, then compare allocation rules — lowest-scores-first, effect-band targeting, uplift-based targeting net of counterfactual — on total expected gain and equity, producing a decision tool a district could actually run. An evaluation-design team could design a rolling-enrollment RCT that estimates effects by prior-support status, so the counterfactual is measured rather than assumed. A qualitative team could document how tutor leads and principals actually decide who gets tutored and where sessions fit, surfacing constraints the model must respect. Skills: causal inference, optimization, education policy, fieldwork.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"The Scaling Dynamics and Causal Effects of a District-Operated Tutoring Program," Matthew A. Kraft, Danielle Sanderson Edwards & Marisa Cannata, EdWorkingPaper No. 24-1030, Annenberg Institute at Brown University, August 2024, (full text ), accessed 2026-08-17; "High-Impact Tutoring: State of the Research and Priorities for Future Learning," Carly D. Robinson & Susanna Loeb, National Student Support Accelerator, EdWorkingPaper 21-384, accessed 2026-08-17 go to source 1 ↗ go to source 2 ↗ go to source 3 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
The primary source is a tier-1 research working paper (Annenberg EdWorkingPaper, August 2024) with experimental and quasi-experimental estimates and extensive implementation data; the NSSA research agenda is an expert-facing priorities document (2021). Effect sizes, cost, dosage, and percentage figures are taken directly from the paper's abstract, implementation tables, and discussion as extracted from the PDF; the "40th–60th" and "50th–70th" percentile bands are the authors' description of Figure 7 and should be checked against the figure. `failure:lab-to-field-gap` captures the WHERE (small-RCT effects not reproduced at district scale, with named infeasible design features); `failure:ignored-context` captures the WHY (deployment/operational — the counterfactual block already provided individualized help, and universal curricula ignored where effects concentrate); the two tags are used for distinct aspects per the taxonomy note. `constraint:coordination` and `stakeholders:multi-institution` were considered (district, foundations, state corps, nonprofit partners) and rejected: a single district controls allocation, and the binding constraints are cost and missing information. `temporal:window` was considered because ESSER funding has expired but rejected as salience rather than a change in the barrier. Duplicate check: no existing brief addresses tutoring; `education-growth-mindset-structural-blind-spot` and `education-brac-play-based-learning-transition-gap` are unrelated interventions. Note the tutoring topic is widely discussed in popular media; this brief is scoped to the allocation-under-counterfactual sub-problem, which is not.
Source type: Self-articulated (researchers evaluating a district program identifying the mechanisms that limit effects at scale)
Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.