education
averaging a moving target
Thirteen States are building "Through-Year" Tests, but no one has a defensible way to turn fall and winter scores into an end-of-year proficiency claim
Problem statement
A "through-year" assessment replaces the single end-of-year state test with several shorter tests spread across the school year, promising less testing time and results teachers can actually use. But federal accountability still requires one annual, comparable proficiency determination per student, and state standards are written as end-of-year expectations. The unsolved measurement problem is how to combine evidence gathered in the fall and winter — when students had not yet been taught much of the year's content, and may since have forgotten or mastered what was tested — into a valid end-of-year claim. As of 2023, thirteen states were designing, piloting, or implementing through-year programs, yet the Center for Assessment's matrix of state designs shows that no state has adopted the model in which every administration samples the full standards and all scores are combined; most states sidestep the question by using only the spring score for accountability and treating fall/winter results as "informational."
Why this matters
Statewide assessment drives what is taught and how schools are judged; if states cannot make the aggregation defensible, they either abandon the instructional benefits (reverting to a spring test with extra interim tests bolted on, which increases rather than reduces testing time — the report's Consideration 9) or make proficiency claims that will not survive federal peer review or legal challenge. The Center for Assessment notes that "the field also lacks definite solutions to the many technical, logistical, and policy challenges that arise as a state moves to a through-year assessment model," and that the research base "has not deepened substantively" since its 2021 convening. Every state that pilots without a solution is spending public money and instructional time on a design it may have to walk back.
What’s been tried and why it hasn’t worked
This is the second wave of interest, not the first: the U.S. Department of Education's 2010 Race to the Top assessment grants explicitly invited designs that based annual proficiency on tests given throughout the year, and the PARCC consortium proposed a multi-session model — then dropped it from its final design "citing concerns about cost, testing time and local control of curriculum." ESSA (2015) reopened the door by permitting "multiple interim statewide assessments" to yield an annual determination, and the Innovative Assessment Demonstration Authority encouraged pilots. Gong's 2021 analysis identifies the logical core: a through-year test used for summative purposes must show its within-year evidence is at least as good as end-of-year evidence, which requires "showing that the student did not change significantly between when the evidence was gathered" and the end of the year — but the whole point of instructionally useful within-year testing is that instruction does change the student, so "effective instructional uses of an assessment should reduce the 'predictive validity' of that assessment with subsequent performance." States have responded by choosing among eight design models (Dadey et al.'s appendix): those that combine within-year scores (Delaware, Georgia's Navvy, Louisiana, Montana) do so only by testing a subset of standards per administration, which ties the test to local scope-and-sequence and threatens comparability; Texas lets earlier scores count only if they help the student; six states (Alaska, Georgia MAP, Kansas, Nebraska, Maine, Virginia) use fall/winter results only to build a common growth scale or seed a multistage adaptive test; and Florida uses them purely for information. Nobody has a principled aggregation rule for the model states actually want.
What would unlock progress
A psychometric and policy framework for weighting time-stamped evidence about a learner who is expected to change — something like an evidence-decay or forgetting-aware measurement model that states the value of a September item response as evidence about June proficiency, together with clear rules about which standards can be certified as "met" mid-year and which must be re-verified. Adjacent fields have relevant machinery: knowledge-tracing models in intelligent tutoring systems explicitly estimate time-varying mastery from sequences of responses; longitudinal item response theory and growth models handle repeated measures; and clinical trial designs with interim analyses have formal rules for combining early and final evidence without inflating error. Also needed is empirical evidence, which the report says barely exists, on how large the "students may have forgotten, or better mastered, the content" effect actually is between administrations.
Entry points for student teams
A team with psychometrics or data-science skills could simulate a state's through-year program using synthetic or public longitudinal item-response data (e.g., interim assessment datasets or ITS logs), compare candidate aggregation rules — spring-only, simple averaging, best-of, evidence-decay weighting, and knowledge-tracing-based estimates — and quantify how often each misclassifies students relative to an end-of-year criterion. A policy-design team could take the eight-model matrix and write the theory-of-action and validity argument for one state's chosen model, specifying exactly which claims each administration supports. A design team could prototype the teacher-facing report that makes mid-year results useful without implying an end-of-year proficiency claim. Relevant disciplines: measurement/psychometrics, statistics, education policy, information design.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"Through-Year Assessment: Ten Key Considerations," Nathan Dadey, Carla M. Evans & Will Lorié, National Center for the Improvement of Educational Assessment (Center for Assessment), March 2023, accessed 2026-08-17; "Why Has It Been So Difficult to Develop a Viable Through-Year Assessment?," Brian Gong, Center for Assessment blog, 2021-02-23, accessed 2026-08-17 go to source 1 ↗ go to source 2 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
The Center for Assessment is a nonprofit that advises state education agencies on assessment and accountability; the 2023 paper is an expert-to-expert technical/policy paper (CC-BY) with a state-by-state design matrix, and the 2021 Gong blog states the validity argument in its sharpest form — hence tier 2 (technical/analyst report by a practitioner body), not peer-reviewed research. The count of "13 states" and the model matrix are as of March 2023 and will have moved; a verifier should check the current state list (Texas and Florida were flagged in the report as not finalized). `failure:theoretical-gap` is applied because the report says the field lacks solutions to the technical challenges and Gong frames a logical contradiction, not because no one has tried; `failure:adoption-barrier` captures the PARCC through-course design being dropped over cost, testing time, and local control. `temporal:window` is the deadline type: states are making design commitments now. `constraint:coordination` was considered (states, districts, vendors, USED peer review) and rejected: stakeholders do not agree on the approach, and the binding constraint is the measurement problem, not willing actors failing to coordinate. Related existing brief: `education-curriculum-assessment-misalignment` (competency curricula vs. recall-based tests) is a different mismatch; this brief is about the temporal aggregation of evidence within a single accountability test.
Source type: Self-articulated (assessment-advisory body articulating an unsolved technical problem in the systems it helps design)
Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.
Related collection briefs (distinct sub-problems, cross-referenced at intake 2026-08-17): `education-early-grade-assessment-global-linking`.