education · family: the data exists but can’t talk to itself
earlygradesoffthemap
The world tests millions of second- and third-graders with EGRA, EGMA, citizen-led and household surveys — and almost none of it can be used to report whether they meet the global minimum proficiency level
Problem statement
SDG indicator 4.1.1 asks every country to report the share of children reaching a minimum proficiency level (MPL) in reading and mathematics at three points: grades 2/3 (4.1.1a), end of primary (4.1.1b), and end of lower secondary (4.1.1c). The early-grade point is where foundational learning is won or lost, yet it is by far the least covered: UIS counts only 31 countries (2013–2017) rising to 34 (2018–2022) with usable grade 2/3 data, versus 85–98 for end of primary. The paradox is that early-grade assessment is abundant — the Early Grade Reading/Mathematics Assessments (EGRA/EGMA), the PAL Network's citizen-led household assessments, and UNICEF's MICS Foundational Learning Module "have been applied globally" — but, in UIS's words, "they cannot be currently used for global reporting, mostly because they were not intended to generate comparable data." No accepted method yet expresses results from these instruments on the global MPL scale.
Why this matters
Without comparable early-grade data, the international system cannot see where foundational-learning gaps open, cannot target support, and cannot track whether early interventions work — UIS notes the data gap "inhibits international efforts to provide targeted support to countries that need it most." The school-age population covered at grades 2/3 is roughly 92–110 million children versus 291–379 million at end of primary (UIS Figure 1), so hundreds of millions of children are invisible at exactly the stage where remediation is cheapest. Countries that already pay for EGRA rounds or host citizen-led surveys are, in effect, generating data that cannot count toward the global goal they signed up for.
What’s been tried and why it hasn’t worked
UIS deliberately anchored 4.1.1 to a concept — the Global Proficiency Framework's MPL descriptors — rather than to a single test score, so that disparate assessments could be aligned to a common benchmark, and three linking routes have been developed. The IEA's Rosetta Stone study builds statistical concordance tables from regional to international assessments, but its first implementation linked ERCE and PASEC to TIMSS and PIRLS — end-of-primary instruments, not early-grade tools. Policy Linking is a non-statistical, judgment-based method (a 5–6-day workshop with 15–20 teacher panelists aligning national items to the GPF and setting benchmarks); it was proposed in 2017, piloted in 2019, revised in 2020, re-piloted 2021–2022, revised again in January 2023, and remains "under piloting phase" — and if its five quality criteria (enough aligned items, nationally representative sample, minimum administration standards) are not met, "the workshop will be considered a capacity building activity," i.e., it yields no reportable number. The third route, calibrated Assessments for Minimum Proficiency Levels, produced AMPL-b (end of primary) in 2021 for six African countries under the MILO project, but AMPL-a for early grades was still under development at the time of writing. The structural obstacles UIS names are that "every country sets its own standards, leading to inconsistent definitions of performance levels," different regions have "different traditions concerning the stringency of proficiency benchmarks," and early-grade tools like EGRA are individually administered oral fluency measures built for program evaluation, not sampled population reporting.
What would unlock progress
A validated, low-cost linking method specifically for early-grade instruments — for example, embedding a short common calibrated module (an AMPL-a-style anchor) inside routine EGRA/EGMA or citizen-led rounds so that statistical linking replaces judgment-only linking, or a psychometric bridge between oral-fluency measures (words correct per minute, item-level EGRA data) and the GPF grade 2/3 descriptors. Adjacent fields have solved analogous problems: clinical outcome measures are cross-walked with common-item equating (PROMIS linking), and labor-force surveys harmonize national instruments to ILO definitions through anchor modules. The UIS paper itself frames the destination as "an international community of practice" to converge on procedures; the missing piece is a demonstrated technical route for 4.1.1a.
Entry points for student teams
A measurement/statistics team could take one publicly available EGRA or citizen-led dataset (USAID's EGRA datasets and PAL Network member data are partly open) and the published GPF grade 2/3 descriptors and produce a documented, reproducible linking study — mapping items to descriptors, estimating the share at MPL under alternative benchmark rules, and quantifying uncertainty. A design team could specify a 10–15-item common anchor module and simulate how much precision it adds to judgment-based linking. A policy team could cost the options for one country (UIS references a Zambia costing paper) and write the reporting pathway. Skills: psychometrics/IRT, survey statistics, comparative education, and some field-instrument familiarity.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"Measuring and Monitoring Learning Outcomes and Skills: What Are the Challenges Going Forward?" (draft, October 2023), UNESCO Institute for Statistics, Session 1 paper for the UNESCO Conference on Education Data and Statistics, 7–9 February 2024, UIS/ESC/10, accessed 2026-08-17 go to source ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
The UIS paper is a UN statistical agency's own gap statement prepared for its technical conference — tier 1 (agency gap document); it is marked DRAFT (October 2023) and its coverage figures should be re-checked against the latest UIS SDG 4 data release. Country counts (31→34 for grades 2/3; 85→98 end of primary; 78→85 lower secondary) and population figures are read from the paper's Figure 1 as extracted from the PDF text and should be verified against the graphic. `failure:ignored-context` is applied in its data/information sub-pattern: the early-grade instruments were designed for program evaluation and were never built to be population-comparable, so the global-reporting use case was not part of their design. `stakeholders:multi-institution` passes the three-criteria test: UIS (indicator custodian), IEA (Rosetta Stone), USAID/RTI (EGRA/EGMA), the PAL Network, and UNICEF (MICS) each own a non-substitutable instrument or method, and no single one can produce comparable 4.1.1a data. `constraint:coordination` was considered and rejected: the actors broadly agree on the goal and the concept-based MPL approach; the binding constraint is the absence of a validated linking method and comparable data, not unwilling parties. `temporal:window` is the deadline type (2030 SDG horizon). Related brief: `education-curriculum-assessment-misalignment` concerns national curriculum-vs-test alignment, not cross-national comparability. Follow-up: UIS's "country's options to report" paper and the Learning Data Compact for cost figures.
Source type: Self-articulated (indicator custodian agency documenting its own measurement gap)
Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.