education · digital · family: the data exists but can’t talk to itself
experimentswithoutfaces
Digital learning platforms can now run cheap randomized experiments on millions of students — but cannot answer "For whom does it Work?" because they Don't hold student demographic data
Problem statement
The Institute of Education Sciences created SEERNet in 2021 to turn five widely used digital learning platforms (ASSISTments/E-TRIALS, MATHia/UpGrade, OpenStax/Kinetic, Canvas/Terracotta, ASU Online) into research infrastructure, so that education experiments could be run faster, cheaper, and at scale, with an explicit mandate for "equity-focused" research. The obstacle SEERNet's own guidance surfaces is that the platforms generally do not know who their students are: "a DLP may not have access to individual student or teacher demographic data (that is, without connecting to other data sources or embedding a survey) but may collect school-level covariates, including locale type … or Title I status." Answering the question policymakers most want answered — whether an intervention narrows or widens gaps by race, income, language status, or disability — requires linking platform records to student information systems held by districts, which means data-sharing agreements, de-identification protocols, and IRB review that a platform-embedded experiment was supposed to make unnecessary.
Why this matters
The promise of platform-based research is replication at scale: hundreds of A/B tests on real coursework instead of one $3M field trial. If that machinery can only report average effects, it will systematically miss heterogeneous effects, and equity claims — the stated purpose of the IES investment — will rest on school-level proxies (Title I status, urban/rural) that blur within-school differences. Researchers also cannot easily measure attrition or compliance for subgroups, and, as the report notes, some platforms do not even provide direct attrition data — it must be reconstructed from activity timestamps. The result is a research infrastructure whose cheapest studies are its least informative for the students the studies are meant to serve.
What’s been tried and why it hasn’t worked
SEERNet's approach so far has been procedural: advising researchers to explore existing data before designing a study, to negotiate with the platform about what can be varied, and to "weigh the tradeoffs and level of effort that may be required to access and integrate additional data, especially if that necessitates further data sharing agreements with educational institutions." Some platforms offer to broker district relationships and coordinate de-identification "such that the researcher never sees personally identifiable information," and some populate rosters from an LMS or SIS, which "could afford possibilities for linking with other sources of data given appropriate permissions" — but each of these is a per-study negotiation, which reintroduces the cost and delay the infrastructure was built to eliminate. Embedded consent-plus-survey collection of demographics works better for platforms serving adults (OpenStax, ASU) than for K–12 platforms where minors cannot consent and districts control the data. The report also flags that IRBs unfamiliar with platform research "may … be more conservative in their review process," and that researchers must find out whether a missing data type reflects "current technical constraints that could be relaxed in the future or if it is so by design." No standard, reusable mechanism exists for equity-relevant covariates to travel with a platform experiment.
What would unlock progress
A standardized, privacy-preserving way for platform experiments to obtain subgroup covariates without per-study bespoke agreements — for example, a district-side service that returns only aggregate subgroup treatment effects (or salted, hashed subgroup labels) computed against the platform's randomization, in the spirit of secure multiparty computation or differential-privacy releases used in census and health data; a model data-sharing agreement and IRB template specific to platform-embedded experiments; or validated methods for estimating heterogeneity from the school-level covariates platforms already hold, with honest bounds. Adjacent precedents: statistical linkage units in health (trusted third parties that link and de-identify), and ad-tech's aggregate-measurement APIs that report conversions by cohort without exposing users.
Entry points for student teams
A software/privacy team could prototype a "trusted linkage" service: the platform exports randomization assignment and outcome by pseudonymous ID, the district holds demographics, and the service returns only subgroup effect estimates above a minimum cell size — testing it on synthetic SIS and platform data and measuring what precision is lost relative to individual-level linkage. A policy/legal team could draft a model FERPA-compliant data-sharing agreement and IRB protocol for platform experiments and pilot it with one district and one open platform (ASSISTments' E-TRIALS is designed to host external researchers). A statistics team could quantify, using public NAEP or state data, how much subgroup heterogeneity is recoverable from school-level proxies alone. Skills: privacy engineering, education law, causal inference, human-subjects ethics.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"Considerations for Conducting Research in Digital Learning Platforms," A. Schellinger, J. Zacamy, J. Roschelle, A. Closser & C. D. Zepeda, Digital Promise (SEERNet, IES-funded network lead), April 2024, ERIC ED657728, accessed 2026-08-17; IES SEERNet program page, accessed 2026-08-17 go to source 1 ↗ go to source 2 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
The Digital Promise paper is the network lead's synthesis of "conversations and informal interviews" with SEERNet research teams — practitioners writing for researchers, on ERIC, funded by IES (R305N210034 per the paper's disclaimer). Tier 2 (network technical report). The `stakeholders:multi-institution` tag passes the three-criteria test: platform vendors hold interaction and outcome data, districts hold demographics, and researcher institutions hold IRB authority — none can produce an equity estimate alone, and the inter-institutional data boundary is the binding barrier. `constraint:coordination` was considered and rejected: the actors are willing (platforms broker district contacts), and removing coordination friction alone would still leave FERPA/consent constraints and platform data architecture in place — the binding constraint is data plus regulatory. `failure:ignored-context` is applied in its data/information sub-pattern: platforms were architected for instruction, not for the covariates research requires. `temporal:newly-created` because platform-as-research-infrastructure is a post-2021 construct (SEERNet, AIMS Collaboratory). Related briefs: `education-essay-scoring-dialect-bias` (subgroup harm from an ed-tech system) and `education-displaced-student-data-portability` (student data that cannot follow the learner) are neighbors, not duplicates. Follow-up: SEERNet's public "Guide to Doing Research on DLPs" and the AIMS Collaboratory "researcher-ready" catalog would sharpen which platforms already support consent-based demographic capture.
Source type: Self-articulated (research-network lead documenting a limitation of its own infrastructure)
Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.