health · digital · family: the regulator demands evidence that cannot exist
approvedwithout proof
Over 1,250 AI medical devices cleared by the FDA — nearly half lack public clinical evidence
Problem statement
More than 1,400 AI-enabled medical devices have been authorized for marketing in the United States (FDA AI-enabled device list, March 2026 update). A cross-sectional study of the 903 devices on the FDA's list through August 2024 (Windecker et al., JAMA Network Open, 2025) found that only 505 (55.9%) reported clinical performance studies at the time of authorization, and 218 (24.1%) explicitly stated that no clinical performance studies were conducted. Most were cleared through the 510(k) pathway based on bench testing or retrospective dataset performance, without prospective clinical validation demonstrating real-world diagnostic accuracy: of the reported studies, only 8.1% were prospective and 2.4% randomized (Windecker et al.), echoing an earlier analysis of 130 FDA-approved AI devices that found evaluations were overwhelmingly retrospective (Wu et al., Nature Medicine, 2021). The evidence bar is low even for the De Novo pathway that first-of-a-kind devices use: among the 63 novel moderate-risk therapeutic devices authorized via De Novo between 2011 and 2019 — therapeutic devices generally, not AI devices specifically — 31% of pivotal studies failed to meet at least one primary effectiveness endpoint, yet the devices were authorized (Johnston et al., JAMA Internal Medicine, 2020). Clinicians integrating AI tools into diagnostic and treatment workflows cannot independently assess whether the device performs as claimed in their specific patient population, clinical setting, and workflow context.
Why this matters
AI-enabled devices are used across radiology (stroke detection, pulmonary embolism triage, fracture detection), cardiology (ECG interpretation), pathology (cancer detection), and ophthalmology (diabetic retinopathy screening) — specialties where clinicians make time-critical decisions based on AI outputs. When independent academic validation studies have been conducted, they frequently find performance below manufacturer claims in real-world settings. If an AI device has a higher false-positive rate in a specific demographic or clinical setting than its clearance data suggests, patients may undergo unnecessary interventions or experience delayed treatment for life-threatening conditions.
What’s been tried and why it hasn’t worked
The 510(k) pathway requires demonstration of substantial equivalence to a predicate device, not independent clinical validation — and for AI devices, the predicate may use an entirely different underlying technology, making the comparison structurally weak. In the same De Novo analysis, 19% of the novel therapeutic devices (12 of 63) were never evaluated in pivotal studies at all (Johnston et al., 2020) — and no equivalent published accounting yet exists for AI devices specifically. The FDA's AI/ML action plan and Predetermined Change Control Plan (PCCP) guidance address how algorithms can be updated post-market but do not address the baseline clinical validation gap. Some academic institutions have begun independent validation studies, but these require access to clinical datasets that raise privacy and cost barriers. Manufacturers consider performance data proprietary and competitive, actively resisting transparency mandates. No standardized reporting framework exists for AI device performance analogous to STARD for diagnostic accuracy studies, though TRIPOD+AI is emerging.
What would unlock progress
A regulatory mechanism that requires minimum clinical evidence thresholds for AI devices affecting clinical decision-making — without requiring the full PMA pathway that would stifle innovation — would close the gap. A standardized, publicly accessible performance reporting framework (building on TRIPOD+AI) that enables clinicians to compare AI devices by clinical setting, patient demographics, and workflow integration would transform purchasing and adoption decisions. Federated validation approaches, where performance is tested across institutional datasets without centralizing patient data, could resolve the privacy-versus-transparency tension.
Entry points for student teams
A student team could build a systematic database of FDA-cleared AI medical devices and their publicly available clinical evidence, extending existing cross-sectional analyses and creating a tool clinicians could use to evaluate devices before adoption. A team with machine learning and clinical skills could conduct an independent validation study of one or more FDA-cleared AI diagnostic devices using institutional or publicly available imaging datasets, documenting performance across demographic subgroups. A design-oriented team could prototype a "nutrition label" format for AI device performance that presents accuracy, sensitivity, specificity, and demographic breakdown in a standardized, clinician-friendly format.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
Windecker D, Baj G, Shiri I, et al., "Generalizability of FDA-Approved AI-Enabled Medical Devices for Clinical Use," *JAMA Network Open*, 2025;8(4):e258052. DOI: 10.1001/jamanetworkopen.2025.8052. Accessed 2026-08-20; Johnston JL, Dhruva SS, Ross JS, Rathi VK, "Clinical Evidence Supporting US Food and Drug Administration Clearance of Novel Therapeutic Devices via the De Novo Pathway Between 2011 and 2019," *JAMA Internal Medicine*, 2020;180(12):1701–1703. DOI: 10.1001/jamainternmed.2020.3214. Accessed 2026-08-20; Wu E, Wu K, Daneshjou R, et al., "How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals," *Nature Medicine*, 2021. Accessed 2026-08-20; MedTech Dive, "AI in medtech is booming. Track new devices here." (tracker of the FDA AI-enabled medical device list; more than 1,400 devices authorized as of the FDA's March 2026 update). Accessed 2026-08-20; Applied Radiology, "Gaps in Clinical Data for FDA-Approved AI-Enabled Medical Devices" (news summary of Windecker et al.). Accessed 2026-08-20 go to source 1 ↗ go to source 2 ↗ go to source 3 ↗ go to source 4 ↗ go to source 5 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
Primary sources include the FDA Map / Applied Radiology cross-sectional analysis (2024), Nature Digital Medicine analysis of De Novo pathway performance (2024), and the FDA's draft guidance on "Artificial Intelligence-Enabled Device Software Functions" (January 2025). The "regulatory-mismatch" failure tag captures the core structural issue: the 510(k) pathway's legal standard is substantial equivalence, not clinical effectiveness, and changing this requires Congressional action. The "worsening" temporal tag reflects that AI device authorizations are accelerating (the count has grown from ~100 in 2020 to over 1,250) while the evidence gap compounds. The "design-proposal" tractability tag reflects that the regulatory and reporting framework changes needed are well-characterized but not yet implemented. Related briefs may include algorithmic bias, clinical decision support, or medical device regulation problems.
Reconciliation 2026-08-20: re-sourced after the citation-drift sweep flagged that the sole citation was a popular Applied Radiology news article while the body carried four precise study statistics. All four numbers were traced to their actual peer-reviewed sources and verified against them: (1) the 55.9%-with-clinical-performance-data figure (505 of 903 devices, plus 24.1% explicitly reporting no studies, 38.2% retrospective / 8.1% prospective / 2.4% randomized designs) is from Windecker D, Baj G, Shiri I, et al., JAMA Network Open 2025;8(4):e258052 (doi:10.1001/jamanetworkopen.2025.8052), a cross-sectional study of the FDA list through 2024-08-31 — the Applied Radiology piece was a news summary of this study, now kept only as supplementary; (2) the "one-third of De Novo devices failed primary effectiveness endpoints" and "one-fifth never evaluated in pivotal studies" figures are from Johnston JL, Dhruva SS, Ross JS, Rathi VK, JAMA Internal Medicine 2020;180(12):1701–1703 (doi:10.1001/jamainternmed.2020.3214) — verified as 31% of pivotal studies failing at least one primary effectiveness endpoint and 19% (12/63) of devices with no pivotal study; the body previously implied these were AI devices, but the Johnston cohort is novel moderate-risk therapeutic devices generally, and the body now says so explicitly; (3) the device count is updated from "over 1,250" to "more than 1,400" per the FDA's March 2026 list update as reported by MedTech Dive's tracker (the FDA's July 2025 update listed 1,247, per the FDA Law Blog — the title's "over 1,250" remains accurate); (4) the retrospective-evaluation claim is additionally anchored to Wu et al., Nature Medicine 2021 (130-device analysis). The first paragraph of these Source Notes predates this reconciliation: its "Nature Digital Medicine analysis of De Novo pathway performance (2024)" does not match any source found — the De Novo numbers come from the 2020 JAMA Internal Medicine research letter above. Genome tags untouched.