health · digital · chemistry
the mouse said it was safe
More than 90% of drugs that clear animal safety testing fail in humans
Problem statement
"Historically, more than 90 percent of drugs that clear animal studies do not receive FDA approval, often due to safety or efficacy issues identified in human trials" (FDA, April 2026). Toxicity is a large share of that attrition but not the largest: analyses of clinical trial data from 2010 to 2017 attribute the 90% failure rate to lack of clinical efficacy (40%–50%), unmanageable toxicity (30%), poor drug-like properties (10%–15%), and lack of commercial needs and poor strategic planning (10%) (Sun et al. 2022). Animal models do not reliably predict human absorption, distribution, metabolism, excretion, and toxicity (ADME-Tox) because of fundamental species differences in drug-metabolizing enzymes, transporter proteins, and organ physiology: the reference concordance study, covering 150 compounds and 221 human toxicity events across 12 pharmaceutical companies, found a true-positive concordance rate of 71% for rodent and non-rodent species combined, "with nonrodents alone being predictive for 63% of HTs and rodents alone for 43%" (Olson et al. 2000). ARPA-H's own framing is that animal models "are expensive and not predictive of all aspects of human physiology," and that "there has been no increase in the frequency of new drug approvals and no decrease in the drug failure rate in 40 years." No computational model currently exists that can reliably predict human drug safety from molecular structure and preclinical data alone, because the multi-organ physiological interactions that produce toxicity (liver metabolism generating toxic metabolites that damage the kidney, for example) are too complex to model from first principles.
Why this matters
ARPA-H puts the average cost of getting one new drug "from discovery, through preclinical testing and clinical trials, and finally to people who need it" at $2 billion. The standard academic estimate is $1,395 million out-of-pocket per approved new compound, $2,558 million once capitalized to the point of marketing approval, and $2,870 million including post-approval R&D (2013 dollars; DiMasi, Grabowski & Hansen 2016) — figures that are per approved drug precisely because the costs of compounds abandoned during testing are loaded onto the survivors. Liver toxicity is, in FDA's words, "the most common cause for the discontinuation of clinical trials on a drug, as well as the most common reason for an approved drug's withdrawal from the marketplace." In the U.S. Acute Liver Failure Study Group's prospective multicenter cohort, 275 of 662 acute liver failure cases (42%) were acetaminophen-induced, rising from 28% of annual cases in 1998 to 51% in 2003 (Larson et al., Hepatology 2005;42:1364–1372); a further 11.1% of 1,198 ALF subjects in the same registry were adjudicated as idiosyncratic drug-induced liver injury (Reuben, Koch & Lee, Hepatology 2010;52(6):2065–2076). The most-cited estimate of the wider toll — 2,216,000 serious and 106,000 fatal adverse drug reactions among hospitalized U.S. patients in 1994 — comes from a meta-analysis of 39 prospective studies (Lazarou, Pomeranz & Corey, JAMA 1998;279(15):1200–1205) and has not been re-estimated prospectively at national scale since. If computational models could predict human toxicity before clinical trials, unsafe drugs could be eliminated earlier, development cost and timeline would fall, and promising drugs that fail in animals but would work in humans could be rescued.
What’s been tried and why it hasn’t worked
Organ-on-a-chip systems (microphysiological systems) use human cells in microfluidic devices to model individual organ responses, but linking multiple organs into a functioning "human-on-a-chip" with correct blood flow ratios and pharmacokinetics has not been achieved at physiologically relevant scales. Computational ADME-Tox models (physiologically-based pharmacokinetic models, or PBPK) can predict plasma drug concentrations reasonably well but cannot predict organ-specific toxicity mechanisms. AI/ML models trained on historical clinical trial data can identify statistical correlations between drug structure and toxicity but lack mechanistic understanding — they cannot explain why a drug is toxic or predict novel toxicity mechanisms not represented in training data. Regulatory appetite has moved faster than the science: FDA's April 2025 Roadmap to Reducing Animal Testing in Preclinical Safety Studies committed to reducing, refining, or replacing the animal-testing requirement using "AI-based computational models of toxicity and cell lines and organoid toxicity testing in a laboratory setting (so-called New Approach Methodologies or NAMs data)," and the agency's April 20, 2026 year-one report records that it has "qualified the first artificial intelligence-based drug development tool," issued draft guidance on reducing non-human-primate testing for monoclonal antibodies, and launched a searchable database of where alternative methods are acceptable. What is still missing is a settled validation standard for computational safety models, which creates a chicken-and-egg problem: models cannot be validated without clinical data, but generating clinical data has required animal testing first.
What would unlock progress
In silico models of human physiology — "digital twins" of human ADME-Tox processes — that integrate mechanistic pharmacokinetic modeling with AI-driven toxicity prediction from molecular structure could replace or supplement animal testing. This requires: (1) comprehensive training datasets linking drug molecular features to human clinical outcomes across multiple organ systems; (2) multi-organ physiological models that capture inter-organ drug metabolism and toxicity cascades; (3) regulatory acceptance frameworks for computational safety evidence that don't require retrospective animal validation. The EU's cosmetics bans have already created sustained regulatory pressure for alternative methods: under Regulation (EC) No 1223/2009, testing finished cosmetic products on animals has been prohibited since 11 September 2004 and testing of ingredients since 11 March 2009, with the marketing ban extended to repeated-dose toxicity, reproductive toxicity, and toxicokinetics endpoints on 11 March 2013 (European Commission, Internal Market — Cosmetics: animal testing).
Entry points for student teams
A student team could build a machine learning model that predicts drug-induced liver injury (DILI) risk from molecular descriptors using the FDA's DILIrank database (DILIrank 2.0 covers 1,336 FDA-approved drugs sorted into four DILI-concern classes; https://www.fda.gov/science-research/liver-toxicity-knowledge-base-ltkb/drug-induced-liver-injury-rank-dilirank-20-dataset), benchmarking against existing QSAR models and testing on a held-out validation set of drugs with known clinical outcomes. A more experimental team could design a two-organ microfluidic chip (liver + kidney) and measure how drug metabolism in the liver compartment produces nephrotoxic metabolites in the kidney compartment. Relevant disciplines: pharmacology, machine learning, biomedical engineering, toxicology, regulatory science.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
ARPA-H, "CATALYST — Computational ADME-Tox and Physiology Analysis for Safer Therapeutics," program page, ARPA-H, "CATALYST program to fast-track safer medicines from lab to patients," December 4, 2025, U.S. Food and Drug Administration, "FDA Achieves Year 1 Goals in Reducing Animal Testing in Drug Development," press announcement, April 20, 2026, Sun D, Gao W, Hu H, Zhou S. "Why 90% of clinical drug development fails and how to improve it?" Acta Pharmaceutica Sinica B 2022;12(7):3049–3062, doi:10.1016/j.apsb.2022.02.002; Olson H, Betton G, Robinson D, Thomas K, Monro A, Kolaja G, Lilly P, Sanders J, Sipes G, Bracken W, Dorato M, Van Deun K, Smith P, Berger B, Heller A. "Concordance of the toxicity of pharmaceuticals in humans and in animals." Regulatory Toxicology and Pharmacology 2000;32(1):56–67, doi:10.1006/rtph.2000.1399; DiMasi JA, Grabowski HG, Hansen RW. "Innovation in the pharmaceutical industry: New estimates of R&D costs." Journal of Health Economics 2016;47:20–33; U.S. Food and Drug Administration, "Liver Toxicity Knowledge Base (LTKB)," Accessed 2026-08-21. go to source 1 ↗ go to source 2 ↗ go to source 3 ↗ go to source 4 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
Related briefs: `health-ai-device-clinical-evidence-gap` (FDA evidence requirements for AI-based tools — CATALYST models would face similar scrutiny); `digital-safe-rl-exploration-guarantees` (safety verification challenges for AI systems making consequential decisions). The `failure:unrepresentative-data` tag is primary — the core problem is that animal data does not represent human physiology. The `failure:lab-to-field-gap` captures the additional challenge that even human cell-based in vitro models don't replicate the multi-organ interactions seen in vivo. The brief is tagged `temporal:static`: the animal-to-human predictivity gap is a persistent, decades-old barrier rather than one measurably accelerating. (The rising complexity of drug candidates — biologics, RNA therapeutics, cell therapies — does make traditional animal models even less predictive, but that is a compounding pressure, not the clean measurable-deterioration-plus-feedback pattern the `temporal:worsening` test requires.) Source-bias note: ARPA-H frames this as a pure technical barrier; the regulatory challenge of getting FDA to accept in silico evidence as a replacement for animal data is equally significant.
Reconciliation 2026-08-21: The citation itself held up — both ARPA-H items on the Source line are real and were fetched: the CATALYST program page (full program name "Computational ADME-Tox and Physiology Analysis for Safer Therapeutics," https://arpa-h.gov/explore-funding/programs/catalyst) and the press release "CATALYST program to fast-track safer medicines from lab to patients," which is dated December 4, 2025 (https://arpa-h.gov/news-and-events/catalyst-program-fast-track-safer-medicines-lab-patients) and announces up to $125 million over 4.5 years across eight performer teams. The drift was in the numbers, all of which were carried past what that agency page supports. (1) The old headline claimed animal models "Fail to Predict Human Toxicity 90% of the Time." The 90% is the overall clinical failure rate, not a toxicity-prediction miss rate; the reference concordance study found animal studies detect 71% of human toxicities with rodent and non-rodent data combined, 63% non-rodent alone, 43% rodent alone (Olson et al., Regul Toxicol Pharmacol 2000;32(1):56–67, doi:10.1006/rtph.2000.1399, record read at https://www.ebi.ac.uk/europepmc/webservices/rest/search?query=EXT_ID:11029269). Title restated to FDA's own wording, "Historically, more than 90 percent of drugs that clear animal studies do not receive FDA approval" (FDA press announcement, April 20, 2026). (2) "Unexpected toxicity being the leading cause of failure in Phase 1 and Phase 2" is wrong: Sun D, Gao W, Hu H, Zhou S., Acta Pharm Sin B 2022;12(7):3049–3062 (https://pmc.ncbi.nlm.nih.gov/articles/PMC9293739/) attributes the 90% to lack of clinical efficacy (40%–50%), unmanageable toxicity (30%), poor drug-like properties (10%–15%), and commercial/strategic reasons (10%) — toxicity is second, not first. Corrected in place. (3) "Each clinical trial failure costs $1–2 billion" could not be sourced anywhere and was removed; it appears to be the per-approved-drug development cost misapplied to a single trial failure. (4) "$2.6 billion... with over 90% of that cost attributable to clinical trial failures" — the $2.6B is a rounding of DiMasi, Grabowski & Hansen, J Health Econ 2016;47:20–33 ($1,395M out-of-pocket, $2,558M capitalized pre-approval, $2,870M with post-approval R&D, 2013 dollars; abstract read verbatim at https://econpapers.repec.org/RePEc:eee:jhecon:v:47:y:2016:i:c:p:20-33), so exact figures were substituted alongside ARPA-H's own $2 billion; the "over 90% of that cost" split is in neither source and was removed. (5) "Drug-induced liver injury alone accounts for 30–50% of acute liver failure cases" conflated acetaminophen overdose with idiosyncratic DILI: replaced with Larson AM, Polson J, Fontana RJ, Davern TJ, Lalani E, Hynan LS, Reisch JS, Schiødt FV, Ostapowicz G, Shakil AO, Lee WM, Acute Liver Failure Study Group, "Acetaminophen-induced acute liver failure: results of a United States multicenter, prospective study," Hepatology 2005;42:1364–1372, doi:10.1002/hep.20948 (275/662 = 42%, rising 28%→51% 1998–2003) and Reuben A, Koch DG, Lee WM, "Drug-Induced Acute Liver Failure: Results of a U.S. Multicenter, Prospective Study," Hepatology 2010;52(6):2065–2076, doi:10.1002/hep.23937 (133 of 1,198 = 11.1% idiosyncratic DILI; https://pmc.ncbi.nlm.nih.gov/articles/PMC3992250/). (6) "Most common cause of post-market drug withdrawal" is confirmed and now quoted from FDA's Liver Toxicity Knowledge Base page verbatim ("the most common cause for the discontinuation of clinical trials on a drug, as well as the most common reason for an approved drug's withdrawal from the marketplace," https://www.fda.gov/science-research/bioinformatics-tools/liver-toxicity-knowledge-base-ltkb). (7) The "2 million serious ADRs / over 100,000 deaths" figures verify to Lazarou J, Pomeranz BH, Corey PN, "Incidence of adverse drug reactions in hospitalized patients: a meta-analysis of prospective studies," JAMA 1998;279:1200–1205, doi:10.1001/jama.279.15.1200 — but they are a 1994 estimate restricted to hospitalized patients (serious ADRs 6.7%, fatal 0.32%), so the scope qualifier and the citation were added rather than the numbers changed. (8) "The cost of drug development could be reduced by an order of magnitude" is an unsourced counterfactual magnitude and was cut. (9) The FDA-regulatory sentence was stale: FDA's April 2025 Roadmap to Reducing Animal Testing in Preclinical Safety Studies (announced April 10, 2025, https://www.fda.gov/news-events/press-announcements/fda-announces-plan-phase-out-animal-testing-requirement-monoclonal-antibodies-and-other-drugs) and the April 20, 2026 year-one report (https://www.fda.gov/news-events/press-announcements/fda-achieves-year-1-goals-reducing-animal-testing-drug-development) now anchor it, including the qualification of the first AI-based drug development tool; the validation-standard chicken-and-egg framing survives and was kept. (10) The EU cosmetics claim verified and was made specific from the European Commission's own page (Regulation (EC) No 1223/2009; finished-product testing ban 11 Sep 2004, ingredient testing ban 11 Mar 2009, marketing ban complete 11 Mar 2013; https://single-market-economy.ec.europa.eu/sectors/cosmetics/animal-testing_en). (11) DILIrank confirmed live and updated to DILIrank 2.0 (1,336 drugs). Genome tags untouched.