food-safety · digital
models that stop at the loading dock
Predictive microbiology can model one processing step at a Time, but the central open model repository cannot chain those models along the supply chain companies are now legally responsible for
Problem statement
Predictive microbiology — models that estimate how fast Listeria, Salmonella or E. coli will grow, survive or die under a given temperature, pH, water activity and time — is one of the food industry's main tools for deciding whether a process is safe. The world's central open repository for these models and their underlying data is ComBase, maintained by USDA-ARS and partners. The problem is that ComBase's models predict single environments in isolation: a user can ask what happens to Listeria in a product held at 8 °C, but cannot link a chilling step to a transport step to a retail-display step to estimate the risk of a real, multi-node journey. Under the U.S. Food Safety Modernization Act, companies are increasingly responsible for exactly those upstream and downstream nodes, and the tool cannot follow them there.
Why this matters
USDA-ARS describes ComBase as "the number-one web-based resource for quantitative and predictive food microbiology," used by regulators, industry and researchers to underpin HACCP plans, shelf-life decisions and risk assessments. Its own program plan lists the gaps: "At present, users cannot produce probabilistic estimations of microbial behavior based on data and models in ComBase. If the internet fails, there is currently no means to access ComBase. Under FDA-FSMA, food companies are increasingly responsible for up-and down-stream nodes that influence food safety. Currently, ComBase has no feature that allows models to be 'linked' to predict outcomes of food processes along a simulated supply chain." Companies without in-house quantitative-risk-assessment specialists therefore either build a bespoke spreadsheet chain of models by hand — error-prone and unauditable — or do not do it at all.
What’s been tried and why it hasn’t worked
ComBase has accumulated tens of thousands of microbial-response records donated by partners and funded projects, plus a suite of growth, survival and inactivation models, and the ARS plan explicitly targets "models that predict pathogen and non-pathogen behavior in complex food systems" and "data that demonstrates how models can be integrated more fully into supply chains (nodes)." Three things have blocked the leap from single-step to chained prediction. First, the models are deterministic point predictions built from laboratory broth and single-matrix data; the plan calls for "probabilistic modeling to balance the deterministic approaches" and notes performance is uncertain "especially in complex food matrices where the intrinsic and extrinsic parameters may change" — errors that compound when models are chained. Second, the interface assumes expert users: "Use and interpretation of ComBase requires a level of technical expertise that is generally lacking by ComBase users, especially industry," so the people who most need chained predictions (small and mid-size processors) are least equipped to build them. Stand-alone quantitative microbial risk-assessment tools do exist that let an expert chain process steps with Monte Carlo variability — FDA-iRISK is the best-known — but they are decoupled from ComBase's open model and data repository and presuppose the QMRA expertise ARS says industry users lack, so the gap is a chained, probabilistic capability inside the repository non-experts actually use, not the absence of any chaining software. Third, "a long-term issue is still data acquisition. Most data in ComBase was donated ... however, this does not keep pace with the needs of industry and government," so the organism–food–condition combinations that occur along real chains are often missing from the database, and there is no offline or stand-alone version for plants with poor connectivity.
What would unlock progress
The enabling move is a model-composition layer: a way to represent a supply chain as a sequence of environmental histories and propagate a distribution (not a point estimate) of pathogen concentration through them, with uncertainty that honestly widens at each step and with clear flags where the underlying data are thin. The adjacent precedent is process-simulation software in chemical engineering and discrete-event simulation in logistics — both chain validated unit-operation models into a flowsheet with propagated uncertainty. ARS's own wish list names the components: probabilistic predictions, a stand-alone version, model linking for process-risk estimates, and training modules for users with different skill levels.
Entry points for student teams
A team could build an open prototype that wraps a small set of ComBase-style growth/inactivation models as composable nodes, lets a user define a temperature–time history across chilling, transport and display, and outputs a Monte Carlo distribution of pathogen levels with per-node uncertainty contributions — validated against published multi-step challenge studies. A second scoped project is a "data-gap map": query ComBase for the organism–food–condition combinations most requested by industry and quantify where chained predictions would rest on extrapolation, producing a prioritized data-collection list. A third is an offline-first, low-expertise interface tested with quality managers at small processors. Relevant skills: food microbiology, statistics/uncertainty quantification, software engineering, human-centered design.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"2021–2025 Action Plan, National Program 108 Food Safety," USDA Agricultural Research Service, Office of National Programs, accessed 2026-08-17; ComBase (Combined Database for Predictive Microbiology), accessed 2026-08-17 go to source 1 ↗ go to source 2 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
Source is the USDA-ARS National Program 108 (Food Safety) Action Plan for 2021–2025, a tier-1 agency research agenda in which ARS states its own program's gaps; the ComBase-specific issues are credited in the plan to the ComBase scientific advisory committee (Dr. Tamplin, University of Tasmania). All quotations above are verbatim from that document (Problem Statement 5 / ComBase discussion and Research Needs). `failure:lab-to-field-gap` is used in its benchmark-to-deployment sense: models validated step-by-step in controlled matrices are not validated when chained across a real multi-node chain, and the plan itself flags performance uncertainty in complex matrices; `failure:not-attempted` was considered and rejected because ARS names an active plan to build linking, probabilistic and stand-alone features — the work is attempted but incomplete. `constraint:technical` is applied narrowly (uncertainty propagation through chained biological models is genuinely unresolved), alongside the primary `constraint:data`. Related collection briefs: `food-safety-food-water-data-fragmentation` (data-siloing across food and water systems) and `digital-food-chain-interoperability-failure` (traceability platforms) — this brief is the distinct predictive-modeling problem, not traceability. Follow-up: check the NP 108 2021–2025 Retrospective Review (published late 2024) for which ComBase initiatives were actually delivered.
Source type: Self-articulated (federal research program stating gaps in a resource it maintains)
Verifier note 2026-08-17: all NP 108 quotations confirmed against the PDF; H1 and Why-It-Matters hedged because the ARS gap statement is specific to ComBase — expert QMRA tools (e.g., FDA-iRISK) already chain process steps probabilistically, so the problem is scoped to the open repository and non-expert users; no evidence found (Aug 2026 search) that ComBase has since added model linking. Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.