Skip to content
problem genome

Data & method

Where these problems come from, and how much to trust them

How a problem gets in

Problems are collected from places where experts talk to each other, not to the public: research agendas, standards-body roadmaps, agency gap reports, conference proceedings, expert interviews, post-mortems. A problem earns a place when it is real (experts say so, in writing), open (nobody has solved it), specific enough to work on, and comes with constraint information — why past attempts failed.

Each problem becomes a brief: the problem statement, why it matters, what’s been tried and why it hasn’t worked, what would unlock progress, and entry points for a student team — every claim cited to its source (see what a finished brief looks like). There are 613 briefs today, and the collection is alive: new ones must pass the verification gate below before they enter.

The six-layer fact-check

Every brief goes through six checks before we call it verified:

  1. The source link works and points at the named document — every URL is fetched and checked.
  2. Every number in the brief is extracted and checked against the source.
  3. The brief says what the source says — no paraphrase drift.
  4. The causal story holds — why-it's-stuck has to survive a skeptical read.
  5. The tags fit their written definitions, with the hardest tags passing an extra decision gate.
  6. The brief is consistent with the rest of the collection — duplicates and too-thin briefs are flagged.

Sources are graded by how close they are to the experts: Tier 1 = a primary expert source (the agency’s own gap report, the standards body’s own roadmap, the paper itself) — 449 of 613 briefs; Tier 2/3 = analyst, conference, or community sources, still fact-checked. A small number of briefs (3 right now) carry a “needs sourcing” flag: the problem looks real but we want a deeper source before we’d stake a semester on it.

  • highTier-1 source — a primary expert source, fully fact-checked
  • mediumTier-2/3 source — analyst, conference, or community source, fact-checked
  • needs sourcingflagged for deeper sourcing — help us source it

How much to trust a tag

Every problem carries genes — structured tags for its field, its binding constraint, why it’s stuck, what breakthrough it needs. Tags are only useful if two people reading the same brief would apply the same ones, so we test exactly that: independent readers re-tag samples of the collection blind, and we measure how often they agree. On the most recent run, two of three blind readers agreed nine times in ten across the tag system. The hardest tags — the ones readers disagree on most — must pass an extra written decision gate before anyone applies them, and they’re drawn with a dashed border everywhere on this site.

The workbenches ring on each brief (which of seven systems a problem engages) was coded the same way: on a shared sample, readers agreed on whether a given workbench is engaged about nine times in ten, but on the single primary workbench only about two times in three — so read the long wedge as our best single call, not ground truth.

The codebook, and how it changes

The full codebook — every gene with its plain-language definition and how to recognize it — lives on this site: start from the notation page and choose any gene. The tag system is a living document: it is reviewed against the data at fixed milestones, tags are merged, split, added or retired based on how they perform, and every change is logged. The most recent review (August 2026, run over the full collection) added four tags — each now under review for its first hundred uses, drawn dashed until it graduates. The 19 families are re-derived after each review, and each family page names how its grouping has grown from release to release.

Who makes this

The Problem Genome is a research project of Natural Artificial Labs at Dartmouth, built for design and engineering education: the collection is the instrument, and classrooms that adopt problems from it feed the research on how people learn to read — and choose — problems worth solving. It is inspired by the Music Genome Project and the Art Genome Project: start with a reasonable set of genes, then let the data teach you better ones.

Take the data

The collection is licensed CC BY-SA 4.0. Use it, teach with it, build on it, sell what you build — with two conditions: credit the collection, and license whatever you build from it the same way, so it stays open for the next person.

Download all 613 problems as JSON → — every brief with its genome tags, source, and body text. Each brief page also offers a per-problem citation in BibTeX.

Cite it as: The Problem Genome Project, Natural Artificial Labs, Dartmouth College, www.problemgenome.com — CC BY-SA 4.0. One limit worth stating plainly: the license covers our briefs, not the sources they cite. Every brief points at documents that keep their own terms — check those before reproducing anything from them.