Inspiration

It's late December 2019. A doctor in Wuhan finishes a fourteen-hour shift and files a report. Eight words: undiagnosed pneumonia, Hubei, request for information. It sits in a database, unread. Thirty-one days later, the world shuts down.

That report existed. The warning existed. Nobody was reading it in time.

This isn't a one-off. In 2014, a health worker in Guinea described a strange fever in a small village — three months before we called it the West African Ebola epidemic. Mpox trickled out of a handful of countries before anyone connected the dots; by then it was in dozens. The first warning almost always exists. It almost always gets lost, because no human team on earth can read every disease report the moment it's filed.

We wanted to find out: is the signal actually there in that first report — and can a machine catch it?

What it does

Tirith reads the earliest report of an outbreak and predicts whether it will escalate into an epidemic or fizzle out — before the official alarm is raised.

It ingests a disease-outbreak report (the kind that lands on a public-health analyst's desk daily), reasons about the epidemiological signals in the narrative, and returns an escalation probability with its reasoning. It's a triage tool for the people drowning in early signals: which of this week's thirty reports is the one that matters?

The dashboard presents active signals ranked by escalation risk, a live global threat map, and an analyst interface you can question directly — including the one that matters most: "Would you have caught COVID?"

How we built it

Data. We started from a public database of ~3,300 WHO Disease Outbreak News reports spanning 1996–2019 (Carlson et al.). The database had structured fields but no narrative text and no outcome labels, so we:

  • Fetched the full narrative of each report from WHO's archive (~2,100 clean reports).
  • Grouped reports into outbreak events and derived an escalate/fizzle label from what actually happened (follow-up reports, case-count growth, geographic spread).
  • Applied a strict anti-leakage cut: the model only ever sees an outbreak's first report. It reads the opening line of the story and predicts the ending.
  • Split the data temporally — train on the past, validate on the future — so no result can leak backward through time.

The finding. Before training anything, we ran a go/no-go test. A frontier model reading only first-report narratives hit ~73% balanced accuracy at predicting escalation — versus ~52% for a model using only structured case-count features. The signal is real: whether an outbreak explodes is written in the language of its very first report, not just the numbers.

The model. We distilled reasoning traces from a frontier model on our training examples and used them as supervised fine-tuning targets, teaching a smaller Qwen 3.5 model (trained on Freesolo) to reason about escalation rather than mimic a template. We then applied reinforcement learning (GRPO) with a reward combining prediction correctness and confidence calibration, plus asymmetric penalties — punishing a missed outbreak more heavily than a false alarm — because in this domain, a false negative is a body count.

Challenges we ran into

  • The data didn't exist where we thought. Our first two data sources fell through — one shipped no narratives, another paywalled its archive and blocked scraping. We found a clean, public path through WHO's own reports only after several dead ends.
  • Preventing leakage was harder than the modeling. A model that "predicts" an outbreak it has already seen the ending of is worthless. The first-report cut and temporal split were non-negotiable, and we verified them by hand.
  • Small models struggle with subtle narrative signal. Our fine-tuned model became strongly safety-biased — it misses fewer real escalations than the frontier model, but over-flags, raising too many alarms. Two independent reward configurations landed at the same operating point, telling us the bottleneck was model capacity and training signal, not reward tuning — which is exactly why we pivoted to reasoning distillation.

What we learned

  • The signal is real and measurable. Outbreak escalation is genuinely predictable from first-report narrative text, and by a wide margin over structured features. That finding stands on its own.
  • Reward design is a scalpel, not a hammer. Weighting missed-escalations too heavily on imbalanced data drove the model straight into "always escalate." GRPO faithfully optimizes exactly what you tell it to — including your mistakes.
  • Honest evaluation beats a flattering number. The temporal split and first-report cut made our numbers lower and truer. We'd rather report a defensible result than an inflated one.

What's next

  • Reasoning distillation at scale — transferring more of the frontier model's analytical ability into a small, cheap, deployable model.
  • Larger base models to close the remaining capacity gap.
  • A live ingestion layer connecting Tirith to real-time public-health feeds, so the next quiet report at the end of a long shift is actually being read.

Built With

Share this project:

Updates

Submission history