Inspiration
One of us studies biotechnology and the other works in machine learning, so we looked for a problem that needed both. Antimicrobial resistance kept coming up. Resistant infections directly caused about 1.27 million deaths in 2019, and part of the problem is time: a lab needs 2 to 3 days to grow the bacteria and test which antibiotics work. Until then, doctors guess, which further increases resistance.
The answer is often already in the bacterium's DNA, and sequencing it is now cheap. The hard part is reading it. Existing tools return a list of gene names that most clinicians can't act on, and they never say how sure they are. That fit the track: the information exists, it just isn't usable yet.
What it does
AMR Lens takes an assembled E. coli genome and, for five common antibiotics (ampicillin, cefotaxime, ciprofloxacin, gentamicin and trimethoprim-sulfamethoxazole), says whether the drug is likely to work.
- Each answer is Resistant, Susceptible or Uncertain, with a probability.
- Each answer lists the genes behind it, so a clinician can check the reasoning.
- When the model isn't confident enough, it says "uncertain, confirm in the lab" instead of guessing.
- An optional short summary, written by Claude from the results, puts it in plain language.
How we built it
Data. We pulled every lab-tested E. coli result from the public BV-BRC database and downloaded the matching genomes, about 12,000 after quality checks. Many labs only reported a raw measurement (an MIC), so we converted those to resistant/susceptible using the EUCAST v16.1 clinical breakpoints.
Biology. Each genome was checked for size, fragmentation and species, then scanned with NCBI's AMRFinderPlus, which finds known resistance genes and mutations. That gave us 331 markers. We then reviewed by hand which genes actually affect which antibiotic.
Machine learning. One logistic regression model per antibiotic, using the markers as inputs. We tried LightGBM too, but it wasn't better, and logistic regression lets us explain every prediction exactly. Conformal prediction sets a confidence threshold per antibiotic so that confident answers keep an error rate of about 2%; anything below it is reported as uncertain.
Testing honestly. Bacteria come in families of near-identical strains, often from the same outbreak. If relatives land on both sides of a train/test split, the model can memorise strains and look better than it is. We grouped the genomes into 807 families using Mash DNA distances and always tested on families the model had never seen.
App. Streamlit, with a card per antibiotic showing the call, probability, evidence genes and the lab result when known. It's deployed on Streamlit Community Cloud with example samples. Uploading your own genome runs locally, since it needs the bioinformatics tools.
Challenges we ran into
- Our first rules were terrible. The auto-generated gene-to-drug map called almost every sample resistant to ampicillin, because of a gene nearly every E. coli carries but that rarely causes resistance on its own. The biology review fixed it.
- A tool that guessed at random. Our strain-typing tool labelled about 280 E. coli as Salmonella. When two species tie on score, it picks one at random. We caught it because the result looked wrong, and a second method confirmed they were E. coli.
- Uncertainty that broke its own promise. Our first confidence thresholds were tuned on a held-out set dominated by one large family. They aimed for 5% errors and delivered 10%. Spreading the held-out data across all families fixed it.
- A laptop and bad Wi-Fi. Scanning 12,000 genomes took two overnight runs. The first one silently failed on 1,500 downloads when the Wi-Fi dropped, so we changed the pipeline to download everything first and scan offline.
- Labels we couldn't trust. In two cases where the model disagreed with the lab, the lab's own measurement agreed with the model. We traced 93 contradictory labels to one study.
Accomplishments that we're proud of
- 95–96% balanced accuracy for every antibiotic on bacterial families the model never saw, and 97–98% accuracy when it's confident.
- Without being told any biology, the model's top genes match the textbook resistance mechanisms.
- Doubling the data across more studies cut "uncertain" answers for cefotaxime from 60% to 12% without giving up any safety.
- In our live demo, a real hospital genome the model never saw, carrying the KPC-2 carbapenemase gene, gets all four of its lab results right.
- We found real errors in a public dataset instead of just training on them.
What we learned
- How you test matters as much as the model. Grouping by family changed what our numbers meant. A model trained on one study did noticeably worse on others.
- More varied data beat a fancier model. Adding data from more studies helped far more than switching algorithms.
- Expert rules are a strong baseline. For some antibiotics, a curated gene list does about as well as our model. The model helps most where resistance depends on combinations of genes, and when it can say it isn't sure.
- "I don't know" is useful. In a medical setting, a model that flags its uncertain cases is worth more than one that's slightly more accurate on average.
- Check the data, not just the model. Some of our "errors" were the lab data being wrong.
What's next for AMR Lens
- Test it with a hospital lab on new samples.
- Add more species (starting with Klebsiella) and more antibiotics.
- Work from DNA sequenced directly from patient samples, so the growing step can be skipped too.
- Reduce the uncertain answers for gentamicin, which is still high at about 49%.
Built With
- amrfinderplus
- biotech
- bv-brc
- eucast
- machine-learning
- python
- seqkits
- streamlit

Log in or sign up for Devpost to join the conversation.