-
-
PhageScout reaches a non-redundant pair in ~15 assays; random & broad-host screening hit the 30-budget ceiling without one. Lower=better.
-
PhageScout finds a diverse pair in ~2 assays — fewest of all five policies, far ahead of random/broad-host. Lower is better.
-
PhageScout ranks phages by belief with uncertainty, tracks strain coverage, and recommends the exact next batch to test.
-
Confirmed, non-redundant shortlist — every phage was experimentally validated in the replay. Triage support, not a prescription.
-
Cocktail tab: a confirmed shortlist, every phage experimentally validated against this isolate; triage support, not a prescription.
-
PhageScout is fully open source — belief updating, acquisition, baselines, and replay in one clean repo.
-
PhageScout
What inspired us: When antibiotics fail against a drug-resistant infection, bacteriophages—viruses that infect bacteria—can become a last-resort treatment option. The problem is that each phage only works against certain bacterial strains, so laboratories may need to test dozens or hundreds of candidates before finding a match.
Most computational systems rank phages once and stop. We wanted to build something closer to how a real laboratory works: every assay result should change what gets tested next.
What does PhageScout do: PhageScout is a closed-loop decision-support system that recommends which phages a laboratory should test next against a patient’s bacterial isolate.
It works in two phases:
- Phase 1: Find one active phage as quickly as possible.
- Phase 2: Find a second active phage with a different host-range or receptor-binding profile.
The goal is to build a confirmed, non-redundant two-phage cocktail using fewer laboratory assays. After every positive or negative result, PhageScout updates its predictions across the remaining phage bank and recommends the next batch.
The dashboard displays:
- Current infection probability for every phage
- Tested negative phages
- Confirmed active phages
- Assays required to find the first hit
- Assays required to find a diverse pair
- A round-by-round decision log
- Comparisons against baseline screening strategies
How did we build it: We designed PhageScout as an active experimental-design system rather than another static phage-host classifier.
Its selection strategy balances three goals:
- Exploitation: Test phages that are likely to work.
- Exploration: Test phages whose results could improve predictions across the bank.
- Diversity: Avoid repeatedly testing highly similar phages.
The app can generate synthetic interaction matrices or accept real laboratory datasets. We processed public E. coli and Klebsiella pneumoniae interaction data into a standard format where rows represent phages, columns represent bacterial strains, and each cell records whether an interaction occurred.
Users can select a strain as the simulated patient isolate, configure the initial model quality, upload optional phage features, and set a testing budget. Each simulated assay reveals one previously hidden interaction.
What makes it innovative: Existing phage-susceptibility systems mainly ask:
Will this phage infect this bacterium?
PhageScout asks:
Given everything we have learned so far, which experiment should the laboratory perform next?
The main innovation is not simply predicting interactions. It is optimizing the full screening sequence toward a practical goal: finding multiple confirmed, non-redundant active phages with as few assays as possible.
What challenges did we face: Real biological datasets use different formats and scoring systems. Some contain strong, weak, and negative interactions, while others include blank cells that mean “not tested” rather than “inactive.”
We created preprocessing workflows that preserve these distinctions instead of incorrectly converting missing data into negative results.
Another challenge was defining phage diversity honestly. Different host-range profiles or receptor-binding proteins may suggest that two phages are non-redundant, but they do not prove independent resistance mechanisms. We therefore present diversity as a prediction rather than claiming the cocktail is evolution-proof.
We also had to balance immediate success with information gathering. Testing only the most likely phages may produce redundant results, while testing only uncertain phages may waste limited laboratory resources.
What are we proud of: We built an end-to-end prototype that can:
- Load synthetic or real phage–bacteria interaction matrices
- Treat any bacterial strain as a patient isolate
- Recommend sequential batches of phages
- Learn from every assay result
- Switch automatically from first-hit discovery to cocktail diversification
- Track assays-to-first-hit and assays-to-diverse-pair
- Compare performance against random and static screening
- Support real E. coli and Klebsiella datasets
The system produces a clear, measurable result: how many assays PhageScout needs compared with simpler screening strategies.
What did we learn: We learned that the most useful machine-learning contribution is not always a more complex prediction model. In this problem, the key question is how predictions guide the next real-world action.
We also learned the importance of separating prediction from confirmation. PhageScout does not prescribe an untested therapy. It prioritizes laboratory experiments, and only confirmed active phages are considered for the final cocktail.
What is next for PhageScout: Our next step is to evaluate PhageScout across every eligible strain in the real E. coli and Klebsiella datasets.
We plan to measure:
- Assays required to find the first active phage
- Assays required to find a valid two-phage pair
- Success rates under fixed testing budgets
- Improvement over random and static screening methods
Future versions could replace simulated starting probabilities with genomic models using bacterial genomes, phage genomes, receptor-binding proteins, and historical interaction data. We also plan to add assay-noise modeling, uncertainty calibration, testing costs, turnaround-time constraints, and experimentally validated cross-resistance information.
Our long-term goal is to connect genomic prediction, phage-bank inventory, and wet-lab testing in one adaptive system that helps laboratories find personalized phage combinations faster.
Built With
- github
- matplotlib
- numpy
- openpyxl
- pandas
- python
- scikit-learn
Log in or sign up for Devpost to join the conversation.