Inspiration
Two years ago, my uncle passed away from a resistant bacterial blood infection. Finding an effective treatment required culturing the bacteria in a lab, which is a process that takes days. He simply didn't have that time.
His story is tragically common. Even in top hospitals, definitive results take 48 to 72 hours, and in much of the world, testing isn't an option at all. As a result, doctors are forced to guess and end up prescribing the wrong initial antibiotic one-third of the time. That single misstep doubles a patient's odds of death and accelerates global antibiotic resistance.
What it does
DNAgen takes a bacterial whole-genome sequence from a bloodstream infection and predicts which antibiotics are likely to work against that specific strain.
For each antibiotic, the AI model predicts the minimum inhibitory concentration (MIC), which is the minimum concentration of the antibiotic to be considered effective against the bacteria, and converts it into a probability of effectiveness. It then ranks the antibiotics and shows which treatments are likely to work and which are likely to fail.
How we built it
We built DNAgen around a multi-task multilayer perceptron trained on public lab data from two databases, BV-BRC and NCBI: roughly 28,000 bacterial genomes across five common bloodstream pathogens:
- E. coli,
- K. pneumoniae,
- S. aureus,
- P. aeruginosa,
- A. baumannii
and 37 antibiotics, totaling about 260,000 susceptibility measurements for training and 47,000 held out.
Our pipeline first confirms the species with Mash, then runs NCBI's AMRFinderPlus to extract resistance related features from each genome, such as mutated genes or genes that express resistance. About 700 resistance gene families and 2,300 resistance point mutations were founded. Since it is a large dataset, we ran the data processing pipeline on an AWS VM and stored the dataset in an S3 bucket.
We then form a sparse input of about 3,000 columns, combined with a learned 16-dimensional species embedding vector. We split the data by genetic cluster, keeping near-identical germ line families together to prevent overfitting, so DNAgen is evaluated on strains it has not seen during training. Strain type, country and year are never used as inputs, so the model learns resistance mechanisms and not outbreaks.
Because the minimum amount of antibiotic effective against that bacteria (mic) labels used in the database are in mg/L base 2 numbers, and the mic levels of the same datapoints can sometimes vary, susceptibility labels are usually limits and not exact values. We formatted the labels to be MIC interval and train with an interval-censored negative log-likelihood. This lets DNAgen learn directly from bounded MIC measurements without imputing a single value.
We used an multilayer perceptron as our model. The network is small (376,000 parameters): a shared trunk of two 256-unit layers feeds 37 output heads, one per antibiotic, so drugs with few measurements borrow strength from drugs with many. We trained with AdamW, dropout and early stopping across 5 cross-validation folds. Validation loss flattens after about 30 epochs and converges in a nice exponential decay graph with very small gap between training and testing loss.
The model then predicts an MIC distribution for each antibiotic, which we compare with the clinical breakpoint (CLSI) to get a probability of susceptibility.
We also used neon as a database to store uploaded genome sequencing information. With the user’s consent, we would also be able to query the data from that backend and iteratively improve the model.
Challenges we ran into
Technical understanding of the concept, the dataset, and designing how the data changes through the pipeline before feeding into the model was a challenge for us. Since there were many antibiotic resistant genes (AMR genes) that can be detected, it took us a while to figure out the method of use spatial embeddings to represent those genes before feeding into the MLP.
Some of our teammates don’t have experience with biology, so we had to take some time learning the biology part of the project in order to deeply understand before starting the project.
Dataset generation took a very long time since while scanning for the AMR genes on a large dataset. We were able to find a dataset from NCBI which already comes with a pre-labelled AMR gene. We pivoted to that dataset instead which saved us some time.
We were deciding between XGBoost model and a simple MLP model. We trained with the XGBoost model first and it had an on par result with the MLP. The only problem is it requires training on multiple models because since the dataset was missing pieces of labels, we have to train multiple models based on the number of labels.
Accomplishments that we're proud of
Building a model that actually knows when it is uncertain:
We're really proud that our model doesn't just blindly guess. On genetic clusters it had never seen before, DNAgen matches lab results 95.5–98.5% of the time when it has a high confidence (over 90% probability). The thing we cared about most was making sure it rarely calls a resistant strain "susceptible", which happens only 0.8–1.4% of the time, since that's the error that could actually hurt a patient. If the evidence is weak, it just says "uncertain".
Certain probability prediction:
When the model reports a 95% probability, it's actually right about 94% of the time on our held-out data for E. coli, K. pneumoniae, and S. aureus. It also successfully ranks a working drug above a failing one about 96% of the time for K. pneumoniae.
Good validation loss:
Our validation loss dropped from around 2.0 to 0.98, flattened at about 30 epochs, and early stopping kicks in before it memorizes anything. The training loss and test loss follows an exponential decay, and barely has a small gap between the two, suggesting the model has a good fit. Plus, the whole model trains in just six minutes on a regular laptop CPU. (Most of the time is processing and downloading the data).
Building it all from scratch in a weekend:
We built the data pipeline, the model, the loss function, the calibration, and the serving layer completely from zero under 24 hrs!!!
Learning the biology as we went: Several of us came into this with no biology background, so we had to spend a lot of time learning the science just to understand what we were building.
What we learned
1) Learned the biological concepts and integrating those concepts with AI + software engineering in under 24 hours. 2) Always make sure to check if there are pre-existing datasets that would save time on data processing. 3) It is always good to nail down the concept and understand what you’re building before actually building. Have a solid pipeline and workflow ready, have the math ready, and building would be way faster.
What's next for DNAgen
We want to implement this software with an Oxford Nanopore handheld sequencing machine to be used in rural / field hospitals without wet labs. This would reduce both the cost and time for saving a patient’s life. We also intend to train a sequencing genome predictor and predict while the machine is sequencing. This would include predicting the genome sequence that includes the mutation and the AMR genes which would yield a faster result.
Built With
- amazon-web-services
- fastapi
- neon
- python
- pytorch
- s3
- tailwind
Log in or sign up for Devpost to join the conversation.