Inspiration
Over 23 million people have taken direct-to-consumer DNA tests, yet the VCF files they receive are unreadable without a bioinformatician. Genetic counselors cost $300+/session and there are fewer than 8,000 certified ones for 330M Americans. Meanwhile, 36% of all ClinVar variants are still classified as "Uncertain Significance" — the most unhelpful answer a patient can get. We wanted to fix that, entirely offline, with no cloud dependency or API key required.
What it does
GeneOracle takes a raw VCF file (from 23andMe, Illumina DRAGEN, or GATK) and produces a plain-language, clinically-grounded genomic risk report across 6 disease modules — all offline on a standard laptop. It classifies single variants using ClinVar + AlphaMissense, computes Polygenic Risk Scores using the PGS Catalog, flags pharmacogenomic drug interactions via PharmCAT, and generates a tiered narrative report using a local RAG-grounded Phi-3-mini LLM. Every claim in the report traces back to a cited database.
How we built it
We built a 5-layer modular pipeline:
- L1 – VCF ingestion and normalization (cyvcf2, PyVCF3)
- L2 – Offline ClinVar + gnomAD SQLite lookup for single-variant pathogenicity
- L3 – VUS resolution using locally-cached AlphaMissense scores (~216M variants) + ML-based ACMG classification (scikit-learn)
- L4 – Multi-disease Polygenic Risk Scoring via PGS Catalog + pgsc_calc
- PharmCAT for PGx star alleles
- L5 – RAG-grounded narrative generation with Phi-3-mini GGUF running via llama.cpp, using ChromaDB as the local vector store
Challenges we ran into
Preparing 14GB dataset was the biggest challenge.
Accomplishments that we're proud of
- First fully offline tool combining single-variant pathogenicity + multi-disease PRS + PGx in one pipeline
- Achieving <5 second lookup time on AlphaMissense despite 216M variant entries, using indexed SQLite
- RAG architecture validated by Yale/Microsoft research (Bioinformatics Advances 2025) - outperforms fine-tuning for genomics at a fraction of the cost
- Report output passes a "no bioinformatics background needed" readability test — Executive Summary, Clinical Detail, and Screening Recommendations as separate tiers
What we learned
- RAG over large genomic databases is more practical and accurate than fine-tuning for factual variant interpretation (Lu & Cosgun, 2025)
- Multi-disease PRS ensembles outperform single-disease PRS by 10–20% AUC (Bedő et al., PLOS Genetics 2023) — worth the extra compute
- Offline-first architecture forces you to be extremely deliberate about data pipelines; every dependency needs a local fallback
- Medical AI needs citation trails, not just answers — users trust reports more when every claim links to ClinVar, gnomAD, or a published score
What's next for GeneOracle
- Expand from 6 to 20+ disease modules using the full PGS Catalog
- Add whole-exome sequencing (WES) support alongside consumer SNP arrays
- Replace Phi-3-mini with a larger local model (Mistral 7B) for richer narratives
- Build a clinician portal with FHIR export for EHR integration
- Add longitudinal tracking — re-run reports as new ClinVar evidence emerges for previously uncertain variants
Log in or sign up for Devpost to join the conversation.