Inspiration

Over 23 million people have taken direct-to-consumer DNA tests, yet the VCF files they receive are unreadable without a bioinformatician. Genetic counselors cost $300+/session and there are fewer than 8,000 certified ones for 330M Americans. Meanwhile, 36% of all ClinVar variants are still classified as "Uncertain Significance" — the most unhelpful answer a patient can get. We wanted to fix that, entirely offline, with no cloud dependency or API key required.

What it does

GeneOracle takes a raw VCF file (from 23andMe, Illumina DRAGEN, or GATK) and produces a plain-language, clinically-grounded genomic risk report across 6 disease modules — all offline on a standard laptop. It classifies single variants using ClinVar + AlphaMissense, computes Polygenic Risk Scores using the PGS Catalog, flags pharmacogenomic drug interactions via PharmCAT, and generates a tiered narrative report using a local RAG-grounded Phi-3-mini LLM. Every claim in the report traces back to a cited database.

How we built it

We built a 5-layer modular pipeline:

  • L1 – VCF ingestion and normalization (cyvcf2, PyVCF3)
  • L2 – Offline ClinVar + gnomAD SQLite lookup for single-variant pathogenicity
  • L3 – VUS resolution using locally-cached AlphaMissense scores (~216M variants) + ML-based ACMG classification (scikit-learn)
  • L4 – Multi-disease Polygenic Risk Scoring via PGS Catalog + pgsc_calc
    • PharmCAT for PGx star alleles
  • L5 – RAG-grounded narrative generation with Phi-3-mini GGUF running via llama.cpp, using ChromaDB as the local vector store

Challenges we ran into

Preparing 14GB dataset was the biggest challenge.

Accomplishments that we're proud of

  • First fully offline tool combining single-variant pathogenicity + multi-disease PRS + PGx in one pipeline
  • Achieving <5 second lookup time on AlphaMissense despite 216M variant entries, using indexed SQLite
  • RAG architecture validated by Yale/Microsoft research (Bioinformatics Advances 2025) - outperforms fine-tuning for genomics at a fraction of the cost
  • Report output passes a "no bioinformatics background needed" readability test — Executive Summary, Clinical Detail, and Screening Recommendations as separate tiers

What we learned

  • RAG over large genomic databases is more practical and accurate than fine-tuning for factual variant interpretation (Lu & Cosgun, 2025)
  • Multi-disease PRS ensembles outperform single-disease PRS by 10–20% AUC (Bedő et al., PLOS Genetics 2023) — worth the extra compute
  • Offline-first architecture forces you to be extremely deliberate about data pipelines; every dependency needs a local fallback
  • Medical AI needs citation trails, not just answers — users trust reports more when every claim links to ClinVar, gnomAD, or a published score

What's next for GeneOracle

  • Expand from 6 to 20+ disease modules using the full PGS Catalog
  • Add whole-exome sequencing (WES) support alongside consumer SNP arrays
  • Replace Phi-3-mini with a larger local model (Mistral 7B) for richer narratives
  • Build a clinician portal with FHIR export for EHR integration
  • Add longitudinal tracking — re-run reports as new ClinVar evidence emerges for previously uncertain variants

Built With

  • alphamissense
  • catalog
  • chromadb
  • clinvar
  • clinvcf
  • cyvcf2
  • gnomad
  • java
  • llama.cop
  • pandas
  • pgs
  • pgsc-calc
  • pharmacat
  • pharmgkb
  • phi-3-mini
  • plink2
  • polars
  • python
  • pyvcf3
  • rag
  • reportlab
  • scikit-learn
  • sentence-transformers
  • snpeff
  • sql
  • streamlit
Share this project:

Updates

Submission history