Inspiration

  • Biopsy slides are read in arrival order, not urgency order, so the urgent one can wait days in the stack.
  • US pathologists fell ~17% (2007–2017) while caseload per pathologist rose ~41%.
  • Cloud AI is a poor fit: slides are 0.3–4 GB each, patient data is sensitive, and maintaining GPU servers is unfeasable for this function.
  • The ASUS GX10 (128 GB unified memory) can run a vision model and an LLM side by side, locally.

What it does

  • Scores every tissue tile of every slide for tumor and builds a heatmap.
  • Measures each suspicious region and ranks the tray by clinical urgency.
  • Nemotron, running locally, drafts the report and answers questions, with every number checked.
  • Live pipeline view, gigapixel viewer, and an overlay of the pathologist's own tumor outline.

How we built it

  • Vision: 96 px tiles at 10× → UNI2-h (681M-param pathology foundation model, frozen) → 1,536-number feature → logistic regression trained on 30k PCam tiles.

  • Regions: tiles above 0.5 grouped; a region needs ≥ 2 tiles and one ≥ 0.9. Area of $n$ tiles:

  • Urgency: category from the largest region's size (> 2 mm macro, > 0.2 mm micro); score rises with size on a log scale.

  • LLM: nemotron3:33b via Ollama. Code writes every number; the model writes only prose; replies with unsupported numbers are rejected.

  • App: FastAPI + OpenSlide + OpenSeadragon; live progress events; all assets bundled locally, so it works offline.

  • Speed: ~135 tiles/s, ~100 slides/hour; UNI2-h (~6 GB) and Nemotron (~33 GB) loaded together.

Challenges

  • Moving Model Weights and Test Data: the GX10 downloaded the public data itself; we worked remotely over Tailscale.
  • LLM issues: empty replies (hidden reasoning) and made-up counts → reasoning off plus number checks.
  • Alignment: heatmaps, outlines and slide tiles all had to share full-resolution coordinates.
  • Tuning: one-tile false spots and bunched urgency scores → two-level thresholds plus size-based scoring.

What we learned

  • A foundation model plus a tiny classifier is powerful, but the real work is the whole-slide pipeline around it.
  • Medical LLMs need guardrails: code supplies the facts, the model writes the prose.
  • Unified memory makes a fully local, multi-model pathology box practical.
  • Keep a truly held-out test set, and state caveats up front.

What's next

  • Validate on slides from other hospitals; add more tissue types.
  • Connect directly to slide scanners and lab systems.

Built With

Share this project:

Updates

Submission history