LesionLens
Decision support for a doctor's second look. LesionLens is an AI tool that flags MS-typical vs. atypical white-matter lesions on FLAIR MRI slices, validated on MRIs from a hospital it never trained on.
Inspiration
Multiple sclerosis is notoriously hard to diagnose from an MRI alone, and the mistake is common. A ScienceDirect study found that nearly 18% of patients referred to two major MS centers had actually been misdiagnosed with the disease. A separate study led by Dr. Andrew Solomon found that 72% of misdiagnosed patients were put on MS therapies they didn't need, and 33% of them stayed misdiagnosed for a decade or longer before anyone caught it.
The most common correct diagnosis hiding behind an MS mislabel is migraine, a condition that disproportionately affects women, and whose MRI lesions can look similar to MS-typical ones on a radiologist's quick visual read.
The root cause here isn't doctors being careless. It's that judging whether a white-matter lesion looks like MS is currently a completely subjective call with no quantified baseline. That's exactly where the 18% comes from, and that's where LesionLens comes in. We wanted to give clinicians something they don't have today: a fast, repeatable second opinion at the precise decision point where that number gets created.
What it does
LesionLens is a web app for clinicians. Upload a FLAIR MRI slice, and it returns:
- An annotated image with every white-matter lesion boxed
- A per-lesion flag: MS-typical pattern vs. atypical/nonspecific pattern
- Lesion count and total area (a burden score)
- A plain-language summary of the findings, generated by Gemini, written into the "findings note" of the results
- A precision/false-positive readout, including performance on a hospital dataset the model never saw during training
It's designed to sit alongside the McDonald criteria a neurologist already uses, not replace their judgment. It is never shown to the patient, and it is never a standalone diagnosis; it's a second look.
Key Features
- RF-DETR-B object detector, fine-tuned on real lesion masks that box every lesion on a FLAIR slice
- A heuristic pattern-scoring layer that encodes established radiological criteria: Dawson's fingers elongation/orientation, periventricular vs. subcortical location, and lesion shape to flag MS-typical vs. atypical
- Burden scoring: lesion count + total area per scan, trackable over time
- Cross-dataset validation: trained on one dataset (MS3SEG), scored (never trained) on a completely separate multi-hospital dataset (MSLesSeg)
- Gemini-generated plain-language explanation for every result
- MongoDB Atlas case history so a clinician can revisit a patient's past scans
How we built it
Flow: Upload → RF-DETR-B inference (boxes + confidence) → heuristic pattern scoring per lesion → burden metrics → Gemini plain-language summary → case saved to MongoDB → React frontend renders the annotated image, burden score, explanation, and a validation panel.
Data. We sourced lesion masks from two public datasets: MS3SEG for training and MSLesSeg for validation only (115 scans across multiple hospitals, sealed away from training entirely). MS3SEG has ~2,000 annotated FLAIR slices available; we trained on a 649-image working subset (507 train / 97 valid / 45 test) to keep iteration fast within the hackathon window. Since we weren't replicating a dense multi-class segmentation model, we converted each dataset's masks into bounding boxes via connected-component extraction, ran everything through OpenCV preprocessing (normalization, CLAHE contrast, consistent resize/pad), and exported to COCO format through Roboflow.
Model. We fine-tuned RF-DETR from its pretrained checkpoint on Roboflow's hosted training rather than training from scratch. We chose this because it's the detection architecture the team already knew, which protected our timeline against the extra heuristic layer we were building on top of it.
Heuristic layer. On top of every detected lesion, we engineered a rule-based classifier using geometric features the radiology literature already uses to separate MS-typical lesions from vascular or nonspecific white-matter hyperintensities: elongation and orientation relative to the ventricles, periventricular vs. subcortical location, and ovoid shape.
Validation. A batch script runs the trained model plus the heuristic layer over the sealed MSLesSeg set and reports precision and per-scan false-positive rate on both held-out MS3SEG and MSLesSeg. This is our strongest evidence that this generalizes past the hospital it trained on.
Backend + AI. One FastAPI endpoint takes an uploaded scan, runs inference, and returns detections, flags, and burden metrics as JSON. That JSON feeds Gemini for the plain-language summary, and MongoDB Atlas stores case metadata, results, flags, and the explanation.
Deployment. The app runs live on a Vultr VPS, with nginx serving the frontend and reverse-proxying the FastAPI backend under systemd, a custom domain (lesionlens.health), and HTTPS via Let's Encrypt/certbot.
Frontend. React + Vite + TypeScript. An SVG overlay color-codes each box by pattern flag, a lesion table lets a clinician click a row to jump to it on the image.
Challenges we ran into
- Roboflow's free plan trains RF-DETR but doesn't allow downloading the weights, so every inference call is an HTTP request to their hosted API, meaning our live demo depends on the Wi-Fi where we are working.
- Neither of our datasets contains confirmed migraine-mimic cases, so there's no labeled data to train a model that tells MS apart from migraine directly. We had to be disciplined about the scope and never claim that our tool detects migraine, since that's a topic judges are very likely to ask about.
Accomplishments that we're proud of
- Shipping a full pipeline (data, detector, heuristic layer, backend, frontend) built as three parallel workstreams that met at a single integration contract, all within hackathon time
- Running real cross-hospital validation (train on one dataset, score on a completely separate one) instead of just reporting in-sample accuracy, which is the check most hackathon medical-imaging projects skip to take the easy way out
- Making every integration including inference, Gemini, and MongoDB degrade properly, so the whole app runs and demos cleanly even with zero API keys configured
- Shipping a real, publicly reachable deployment, not just a local demo, live on Vultr behind HTTPS at a real domain
What we learned
- The radiological features that separate an MS-typical lesion from a nonspecific one such as Dawson's fingers, periventricular vs. subcortical positioning, and lesion morphology and how to turn that domain knowledge into a rule-based classifier
- Why precision and per-scan false-positive rate are the best headline metrics for a clinical detection tool, not accuracy
- How to avoid overclaiming what our model can do. We drew a fine line between claiming that "we encode diagnostic criteria for MS-typical lesions" and "we detect migraine."
- How to keep three people working across three directories (data/, ml/, backend+frontend/) without breaking anything when it is time to integrate
What's next for LesionLens
- Expand our input to read in an entire MRI series rather than standalone images, and change our intelligence to factor in entire series rather than detections in one brain
- Validate across more hospitals and scanners to get a real sense of how the pattern flag holds up outside two research datasets
- Partner with an MS clinic to see the tool in action, and to scope what it would actually take before ever suggesting it distinguishes MS from migraine specifically
- Surface burden trend alerts per patient on the case view itself. Right now a clinician has to open History and notice a rising trend; the next step is flagging it to them automatically.
- Move inference off a hosted API and onto downloadable weights, so the tool doesn't depend on an internet connection in a clinical setting

Log in or sign up for Devpost to join the conversation.