Inspiration

A vessel is more than a bright shape in a CT scan. Where does it actually join the aorta? Is it one common trunk or two separate openings? Does the proposed branch continue into a supported lumen, or follow an artifact?

For the Toralis Labs challenge, we wanted an answer that could be inspected. Our goal was to connect automatic branch discovery to the image evidence behind each prediction, on the kind of offline CPU laptop the challenge requires.

What it does

Branchseed accepts a CT volume and a mask of the parent aorta, discovers a variable number of candidate direct daughters, and exports an instance for each accepted prediction:

  • its opening at the aortic wall;
  • a seed 5 mm along the proximal vessel;
  • the estimated local seed radius;
  • a unit direction vector and a parent link.

Our Aorta Explorer makes those outputs tangible. Orbit the parent vessel, select a branch, inspect linked CT planes, see where origins sit on the wall map, and travel through the reconstructed lumen. Export the same prediction as machine-readable JSON. The interface keeps the scan, geometry and measurements connected rather than leaving a reviewer with an isolated score.

How we built it

The submitted detector combines classical image processing and geometry with a small candidate filter. We validate physical coordinates, crop around the parent, work on a 1 mm grid, estimate scan-relative contrast, enhance tubular structures at multiple scales, propose wall contacts and trace supported proximal paths. Two proposal passes run the same detector, a strict profile and a looser review profile, both at native contrast scale 0.9. Connection, length, size and shared-path checks reject crop caps and duplicate openings, and the origin-diameter policy uses the judge's 2 mm minimum with an explicit allowance for native-voxel uncertainty.

Each surviving candidate gets 13 geometric and intensity features and a score from a small logistic model fitted on synthetic vascular phantoms. Candidates scoring below 0.15 are dropped, strict survivors are kept first, and review survivors within 3 mm of a kept origin are merged. Scoring before merging stops a weak review copy from displacing a strong strict one.

Python, SimpleITK, NumPy, SciPy and scikit-image power the detector. TypeScript and Three.js power the Explorer; Vite builds its local assets. The Python server serves the built website and analysis API. The judge package ships the detector and its bundled model with ten pinned wheels for offline Windows installation. After setup, neither inference nor the Explorer needs a hosted model, API key, CDN or GPU.

We also built a research pipeline: synthetic vascular phantoms with analytic references, candidate-feature classifiers, physical CT patches, a small CNN, model blends and case-separated selection. Only the logistic candidate filter made it into the submission; the rest informed the decision.

Challenges we ran into

Direct connectivity is harder than finding bright voxels. Thin or curved vessels can lose contrast through partial volume, nearby structures can look tubular, and relaxing a connection check can turn a missed branch into several false positives. One opening must remain one instance even when its trunk splits downstream.

The reference data also required care. The judge approved 19 AI-assisted targets across five development cases, but the annotations may omit eligible branches. We had already inspected these cases. Repeatedly choosing the best score would not establish performance on new patients, so we tracked exposure, replayed frozen evidence and kept real-reference agreement separate from synthetic regression results.

Accomplishments that we're proud of

We built a complete path from new CT/mask inputs to valid physical-coordinate JSON and an interactive evidence viewer. The packaged judge application completed all 25 supplied scans with network access denied, averaging 11.6 seconds per case, with a 52.7-second maximum and 1499 MiB maximum sampled process-tree RSS under four-core Linux affinity. These are development measurements; organizer-Windows timing remains to be measured.

We also made the performance claims inspectable:

Evidence Submitted result What it measures
Five reused reference cases, 19 targets Precision 0.778, recall 0.737, F1 0.757; count MAE 1.8 Local one-to-one ostium agreement at 3 mm
Twenty-four synthetic topology cases 47 TP / 11 FP / 2 FN, F1 0.879 Procedural topology regression at 3 mm
Twenty-five supplied scans 25 completed, byte-identical to a full-checkout replay Execution and output coverage, not detection accuracy

The local reference result still includes 5 missed targets and 4 unmatched predictions, and the same five cases chose the 0.15 threshold. Synthetic F1 is not real-patient accuracy: the filter admits eleven synthetic false positives, four of them in negative controls, where the strict baseline had none. Neither result establishes hidden-test or clinical performance. We would rather show the evidence and its limits than attach an unsupported "accuracy" percentage to the product.

What we learned

A more complicated model is useful only if its evidence supports the change. Our frozen selection replay covered 280 variants, with 180 eligible for the case-separated comparison. The retrospective random-forest winner scored higher on the reused cases but lacked independent promotion evidence, and the case-separated selection composite was less reliable than fixed strict.

Order matters as much as the model. Merging proposals before scoring let a low-scoring review copy displace a strong strict one; scoring first and merging after raised local F1 from 0.533 to 0.757 and lowered daughter-count error from 2.0 to 1.8. A separate guarded connection recovery reached F1 0.581 but worsened count error to 2.2, so it stays an explicit strict-only option. The unfiltered strict baseline remains one flag away.

What's next for Branchseed — Aorta Explorer

The next step is complete expert review of missed origins and unlisted predictions, followed by evaluation on previously unused patients with different contrast and slice thickness. Hardening the candidate filter against the synthetic false positives is the next engineering target, alongside better candidate generation, curved-wall recovery and origin-size calibration. We also need measurements on the organizer's actual Windows laptop.

Branchseed is a research prototype for branch discovery and visual verification. It does not segment the parent automatically, reconstruct the full distal tree or provide a clinically validated diagnosis.

Credits: Toralis Labs for the challenge and the supplied CT/mask data; SimpleITK, NumPy, SciPy, scikit-image, Three.js, Vite and the other open-source libraries we depend on; the Manrope and Fraunces fonts under their bundled licences; Devin and Claude Code for AI-assisted development. The showcase uses recorded interface footage and an original synthesized soundtrack.

Built With

Share this project:

Updates

Submission history