Inspiration
When radiologists read a follow-up chest X-ray, the prior film and its report are usually open beside it. If that report says "No pleural effusion," a slightly blunted costophrenic angle on today's film is easy to dismiss. The Kim-Mansfield classification of radiology errors names both failures: type 7, reading the current film without properly comparing it to the prior, and type 12, satisfaction of report, where an earlier report steers the read.
Most AI tools take the images and the report together, so the model can be anchored the same way. We wanted a system that commits to what it sees in the pixels before anything written can influence it, then sends any conflict to a person. We called our product Pleural Sight, since the job is comparing two (plural) moments and providing insight on pleural effusion.
What it does
Pleural Sight is a second-reader worklist for follow-up chest X-rays. Each case is an earlier film, a current film, and optionally a report for each.
- A plain-Python gate rejects pairs that can't be compared: byte-identical files, and AP vs. PA projections (portable AP films spread fluid out and magnify the heart). These return Cannot compare without calling a model.
- A Blind Reader gets the two films as FILM 1 and FILM 2, with no dates and no report, and labels each one present, absent, uncertain, or not assessable, with evidence.
- A Report Reader, which never sees the images, extracts one effusion claim per report. Each supporting quote must appear verbatim in the report.
- If a stable, definite blind read contradicts a report, an Investigator takes one targeted, still report-blind second look at the disputed film.
- Ordinary code assigns the status, and the worklist puts the cases that most need a person first.
- A clinician approves, overrides with a required reason, or escalates. Every action is logged.
The worklist order is:
$$ \textit{Disagreement} \prec \textit{Cannot compare} \prec \textit{Unstable read} \prec \textit{Image only} \prec \textit{Agrees with report} $$
The viewer syncs both studies in side-by-side, flicker and swipe modes. It shows a plain-language conclusion such as "Possible new fluid in the current study," and keeps the evidence and the agent trace in expandable panels. We left out heatmaps and confidence percentages because nothing calibrated sits behind them.
How we built it
The backend is FastAPI, with SQLite for intake and the audit log. OpenSwarm launches a fresh agent for each role, and each agent gets a run-specific MCP connector that exposes exactly one tool with no arguments. The tools call Gemini over HTTP with strict JSON schemas validated by Pydantic. The frontend is vanilla JS and CSS with no third-party scripts. Images come from the public NIH ChestX-ray14 dataset; a script downloads only the 37 films we use.
upload ─► comparability gate (Python, no model)
├─► Blind Reader ─► read_pair ─► Gemini (films only, read twice)
└─► Report Reader ─► read_reports ─► Gemini (text only, quotes checked)
│ contradiction?
└─► Investigator ─► reassess ─► Gemini (once per run)
─► verdict rules (Python) ─► worklist ─► clinician sign-off
Keeping the reader blind
Agents never touch the data. Each run writes a manifest with the image paths, report text and gate result, and only the MCP process reads it. The Blind Reader's whole job is to call read_pair({}). The tool sends the image bytes to Gemini and returns a summary such as
{"reading_a": "new", "reading_b": "new", "order_swap_consistent": true, "stored": true}
so image bytes and report text never enter an agent prompt. The server rejects any arguments and any call from the wrong role, so an injected "also call read_reports" fails. Agents share results through a SQLite stages table, and each connector is deleted when its agent finishes.
Testing for position bias
Vision models sometimes answer differently depending on which image comes first, so the Blind Reader reads every pair twice. Let \( f(x_1, x_2) \to (s_1, s_2) \) be one read of two films:
$$ R_A = f(\text{prior}, \text{current}), \qquad R_B = \sigma\big(f(\text{current}, \text{prior})\big) $$
where \( \sigma \) maps the second read back to prior/current positions. The read is stable only if
$$ R_A^{\text{prior}} = R_B^{\text{prior}} \quad\wedge\quad R_A^{\text{current}} = R_B^{\text{current}}. $$
Otherwise the case is marked Unstable read; we never average the two. A stable read becomes a transition label:
$$ T(s_p, s_c) = \begin{cases} \text{absent} & (\text{absent}, \text{absent}) \\ \text{new} & (\text{absent}, \text{present}) \\ \text{resolved} & (\text{present}, \text{absent}) \\ \text{persistent} & (\text{present}, \text{present}) \\ \text{indeterminate} & \text{otherwise} \end{cases} $$
Models observe, code decides
For a report claim \( c \) about study \( k \), with blind state \( v_k \):
$$ \text{verdict}(c) = \begin{cases} \text{no relevant claim} & c = \text{no relevant claim} \\ \text{uncertainty} & v_k \notin \lbrace\text{present},\text{absent}\rbrace \ \lor\ c = \text{uncertain} \\ \text{agreement} & v_k = c \\ \text{contradiction} & \text{otherwise} \end{cases} $$
The Investigator runs only after a contradiction from a stable, definite read. The MCP server checks that condition again itself, and a primary key on the stage name allows reassess to run once per case.
One pipeline serves two engines. The OpenSwarm engine launches real agents; the headless engine calls the same tools in-process, for batch evaluation or when OpenSwarm isn't running. Both return the same result shape.
For evaluation we wrote three single-call baselines to run on a 19-pair NIH pilot set (16 PA-to-PA pairs, four per transition, plus 3 AP-vs-PA): a naive prompt, a careful prompt that may abstain, and one that also sees the current report, to measure anchoring directly.
Challenges we ran into
Our first full runs failed because agents finished without working. OpenSwarm sessions reported completed, but no tool result had been stored; one agent described what it planned to do and stopped. We only accept a stored MCP result as proof of work, so the run failed instead of using a chat answer. We tightened the role prompt, set the one tool to always allow, and now check at discovery time that each connector exposes exactly that tool.
The free Gemini tier allows 20 requests per model per day. An image-only case uses 2 requests, a case with reports uses 3 to 5, and the evaluation needs about 100. That leaves
$$ \left\lfloor \frac{20}{2} \right\rfloor = 10 \text{ image-only cases per model per day} $$
before anyone rehearses the demo. After an HTTP 429 in the middle of a test, we added a fallback chain across six models (each has its own quota), a rotating key pool, a response cache keyed on a SHA-256 of the request, and error messages that say how to fix the problem. A consumer Google AI Pro subscription doesn't raise API limits.
The Report Reader kept treating the earlier study's report as history and returning no claim. We had to tell it that what a report says about its own study is a claim for that study, and that only mentions of even older exams are history.
A review caught the UI saying "Images and reports agree" whenever nothing was contradicted, including hedged reports and reports that never mentioned fluid. The summary now comes from every claim's verdict, and uncertainty or silence shows a neutral state that asks for review.
Cancelling during an OpenSwarm request could still create a connector or start an agent on the server. Every create and launch call is now wrapped in asyncio.shield: on cancel we wait for the response, delete what was created and stop any session. A marker file makes the tools reject late results.
Three of us built the workspace UI, the backend and the animated landing page in parallel, and merged them into one app overnight.
NIH labels are mined from reports by NLP, so they carry the same anchoring bias we were trying to catch. We show them as reference labels and never put them in a prompt.
Accomplishments that we're proud of
- The data boundary lives in code. The image reader never receives report text, and each agent can reach only one tool, with no arguments.
- Position bias shows up as an Unstable read status a clinician can see.
- Every report quote is checked character by character against the source, and a mismatch fails the run.
- The gate refuses AP-vs-PA pairs before any AI runs. "These X-rays were taken from different sides, so we won't compare them" is our favorite moment in the demo.
- Failures such as an exhausted quota, an agent that skipped its tool, a scanned PDF or a 64-pixel image produce a specific message. Our design rules ban fake progress bars and invented percentages.
- The test kit has seven demo folders, including one of files that should be rejected, plus five cases aimed at hard report reading: fluid that resolves while other changes remain, a hedging report, a report that never mentions fluid, and a current report that mentions a past negative.
What we learned
- Two narrow readers that can't see each other's inputs, joined by a few lines of
ifstatements, were easier to debug and audit than one prompt that sees everything. - Once Gemini only had to fill a strict schema, deciding what counts as a disagreement became ordinary, unit-tested code.
- An agent saying it finished means nothing until its tool has stored a result.
- With \( n = 19 \), a 95% interval on an accuracy near 50% has a half-width of \( 1.96\sqrt{0.25/19} \approx 0.22 \). The pilot shows how the system behaves; it can't measure accuracy.
- Caching, fallbacks, and replays decided whether we could demo at all.
- "No fluid detected in either study" belongs in the headline.
absent_bothbelongs one click deeper.
What's next for Pleural Sight
- A held-out, clinician-annotated evaluation with patient-level separation, reporting image sensitivity and specificity, report extraction accuracy and workflow completion separately, with confidence intervals.
- Suitability checks that send non-chest images and low-quality films to not assessable instead of producing an effusion call.
- The same blind-read, report-read, compare pattern for pneumothorax, cardiomegaly, and line or tube position.
- DICOM and PACS integration, so view, chronology and patient identity come from metadata instead of checkboxes.
- Authenticated reviewer identity on sign-off, and OCR for scanned reports.
- Calibration, so any number we eventually show means something.
Research prototype. Not for clinical use. Every flag needs clinician review.
Log in or sign up for Devpost to join the conversation.