Inspiration

A discharge summary is written by a clinician, for clinicians. The patient gets handed a copy as a side effect.

"Patient presented with acute exacerbation of CHF, diuresed with IV furosemide, discharge on furosemide 40mg PO daily, hold if SBP <100, f/u cardiology 7–10 days."

Nobody outside medicine can act on that. So the medication gets taken wrong, the warning signs get missed, the follow-up never happens — and the patient is readmitted. Hospitals are financially penalised for exactly that readmission.

The obvious move is "ask an LLM to simplify it." That's also the dangerous move: a fabricated dose in a patient handout is worse than no handout. So the interesting part of this project was never the translation. It was the proof.

What it does

Discharge summary in → plain-language patient page out, plus evidence that nothing was invented.

The output page has four sections:

  • What happened
  • Your medicines
  • Warning signs — split into call the clinic vs go to the ER
  • Follow-up

Every item is traced to a verbatim source_line from the original note. The UI shows that trace alongside the plain-language text, and it shows the agent's run timeline — extract → review sources → draft → verify → render — so a reviewer can see what the agent did, not just what it produced.

When the source can't be safely translated, it stops. It does not guess.

Positioning, stated plainly: a nurse or discharge coordinator runs this, reviews the output, and hands it to the patient. Clinician-side, not patient-side. The agent makes no clinical decisions — the treatment plan already exists, and is being translated, not authored.

All data in the demo is synthetic. No real patient data was used.

How I built it

One agent. No sub-agents — the branching is genuinely a graph, and sub-agents would have added orchestration bugs without adding capability. "I didn't need the complexity" was a deliberate call.

The LangGraph pipeline:

discharge summary
      │
      ▼
  [extract] ─────────► grounded JSON, verbatim source_line on every item
      │
      ▼
[review_sources] ────► deterministic safety rules on the extracted evidence
      │                   contradictory doses · non-positive doses ·
      │                   unsupported schedule vocabulary
      ▼
   [draft] ──────────► plain-language patient page
      │
      ▼
  [verify] ──────────► every claim rechecked against the extracted JSON
      │                   fails? rewrite once, then re-verify the whole page
      ▼
   clean ──► patient page        still unverifiable, or a source conflict
                                            │
                                            ▼
                                ESCALATE — stop, show the evidence,
                                           flag for human review

Two design decisions that carry the whole project:

Extraction runs exactly once. A rewrite cannot quietly change the evidence it is being judged against. Without this, the verifier can be talked into agreeing with itself.

Verification re-checks the entire page after a repair, including sentences kept from the first draft — not just the sentence that failed.

Stack: Python · LangGraph · Anthropic SDK (claude-sonnet-5) · FastAPI · React 19 + Vite + Tailwind · pytest. No database, no auth — it runs from a laptop.

Built solo, in gated chunks: Claude Code built one tool at a time, and each chunk was handed to a second model with one narrow question — "what breaks this?" — which generated adversarial inputs and ran them, rather than reading the code. Only demo-breaking findings got fixed.

Challenges I ran into

Making the safety check strict without making it cry wolf. The clean test case moves a patient from IV furosemide 40 mg BID to furosemide 40 mg PO daily. That's a normal route change, not a contradiction — but a naive "same drug, two mentions" check escalates on it, and would escalate on nearly every real discharge summary. A safety net that fires constantly gets switched off. Getting case 1 to run clean while case 2 still escalates was the hardest single constraint in the build.

Deciding what the agent is not allowed to infer. Alendronate 70 mg PO weekly is a perfectly valid prescription, but "weekly" sits outside the formatter's supported schedule vocabulary. The tempting fix is to have the model phrase it. The correct fix is to escalate — because the failure mode of guessing patient-facing wording is silent and unbounded.

Accomplishments that I'm proud of

The escalation path. It's roughly ten lines of logic and it's the most grown-up part of the demo — it's the part that says this was built by someone who thought about being wrong.

Concretely, on the seeded-flaw case the same drug appears at two different doses for the same phase of care — 40 mg in the medication list, 20 mg in the instructions. Nothing in the note says which is right. The agent doesn't pick one. It stops and asks for a human.

What I learned

In a clinical context the verifier is the product. Plain-language rewriting is the easy half and largely solved; the half that decides whether anyone can actually deploy this is the traceability and the refusal to guess. I spent the build accordingly.

What's next for ClearDischarge

Deliberately out of scope today, and named as such rather than omitted:

  • Plain-language drug information from openFDA, held to the same verification standard
  • Readability scoring against a target reading level
  • Translation into the patient's own language
  • Memory across a patient's admissions
  • Push into the patient portal, and follow-up appointment booking
  • A clinician review queue, so escalations land somewhere rather than just stopping

Built With

Share this project:

Updates

Submission history