-
Clinical Assessment: entering SOFA inputs for a septic shock case, live five-agent pipeline on the right.
-
Sepsis Bundle: five agents converge on one line — Physician Decides. AI proposes, physician confirms.
-
Device exposure, risk modifiers, and allergies feed the governed assessment — one button, physician still decides.
-
Governed result: SOFA 16, antibiotic class + dose proposed, governance PASS — all computed independently.
-
Plain-language rationale for the antibiotic class, plus a 48h reassessment deadline — raw JSON stays one click away.
-
Culture result comes back — the de-escalation engine checks susceptibility and proposes the narrowest safe option.
-
Result: narrower option found — Ceftriaxone over Piperacillin-tazobactam, governance PASS, 48h reassessment set.
-
Live audit trail — every assessment and governance decision, timestamped and hash-chained, pulled straight from disk.
-
Aggregate, de-identified stats plus deterministic outbreak signals — no patient identifiers ever rendered here.
Inspiration
As a physician, I kept seeing the same gap: sepsis kills through delay. The Surviving Sepsis Campaign's Hour-1 Bundle exists precisely because every hour without appropriate antibiotics and fluid resuscitation raises mortality — yet in real ICU workflows, the SOFA score, the antibiotic decision, the culture follow-up, the guideline check, and the outbreak pattern across patients all live in different heads, different papers, different moments. I wanted these to converge into one governed channel instead of one overworked clinician juggling five mental checklists at once — without ever taking the final decision away from the physician.
What it does
Sepsis Bundle is a five-agent clinical decision support system for sepsis and septic shock:
- Sepsis Bundle Agent — computes the SOFA score deterministically from entered clinical values (respiratory, coagulation, liver, cardiovascular, CNS, renal) and tracks Hour-1 Bundle completion.
- Antibiotic Specialist Agent — proposes an empirical antibiotic class (not just a drug name) from a versioned knowledge base, with renal dose adjustment and MRSA/MDR/anaerobic/fungal risk modifiers. Once a culture result comes back, a deterministic de-escalation engine checks whether the current regimen is still covered, raises a resistant-organism alert, or identifies a genuinely narrower-spectrum option — ranked by an actual spectrum table, not just "any drug the lab marked sensitive."
- Guideline Surveillance Agent — watches for antibiotic guideline changes and raises them as pending reviews with an SLA. Nothing activates automatically; every change requires physician/pharmacist approval before it becomes the active knowledge base.
- Governance Agent — the safety boundary running underneath all four other agents. It independently recomputes every deterministic decision and compares it against what's about to be shown; any mismatch is blocked before it reaches the physician. Unknown findings fail closed by design — a property that caught my own bugs twice during development.
- Memory & Analytics Agent — aggregate, de-identified longitudinal statistics, plus an outbreak-style pattern layer: deterministic organism clustering and device-exposure association (e.g. "3 of 3 E. coli cases shared a urinary catheter >48h"), grounded directly in the peer-reviewed infection-control literature on outbreak detection (frequency-alert methods, their known specificity limits versus genomic typing) rather than invented from scratch. A bounded LLM narration layer summarizes these statistics in plain language for a physician/infection-control reviewer — it is only allowed to describe numbers that were actually computed, must always offer multiple possible explanations rather than settling on one (a documented failure mode in published LLM infection-control benchmarks), and must say so plainly when a sample is too small to trust.
Every case, every governance decision, and every guideline approval is written to an append-only, hash-chained audit trail. Real authentication (PBKDF2-hashed passwords, server-side sessions) protects every clinical endpoint — this isn't a public demo toy. The system's guiding line, shown on every screen: AI proposes, physician decides.
How I built it
FastAPI backend, deterministic Python rules engines for every clinically consequential decision (SOFA, antibiotic selection, de-escalation, governance validation) — Claude is used only for narration, describing a decision that was already made, never for making it. Vanilla JS/HTML frontend (no build step), deployed on Render and Vercel. 45 automated tests cover the deterministic core, and every feature in this write-up was additionally verified end-to-end against the live deployment, case by case, with hand-computed expected SOFA scores and antibiotic outputs checked against what the running system actually returned.
Challenges I ran into
The most valuable part of this build was the number of real safety bugs that surfaced under scrutiny rather than in production:
- A vasopressor dropdown that was hardcoded to "none" in the frontend, meaning the cardiovascular SOFA component could never score correctly for septic shock patients actually on vasopressors — the exact population the score matters most for.
- A urine-output field ambiguous between a 24h total and an hourly rate, silently misread as whichever the code assumed.
- A de-escalation governance rule that, on first pass, blocked a valid resistant-organism alert from ever reaching the physician, because the finding wasn't yet registered in the policy matrix — the system's own "fail closed on the unknown" design caught my own gap before a user ever would have.
- Early instinct to let the Memory Agent's LLM layer freely "guess" an infection source from statistics — reading the actual outbreak-detection literature first showed this exact failure mode already documented (a model jumping to one diagnosis from one data point without considering alternatives), and reshaped the design into multiple-hypothesis, explicitly-bounded narration instead.
Each of these taught the same lesson: in clinical software, a UI that looks complete and a UI that is safe are not the same claim, and the gap between them is exactly where governance has to live.
What's next
Wiring a licensed guideline source into the surveillance agent (currently a scaffold pending institutional licensing), pharmacist/ID-physician review of the antibiotic knowledge base's drug-class and spectrum-ranking scaffold, adding ward/unit-level data to sharpen the outbreak pattern signal (the single highest-impact addition per the infection-control literature), and a persistent-disk deployment so audit history and pattern data survive across releases in production.

Log in or sign up for Devpost to join the conversation.