Inspiration
Every year, emergency doctors face a race against time. Somewhere buried in a patient's history is the one fact that could change everything — a hidden allergy, a critical medication, a warning sign from years ago. But uncovering it can take fifteen minutes. Fifteen minutes a doctor doesn't have.
AI has been brought into healthcare before, and too often it's made things worse — generating answers that sound confident but were never actually in the patient's record. In medicine, a plausible lie is more dangerous than no answer at all.
We built GRUS to solve this properly: an emergency medicine AI agent that reads a patient's entire record and produces a verified clinical brief in seconds — where every claim traces back to a real source, and where missing information is reported honestly instead of guessed.
What it does
GRUS reads a patient's full medical history and generates a sourced, verified clinical brief in under ten seconds. Every fact links directly to the database row it came from. If something is missing — a vital sign, a lab result, a home medication — GRUS says so plainly instead of filling the gap.
Beyond summarization, GRUS:
- Runs fifteen validated clinical decision scores (PERC, HEART, qSOFA, KDIGO, and others) — computed deterministically, never guessed. If criteria are missing, it asks rather than assumes.
- Runs three machine learning risk models (acute kidney injury, transfusion likelihood, electrolyte crisis) trained on 20,000 MIMIC-IV admissions using XGBoost on SageMaker.
- Lets doctors ask direct questions through a grounded chat assistant that only answers from retrieved, cited facts.
- Supports live patient registration, with vitals and labs streamed in real time via EventBridge and Lambda — the same pipeline a real hospital monitor feed would use.
How we built it
Data layer. We used MIMIC-IV, a real de-identified clinical database of over 400,000 admissions. We explored the data first, narrowing it down through anticoagulation, trauma, and critical lab values to find cases that would genuinely test the system. An ETL pipeline loads this into Aurora PostgreSQL — and along the way, twelve distinct structural problems surfaced in the raw data (messy drug names, corrupt timestamps, repeated dose entries), each becoming a deliberate schema decision.
Retrieval. Clinical notes are split by clinical section — not by character count — so a medication list never gets cut mid-entry. Each section is embedded using Amazon Titan and stored directly in Aurora using the pgvector extension. We run both semantic and keyword search together: in one real case, the passage naming a patient's reversal agent scored just 0.185 on vector similarity — keyword search found it instantly. Hybrid retrieval is not a hedge here; it's necessary.
Agent orchestration. Five agents, built on the Strands Agents SDK and containerized for Amazon Bedrock AgentCore, run as a directed graph. Three run in parallel — a Retriever pulling facts, a Reconciler checking for contradictions between coded and written data, and a Risk agent evaluating both deterministic rules and ML predictions. All three converge on a Verifier, which confirms every claim traces to a real database row before anything is shown. Only then does a Composer write the final brief.
Four agents run on Qwen 3 32B for fast structured extraction. The Verifier runs on Amazon Nova Pro — judging whether a claim is genuinely supported requires deeper reasoning than simply finding it.
Machine learning. We trained seven candidate risk models on SageMaker. Only three shipped. Two were rejected for being too rare to learn reliably. One was rejected because white cell count carried 48% of its weight — it was restating a current value, not predicting a future one. One was rejected for firing on 61% of all patient-hours — noise, not signal. The three that shipped: AKI risk (AUC 0.907), transfusion likelihood (AUC 0.892), and electrolyte crisis (AUC 0.818), each registered in SageMaker Model Registry behind a manual approval gate.
Backend and deployment. A FastAPI backend runs on AWS Elastic Beanstalk, serving a Next.js frontend deployed on Vercel. Live patient vitals are fed through Amazon EventBridge and Lambda, writing directly into Aurora exactly as a real hospital monitor integration would.
Challenges we ran into
- Row-type mismatches across the API — several endpoints expected
tuple-indexed database rows while others used dict rows, causing
silent
KeyErrorfailures that took careful tracing to isolate. - AWS Marketplace billing — Anthropic models on Bedrock require a payment method our account's e-mandate card didn't support. We switched our reasoning model to Amazon Nova Pro, which bills directly through standard AWS billing.
- Packaging for deployment — Windows zip tooling silently corrupted nested folder structures multiple times during Elastic Beanstalk deployment, requiring us to move model configuration files to S3 and fetch them at runtime instead.
- Leakage in model training — a label we'd dropped for being too rare was still present in the feature set for another model, quietly inflating its accuracy. Removing it dropped the AUC from 0.913 to 0.892 — a more honest number.
- KDIGO staging bug — our first implementation computed a patient's kidney injury baseline as their lowest recorded creatinine, which silently hid every case where a patient's kidneys recovered. Fixed by anchoring baseline to the admission value instead.
What we learned
That the deepest engineering problem in clinical AI isn't reasoning — it's honesty. A model that can say "I don't know" and mean it, that never fills a gap with something plausible, is harder to build than one that always has an answer. Every fix in this project moved us toward a system that's not just capable, but auditable.
What's next for GRUS — Hospital Emergency Decision Support System
- Full AgentCore production deployment with managed scaling
- Real hospital monitor integration replacing the EventBridge simulation layer
- Expanding the clinical score library beyond the current fifteen
- Formal clinical validation with practicing emergency physicians
Built With
- amazon-bedrock
- amazon-nova
- amazon-titan
- amazon-web-services
- aurora
- aws-lambda
- bedrock-agentcore
- docker
- elastic-beanstalk
- eventbridge
- fastapi
- mimic-iv
- next.js
- pgvector-aurora
- postgresql
- python
- qwen
- react
- sagemaker
- strands-agents-sdk
- tailwindcss
- typescript
- vercel
- xgboost


Log in or sign up for Devpost to join the conversation.