Inspiration
When a mass-casualty event hits, the hard question isn't medical, it's logistics: where does everyone go? Baltimore ERs already run near capacity on an ordinary night. The records behind those decisions often disagree, because roughly half of hospital record text is copied forward from older notes, so stale facts look current and nobody re-checks them in a surge. The NHS's own bed-allocation pilot at Kettering built a greedy allocator and reported it couldn't adapt to changed ward layouts or flu peaks. We wanted something that re-plans live, shows its reasoning, and stops at the one thing software shouldn't guess: whose record is right.
What it does
EmerFlow has three screens. The EMS map, our homepage, shows how full every Baltimore ER is right now from live MIEMSS crowding data, ranked by drive time plus the wait on arrival rather than distance alone, with a what-if simulator that drops a mass-casualty incident on the map and spreads the casualties across hospitals. The Hospital Swarm board is the staff command view, where ten department AI agents covering the ER, ICU, step-down, surgery, staffing, imaging, X-ray, lab, blood bank and EMS each see only their own unit and negotiate with a coordinator agent every few minutes, and every patient row shows what's wrong, what they need, where they're going and who decided. DeepChart is the doctor's portal: when a move depends on a fact the records disagree about, the move pauses and the doctor sees every hospital's version side by side with sources and dates, and the software never says which one is right. The rule throughout is that AI talks, code counts, and a human approves the big moves.
How we built it
The backend is Python and FastAPI over a simulated hospital with nine units, nurse ratios, CT, X-ray and lab queues, blood stock and surgery, all driven by a clock. Eleven Gemini agents, ten departments plus a coordinator, run on gemini-3.1-pro-preview through Vertex AI, each with its own persona, lens and temperature, answering in structured JSON behind per-call timeouts and a circuit breaker. Code owns everything countable: a fast lane that places critical patients instantly, a validator, the records check and a five-level escalation ladder, and every move in the system goes through a single commit path. The front end is Next.js 16 and Tailwind over a Server-Sent Events stream, including a three.js view of each agent round and a 3D graph of what every agent currently remembers. It all deploys as one Cloud Run service, and because every live Gemini run is recorded, replay mode can play one back with the network off.
Challenges we ran into
Latency was the first wall: ten departments in parallel plus a coordinator takes 10 to 25 seconds per round, so we made code place the obvious cases instantly and send only genuine contention to the agents. The coordinator kept proposing impossible moves until we stopped asking it for beds and handed it a menu of what was allowed per patient, which is where the rule that LLMs never do bed arithmetic came from. Agent memory turned out to poison prompts with stale bed counts, so we banned counts from memory entirely, wrote notes in plain words instead of unit codes, and added supersede rules so "couldn't move" never sits beside "moved". Cost was a real constraint too, since a naive loop burns tokens with nobody watching, so the clock only runs while a browser is actually holding the event stream open. Hardest of all was making sure a critical patient never waits on paperwork: a records conflict flags a resus move, but it never blocks it.
Accomplishments that we're proud of
The agents genuinely argue, and ICU objecting that a plan leaves no reserve bed is real model output rather than a script. We can measure the safety claim instead of asserting it, because we planted the record conflicts ourselves and kept an answer key: on a 25-patient surge the system flags 11 of 11 conflicts and 11 of 11 lookalike patients with no false alarms. Coordination itself proved worth a great deal, since on the same seeded surge departments acting alone averaged 18 minutes to a bed with a worst case of 141 and five people still waiting, while coordinated it was 1 minute on average, 12 at worst, and nobody left waiting, though we should be clear that this compares coordination against none rather than AI against rules, because offline the swarm and the rule-based planner run the same logic. On top of that we have a full wifi-off demo path, 93 passing tests, and a deployed live site.
What we learned
Put every number in code and every judgement call in the model. The agents earn their place as negotiators and explainers rather than calculators, and the moment we let one near arithmetic it invented beds. The most valuable feature turned out to be knowing when to stop, because pausing a move and handing a doctor two conflicting records is more useful than any confident answer we could have generated.
What's next for EmerFlow
We want to read real bed-tracking, scheduling and staffing feeds read-only, which is a small step because each agent already looks at the world through one narrow view function, so the agents themselves wouldn't change. From there we'd pull outside records over a health information exchange in FHIR instead of our synthetic three-hospital set, and run enough live Gemini rounds to say honestly whether negotiation beats the rule-based ladder. The goal after that is a shadow-mode pilot running alongside a real bed-management team.
Built With
- docker
- fastapi
- github-actions
- google-cloud
- google-cloud-run
- google-cloud-vertex-ai
- google-gemini
- leaflet.js
- miemss-edas-api
- next.js
- pydantic
- pytest
- python
- radix-ui
- react
- react-force-graph-3d
- react-leaflet
- server-sent-events
- shadcn-ui
- tailwindcss
- three.js
- typescript
- uvicorn
- vite
Log in or sign up for Devpost to join the conversation.