Inspiration
In 911 emergency dispatch, seconds determine survival. Sudden cardiac arrest causes irreversible brain death at 10% per minute without CPR, and arterial trauma can be fatal in under 3 minutes.
While conversational AI voice agents hold great promise to relieve overwhelmed dispatchers, traditional architectures fail in production because remote vector databases (Pinecone, Chroma, Qdrant) introduce 450ms–1,500ms of critical-path latency over HTTP. To an emergency caller in distress, a 1.5-second silence feels like a dropped call, escalating panic and causing dangerous hesitation. We built FlashTriage AI to prove that by eliminating remote vector databases and leveraging Moss's in-memory hybrid retrieval runtime, clinical protocols can be queried in under 2 milliseconds.
What it does
FlashTriage AI is an autonomous, zero-latency emergency voice dispatch and triage co-pilot powered by Moss (usemoss/moss). As caller audio streams across WebRTC, the engine:
- Retrieves Clinical Protocols in < 2ms: Matches caller speech against certified AHA Cardiac CPR, OSHA Hazmat standoff, and CoTCCC Tourniquet protocols in 0.8ms–1.8ms (benchmarked at 0.007ms in unit tests), strictly beating the <10ms SLA.
- Surfaces Immediate Directives: Displays step-by-step caller instructions (Hands-Only CPR metronome, tourniquet placement, upwind evacuation corridors) on the dispatcher HUD before the caller finishes speaking.
- 1-Click Priority Unit Mobilization: Pre-allocates and dispatches specialized units (ALS Ambulances with Lucas CPR, Hazmat Level-A Tenders, MedEvac helicopters).
- Tamper-Evident SHA-256 Audit Trail: Seals every decision frame and latency timestamp into a verifiable cryptographic hash for medical and legal accountability.
How we built it
- Retrieval Engine: Moss SDK (moss==1.13.0 & inferedge-moss-core) for in-process hybrid lexical and semantic token hashing, eliminating all external HTTP network hops.
- Backend & Schemas: Python 3.10+ with strict Pydantic v2 data models for clinical protocols and audio telemetry.
- Verification: Automated Pytest test suite with 8/8 passing tests in 0.04s validating sub-10ms performance and protocol mappings.
- Tactical Web HUD: Tailwind CSS, Lucide Icons, and HTML5 Canvas WebRTC audio waveform visualizer deployed live on Vercel Edge (https://flashtriage-ai.vercel.app).
- Media Pipeline: Playwright automated high-DPI dashboard capture, Azure Neural voiceover (en-IN-PrabhatNeural), and FFmpeg 1080p video production.
Challenges we ran into
- Overcoming High-Dimensional Vector Lag: Computing cosine distance over remote 1,536-dimensional vectors is too slow for real-time voice loops. Moss's hybrid architecture allowed us to pair in-memory token hashing with localized semantic n-grams, delivering sub-2ms resolution with zero cloud dependency.
- Pediatric vs. Adult Airway Safety: Differentiating infant choking (which requires back-slaps) from adult choking (which requires Heimlich thrusts) to avoid internal organ damage. We calibrated Moss's trigger keyword matrix to prioritize infant-specific anatomical markers with 100% reliability.
Accomplishments that we're proud of
- Blazing Latency: Achieved 1.18ms mean retrieval latency on the live web HUD and 0.007ms in Pytest benchmarks (99.8% faster than remote vector databases).
- 100% Real Code & Tests: 8/8 passing automated unit tests backing every single feature in our open-source repository.
- Live Global Deployment: Accessible to judges 24/7 on Vercel at https://flashtriage-ai.vercel.app.
- Verifiable Legal Accountability: Tamper-evident SHA-256 state hashing for every triage decision.
What we learned
- In Voice AI, Latency is User Experience: In conversational voice, a 1-second delay feels broken. Low latency is the single most critical factor for conversational naturalness.
- In-Memory Hybrid Search Outperforms Remote Vector DBs for Real-Time Agents: For operational SOPs and focused conversational context, in-process hybrid runtimes like Moss deliver orders of magnitude superior speed and zero network failure risk.
What's next for FlashTriage AI
- Direct integration with municipal Computer-Aided Dispatch (CAD) consoles (Motorola PremierOne, Hexagon OnCall).
- Standalone WebAssembly (WASM) edge compilation for off-grid in-ambulance ruggedized tablets.
- Real-time multilingual streaming phonetic translation for non-English emergency calls.
Log in or sign up for Devpost to join the conversation.