Mayday

The minutes before the ambulance, coached.

Inspiration

When someone collapses or starts bleeding badly, the people around them are the only help there is until the ambulance arrives, and that takes seven minutes in most of the country and thirteen in rural areas. Almost none of them are trained. The 911 dispatcher coaching them by voice cannot see what they are doing, and they cannot tell whether they are doing it right. So the minutes that decide whether the person lives are spent panicking, waiting, or doing the right thing wrong.

A bystander who starts CPR gives someone two to three times the chance of survival. Only about 40% of victims get CPR before the ambulance arrives, and fewer than 10% survive a cardiac arrest outside a hospital.

The knowledge that would save that person is public and free, from the American Heart Association, Stop the Bleed and the Red Cross. What is missing is real-time support for the person standing there, and that person is the one who could be saving a life. We wanted to put that support in the phone already in their pocket.

What it does

Mayday is an AI emergency dispatcher with eyes. A bystander opens it on a phone, says what is happening or taps one button, props the phone up, and it coaches them through the correct first-aid protocol out loud, one step at a time with a picture for each, until EMS arrives.

The camera watches them work. It counts chest compressions against a metronome at 110 a minute and says "Faster. Push with the beat." when they drift. It notices when their hands leave a wound and says "Don't let go!" within about two seconds. When it cannot see, it says so and keeps coaching by voice.

CALL 911 is always on the screen. While the call is open, the app shows a situation report to read aloud, rebuilt every second: where, what, how long, when CPR started, the average rate, the longest pause. When the paramedics arrive, the handoff screen gives them the headline numbers, a timeline, and a QR code that carries the report itself, so it scans with no signal.

Two emergencies ship: cardiac arrest with hands-only CPR, and severe bleeding, with choking captured as data. Every spoken instruction is a line a human transcribed from the published guideline. No AI picks, orders, or invents a step.

How we built it

Everything in the coaching loop is browser code, so it keeps working with the network off. Four parts on the phone do the work: the Eyes, the Ears, the Brain and the Voice.

The Eyes are Google's MediaPipe Pose and Hand Landmarker running as WebAssembly in the page, 33 body and 21 hand landmarks per frame with two poses tracked. From those we extract compression peaks and rate, recoil, hands on the wound, and a bounding box with posture and stillness for each person in view.

The Ears are the Web Speech API used as keyword spotting, with phrase cues for how people talk in a panic.

The Brain is a hand-rolled state machine engine, 279 lines with no dependencies, that walks four scripts written as plain data, one per emergency, each medical step citing the guideline page it came from.

The Voice is a three-priority queue that speaks through the browser's own synthesis and keeps the beat on the Web Audio clock.

Every event goes into a log, and the situation report and the paramedic handoff are folded out of it. The app installs as a PWA and caches its models, so it opens with wifi off.

The cloud sits behind a small key proxy on the same origin and behind a flag, and the app is complete without it:

  • Gemini (gemini-3.6-flash, Google AI Studio) sees one photo at the start with a fixed prompt that asks for JSON only: a label from a closed list, a confidence, one sentence about what is visible, and the patient's bounding box on a 0 to 1000 grid. The label becomes a spoken question the person confirms; it never moves the app on its own.
  • ElevenLabs gives the coach a human voice, Brian on Flash v2.5, with every script line cached at launch so it plays with wifi off and a 0.8 second fallback to the browser's voice. The 911 call-taker is a second voice, Sarah, and with the live agent on it is an ElevenLabs Agents conversation over a WebSocket that hears the mic and is instructed to keep the person following the coaching.
  • Grok (xAI, grok-4-1-fast-non-reasoning) reads a panicked sentence the app has no phrase for and answers with the number of the on-screen button or cited answer it meant, so it cannot invent a step. It also rewords one canonical line for the moment, and every rewording is checked against the step's required words before it is spoken.

Five rules hold it together: authority is deterministic, reflexes are local, cognition is episodic, fail loud and never wrong, and a human dials 911.

Challenges we ran into

The hardest problems were in computer vision: getting a phone to see an emergency, measure what a bystander is doing, and know when it cannot trust what it sees.

  • Measuring compressions from a phone camera. The signal is the up-and-down of the helper's shoulders, and it is noisy: two people in frame, a handheld phone, bad light. We track two poses at once and pick the rescuer by geometry, the one kneeling with hips below shoulders, smooth the signal, and count peaks with a hysteresis detector so a wobble does not count as a push. Recoil is estimated from the trough between pushes and we call it an estimate.
  • Classifying the emergency from what the camera sees. A person lying flat is a cue, not a diagnosis, so the on-device detector, torso within 25 degrees of horizontal for two seconds, only ever earns a question. For the unclear case we send one photo to a vision language model with a closed label list, because an open-ended model under a bad picture invents things. Anything outside the list becomes unclear, and a sentence that starts giving advice is dropped.
  • Bounding boxes that mean something. Boxes with posture and stillness for every person in view come from the pose landmarks on the device; the patient's box from Gemini comes back on a 0 to 1000 grid that had to be mapped onto the live camera, and another provider answered in pixels instead, so the parser reads both.
  • Knowing when to stop trusting the picture. Cover the lens, step too close, lose the light, and every number must go blank and the app must say so, rather than coach off a frozen value. That blind gate took as much care as the measurement itself, and the handoff prints that time as not measured, never as a pause.
  • Two models per frame froze the phone. Pose and hand tracking together during triage stalled the camera. Hand tracking now runs only in the bleeding states, and the person-down cue rides on the pose we already measure.
  • Gemini's thinking ate the answer. Gemini 3.x reasons before it replies and the reasoning tokens came out of our budget, so replies arrived truncated or empty. Turning reasoning off for a classification call took it from 3.4 seconds to under one.
  • Two intent routers, built in parallel. Two of us built the same feature five hours apart without knowing, one on Grok and one on Gemini. We kept the measured one and made xAI a provider row in the proxy, so the same route runs on either.
  • Phones do not want to listen. iPhone speech recognition never sends a final result and a home screen app gets no recognizer at all, so we settle interim results into sentences and every voice command has a button.

Accomplishments that we're proud of

Four people, one weekend, and a product that coaches a real person through CPR from a phone propped on the floor.

  • The team. We split the work into four parts, Eyes, Brain, Mouth and Face, agreed on the seams between them in the first hour, and then nobody waited on anybody. Four modules built in parallel met in the middle and worked.
  • A live compression rate read from the camera on a phone, corrected out loud within a second, with the whole loop running on the device and surviving wifi off.
  • Learning classification the hard way: a person-down detector from pose geometry, a vision language model held to a closed label list, bounding boxes mapped from a model's grid onto a live camera, and a blind gate that knows when not to trust any of it. None of us had built that before Friday night.
  • Four AI services that each make the app better and none of which can put a word in its mouth: the label is a question, the rewording is checked, the voice reads the script.
  • The script-as-data design: adding an emergency is adding one cited file, and the engine never changes.
  • A handoff report that tells a paramedic what the bystander did, with time the camera could not see printed as not measured, because we will not tell a paramedic something we did not observe.

What we learned

We learned more in one weekend than we expected to, and the main lesson is how much there is still to learn.

  • Classification is a discipline, not a call. Choosing a closed label set, deciding what a model may and may not say, mapping its answer onto a live camera, and deciding when to ignore it entirely was most of the work, and none of it was the API call.
  • Where to draw the line for AI in a safety-critical loop: seeing, hearing and speaking are fair game; choosing the next medical step is not. Once that line was clear, adding three cloud services in one night stopped being scary.
  • Agree on the seams first. Four people owning four modules with typed interfaces between them meant nobody waited on anybody, and the merge at the end was boring.
  • The browser is enough. Pose tracking, audio scheduling, speech in and out, offline caching and a QR handoff all ran from one page with no native code and no app store.
  • Phones are hostile to good intentions: speech recognizers that never finish a sentence, two models that freeze the camera, certificates that need a password at 2 AM. Every one of those taught us something we now know to check first.

What's next

  • More scripts. Choking with the camera cue enabled, then stroke, seizure and overdose, each one file with its guideline.
  • More that the camera can honestly tell a bystander: breathing rate from chest motion and pain cues from the face, offered as questions for the person, never as diagnoses.
  • An on-device vision model, so the scene assessment is local too and nothing leaves the phone.
  • A real deployment path: the software-as-a-medical-device pathway, and dispatch integration through a platform such as RapidSOS so the situation report reaches the real call-taker.

Team

  • Emmanuel Adedeji
  • Bryce Biyeba
  • Ricky Chen
  • Israel Ogwu

Built With

  • elevenlabs
  • gemini-api
  • google-ai-studio
  • grok
  • mediapipe
  • pwa
  • react
  • typescript
  • vite
  • web-audio-api
  • web-speech-api
  • webassembly
  • xai
Share this project:

Updates

Submission history