Inspiration
My first hackathon was at Johns Hopkins. Getting there meant driving over 16 hours from my college town. Classes left no room for long rest stops, so we pushed through.
My friends and I were sure we could handle it. We were wrong. That false confidence put us in dangerous situations many times on that road.
I kept thinking about the Autopilot in my dad's Tesla back home. That kind of safety net comes with the car. We wanted one that comes with the phone you already have: something that reads physiological signs of tiredness, warns you in different ways as it gets worse, and keeps the people in your life accountable for each other, no matter what you drive.
What it does
A phone sits on the dash. It watches the driver's face, scores how risky the driver looks right now, and talks to them. The driver never touches the screen.
See. The front camera reads pulse, breathing, blinks, expression, eye closure, and yawns. GPS and the accelerometer add speed against the posted limit, hard brakes, and swerves.
Think. Every 10 seconds the backend turns those signals into a risk score from 0 to 100 and picks one of four tiers.
| Score | Tier | What happens |
|---|---|---|
| under 40 | 0 | Logged only |
| 40 to 69 | 1 | Calm voice check-in |
| 70 to 84 | 2 | Firm voice warning |
| 85 and up | 3 | Urgent voice, alarm, and the family is told |
Act. The voice gets more urgent as the tier rises. "Want me to find a rest stop?" Say "yeah" and navigation opens. Say "I'm fine" and it backs off and learns that warning was a false alarm.
Family. At the top tier the agent posts to the family chat over iMessage. Guardians get the location. Family can text back, and the agent reads the message to the driver out loud. The driver answers by voice: "tell her I'm stopping in ten." If the driver is falling asleep, the agent asks the family to roast them awake and reads the roasts aloud right away.
After the drive. The trip gets a report card with a safety score, a letter grade, and five category meters. The family chat gets it as an image. The location is never on it.
Some rules do not wait for the score. Eyes shut for 1.5 seconds is a microsleep and goes straight to tier 3. Kids in the car raises every alert by one tier.
How we built it
Sensing, on the phone (Kotlin, Android). Presage SmartSpectra gives pulse, breathing, blinking, expression, and 478 face landmarks from the camera alone. No wearable. Presage has no "eyes closed" or "yawning" output, so we compute them from the landmarks: eye aspect ratio for closure, mouth aspect ratio for yawns. The phone only senses. It sends raw signals and scores nothing.
Risk engine and decision tree, on the backend (TypeScript). The math is in the next two sections.
The risk score, step by step
10 s window
│
▼
baseline ──► smooth ──► six levels ──► weighted sum ──► score R ──► decision tree
(first 60 s) (3 windows) x_i in [0,1] z = Σ w_i x_i 0 to 100
Step 1. Baseline. The first 6 windows only learn this driver's resting heart rate and breathing rate. Tier stays 0.
Step 2. Smoothing. Numeric signals are averaged over the last 3 windows. The longest eye closure is not, so a microsleep is never averaged away.
Step 3. Six levels. Each raw signal is scaled to 0 to 1, where 1 means "fully present". Write \(\mathrm{sat}(v) = \min(\max(v, 0), 1)\).
| Level | Formula | Reaches 1 at |
|---|---|---|
| Drowsy | \(0.4\,\mathrm{sat}(\tfrac{\text{eye closure}}{0.30}) + 0.2\,\mathrm{sat}(\tfrac{\text{yawns}}{3}) + 0.2\,(\text{low engagement}) + 0.2\,\mathrm{sat}(\tfrac{\text{breathing drop}}{4})\) | eyes 30% closed, 3 yawns, breathing down 4 per minute |
| Agitated | \(0.6\,(\text{stress}) + 0.4\,\mathrm{sat}(\tfrac{\text{pulse} - \text{baseline}}{25})\) | pulse 25 bpm over baseline |
| Speeding | \(\mathrm{sat}(\tfrac{\text{mph over limit}}{20})\) | 20 mph over |
| Phone | 0 or 1 | phone in hand |
| Distracted | \(\mathrm{sat}(\tfrac{\text{seconds gaze off road}}{4})\) | 4 seconds |
| Erratic | \(\mathrm{sat}(\tfrac{\text{hard brakes} + \text{swerves}}{3})\) | 3 events |
A missing signal contributes 0.
Step 4. Weighted sum. Each level is multiplied by the natural log of that factor's crash odds ratio (SHRP 2, AAA). Adding log odds is the same as multiplying odds, which is how a logistic model combines risks.
$$z = \sum_i \ln(\mathrm{OR}_i)\, x_i + s \qquad R = 100 \cdot \frac{\min(z,\ \ln 50)}{\ln 50} \cdot m$$
Here \(s\) is the sleep term, \(m\) is the context multiplier, and \(\ln 50 \approx 3.91\) caps the combined odds at 50 times. \(R\) never goes above 100.
| Factor | Odds ratio | Weight \(\ln(\mathrm{OR})\) | Most it can add to \(R\) alone |
|---|---|---|---|
| Speeding | 12.8 | 2.55 | 65 |
| Agitated | 9.8 | 2.28 | 58 |
| Phone in hand | 3.6 | 1.28 | 33 |
| Drowsy | 3.4 | 1.22 | 31 |
| Distracted | 2.0 | 0.69 | 18 |
| Erratic | 2.0 | 0.69 | 18 |
The last column is \(100 \cdot \ln(\mathrm{OR}) / \ln 50\). It shows the problem: a fully drowsy driver scores 31, which is under the tier 1 line of 40. That is why the decision tree has overrides.
| Sleep last night | Crash odds | Sleep term \(s\) |
|---|---|---|
| under 4 hours | 11.5 times | 2.44 |
| 4 to 5 hours | 4.3 times | 1.46 |
| 5 to 6 hours | 1.9 times | 0.64 |
| 6 to 7 hours | 1.3 times | 0.26 |
| 7 or more, or unknown | 1 | 0 |
| Context | Effect on \(R\) |
|---|---|
| Kids in the car | +15% |
| Low driving experience | +15% |
Step 5. Drowsy or reckless. The sum is split in two. The larger side is the dominant cause and picks the spoken line.
drowsy side = w_drowsy · x_drowsy + sleep term
reckless side = agitated + speeding + phone + distracted + erratic (each w_i · x_i)
dominant = whichever is larger
Per-driver learning. Each weight carries a multiplier per driver, clamped to 0.5 to 1.5.
| Feedback | Dominant factor's weight |
|---|---|
| "I'm fine" (false alarm) | times 0.95 |
| Confirmed alert | times 1.05 |
Worked examples.
| Drowsy driver | Reckless driver | |
|---|---|---|
| Signals | eyes 25% closed, 2 yawns, slow breathing | 15 mph over, stressed, 2 swerves |
| Levels | drowsy 0.70 | speeding 0.75, agitated 0.60, erratic 0.67 |
| \(z\) | \(1.22 \times 0.70 = 0.85\) | \(1.91 + 1.37 + 0.46 = 3.74\) |
| \(R\) | \(100 \times 0.85 / 3.91 \approx 22\) | \(100 \times 3.74 / 3.91 \approx 96\) |
| Tier from score | 0 | 3 |
| Final tier | 2, by the drowsy override | 3 |
The decision tree, step by step
A pure function: state and window in, new state and actions out.
new 10 s window
│
first 6 windows? ── yes ──► learn baseline, tier 0
│ no
▼
score R ──► raw tier R < 40 → 0 40-69 → 1
│ 70-84 → 2 R ≥ 85 → 3
▼
HOLD tier = min(this window, last window)
│
▼
OVERRIDES raise the tier to a floor (table below)
│
▼
KIDS IN CAR any active tier + 1, max 3
│
▼
COOLDOWNS mute repeats, keep the tier
│
▼
tier + actions
| Override | Condition | Result |
|---|---|---|
| Microsleep | eyes shut 1.5 s or longer | tier 3 at once, skips the hold |
| Drowsy, short | drowsy level 0.6 or more for 3 windows (30 s) | at least tier 2 |
| Drowsy, long | drowsy level 0.6 or more for 12 windows (2 min) | tier 3 |
| Stuck at warning | tier 2 for 12 windows (2 min) | tier 3 |
| Step | Why it exists |
|---|---|
| Hold | One noisy window should not trigger an alert |
| Overrides | Drowsiness scores low in the model but is deadly in practice |
| Kids raise | Applied last, so a child in the back makes everything stricter |
| Cooldowns | A system that nags gets turned off |
| Tier | Voice | Listens for a reply | Family | Cooldown |
|---|---|---|---|---|
| 1 | calm nudge | 5 s | never | voice once per 2 min |
| 2 | firm warning | 5 s | only if this driver's threshold was lowered | voice once per 2 min |
| 3 | urgent, plus alarm | no | yes | family once per 10 min |
Learning. Within tiers 1 and 2, a LinUCB contextual bandit picks which intervention to speak. The reward is whether the driver's state improves 120 seconds later. It never changes the tier.
The score needed to text family also adapts to the driver's history:
$$T = \mathrm{clamp}\big(85 + 0.5\,(\text{careIndex} - 75) + \text{learnedShift},\ 60,\ 95\big)$$
Voice (ElevenLabs). One voice, four delivery settings. Lower stability and higher style make it sound more urgent. Alert lines are templates, not LLM output, so they are instant and predictable. Everything spoken to the driver goes through one queue, and the queue pauses for 10 seconds after a hard brake.
Family agent (Photon Spectrum). An iMessage agent that classifies each message as a question, a message for the driver, a roast, or chatter. It answers from a pre-filtered fact list, so the model cannot leak what it was never given.
Storage (Tiger Data). Every 10 second window lands in a TimescaleDB hypertable, with continuous aggregates, 7 day retention, and compression. The report card is stored at trip end so it survives after the raw windows expire.
Report card.
$$\text{penalty} = 0.3\,\overline{R} + 0.2\,R_{p90} + 30\,f_{\text{tier}\ge 2} + 30\,f_{\text{tier}3} + 10\,\min(\text{microsleeps},\ 3)$$
$$\text{score} = \mathrm{clamp}(100 - \text{penalty},\ 0,\ 100)$$
Challenges we ran into
Presage does not detect yawns or closed eyes. We had to build both from raw face landmarks, then track unbroken closed-eye runs across camera frames to catch a microsleep.
The research under-weights drowsiness. In the odds ratio model, speeding weighs 2.55 and drowsiness only 1.22. A driver with heavy eyes and two yawns scores about 22 out of 100. That is tier 0, right before they fall asleep. The score alone was not enough, so we added override rules that fire on eye closure and sustained drowsiness no matter what the score says.
Missing data is not calm data. When the camera loses the face, an average reads as "fine". We drop those frames, and after three faceless windows the engine switches to scoring speed and motion only.
Smoothing hides the worst moment. Averaging over three windows damps noise, but it also averages away a 1.6 second eye closure. Numeric signals are smoothed. The longest eye closure is not.
Not nagging. A system that warns too often gets turned off. Holds, cooldowns, and per-driver false alarm feedback all exist for that reason.
Talking over an emergency. Family messages, roasts, and alerts all compete for the driver's ears. One queue, one item at a time, alerts first.
Accomplishments that we're proud of
- A risk score built on published crash odds ratios, not numbers we made up.
- A full three-way conversation: the AI, the driver, and the family, all hands-free on the driver's side.
- A backend that is tested end to end with a fake phone client, so the whole loop runs without a car.
- It fails soft. No face: score motion only. LLM down: keyword fallback. TTS down: the phone speaks with its own voice.
- Privacy by design. Location goes to guardians only. API keys stay on the backend. A driver can set sharing to "never", and the agent asks by voice before telling anyone.
What we learned
- A model that is right on average can be wrong at the one moment that matters. Rules have to sit on top of the score.
- "High pulse" means nothing without a baseline. Every driver is compared to their own first 60 seconds.
- No evidence is not evidence of danger, and it is not evidence of safety either.
- How a warning sounds matters as much as when it fires.
What's next for HiWay-67
- Verify the full loop on a device in a real car. The backend is tested end to end. The phone app builds and its core logic is unit-tested.
- Check the drowsiness odds ratio against the source paper and replace the placeholder weight for erratic driving.
- Fit the weights from data instead of setting them by hand.
- Nod detection, gaze, and phone-in-hand sensing on the phone.
- A pre-trip check: hours slept and how rested the driver feels.
- Weather and road conditions as a risk factor.
- Weekly trends across trips.
- Persist trip history and contact preferences so they survive a restart.
Built With
- android-studio
- canva
- codex
- eleven-labs
- figma
- imessage
- mapbox3d
- node.js
- photon
- postgresql
- presage
- tigerdata

Log in or sign up for Devpost to join the conversation.