Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for Watch Tower
Inspiration
Falls are one of the most common — and most dangerous — risks for people aging in place or living with a visual impairment. The tools built to help (apps, wearables with push notifications) all share the same weak point: they assume someone is looking at a screen. For a lot of the people who need help most, that assumption fails.
We wanted to build something that changes the channel entirely — instead of a notification nobody checks, a phone that actually rings.
What it does
Watchtower watches a room through a webcam using a custom-trained YOLO
model that classifies frames as fall or non-fall. When a fall is
detected — confidently, consistently across multiple frames, and outside
a cooldown window to avoid duplicate alerts — it uses CALL-E to:
- Call a designated caregiver and describe what was detected (room, time), then ask for a clear spoken decision: dismiss as a false alarm, or escalate.
- If the caregiver escalates — or CALL-E can't get a clear answer at all — place a second call to a secondary contact, so a real human always makes the final call to emergency services. Watchtower never contacts emergency services on its own.
- Log every event to a local database and surface live status through a dashboard (built-in HTML, plus an optional Streamlit version) showing the video feed, current state, and event history.
How we built it
- Computer vision: a custom-trained YOLOv8 model (
best.pt) run throughultralyticsandsupervision's ByteTrack, served over a FastAPI streaming endpoint. - Detection logic: confidence thresholding, consecutive-frame confirmation, and a cooldown window to keep false positives and duplicate calls low.
- CALL-E integration: the Python SDK (
calle-ai), usingcreate_and_waitwith a structuredresult_schemaso CALL-E returns a cleandismiss/escalate/unknowndecision instead of a raw transcript we'd have to parse ourselves. - Safety design: modeled on this repo's own
deployment-approval-callpattern — CALL-E is only ever allowed to call two configured human phone numbers, never emergency services directly, and any ambiguous result fails toward escalation rather than silence. - Persistence & UI: SQLite for event history, a built-in dashboard, and a separate Streamlit frontend that polls the FastAPI backend for a live status view.
Challenges we ran into
- Transient network timeouts during CALL-E's result-polling step looked like call failures but weren't — the call itself often succeeded while only the status check timed out. We added retry logic with backoff rather than treating every timeout as a hard failure.
- CPU-only inference was too slow for smooth video — frame skipping, a smaller inference resolution, and a lower capture resolution got it back to a usable frame rate without a GPU.
- Environment setup on Windows — PowerShell's execution policy blocked venv activation, and OneDrive-synced folders caused venv corruption; keeping the venv outside the OneDrive-synced repo folder fixed it.
Accomplishments that we're proud of
Getting a real, end-to-end loop working: a simulated (and later a real) fall in front of the camera resulting in an actual phone call, a real spoken decision, and — when escalated — a second real call, all logged and visible on a live dashboard.
What we learned
How much of building a reliable phone-call agent is about handling the edges: ambiguous answers, network hiccups, false positives — not just the happy path of "detect fall, make call." Designing for graceful degradation (retry, fail-safe-toward-escalation, never crash the video stream) mattered as much as the detection model itself.
What's next for Watchtower
- Multi-room support with per-camera process instances
- Encrypting the local event log for real deployments
- Exploring a lightweight on-device fall model for lower latency without a GPU
Log in or sign up for Devpost to join the conversation.