Inspiration
Every parent has looked away for one second. That second is all it takes: falls are the leading cause of injury-related emergency visits for children under five and the leading cause of brain injury in children aged 0–4, drowning is the leading cause of death for children aged 1–4, and falls are the leading cause of construction worker deaths. Almost every safety system on the market detects an accident after it happens. We wanted one that predicts it, early enough for someone to step in.
What it does
Foresight watches a room through one camera and turns every frame into a skeleton of 33 body points. Several prediction engines run on that skeleton in real time:
- AI fall prediction: a model we trained reads the last second of movement and outputs a fall-risk score 15 times per second. When risk stays high, a soft chime warns before the fall.
- Fall confirmation: a fast drop followed by lying flat (or landing upside down) sounds an alarm. If the person stands back up for 10 seconds, it clears itself as "recovered".
- Danger zones: a guardian draws boxes around hazards (stove, socket, bathtub, stairs, crib, machinery, drop zones). Foresight predicts approach (seconds until someone reaches a zone) and reach (seconds until a hand touches a stove or socket), and detects climbing out of cribs or onto furniture.
- Not moving: someone lying still too long, or staying still in water, raises an emergency. This covers silent drowning, unconsciousness and worker collapse.
- Home and Worksite modes: falls are watched for everyone in both. Home zones use strict rules for young children; Worksite zones cover electrical panels, machinery, edges and drop zones.
- Dashboard + phone alerts: a web dashboard (live status, live skeleton over the room, today's timeline, activity log, zone setup) and Telegram alerts with a skeleton image.
- Feedback loop: guardians mark every alert as Real or False alarm; those labels become training data for the next model.
Private by design: video is processed on the device and discarded. Only skeleton points and events leave it.
How we built it
- Data: we started from the Oops! dataset (20,700 real "fail" videos). We filtered clips by their human-written descriptions, extracted skeletons with MediaPipe Pose, removed clips with poor tracking, scene cuts or no real fall motion, and manually spot-checked kept and rejected clips, which changed our cleanup rules. Result: 1,045 verified fall clips + 641 normal clips.
- Model: a small temporal convolutional network (PyTorch) on 1-second skeleton windows (15 fps), with train / validation / test splits by source video so the test set is never seen during training or tuning.
- Calibration: we chose the warning threshold by simulating the device's real alert logic (smoothing, half-second hold) on the validation set.
- Zones: a single camera can't see depth, so distance is judged from where a person's feet touch the floor; a hand held in front of the lens won't trigger a zone, but walking up to the stove will.
- System: Python device program (OpenCV + MediaPipe + PyTorch), Flask + SQLite server, HTML/JS dashboard, Telegram Bot API alerts. Designed for a Raspberry Pi 5 + USB webcam; the prototype shown runs on a ThinkPad, with Pi deployment in progress.
Results (test clips the model never saw)
- Warned before 39% of falls
- Warned at least 1 second ahead for 23% of falls
- False warnings on 12% of normal clips
Predicting accidents is much harder than detecting them, and we want to be honest about that. Adding velocity features did not significantly help; the main limit is that YouTube clips start only a few seconds before the fall.
Challenges we ran into
- Fall datasets are built for detection, not prediction: clips start just before the fall, leaving little lead-up to learn from.
- Messy real-world video: shaky cameras, multiple people and scene cuts; our spot-checks showed the tracker sometimes followed the wrong person.
- Depth from one camera, solved with the floor-position approach.
- Real bugs found by testing: a real ladder fall logged as "prevented", a person landing upside down that the alarm missed, "ghost" skeletons with no one in the room, and an alarm that never cleared. Each one led to a fix.
Accomplishments we're proud of
A complete working system, from raw data to a phone alert, with an honest evaluation on unseen clips, and a design that keeps families' video private.
What we learned
Prediction is a different problem from detection, data quality matters more than model tweaks, and honest limits make a safety system more trustworthy.
What's next
- Real, consented child data through the guardian feedback loop
- Multi-person tracking
- Depth and infrared (night-vision) cameras; a thermal sensor so stove alerts fire only when it's hot
- AI-suggested danger zones and automatic re-alignment when the camera moves
- Full Raspberry Pi deployment
Credits
Accident footage: Oops! dataset (Epstein, Chen & Vondrick, CVPR 2020, CC BY-NC-SA 4.0, non-commercial research use). Statistics: CDC, U.S. CPSC, Journal of Safety Research, Injury Epidemiology, Children's Mercy Kansas City.
AI assistance: we used Claude (Anthropic) as an AI coding assistant. It wrote most of the code from our design decisions; we directed the design, collected and spot-checked the data, ran every experiment, tested the system and found the bugs that shaped it. Details are in the README.
Log in or sign up for Devpost to join the conversation.