Inspiration
In the United States the medical costs attributable to falls were about $50.0 billion in 2015, and Medicare paid about $28.9 billion of that for the nonfatal ones (Florence et al., J Am Geriatr Soc, 2018). CMS lists falls and trauma among Hospital-Acquired Conditions, so a fall with injury on a ward already carries no additional payment, and on a UK dementia ward the rate ran at 5.4 per 1,000 occupied bed days before a quality improvement programme brought it to 1.4.
Everything a ward has for this fires too late: a pressure pad knows nothing until the weight has gone, and a fall detector fires after the fall. Standing up from a seat happens in four phases, and in the first two the person is still supported. A camera watching the upper body move over the feet can see those two, and we wanted to call in that window.
The other constraint was consent. Older adults' dominant objections to assisted-living cameras are that they are privacy-invasive, obtrusive and stigmatising (Tham et al., Innov Aging, 2024), and acceptance is lowest for intimate activities such as changing clothes and showering (Maidhof et al., Front Public Health, 2023). So the product keeps no pictures at all.
What it does
Preempt watches a hospital or care-home room and raises a call before a fall.
- Bed and chair exit. It calls when the shoulders arrive over the feet and the hips start to rise, and checks which way the hips travel so sitting down raises nothing.
- Unsteady gait. Sway about the fitted walking line and a hand held out to a wall, over four seconds of walking.
- Hazards on the route. A walking frame out of reach, clutter between bed and door, a wet-floor sign, all in metres on the floor plane.
- Escalation. A quiet prompt in the room, then the nurses' station, then an urgent call when someone is on the floor. An unanswered prompt becomes a station call by itself.
- It says when it cannot see. Too dark, lens blocked, camera moved, nobody in view: each named with its remedy, none of them a clinical call.
The frame is destroyed as soon as the pose is read. A 1280x720 frame is 2.76 MB and the seventeen keypoints kept from it are 408 bytes. The evidence a nurse sees is a stick figure drawn on a blank canvas, and a test checks by reflection that no rendering function accepts an image.
Results, real footage first
| Track | Result |
|---|---|
| UR Fall, 21 real sequences, full pipeline | 10 of 12 falls reach "on the floor", median 0.50 s behind the labelled frame (0.13 to 0.97 s). 2 of 10 activity sequences with no lying frames falsely said "on the floor". |
| CDC chair-stand footage, oblique shot, room set up by hand | 3 of 3 stands called, 0.08, 0.42 and 0.41 s before upright, 0 false calls (3 before the sit-down fix). |
| CDC chair-stand footage, front shot, two people in view | 0 of 3 stands called. The tracker follows the larger box, which is often the clinician. |
| CDC Timed Up and Go, panning camera | Refused with VIEW_UNUSABLE (camera moved), no clinical call. |
| Synthetic decision layer, 60 sequences | 25 of 25 calls, 0 false alarms in 1.23 hours of quiet, median lead 4.60 s over 10 bed exits and 4.47 s over all 15 exits. |
The lead-time figures come from synthetic sequences, because no real dataset labels the moment a person began to stand. A 30-second chair-stand test is fast repeated standing with no slow preparation, so the slow bed exit behind the 4.6 s figure is not in that footage.
How we built it
OpenCV 5.0.0.93, pinned in three places and asserted at build time, because an
unpinned pip install opencv-python resolves to 4.14.x.
- Detection and pose. YOLOX-tiny and RTMPose-t, both Apache-2.0, both through
cv2.dnnwith no second inference runtime. We avoided Ultralytics because AGPL-3.0 would make a hosted demo a source-disclosure obligation. - The new DNN engine.
cv2.dnn.ENGINE_NEWran YOLOX-tiny at 27.46 ms against 44.01 ms forENGINE_CLASSIC(1.60x) and RTMPose-t at 7.83 ms against 9.30 ms (1.19x), with keypoints agreeing to within a pixel. - Floor geometry.
findHomographygives the floor plane, and a point imaged past its horizon is refused. NotsolvePnPwithSOLVEPNP_IPPE_SQUARE: it returns an ambiguous solution and reported 2 degrees for a true 20 degree tilt. With a rescaled vertical vanishing point, heights come back to under 1 cm. - The rest.
cv2.KalmanFilter,cv2.fitLine,cv2.dft,cv2.calcHist,cv2.Laplacian,ORBwithBFMatcher,absdiffwith Otsu,morphologyExandconnectedComponentsWithStats,inRangewithapproxPolyDP, andcv2.FontFace. - Deployment. AWS App Runner in us-east-2, 2 vCPU and 4 GB. Both models are
inside the image, so there is no network at runtime and
/healthzanswers two seconds after start.
Challenges we ran into
Real footage found three flaws the synthetic track never showed. We ran public-domain CDC falls-prevention training video of an older woman doing chair stands.
- The live service had one room built in, and an upload could not bring its own, so her feet fell inside the built-in bed zone, every stand read as "supported" and all three were missed. An upload now carries its own room, checked by the same parser as the command line, drawn over the first frame before anything runs, and the result says whether the camera height was measured or assumed.
- With a room set up for that clip, all three stands were called — and so were three sit-downs. Lowering onto a seat leans over the feet with bent knees, the same posture as getting up, so no threshold separates them; the direction the hips travel does. The sit-down calls are gone, the three stands still fire at the same instants, a regression test replays the clip's keypoints and fails on the old engine, and the synthetic evaluation lost no detection and no lead time.
- Standing still was reported as walking in 36 of 40 upright readings. The camera sat at hip height, where one pixel of keypoint noise moved the hips' floor shadow tens of metres. Preempt measures that conditioning every frame now and uses the feet when it is bad; walking readings during still stands fell to 0 of 40.
"On the floor" was wrong twice, and both versions failed silently. Asking whether the head was low in metres fired zero times on twelve real falls, because a lying person's head is a body length sideways from the feet. Measuring the body's length flattened onto the floor gave about 2.2 m for both standing and lying. What works is where the head lands when back-projected as if it were on the floor: 0.99 to 1.69 m from the feet for someone on the floor, 2.5 to 5.3 m for someone standing.
Two textbook methods failed on real footage. Step-time variability, the standard clinical gait statistic, measured 0.26 to 0.52 at 15 Hz on walks that are steady by construction — the size of the effect it was meant to detect — so it is reported to a clinician and scores nothing. Crossing head lines with foot lines to find the horizon gave a 24-degree tilt on the UR Fall camera and no real solution for the focal length; the module searches for a level horizon instead.
Accomplishments that we're proud of
- 10 of 12 real falls recognised on the floor, from pose alone.
- On real CDC footage, 3 of 3 stands called with no false calls, after two fixes that cost nothing on the synthetic track.
- A privacy guarantee enforced by tests: zero camera bytes written, zero frames
kept, no code path from a camera frame to a saved file, and a real
camera movedrefusal on the panning clip.
What we learned
Synthetic data proves the decision logic and hides everything else: a baked-in room, sit-downs and a camera at hip height could not appear in a ward whose camera sat 2.55 m up in a corner. And grade a sequence per event, not per sequence: one where somebody stands and sits straight back down holds one event to call and one not to, and the false call on the sit-down hid behind the correct call on the rise.
What's next
- A detector trained on people lying down. The binding limitation: of 561 ground-truth lying frames a person was detected in 93 per cent but only 68 per cent gave six usable joints, and the two missed falls are pose collapsing as the person lands.
- More than one person. Preempt tracks the largest box, so with a clinician beside the patient no stand in the front CDC clip is seen start to finish.
- Crouching, which two activity sequences had called as lying, and measured rooms: the CDC rooms were set up by hand with an assumed camera height.
Footage credits
Both source videos are works of the US federal government and in the public domain. Their use implies no endorsement by CDC; no CDC logo or end card appears, and faces are pixelated. Cuts are re-encoded with the audio removed.
- "30-Second Chair Stand Test", CDC STEADI, public domain (
PD-USGov-HHS-CDC), https://commons.wikimedia.org/wiki/File:30-Second_Chair_Stand_Test.webm — the front shot (37.95 to 50.35 s) and the oblique shot (50.42 to 63.25 s). - "The Timed Up and Go (TUG) Test", CDC STEADI, public domain
(
PD-USGov-HHS-CDC), https://commons.wikimedia.org/wiki/File:The_Timed_Up_and_Go_(TUG)_Test.webm — one cut (47.68 to 64.85 s). - UR Fall Detection Dataset, Kwolek and Kepski, University of Rzeszów, CC BY-NC-SA 4.0, http://fenix.ur.edu.pl/~mkepski/ds/uf.html — evaluation only, no imagery shown.
Log in or sign up for Devpost to join the conversation.