Inspiration
Why call it OK? Turn your head sideways: it looks like a person lying down.
Our team all has a background at Motorola Solutions, and we wanted to use our learned technical experience to solve a problem which is near and dear to our hearts.
The young, the old, and the disabled, all require special care and love. However, in our time, life gets busy and things like work and school impact the availability we have to express this for the important people in our life. Attempting to take care of someone while you yourself are stretched thin is demanding and you won't be able to give them the care that they need. Hospitals see this too, with labour shortages and patient overflow, and so we wanted to create a tool that eases the workload for caretakers without compromising.
What it does
OK is hardware agnostic, working on all modern cameras. It leverages the RTSP protocol to stream data into an inference model that detects critical events. These events are configurable based off sound, pose, and presence. OK works across multiple cameras in parallel, allowing scalability.
The currently implemented events are as follows:
- Fall detection
- Lack of movement
- Loss of vision in blind spot
- Loud noise detection
When one of the above events is triggered, OK's dashboard displays an alert telling observers what happened, where it happened, and when it happened.
Why it's different
OK is privacy first. No recordings, and no personally identifying information is saved. All analysis is local and done on-the-fly.
As described previously, OK works with all modern cameras. You can easily drop-in OK to your existing camera network, and have it working within 10 minutes without configuration.
If you require specific tuning, OK is easily configurable and features a modular event system, allowing new events and fine tuning existing events.
How we built it
OK uses MediaMTX to convert RTSP into HTML-friendly format over WebRTC for low-latency streaming into our inference model and stream viewer.
Each camera receives a dedicated inference model to keep latencies low, where a Docker container is spun up upon the addition of a camera to the network.
Our inference model itself uses MediaPipe Pose to perform pose detection on the CPU, allowing us to track posture, falls, and absences. PyAV is used for audio analytics, giving us data on a camera's noise levels in real-time.
Our web application uses FastAPI for the backend, hooking into TiDB to store camera and event information. Our frontend uses Vue.js and Pinia.
Challenges we ran into
The challenges we ran into include:
- Telling a fall from everything else: sitting down fast, bending to pick something up and lying down on purpose all look like falls in a single frame. A fall turned out to be a sequence of events over time, so we built a state machine with time windows instead of a per-frame classifier.
- Perspective: when someone falls toward the camera, their body looks shorter, which threw off every measurement scaled to body size. We fixed this by measuring torso length only while the person is upright and using the median of those readings as the body unit.
- Camera angle: from our webcam's position, the torso angle often never looked horizontal. Using several independent "went horizontal" signals (torso angle, bounding-box shape, head at hip level) meant detection still worked.
- WebRTC in different browsers: Firefox ignores loopback addresses for WebRTC connections, so MediaMTX had to advertise the machine's local network IP as well.
- Integration: implementing multiple services is easy, putting them together is not. It was a challenge to have everything working as one.
Accomplishments that we're proud of
OK is a project we're proud of! Some of our achievements include:
- Low-latency alerts means someone falling in front of a camera will alert the dashboard within a second.
- No recordings and no identifying information. The privacy guarantee comes from how the system is built, not just from a policy.
- Every camera is watched in parallel, so an alert from one room shows up even while a caretaker is viewing another.
- It works with any standard RTSP camera and fits on top of an existing camera network.
- Fall rules that are easy to understand and tune: every threshold is in one config file, measured in body units, and every alert records why it fired (peak velocity, maximum torso angle, which signals triggered).
What we learned
When tackling OK, we learned:
- Pose estimation is the easy part. Turning noisy keypoints into a reliable "this person fell" decision is the real work, and it means designing for time, hysteresis and missing data.
- Real-time video has a lot of moving parts (RTSP, WebRTC, ICE candidates, codecs), each with its own ways to fail.
- Building for healthcare changes priorities. A missed alert is far worse than an extra one, alerts must stay visible until someone acknowledges them, and privacy can't be added later.
What's next for OK
Over the course of the project, our vision for OK grew. Unfortunately, the time constraints do not allow us to achieve everything we wanted to do. In the future, we would want to implement:
- Offer the backend as a hosted service, so a user without spare servers gets the compute and still keeps its data private.
- Phone notifications for users, so alerts reach them away from their screens.
- Audible alerts, and acknowledgements shared across every user's screen.
- Tracking several people per camera, and thresholds tuned on a larger set of real fall recordings.
Built With
- fastapi
- mediapipe-pose
- naive-ui
- python
- typescript
- vue
Log in or sign up for Devpost to join the conversation.