Inspiration
On 9 September 2026 the UK Health and Safety Executive published this, about a worker in Knowsley whose arm had been pulled into a moving conveyor:
"There was no guard in place and no emergency stop button in the area. Working alone at the time, there was nobody nearby to see or hear what had happened. In an effort to raise the alarm, he repeatedly waved at a CCTV camera in the hope that someone monitoring the system would spot him and come to his aid, nobody did."
The camera was installed, pointed at the right place, and recording. He waved at it, again and again, and nothing happened, because the only thing that could turn those pixels into a response was a person at a monitor, and nobody was at the monitor.
In the United States, OSHA recorded 5,283 fatal work injuries in 2023, and Machine Guarding (29 CFR 1910.212) is on its FY2024 list of the ten most-cited standards. Inspection is a control that works on inspection days: OSHA has inspected one builder seven times since 2023, and all seven found a fall protection violation.
What it does
Keepout watches one fixed camera over one machine.
- It learns the danger zone. An operator draws it once on a reference frame, or Keepout proposes one from where the machine moves.
- It detects entry. A person entering the zone while the machine runs raises an alert on that frame, with the frame attached as evidence.
- It checks the fixed guard. If a guard is open, removed or covered, it says so.
- It escalates by context. Entry with the machine stopped is a note, because a changeover is routine work. Entry while it runs is an alert. A person down and motionless is the highest level, and it stays raised until a human acknowledges it.
Before any of that, it checks whether it can see. If the camera has moved, the lens is blocked, the lights are out or the feed has frozen, it stops answering and names the problem. A zone drawn on one view means nothing on another, and a quiet "nobody in the zone" from a blind camera is the worst thing this product could say.
There is no face recognition and no identity. Before an evidence frame is stored, the head of every person detected in it is blurred, and so is every face the YuNet detector finds. The limit, stated rather than implied: a person no detector finds is not blurred.
You can bring your own footage: upload a clip, draw the danger zone and the machine region on a frame of it, and run it on the instance.
How we built it
OpenCV 5.0.0, pinned as opencv-python-headless==5.0.0.93. The pin matters:
OpenCV 4.14.0 came out after 5.0.0, so an unpinned install resolves to 4.x. The
version is asserted at import and during the Docker build, and shown at /version
and in the UI.
- People: YOLOX-tiny (Megvii, Apache-2.0), official ONNX export, through
cv2.dnn. OpenCV 5 removedreadNetFromDarknet, so everything is ONNX. We ruled out the Ultralytics models because AGPL-3.0 section 13 would make a hosted demo a source-disclosure obligation. - Tracking is our own, because OpenCV 5 removed
cv2.TrackerCSRT,cv2.TrackerKCFandcv2.legacyfrom the main wheel. Detections are matched by exact Hungarian assignment with a constant-velocity predictor. Exact over greedy, because a greedy swap between two workers can pin an incident, and its evidence frame, on someone who never entered. It costs 0.05 ms a frame. - Machine state is motion energy inside the machine polygon with detected people masked out. Without the mask, a worker walking past a stopped machine makes it look like it is running.
- The guard check compares Canny edges, after CLAHE equalisation, and a masked normalised cross-correlation against a learned reference, and takes the weaker score.
- The view checks use
cv2.findTransformECC,cv2.phaseCorrelate, Canny edge density, Laplacian variance and a frame-difference test. - Privacy: YuNet (MIT) runs on each stored frame at its own size and at 3x, on top of blurring the head region of every person box. The container checks YuNet's sha256 at build time and refuses to start without it.
On AWS: a linux/amd64 container in ECR, served by App Runner in eu-west-2 at
4 vCPU / 8 GB with managed TLS. Evidence frames go to S3 under an IAM role scoped to
s3:PutObject on one prefix, and logs go to CloudWatch as JSON. Every bundled clip
is analysed when the image is built, so a cold container shows a real alert with its
evidence in under two seconds; "Run this clip live" re-runs the same code on the
instance.
Challenges we ran into
The detector cannot see people lying down. Rotating real person crops through body angle, YOLOX-tiny's detection rate falls from 1.00 at 40 degrees to 0.00 at 50. A person on the floor is near 90. With no box there is no aspect ratio, so the highest escalation level would never fire for the case it exists for. So Keepout stopped relying on seeing them: a confirmed track seen well inside the zone that disappears without crossing the boundary is escalated on its own. On 40 seconds of crowded real walkway footage that rule now produces 0 false criticals, with a minimum track lifetime, an overlap test judged at the moment the track was lost, and re-detection under a new id recognised. Zero on one clip is not a rate, and the rule does not belong on a camera pointed at a crowd.
A frozen feed is not byte-identical, and edges are not lighting-invariant. A maximum-difference freeze test never fires, because an H.264 round trip of a truly frozen feed still differs by up to 7 grey levels; the mean separates frozen from live-but-static by about fifteen times. And dimming a scene 38% with the guard still in place drops raw edge correlation from 1.00 to 0.19, because gradients fall below the fixed Canny thresholds — CLAHE before Canny is what makes the guard check survive it.
A reference frame has to be checked before it is trusted. A Commons time-lapse opens with a fade from black, and a black reference refuses every subsequent frame. A reference is now only taken from a frame that passes the frame-level checks, stays provisional until five frames agree with it, and is replaced if 25 consecutive frames disagree with it and agree with each other. On that file it is the difference between 0 and 1,764 usable frames out of 2,689.
App Runner vCPUs are slow for this. The same ONNX model in the same cv2.dnn
call takes 14 ms a frame on a 22-thread workstation and 386 ms on App Runner's 4
vCPU. The whole pipeline is 440 ms a frame there, about 0.19x real time at 12 fps.
One service cannot keep up with one live camera, which is why the demo is computed
at build time.
Small people are missed. On a warehouse clip where workers are about 5% of frame height, YOLOX-tiny tracked nobody. One detector pass is reliable from about 15% of frame height; tiled detection reaches about 8% at five times the cost. Tiling is an option, and a periodic tiled probe warns when the single pass is missing people.
Accomplishments that we're proud of
On the labelled set — composited, real people cut out of OpenCV's vtest.avi
pedestrian clip and placed into a rendered machine cell:
- 819 of 819 frames with somebody in the zone were detected, 95% Wilson interval [0.9955, 1.0].
- 4 of 4 labelled entries alerted on the frame the boundary was crossed, all at the correct level.
- 0 incidents raised across 228 frames where the view was unusable, with all four kinds of blindness detected.
On real fixed-camera construction footage from Wikimedia Commons, which has no labels and so counts false alarms rather than misses:
- The Malta pump crew raises 16 alerts for about nine people in the zone, 1.8 per person, with no false criticals.
- A handheld warehouse clip is refused as
camera_movedinstead of producing confident answers about a view that keeps changing. - Each defect real footage exposed is pinned by a regression test in
tests/test_real_footage_regressions.py.
And 143 tests, most of them sequences whose answer is known by construction.
What we learned
The number worth chasing is the one that argues against you. The rotation sweep, the courtyard false-critical count, the App Runner throughput figure and the real construction clips each changed the design. Synthetic clips told us the code did what we wrote. Real clips told us what we had written wrong.
What's next
- Labelled footage from a real factory. It is the only thing that will give a false-alert rate worth quoting; nobody fell in the Commons clips, so every critical they raised was false and none of them can measure a miss.
- Camera-to-handset latency measured end to end, as a distribution. Our 0 ms figure starts at the decoded frame and leaves out the stream.
- A fix for a hard cut to a similar-looking camera. On the one real example, 165 of 1,085 frames after the cut were wrongly accepted; a fix based on ECC correlation is measured but not built.
Footage credits
Shown in the film, all via Wikimedia Commons:
- "Malta - Mdina - Lorenzo Calleja ditch - Il-Foss tal-Imdina (construction) 01" by Frank Vincentz, CC BY-SA 3.0. Fourth shot (74.35 to 104.6 s), re-encoded; faces pixelated in the film.
- "Prefabricated house construction" by H. Raab, CC BY-SA 3.0. First camera position (0.6 to 47.5 s), re-encoded; a time-lapse.
- "Amazon warehouse BHX4 loading docks 1" by domdomegg, CC BY 4.0. Shown as the handheld clip Keepout refuses. Its appearance implies no endorsement by Amazon.
Used for evaluation only, no imagery shown: vtest.avi, OpenCV sample data,
Apache-2.0, the source of the person crops in the composited clips and of the
courtyard clip; and "Amazon warehouse BHX4 loading docks 2" by domdomegg,
CC BY 4.0, used to test tiled detection on small people.
Built With
- javascript
- numpy
- opencv
- opencv-python-headless
- python
- rapidapi
- yolox

Log in or sign up for Devpost to join the conversation.