Inspiration
Every year, thousands of construction workers in India are injured or killed in incidents where the first link in the chain was something simple: a missing helmet, a missing hi-vis vest. Site safety officers do walkthroughs with clipboards, but a walkthrough covers a site for minutes a day — violations in between go unseen, and when something does go wrong there is rarely timestamped evidence of what led up to it. We wanted to give every site a tireless safety inspector: a camera feed that never blinks, that doesn't just see violations but does something auditable about them — log, archive, alert, or escalate to a human. And we wanted to build it on the two things this competition is about: OpenCV 5 doing the real vision work, and AWS doing the real cloud work.
What it does
SiteSentry is an agentic vision safety inspector for construction sites. Point it at a site photo, a video clip, or a webcam. OpenCV 5 detects every person, tracks them across frames with a Kalman filter, and checks each worker's helmet and vest using a fine-tuned PPE detection model — all running natively in the new OpenCV 5 DNN engine, no ONNX Runtime needed. It also watches restricted zones (polygon ROIs + background subtraction) and flags suspected falls. Then the agent loop takes over — this is the part we're proudest of. Vision evidence flows into an auditable policy table that decides what happens next: first helmet offense → logged to DynamoDB, evidence frame archived to S3; repeat offender → SNS email alert fires to the site supervisor; third offense, or a suspected fall → the agent recommends a work stoppage — but it cannot act alone. The recommendation lands in a supervisor queue, and a human approves or rejects it in one click. Every step is visible in a live agent trace panel: perceived evidence → policy decision → action. For trust, AWS Rekognition runs DetectProtectiveEquipment as an independent second opinion on flagged frames, and the dashboard shows the agreement matrix between our OpenCV 5 pipeline and Rekognition — disagreements included. Faces are blurred by default, nobody is identified, and evidence auto-expires after 30 days.
How we built it
Vision — OpenCV 5.0.0: a YOLOv8n person detector and the SafetyVision YOLOv8s PPE detector (13 classes: Hardhat, NO-Hardhat, Safety Vest, NO-Safety Vest, Fall-Detected, Person, …) run natively in cv2.dnn on the new DNN engine (ENGINE_NEW on 5.0.0), with an FP16 blob path and automatic FP32 fallback; engine selection detected at runtime. Kalman-filter multi-object tracking, BackgroundSubtractorMOG2 + pointPolygonTest zone intrusion, 30-frame aspect-ratio/motion heuristic for fall candidates, HarfBuzz annotation, Haar-cascade face blurring on by default. Agent — a perceive → plan → act loop; the policy table is a plain auditable data structure (severity, escalation, approval requirements per event type × offense count); MCP-style tools (perceive_frame, query_events, archive_evidence, request_approval). Cloud — AWS ap-south-1 Mumbai: Flask on EC2 t4g.micro (Graviton2), S3 evidence archive with 30-day lifecycle, DynamoDB incident log, SNS email alerts, Rekognition second opinion, CloudWatch latency metrics, least-privilege IAM, no hardcoded credentials. One scripts/deploy.sh reproduces the stack; requirements.txt fully pinned. Honest engineering: every AWS client auto-degrades to a local JSONL/file store when credentials are absent (zero spend), every local output labeled as local; the public demo page is a clearly-marked SIMULATED showcase and the rules allow an arranged live screen-share of the real Flask app for judges.
Challenges we ran into
OpenCV 5's DNN engine API moved between 5.0.0 and later 5.x (ENGINE_NEW → ENGINE_OPENCV) — we detect the enum at runtime. The YOLOv8n ONNX download we planned on vanished, so we folded person detection into the PPE model's own Person class: one forward pass yields persons + PPE items + falls, halving inference cost. PPE accuracy is uneven (the model card admits NO-Safety Vest recall is weak at 0.431) — we set conservative thresholds, documented failure modes honestly, and used Rekognition as a second opinion. Surveillance ethics: face blur default-on, no identity recognition, human approval for consequential actions, 30-day retention — built as features, not boilerplate.
Accomplishments that we're proud of
A genuinely agentic loop where vision changes behavior — the competition literally lists "safety monitoring, and human-in-the-loop operations" as the example, and that's what we built. Real OpenCV 5 depth: new DNN engine, FP16 path, Kalman tracking, MOG2 zones, HarfBuzz annotation — measured per-stage latencies. A second-opinion architecture (OpenCV 5 vs Rekognition agreement matrix) turning disagreement into evaluation evidence. Full reproducibility: pinned deps, one-command AWS deploy, smoke test, offline-capable test suite, local-fallback for every cloud call.
What we learned
Application novelty beats algorithm novelty — construction sites needed a decision loop around the detector, not a new detector. Failure cases are a feature: the rubric asks for them and honesty about NO-Safety Vest weakness made the policy design obviously correct. OpenCV 5's DNN engine is ready for modern YOLO models without ONNX Runtime — but handle the 5.0.0-vs-later API difference. Design for zero spend first: local-fallback clients let us build and demo the whole cloud architecture without an AWS bill.
What's next for SiteSentry
COOL benchmark track (deploy on the Cloud-Optimized OpenCV Library AMI on Graviton, publish measured before/after latency); pose-based fall confirmation; multi-camera site view with homography site-map overlay; on-device edge variant syncing to S3; pilot with a real contractor's safety team with informed consent and audited safeguards.
Built With
- ai-agents
- amazon-web-services
- boto3
- computer-vision
- dynamodb
- flask
- opencv
- python
- rekognition
- s3
- sns
- yolo

Log in or sign up for Devpost to join the conversation.