Inspiration
A store owner we know has 14 cameras. Fourteen. And he still finds out about theft the next morning, from a manager, not from the system that's supposed to be watching.
That's when it clicked for us: the cameras aren't the problem. They're everywhere already — stores, warehouses, apartment buildings, all recording 24/7. The problem is that recording isn't the same as watching. Nobody has time to sit and stare at fourteen video feeds. So the footage just piles up, unused, until something goes wrong and someone has to scrub through hours of it after the fact.
Every "solution" we found for this meant ripping out the cameras and buying a whole new hardware ecosystem — thousands of dollars, for a business that just wanted its existing setup to be smarter.
So we asked a simpler question: what if the intelligence lived in software, not in new boxes bolted to the wall?
What it does
Netra AI plugs into cameras a business already owns and makes them actually useful — no new hardware, no rip-and-replace.
You connect it, it auto-discovers your cameras, and you draw zones on the live feed — loading dock, checkout counter, no-entry after 9 PM. That's it. From there:
- It watches, so you don't have to. Someone in a restricted zone after hours? You get an alert in seconds, with the clip attached — not a discovery the next morning.
- You can search video like you'd search Google. Type "anyone near the loading dock after 10 PM last week" and get back actual clips, instead of dragging a timeline for an hour.
One thing we were firm about: no facial recognition, no biometric tracking. This system understands events, not identities. That line mattered to us.
How we're building it
Here's the one decision that shaped everything else: you cannot run heavy AI on every frame of every camera, all the time. Do that math at even 50 cameras and the cost breaks the whole idea. So instead of one pipeline, we built two — sharing the same foundation, so nothing gets computed twice.
The foundation. Every feed starts with dumb, cheap motion detection. Only frames with actual movement get passed to an object detector (a quantized YOLO model, light enough to run on edge hardware like a Jetson). Detected objects get tracked across frames with ByteTrack — so the system reasons about a person walking through frame, not a thousand individual images of them. That tracking data gets saved either way, whether or not anyone ever looks at it.
Pipeline one: alerts, real-time. Zone and rule checks run directly on that tracking data — cheap, just geometry. Only the genuinely unclear cases get escalated to a vision-language model, and even then just a few key-frames, never a full clip. Anything critical — weapon, fire, a fall — skips the queue and alerts immediately, verified after the fact. Speed beats certainty when it matters.
Pipeline two: search, lazy. This is the part we're most proud of. We don't caption or embed footage as it's recorded — we do it the first time someone actually searches that time range, using the tracking data to narrow down what's even worth looking at. After that, it's cached. First search over a window is a little slower. Every search after that is instant.
Built with Python for the pipeline logic, Qdrant for vector search, a small classifier for cheap first-pass decisions, and VLM calls treated like a scarce resource — used sparingly, on purpose, not by default.
Challenges we ran into
Our first instinct was wrong, and it cost us half a day to find out. We wanted to caption everything, all the time, so search would just work. We prototyped it, ran the numbers on cost at scale, and it fell apart immediately. That failure is the reason the two-pipeline design exists at all — we had to unlearn the "just throw AI at it" instinct before the real architecture could show up.
Latency was its own war. Cloud round-trips were too slow for anything under 5 seconds, full stop — which meant detection and rules had to run at the edge, on hardware that doesn't scale the way a server does.
Then there was noise. Wind, shifting shadows, a light flicking on — motion detection saw all of it as "something happened." Tuning that first filter so it caught real events without flooding the pipeline with junk took more iteration than anything else we built.
And honestly — 36 hours is not much time to build, tune, and demo a two-pipeline vision system. We rationed sleep about as carefully as we rationed our VLM calls.
Accomplishments that we're proud of
Less than 1% of everything the cameras see ever reaches the expensive part of our pipeline. That number is the whole project, honestly — it's the difference between a system a small business could actually afford and one that only makes sense with enterprise money behind it.
We're also proud of what we didn't build. It would've been easy to promise facial recognition, or "detects any crime, automatically" — impressive-sounding, and either legally reckless or just untrue. We scoped it to what we could build and stand behind: event detection, zone rules, honest search. Biometrics stayed out, on purpose, until there's real legal review behind it.
What we learned
A model that performs well in a demo and a system that holds up at 2 AM, on real hardware, in bad lighting, with no cloud round-trip to lean on — those are two very different things. We didn't fully understand that gap until we were standing in it.
We also learned that sharing work across pipelines is one of the highest-leverage moves available. The second we realized alerting and search could sit on top of the same tracking layer instead of duplicating detection twice, half the problems ahead of us got easier.
What's next for netra.ai
Action and pose recognition is next — teaching the system to recognize a fall or a physical altercation, not just presence in a zone. That opens doors we care about: eldercare, campus safety, places where seconds matter.
After that, real multi-site scale — which means rethinking parts of this around a proper message queue instead of per-site processing.
But mostly, we want a real pilot. A real store, real footage, real false-positive rates — the kind of data a hackathon demo can't give us, and the only thing that tells us if this actually works in the wild.
Built With
- ai
- bytetrack
- clip
- computer-vision
- machine-learning
- onvif
- qdrant
- rtsp
- security
- vlm
- yolo
Log in or sign up for Devpost to join the conversation.