-
-
ThirdEye detecting and tracking people and objects in real time on-device.
-
Multiple active incidents shown at once, including restricted-zone activity and an unattended object.
-
Backpack 1 linked to Person 2 with last-seen details, event history, and the object-moved alert.
-
NVIDIA Jetson Orin Nano with the Arducam CSI camera powering ThirdEye’s local computer vision pipeline.
-
ThirdEye’s end-to-end pipeline from camera capture and detection to memory, event reasoning, incident recording, and the live dashboard.
Inspiration
Most security cameras are good at recording video, but not at understanding what actually happened.
They can tell you that motion occurred or that a person was detected, but they usually do not maintain context over time. If someone leaves a backpack behind, another person picks it up, or an object disappears from where it was last seen, a traditional camera often leaves the human operator to manually review footage and connect the dots.
We wanted to build a system that could do more than detect objects frame by frame. Our goal with ThirdEye was to create an edge AI system that maintains memory of a scene, tracks relationships between people and objects, understands meaningful events, and only records footage when something important happens.
We also wanted the system to work locally on dedicated hardware instead of depending on cloud inference, making it faster, more private, and more resilient to unreliable internet connections.
What it does
ThirdEye is a real-time edge AI security and event-intelligence platform running on an NVIDIA Jetson Orin Nano.
A live camera feed is processed locally to detect and track people and portable objects such as backpacks. Instead of treating every frame independently, ThirdEye builds a persistent world state that remembers what entities have been seen and how they relate to one another.
ThirdEye can currently detect and reason about events such as:
- people entering and leaving the scene
- restricted-zone entry
- dwelling in a monitored area
- occupancy and crowding
- backpacks and other portable objects
- associations between people and objects
- unattended objects
- objects moving from their previous location
- objects leaving the scene with a person
- last-seen locations and event history
ThirdEye also includes event-triggered incident recording. Rather than continuously storing hours of footage, the system keeps a short rolling buffer and automatically saves video from before and after a meaningful event occurs.
This allows the system to preserve the context surrounding an incident without requiring constant recording.
How we built it
ThirdEye runs locally on an NVIDIA Jetson Orin Nano using a CSI camera and a real-time computer vision pipeline.
The core architecture is:
Camera
→ YOLO object detection
→ tracking
→ persistent identity/world state
→ spatial and relationship reasoning
→ event engine
→ SQLite event history
→ incident recorder
→ live dashboard
We use Ultralytics YOLO for real-time object detection and built our own tracking, identity, world-state, event, relationship, memory, and recording layers around it.
A major design decision was separating raw tracker IDs from logical identities. Tracker IDs can change when an object briefly disappears, so ThirdEye maintains its own persistent entity layer that allows logical people and objects to survive short detection gaps and maintain their history.
For portable objects, the system also maintains last-known state so an object is not immediately forgotten just because the detector temporarily loses it.
Events are stored locally in SQLite, including timestamps, entity relationships, last-seen information, and incident metadata.
For incident recording, ThirdEye keeps a rolling pre-event buffer in memory. When an important event occurs, it saves footage from several seconds before the event and continues recording afterward, creating a compact incident clip instead of continuously storing video.
The system also includes two interfaces:
- a simplified judge/demo dashboard for real-time monitoring
- a debug view with deeper tracking, timing, and event information
Everything required for the core demo runs locally on the Jetson.
Challenges we ran into
One of the biggest challenges was realizing that object detection alone was not enough.
YOLO could detect people and backpacks, but real-world detections are imperfect. A person might disappear for a few frames, a backpack could temporarily be classified as a handbag or suitcase, or an object could stop being detected entirely depending on lighting, angle, or occlusion.
That meant we had to build logic above the detector to create continuity over time.
We ran into issues such as:
- one real person being assigned multiple IDs
- temporary duplicate detections creating ghost people
- one backpack being classified as multiple different objects
- a backpack disappearing from detection even while still physically visible
- associations between people and objects breaking when detections were missed
- tracker IDs changing when someone left and returned to the scene
We addressed these problems by adding duplicate filtering, object-family normalization, persistent logical identities, archived entity states, last-seen memory, relationship recovery, and more conservative event reasoning.
We also had hardware-specific challenges on the Jetson, including CSI camera configuration, JetPack compatibility, CUDA/PyTorch setup, OpenCV with GStreamer support, power limits, and over-current throttling during heavier model experiments.
Another major challenge was presentation. Our early dashboard exposed too much technical information at once, so we redesigned it to focus on the events a user actually cares about while keeping a separate debug view for development.
Accomplishments that we're proud of
We are especially proud that ThirdEye evolved from a basic real-time detector into a stateful event-intelligence system.
Instead of only answering:
"What is in this frame?"
ThirdEye can begin answering questions like:
"Who was this backpack previously associated with?"
"Where was this object last seen?"
"Did this object move after the original person left?"
"What happened immediately before this incident?"
We also achieved real-time performance directly on the Jetson. During testing, the full pipeline ran at roughly 20–24 FPS while performing detection, tracking, event reasoning, dashboard rendering, and incident recording.
Our event-triggered recorder was able to save footage from approximately five seconds before an event and continue recording afterward without creating a measurable FPS penalty in our testing.
We are also proud that the core system does not depend on cloud inference. ThirdEye continues processing locally even if internet connectivity is unavailable.
What we learned
The biggest thing we learned is that building useful computer vision systems requires much more than choosing a good detection model.
A detector gives you observations. A useful product needs memory, state, relationships, uncertainty handling, and logic across time.
We learned a lot about:
- edge AI development on NVIDIA Jetson
- JetPack and CUDA
- real-time video pipelines with GStreamer and OpenCV
- GPU inference
- object tracking and identity persistence
- managing imperfect detections
- designing event-driven systems
- balancing performance with model complexity
- local data storage and incident recording
- building user interfaces that communicate complex AI behavior clearly
We also learned the importance of being conservative with AI conclusions. ThirdEye intentionally distinguishes between detecting an association and claiming ownership, and it avoids making strong claims when the evidence is ambiguous.
What's next for ThirdEye
The next step is improving long-term identity and object persistence.
We are experimenting with lightweight visual re-identification so ThirdEye can recognize when the same person or object leaves the scene and later returns, while keeping the system anonymous and avoiding facial recognition.
We also want to improve portable-object tracking when objects become partially occluded or difficult for the detector to see.
Beyond that, we envision ThirdEye supporting multiple cameras, configurable security policies, better incident search, and natural-language queries such as:
"Where was Backpack #1 last seen?"
"Who was near it before it moved?"
"Show me the incident where it left the room."
The long-term goal is to move security systems from passive video recording toward local, privacy-conscious systems that understand and remember what is happening in the physical world.



Log in or sign up for Devpost to join the conversation.