Inspiration
AI agents are increasingly moving from systems that simply generate answers to systems that can perceive, decide, and act in the physical world.
That transition creates a fundamental security problem: what happens when an autonomous agent acts on visual information that is incomplete, degraded, ambiguous, manipulated, or simply wrong?
Traditional computer vision pipelines usually focus on answering what is visible. CiberIA VisionGuard explores a different question:
Should an AI agent trust what it sees enough to act?
The project was created as an application of the CiberIA cognitive cybersecurity approach to Agentic Vision, introducing a trust and verification layer between visual perception and autonomous action.
What it does
CiberIA VisionGuard is a cognitive security layer for Agentic Vision systems built around OpenCV 5.
Instead of allowing visual perception to trigger an action directly, VisionGuard evaluates the reliability of that perception through a Visual Cognitive Trust Score (VCTS).
The system follows a security-oriented agentic loop:
PERCEIVE → ASSESS → VERIFY → RE-OBSERVE → ACT / HUMAN REVIEW
OpenCV provides the visual perception and image-processing capabilities required by the agent. VisionGuard then evaluates multiple signals associated with the quality and consistency of that perception.
Depending on the resulting trust level, the agent can:
- ACT when visual evidence is sufficiently reliable.
- VERIFY when additional analysis is required.
- RE-OBSERVE using additional OpenCV processing when uncertainty remains.
- HUMAN REVIEW when autonomous action cannot be justified safely.
This makes uncertainty an explicit part of the agent's decision process instead of silently treating every perception as trustworthy.
How I built it
VisionGuard was implemented as a modular Python application with OpenCV 5 at the core of the perception pipeline.
The architecture separates several responsibilities:
- Vision Engine — processes visual input using OpenCV.
- VCTS Trust Engine — estimates the reliability of the visual evidence.
- Agent Orchestrator — decides whether to act, verify, re-observe, or escalate.
- Human Review Layer — provides a safe fallback for uncertain situations.
- Evidence and Trace Layer — records the agent's perception and decision path.
- API and Web Interface — expose the system through FastAPI and a lightweight interactive interface.
The cloud deployment is designed for AWS and uses infrastructure-as-code for reproducibility.
The AWS architecture includes EC2 running on AWS Graviton, together with services such as Amazon S3, DynamoDB, IAM, and CloudWatch for evidence handling, traceability, permissions, and observability.
A dedicated deployment path uses Cloud Optimized OpenCV (COOL) for AWS Graviton, allowing VisionGuard to run OpenCV workloads on an optimized ARM environment.
The repository also includes a benchmarking framework designed to compare equivalent OpenCV workloads between a standard OpenCV environment and COOL on the same AWS infrastructure.
Agentic Vision
VisionGuard is not simply a computer vision application.
OpenCV is part of the agent's reasoning and action loop.
When perception confidence is insufficient, the system can invoke additional visual analysis rather than immediately accepting the first observation. The resulting evidence changes the agent's internal trust assessment and therefore changes what it does next.
This creates a closed-loop process where:
visual evidence influences trust → trust influences tool use → tool use generates new visual evidence → new evidence influences the final action.
The objective is to make autonomous visual agents more cautious, observable, and controllable when perception becomes unreliable.
Security by design
Security is treated as part of the architecture rather than as an external layer.
VisionGuard includes:
- explicit uncertainty handling;
- bounded autonomous actions;
- human escalation;
- decision traces;
- cloud observability;
- controlled error disclosure;
- dependency vulnerability monitoring;
- secret scanning;
- static security analysis with CodeQL.
During development, GitHub CodeQL identified an information-exposure issue in exception handling. The finding was remediated through a dedicated security fix and validated again by CodeQL before being merged.
This provides a concrete example of the secure-development lifecycle used for the project.
Challenges I ran into
One of the main challenges was defining a meaningful boundary between perception confidence and permission to act.
A computer vision system can always return some result, but an autonomous system needs another mechanism to determine whether that result provides enough evidence for a consequential action.
Another challenge was making verification genuinely agentic. Re-processing an image alone is not enough: additional OpenCV operations must be selected because uncertainty exists, and their results must influence the subsequent decision.
Finally, benchmarking optimized computer vision requires methodological discipline. Performance claims must compare equivalent workloads under controlled conditions, which is why VisionGuard includes a reproducible OpenCV-versus-COOL benchmark rather than relying on theoretical performance claims.
Accomplishments that I'm proud of
The project demonstrates a practical architecture in which visual trust becomes a first-class security signal for AI agents.
In particular, VisionGuard combines:
- OpenCV 5 visual perception;
- an explicit Visual Cognitive Trust Score;
- adaptive verification and re-observation;
- autonomous versus human-reviewed decision boundaries;
- traceable agent decisions;
- AWS-native deployment;
- AWS Graviton execution;
- COOL integration and benchmarking;
- security scanning and dependency monitoring.
The result is not only a vision pipeline, but an experimental cognitive security control layer for Agentic Vision.
What I learned
Building VisionGuard reinforced an important principle: better perception alone does not necessarily produce safer autonomous systems.
Agentic systems also need mechanisms to reason about the reliability of their own perception, recognize uncertainty, collect additional evidence, and refuse autonomous action when the available evidence is insufficient.
OpenCV can therefore play a role beyond initial perception: it can become an active tool that an agent repeatedly invokes to reduce uncertainty before making a decision.
What's next for CiberIA VisionGuard
The next stage is to extend VisionGuard from the current demonstrator into a broader cognitive security framework for visual autonomous systems.
Future work includes:
- additional visual degradation and adversarial scenarios;
- richer trust signals for VCTS;
- temporal reasoning across video frames;
- multi-camera evidence correlation;
- expanded OpenCV tool selection;
- stronger policy-based action boundaries;
- integration with robotics and edge AI;
- additional human-in-the-loop workflows;
- broader benchmarking on AWS Graviton;
- evaluation of VisionGuard against real-world Agentic Vision failure scenarios.
The longer-term objective is straightforward:
Before an AI agent acts on what it sees, it should have a defensible reason to trust what it sees.

Log in or sign up for Devpost to join the conversation.