Inspiration
The idea for SightOps came from a very ordinary problem at home.
I live about an hour away from my parents. One day, my mother called me and told me that their dishwasher had stopped working. She wasn't sure what had gone wrong or what she should do next.
Since I couldn't immediately travel there, I asked her to send me photos of the dishwasher — the control panel, the indicators, and different parts of the appliance.
The troubleshooting process became a conversation:
"Send me a photo of the panel."
I inspected it.
"Can you take another photo closer to the display?"
I inspected that.
"Now check this part and tell me what you see."
We gradually collected enough evidence to understand the problem and resolve it.
Afterwards, I realized that what we had just done was essentially an active visual troubleshooting loop:
Observe the problem
↓
Inspect visual evidence
↓
Form a hypothesis
↓
Need more information?
↓
Ask for another view
↓
Inspect again
↓
Understand the likely problem
↓
Recommend an action
That led to a simple question:
Why couldn't an AI agent do this instantly?
Imagine my mother pointing her phone at the dishwasher and saying:
"It's not working. Can you figure out what's wrong?"
Instead of immediately guessing from one image, an AI agent could inspect the appliance, identify its visible state, read indicators or gauges, determine whether the available visual evidence is sufficient, and then guide her:
"I can see a warning indicator, but I can't clearly see the display. Move the camera closer and slightly to the right."
The agent could inspect the new view and continue troubleshooting step by step.
This could be particularly useful for elderly people living away from their children or technically experienced family members.
But while thinking about the problem, we realized that the same situation occurs in many other environments.
A novice technician may encounter equipment they have never seen before while the senior engineer is unavailable.
A maintenance worker may be standing in front of an unfamiliar control panel.
A field engineer may need expert guidance at a remote site.
A junior data-center technician may see an abnormal indicator but not know which component should be inspected next.
In every case, the fundamental problem is the same:
The equipment is physically in front of someone, but the expertise required to understand it is somewhere else.
That became the inspiration for SightOps — an Agentic Visual Reliability Engineer.
Our goal is not simply to create another system that recognizes objects or describes images.
We want SightOps to behave more like the troubleshooting conversation I had with my mother:
look at the evidence, recognize uncertainty, ask to see something else, investigate further, and only then recommend or take an appropriate action.
That is why SightOps is built around:
See → Measure → Reason → Investigate → Decide → Act
What it does
SightOps turns an ordinary camera into an agentic visual troubleshooting system.
The camera could initially be something as simple as:
- A smartphone camera
- A webcam
- An IP camera
- An existing monitoring camera
Later, the same architecture can support controllable pan/tilt cameras, robotic cameras, drones, and other physical inspection systems.
The user shows SightOps the equipment or appliance that needs troubleshooting.
OpenCV 5 then acts as the system's visual measurement layer.
Depending on the equipment being inspected, SightOps can perform operations such as:
- Detecting equipment regions
- Finding control panels
- Correcting camera perspective
- Detecting LED and indicator states
- Reading analog gauges
- Identifying switch positions
- Detecting visible state changes
- Tracking regions between frames
- Selecting useful frames from video
- Detecting blur or poor image quality
- Estimating whether visual evidence is reliable enough to use
The important difference is that the OpenCV result does not have to be the end of the pipeline.
Suppose the camera sees:
Warning indicator: RED
Confidence: 0.98
Pressure gauge:
Reading uncertain
Confidence: 0.41
SightOps does not need to guess what the gauge says.
The agent can reason:
The warning indicator is reliable,
but the gauge measurement is not.
I need additional visual evidence.
It can then take an action:
"Move the camera closer to the pressure gauge."
If SightOps is connected to a controllable camera, the agent could instead invoke a camera-control tool and reposition the camera automatically.
OpenCV analyzes the new observation:
Warning indicator: RED
Pressure: 87 PSI
Gauge confidence: 0.96
Now the agent has better evidence.
It can correlate multiple observations and decide what should happen next.
Depending on the deployment, SightOps could:
- Explain the likely problem
- Ask the user to perform another visual check
- Retrieve troubleshooting information
- Store visual evidence
- Compare the current state with previous inspections
- Create an incident
- Notify an operator
- Escalate to a technician
- Request human approval
- Trigger a safe automated workflow
This creates a real:
Perception → Decision → Action → New Perception
loop.
The vision result changes what the agent does next.
How we built it
SightOps is designed around three clearly separated responsibilities:
Vision → Reasoning → Action
1. OpenCV 5 — Visual perception
OpenCV 5 performs the core image and video analysis.
We deliberately do not want the entire system to depend on sending every camera frame to a large vision-language model.
For many physical inspection tasks, the important information is measurable.
For example:
Is this LED red or green?
What angle is this gauge needle?
Has this switch changed position?
Is this image too blurry to trust?
Did this region change since the previous inspection?
Those are exactly the kinds of operations where traditional computer vision remains extremely useful.
The planned OpenCV pipeline includes:
Camera calibration
Correct lens distortion and establish consistent geometry.
Perspective correction
Transform angled views of equipment panels into normalized views before attempting measurements.
Region-of-interest detection
Identify relevant components such as gauges, LEDs, displays, switches, or equipment panels.
Indicator detection
Use color-space transformations, thresholding, morphology, contour analysis, and configured regions to determine visual indicator state.
Analog gauge measurement
Detect gauge geometry and estimate needle orientation before converting the angle into a measurement.
Change detection
Compare current observations with previous frames or inspection states.
Image-quality analysis
Evaluate blur, exposure, visibility, occlusion, and geometric quality.
These measurements are converted into structured evidence.
For example:
component: pressure_gauge
value: 87
unit: PSI
confidence: 0.96
component: warning_led
state: RED
confidence: 0.98
The confidence measurement is particularly important because it allows the agent to reason about whether it should trust the observation.
2. Agentic reasoning
The next layer is the inspection agent.
Instead of asking a model to independently interpret every image, the agent receives structured evidence from the vision pipeline.
OpenCV capabilities can be exposed as tools such as:
inspect_panel()
read_gauge(region)
detect_indicator(region)
check_image_quality(region)
compare_with_previous(region)
inspect_region(region)
request_new_view(region)
The agent decides which tool should execute next.
For example:
OBSERVATION
warning_led = RED
confidence = 0.98
pressure_gauge = UNKNOWN
confidence = 0.34
The agent might decide:
DECISION
Insufficient evidence to diagnose.
ACTION
request_new_view(
target="pressure_gauge",
instruction="Move closer and reduce viewing angle"
)
The new image returns to OpenCV.
This makes computer vision part of the agent's tool-use loop, rather than merely preprocessing an image before an LLM call.
3. Active perception
One of the most important concepts behind SightOps is active perception.
Most vision applications effectively ask:
"What can I determine from this image?"
SightOps can ask a different question:
"What should I look at next to understand this problem?"
If the evidence is unreliable, obtaining another observation becomes an explicit agent action.
Initially, that may mean instructing a human:
"Please move the camera closer to the display."
But the same architecture can control a pan/tilt camera:
Agent
↓
inspect gauge
↓
OpenCV
↓
confidence = 0.37
↓
Agent
↓
move camera 12° right
↓
New frame
↓
OpenCV
↓
confidence = 0.96
The physical observation process itself therefore becomes agentic.
4. AWS architecture
AWS provides the cloud execution, reasoning, storage, workflow, and observability layers.
Our planned architecture includes:
AWS Graviton + COOL
We intend to run supported core computer-vision workloads using the Cloud-Optimized OpenCV Library (COOL) on AWS Graviton/Arm infrastructure.
The same reproducible workload will also be executed against an appropriate standard OpenCV 5 baseline.
We plan to measure:
- Processing latency
- Frames per second
- Throughput
- CPU utilization
- Estimated processing cost
This will allow us to evaluate whether the optimized Arm vision path provides measurable advantages for SightOps.
Amazon Bedrock
Amazon Bedrock provides the model/reasoning layer where appropriate.
The model does not replace OpenCV.
Instead:
OpenCV
↓
produces visual evidence
↓
Agent / Bedrock
↓
reasons about evidence
↓
chooses next tool/action
Amazon S3
S3 stores:
- Inspection frames
- Cropped evidence
- Evaluation datasets
- Incident images
- Reproducible test inputs
Amazon DynamoDB
DynamoDB maintains:
- Inspection state
- Equipment/component state
- Confidence values
- Agent workflow state
- Incident metadata
AWS Lambda
Lambda executes event-driven actions such as:
- Creating incidents
- Sending notifications
- Updating workflows
- Triggering integrations
- Executing simulated remediation
Amazon CloudWatch
CloudWatch provides observability across the system, including:
- OpenCV processing latency
- Agent tool calls
- Errors and retries
- AWS resource utilization
- End-to-end inspection traces
The complete flow is:
Camera
↓
OpenCV 5 / COOL
on AWS Graviton
↓
Visual measurements
+ confidence
↓
Inspection Agent
↓
Amazon Bedrock
↓
Decision
│
├── Invoke another OpenCV tool
├── Ask user for another view
├── Reposition camera
├── Store evidence in S3
├── Update state in DynamoDB
├── Trigger Lambda action
└── Request human approval
↓
Observe again
Challenges we ran into
One of the biggest challenges is that real-world vision is messy.
A gauge that is easy to read directly from the front may become difficult when viewed at an angle.
Reflections can change the apparent color of an indicator.
Poor lighting can make red and orange LEDs difficult to distinguish.
Motion can blur an image.
Part of a gauge may be blocked.
A camera may be too far away.
This means the system cannot simply assume:
OpenCV produced a number, therefore the number must be correct.
Instead, we need to treat visual confidence as part of the evidence.
That creates another difficult problem:
When should the agent trust the current evidence, and when should it investigate further?
Too aggressive and SightOps may constantly ask for more images.
Too permissive and it may make decisions based on unreliable observations.
Finding the correct balance between deterministic computer vision, confidence thresholds, agent reasoning, and human escalation is an important part of the project.
Another challenge is deciding where OpenCV ends and generative AI begins.
It would be technically easier to send every frame directly to a powerful multimodal model.
But that introduces cost, latency, reproducibility, and explainability concerns — and it wastes many things traditional computer vision already does extremely well.
We therefore separate:
MEASUREMENT
↓
OpenCV
from
REASONING
↓
Agent / model
Another challenge is safe action.
There is an enormous difference between:
"Send another photo."
and:
"Turn off this machine."
SightOps therefore treats actions according to risk.
Low-risk perception actions can happen autonomously.
Consequential actions can require explicit human approval.
Finally, we need to ensure that our COOL benchmarking is genuinely reproducible. We will use identical inputs and pinned configurations rather than claiming performance improvements before collecting actual measurements.
Accomplishments that we're proud of
The part of SightOps we are most excited about is not individual object detection or gauge recognition.
It is the architecture connecting uncertainty to action.
Consider:
OBSERVATION #1
Warning LED = RED
confidence = 0.98
Gauge = UNKNOWN
confidence = 0.41
Instead of forcing a diagnosis:
DECISION
Evidence insufficient.
That changes the next tool call:
ACTION
Acquire better gauge view.
The new observation produces:
OBSERVATION #2
Warning LED = RED
confidence = 0.98
Pressure = 87 PSI
confidence = 0.96
That changes the operational decision:
DECISION
Abnormal condition confirmed.
ACTION
Create incident.
ACTION
Request human approval
for remediation.
This is the behaviour we wanted when we started thinking about the dishwasher incident.
An experienced person does not simply look once and magically know the answer.
They investigate.
We are trying to give an AI agent that same capability.
We are also proud that the architecture can begin with an extremely accessible use case — helping someone troubleshoot a household device with their phone — while using the same fundamental architecture for professional equipment inspection.
What we learned
One of our biggest lessons is that traditional computer vision and generative AI are complementary rather than competing technologies.
OpenCV is extremely effective when the question is measurable:
Where is the gauge?
What angle is the needle?
What color is the indicator?
Did something move?
Is this frame blurred?
The agent becomes valuable when the question changes:
Is this evidence sufficient?
What should I inspect next?
Do these measurements collectively indicate a problem?
Should I continue investigating or escalate?
Does the next action require human approval?
That separation gives SightOps both deterministic visual measurements and flexible reasoning.
We also learned that uncertainty can itself become an actionable signal.
Instead of hiding a low-confidence prediction behind a final answer:
$$ \text{Low visual confidence} \rightarrow \text{New perception action} \rightarrow \text{Better evidence} \rightarrow \text{Safer decision} $$
That idea is central to SightOps.
Another important realization was that an AI agent does not necessarily need an expensive robot to interact with the physical world.
A smartphone and a human holding it can already create an interactive perception loop.
A $10 pan/tilt camera can take that one step further.
The intelligence can therefore be separated from expensive specialized hardware.
What's next for SightOps
Our immediate objective is to build a reproducible competition demonstrator containing visual elements such as:
- Status LEDs
- Analog gauges
- Switches
- Warning states
- Multiple inspection regions
We will deliberately introduce faults and difficult viewing conditions.
SightOps will need to determine whether the evidence is sufficient, obtain another observation when necessary, and reach the correct final action.
We plan to evaluate scenarios involving:
- Poor lighting
- Blur
- Glare and reflections
- Partial occlusion
- Extreme camera angles
- Ambiguous indicator colors
- Ambiguous gauge positions
- Camera movement
- Network failures
- Model/tool failures
We will measure not only whether the final diagnosis is correct, but also whether the investigation process itself works.
Metrics will include:
- Gauge-reading error
- Indicator classification accuracy
- Visual-confidence behaviour
- Active-perception success rate
- End-to-end task success
- False escalation rate
- Missed escalation rate
- Processing latency
- Throughput
- CPU utilization
We also intend to deploy supported OpenCV operations through COOL on AWS Graviton and benchmark them against an appropriate standard OpenCV 5 baseline.
Beyond the competition
The first inspiration came from helping my mother with a dishwasher.
That remains an important future direction.
A consumer version of SightOps could help people troubleshoot:
- Dishwashers
- Washing machines
- Air conditioners
- Water purifiers
- Refrigerators
- Electrical panels
- Home appliances
- Other household equipment
A user could simply point their phone at a device and say:
"This isn't working. Help me figure out why."
SightOps would visually investigate the problem rather than immediately producing a generic troubleshooting checklist.
The same technology can then extend into professional environments:
- Manufacturing
- Data centers
- Warehouses
- HVAC and building management
- Renewable-energy installations
- Telecom infrastructure
- Utilities
- Remote field maintenance
A technician could arrive at unfamiliar equipment, point a camera toward it, and have SightOps act as an always-available visual troubleshooting partner.
Over time, we envision integrations with maintenance and incident-management platforms so that the agent can move naturally from:
SEE
↓
DIAGNOSE
↓
VERIFY
↓
DOCUMENT
↓
ESCALATE
↓
ACT
The long-term vision for SightOps is therefore larger than equipment monitoring.
We want to give AI agents reliable eyes for the physical world.
Not merely AI that can describe what it sees.
AI that knows when it needs to look again, what it should look at next, when the evidence is strong enough to make a decision, and when a human should remain in control.
Built With
- active-perception
- agentic-ai
- amazon-bedrock
- amazon-cloudwatch
- amazon-dynamodb
- amazon-web-services
- aws-graviton
- azure
- computer-vision
- cool
- docker
- image-processing
- machine-learning
- opencv
- python
- rest-api
Log in or sign up for Devpost to join the conversation.