Inspiration

Honestly? We got tired of watching computer vision demos that end at "look, it detected a crack!" Cool. Then what? A human still has to walk out there with a flashlight and figure out what to do about it.

That gap bugged us. Detection is basically a solved problem. Decision-making isn't. Most "AI inspection" tools are just detectors with a nicer UI. They see something, they highlight it, they stop. The person still has to do all the thinking.

We wanted to build the part that comes after. What if the system looked at what it found, thought about it for a second, and actually 'did' something — zoomed in, re-scanned, flagged a human, logged it, moved on?

That's AXON. The name is from neuroscience — an axon is the part of a neuron that carries the signal from sensing to acting. That felt right. We're not building a better eye. We're building the nerve.

What it does

AXON looks at video, finds defects, and then decides what to do about them. That last part is the whole point.

It runs in a loop. OpenCV 5 watches the frames and finds problems. Those results get handed to an agent as tools it can call — detect, zoom, rescan, compare, escalate. The agent looks at what the vision system found and picks the next move. Low confidence? It rescans. Serious defect? It pings a human. Two frames disagree? It compares them.

Then AWS does the work — Lambda runs the tools, S3 holds the evidence, SNS/SES wakes up a human when needed, CloudWatch logs everything so we can actually see why the agent did what it did. There's a live dashboard too, so you can watch the whole thing happen in real time.

It's not a chatbot. It's not "explain this image to me." The vision output literally changes what the system does next. That was the hard part, and it's the part we're most proud of.

How we built it

Python and OpenCV 5 for the vision side. COOL on AWS Graviton for the core workload — that's the Arm-accelerated version of OpenCV, and it's genuinely faster and cheaper for what we're doing. Amazon Bedrock for the agent's reasoning. Lambda, S3, CloudWatch, SNS/SES, API Gateway and WebSocket for the plumbing. MCP to expose our OpenCV operations as tools the agent can call. Docker, FastAPI, a small React dashboard. AWS CDK for infra.

The vision pipeline is pretty standard stuff — preprocessing, ROI detection, defect detection, temporal analysis across frames. What made it interesting was turning all of that into tools. Instead of one function that returns "defect found," we have five callable operations with structured outputs. That's what lets the agent actually reason instead of just narrating.

For the escalation logic, we ended up with a simple scoring function. Each defect gets a score based on how confident the model is, how severe the defect class is, and how long it's been there across frames:

S(d)=cd.wd.(1+alpha.deltat(d))

If it crosses the escalation threshold, a human gets pinged. If it's in the middle band, we rescan. Below that, we log it and move on. Nothing fancy, but it's tunable and documented, which means it's reproducible — and that turned out to matter a lot for judging.

We also added a check that fails loudly if the COOL workload silently falls back to x86. That one bit us early. We thought we were benchmarking Arm and we weren't.

Challenges we ran into

The "agentic" bar is higher than it sounds. Our first version was a detector that printed a result and asked an LLM to describe it. We thought that counted. It absolutely does not — the rules are clear that the vision output has to change what happens next. We had to throw out the first design and rebuild around callable tools with structured outputs. That was the single biggest change we made.

Getting COOL running on Graviton was a slog. The dependency chain had to be rebuilt, and we spent way too long not realizing we were still running the x86 path underneath. Added an architecture assertion after that. Never again.

Fair benchmarks are annoying. Same instance type, same inputs, same warm-up, same I/O. We built a small harness that runs both configs and spits out one comparison table. Tedious, but now we have numbers we actually trust.

Debugging an agent is weird. When a model makes a decision, "what" is easy and "why" is hard. We ended up tracing every tool call through CloudWatch so we could see the full chain. Turns out that's also how we evaluate whether failure handling works.

Knowing when to stop. We could've built fifteen tools. We shipped five that work and let the agent escalate when it's unsure. That was harder than it sounds.

Accomplishments that we're proud of

The loop actually works. Not "we claim it's agentic" — we have traces showing the vision output changing the next tool call.

The Arm numbers are real. COOL on Graviton, reproducible benchmarks, comparison against x86. We didn't fudge it.

Human-in-the-loop is a feature, not a cop-out. The system knows when it doesn't know, and asks. We think that's the right design and we're not embarrassed about it.

We documented the failure cases we didn't fix. That felt important. A judge should be able to trust the numbers we did report.

And honestly — it runs. You can upload a video with a defect, watch AXON find it, decide to rescan, escalate to a human, and log the whole trace. Live. Under 30 seconds. That still feels good.

What we learned

-->"Agentic" is a real bar, not a marketing word. You can't just put a chat window on top of a detector and call it done.

-->Perception is the easy part. The decision layer is where all the actual engineering lives. And it lives or dies on the tool interface, not the model.

-->Arm isn't a compromise. We assumed COOL on Graviton would be slower. It wasn't. That changed how we think about deployment.

-->Reproducibility is a feature. Pinned deps and fixed seeds caught two bugs we would've shipped.

-->Restraint is engineering. Five tools that work beats fifteen that half-work.

What's next for AXON

Short term, we want to expand the tool set — depth estimation, multi-camera fusion, stereo reconstruction. OpenCV 5 has the modules for it.

Also want to build an edge version that runs on-device for places without reliable connectivity. Pipelines, farms, offshore rigs. That's where inspection is hardest and where a cloud-only solution falls apart.

Mid term, domain packs. Pre-tuned severity weights and thresholds for manufacturing, agriculture, infrastructure, safety monitoring. Right now you'd have to tune AXON yourself, and that's a barrier.

We also want the human escalations to feed back into the model as labeled data. The more AXON asks for help, the better it gets. That feels like the right loop.

Longer term, we want to open-source the MCP tool interface so other people can build perception tools on top of AXON. And eventually — this is the dream — AXON as the perception layer for robotic inspection arms and drones. Actually closing the loop from pixels to physical action.

The whole idea is one sentence: perception that acts. We built the axon. Next we build the rest of the nervous system.

Built With

Share this project:

Updates

Submission history