Inspiration

Every prior OpenCV competition winner analyses one frame, or one session. None of them remember.

But what matters in industrial inspection is not what a panel looks like today — it is what changed since the last time anyone looked at it. Photovoltaic degradation is the clearest case: potential-induced degradation costs affected modules around 15% a year and is partially reversible if caught before saturation, soiling costs 5–20% of annual energy, and a utility-scale plant carries roughly 2,900 modules per MW. Nobody can look at them all, twice.

The same holds for any asset photographed again and again: rust spreading on a steel tank, a crack opening in a concrete wall.

So the question was never "can a model find a crack". It was: can an agent keep the memory of an asset across inspections, and act on the difference?

What it does

afterimage inspects physical assets from still photographs — solar panels, metal structures and concrete — and compares every capture against the memory of the same asset. Each branch is decided by a number, and the number is on screen:

  • It gates its own input. blur_variance under 100, or a frame too dark, too bright or clipped, and the capture is sent back for a recapture instead of being scored.
  • It knows which asset it is looking at. Leave the asset ID empty and it votes among the stored baselines; a split or thin vote ends unidentified instead of a guess.
  • It anchors to memory. OpenCV 5 Features (ALIKED + LightGlue) aligns the new capture to the stored baseline. Under an inlier_ratio of 0.90 it retries with ORB, and under 0.30 it refuses the asset rather than comparing the wrong one.
  • It looks closer on its own. A mean_delta between 30 and 35 is uncertain, so it crops the region and rescans it at twice the resolution on each axis instead of reporting a maybe.
  • It escalates to a human. A severity score at or above 0.4 waits in an approval queue before anything is written to memory. A rejection keeps the reviewer's reason; an approval is followed by reading memory back to confirm the baseline that was written.

Every decision is a trace event carrying {metric, value, threshold, branch}, hash-chained so an edit or a removed event is detected, and any run can be replayed from it. A language model orders the tool calls over MCP and phrases the result; it cannot move a threshold, skip the mandated next tool, or submit a branch the evidence did not produce.

A judge can walk it in a minute without a photo of their own. /app opens on bundled samples — five sets of four, each titled with its use case: solar panels (a hot spot, delamination, a cracked glass), a metal structure (rust on a steel tank) and concrete (a cracked wall). Taken in order, each set runs the first baseline, a recapture, a human-gated defect and a foreign object over a demo asset of the judge's own, and every result links to the next photo. The human gate asks "Is this real damage?" and shows the changed region zoomed, before and now, beside Confirm damage and Dismiss. The interface reads in English or Spanish, in plain or technical language, in a light or dark theme, and works by keyboard.

How we built it

Six perception tools — identify, quality, alignment, diff against memory, crop-and-rescan, severity — run in an arm64 OpenCV 5 container on AWS Lambda (Graviton) and are exposed over MCP. The loop opens an in-process MCP session to them, and the same server answers any outside MCP client over stdio.

Memory is a single DynamoDB table that returns an asset's whole history in one query, with images in S3 and baselines that are superseded rather than overwritten. The policy lives in one file: perception reports numbers, and how they combine and what threshold each meets is the policy's business and nothing else's.

Delivery is CloudFormation from GitHub Actions over OIDC with a permissions boundary, exact pins, an image retention policy, and a calibration check that takes the service out of rotation if the alignment stack stops answering as expected. A failed run can be retried into a new run without touching the original record. An inspection takes 20.4 s and bills $0.0005 at list price.

Challenges we ran into

OpenCV 5 broke compatibility with everything the models know: Features2D is now Features, the C API is gone, ML and G-API moved to contrib. Every API had to be verified against the 5.x docs rather than recalled. Two of them carry the alignment and do not exist in OpenCV 4 — cv2.ALIKED and cv2.LightGlueMatcher — so the project cannot be ported back to 4.x by changing an import. The OpenCV 5 DNN engine has no GPU support, so the system was designed for CPU on Graviton from the first commit instead of being ported later.

The harder problem was honesty. It is easy to build a demo where the agent looks decisive. It is much harder to make every decision reconstructible — which is why the causal link is a field in the trace, not something a judge has to infer by reading several events in order:

{"input_metric": "inlier_ratio", "value": 0.0385, "threshold": 0.3, "branch": "unrecognized_asset"}

Accomplishments that we're proud of

29 scenarios, 18 of them on licensed photographs of real photovoltaic modules. 24 pass every assertion: branch, defect class and required decision path.

Metric Score
Branch accuracy 0.8621 (macro F1 0.8624)
Defect macro F1 0.8542 (precision 0.75 or better on every evaluated defect class)
Mean IoU 0.7875
Scenarios passed 24 / 29
Real photographs passed 14 / 18 (95% Wilson interval [0.5478, 0.91])

The five failures are published, each traced to a root cause rather than explained away. The test suite fails if any headline figure drifts from the evaluation artefact, so the README cannot quietly disagree with the numbers.

What we learned

The interesting part of an agent is not the reasoning, it is the refusal. An unusable capture sent back, an unrecognised asset refused instead of guessed, a severe finding that waits for a human — those are the behaviours that make the loop the product rather than the demo.

And a verification gate catches what review does not. Several bugs shipped past careful reading and were caught only by a check that ran the real thing.

What's next for afterimage

The published failures are the roadmap: a coverage gate that cannot take one global default, an exposure gate firing before severity is ever assessed, a severity score that ignores the class the classifier just produced, and a soiling rule calibrated on a synthetic panel that does not survive real texture. Beyond that: field validation on plant imagery outside the committed dataset.

Known limitations: the published effectiveness applies to the committed dataset only and is not field accuracy; the evaluation measures solar panels only, and the metal and concrete sets are demonstrations with a synthetic defect on one photograph each; the trace's hash chain detects ordinary editing but is not a signed audit log; the public deployment has no login and is a bounded demonstration.

Built With

Share this project:

Updates

Submission history