-
-
Full view. OpenCV and ARPI disagree. ARPI finds the tear (break between digit 10-11, +1.4 modules) and its top candidate matches the print.
-
Simple view, the cascade in two lines: what the standard scanner said and what ARPI is doing.
-
The start page.
-
The failure case. Tape over a third of the code, 32 codes fit. ARPI shows the strongest three and asks for a list. The true code is second.
-
Torn and distorted, three digits unreadable from the bars. With the known-code list only one code fits.
-
The first digit is under a piece of paper. Recovered from the parity pattern of the left half.
Everyone knows the pain. You try to read a barcode with a phone, and it does not read. A label that has been through a warehouse, a torn corner, a strip of tape over it. The scanner beeps, you try again, and after a few tries you type the digits by hand.
ARPI (Agentic Reconstruction of Partial Identifiers) reads those codes. When it can read the bars, it does. When part of the code is gone, it works out what the code must have been, shows its working, and shows the options instead of picking one. EAN-13 and UPC-A for now.
Inspiration
1D barcodes are old technology, and for the person holding the scanner the process has hardly changed. I know barcodes. My bachelor's thesis in 2006 was "Implementation of Barcode Functionality in ERP-system": EAN.UCC code structure and a wireless data collection system for a printing ink plant. The handheld reader's fallback back then was the keyboard. It still is.
An EAN-13 has twelve data digits and one check digit, and that is the whole defence:
$$\sum_{i=1}^{12} w_i d_i + d_{13} \equiv 0 \pmod{10}, \quad w_i = 1, 3, 1, 3, \dots$$
When part of the code is damaged, a normal reader fails the checksum and gives up. Or worse, it finds a wrong code that passes the checksum anyway. A locally damaged code does that roughly once in a hundred, and the output can be a drug, a part or a price.
But a damaged code is not a random string. Each digit is seven modules with a fixed shape. The guard bars are fixed. The left half hides the 13th digit in an odd/even parity pattern. The leading digits must be a real GS1 prefix. And in a real warehouse the code is one of the few thousand that exist in the company's ERP. Stack all that and two destroyed digits usually come down to a handful of candidates, often one. With modern tooling it is actually a very simple problem.
The idea is not new. Damaged-barcode reconstruction is patented in several places (Intermec 1998, self-checkout systems that offer ranked candidates, parcel networks that match against shipper and date, and more recently Brady's US12001916 and Socure's US11875259). None of it is something you can download and try. I did not find an earlier system that puts machine vision, a deterministic constraint search and an agent that asks for a better photo together, and keeps all of it invisible to the user.
What it does
Open the page on a phone, point at one code, hold still.
- OpenCV's own barcode decoder goes first. It is fast and reads the easy ones.
- When OpenCV cannot read the code, ARPI takes it. When OpenCV does read it, ARPI checks the read against the bars, because OpenCV is sometimes confidently wrong.
- Every frame adds evidence. Glare and blur move when the phone moves, so the next frame fills in what the last one could not see.
- A small agent decides after each frame: stop, or ask for a different shot ("right half confidence 0.38, other half 1.00"). Every request cites the measurement behind it.
- The result is one of four things: a clean read, the only code in your list that fits, ranked candidates, or what was readable per digit and nothing more.
The rule I care most about: ARPI never asserts a code it had to reconstruct. A reconstruction comes back as options with their fit, and the digits they differ in are marked. When a list leaves only one code, the person confirms it. The system selects, the human confirms.
Optional: load a CSV of the codes that exist in your system. That is the step that turns "probably this one" into "only this one fits".
How I built it
Vision, OpenCV 5. Localise the symbol by gradient orientation coherence (bars have a strong directional signature). Rectify it. Read 32 scanlines. On each line, fit a module grid with offset, width and a perspective bend, and score the fit by asking "does every digit look like some legal digit" rather than "where are the edges". That is why it survives destroyed guard bars. A piecewise warp handles wrinkles and torn pieces that were put back a bit wrong. The output is a confidence map: 95 modules, each with a bar probability and a confidence.
Printed digits. The digits under the bars are read with OpenCV DNN and a CRNN text model. The bar grid already says where digit 7 sits, so each digit is cut out by geometry, no text detection needed.
Reasoning. Every source (bars, printed digits, more frames, the code list) becomes the same thing: a log-likelihood for every digit at every position. They add up inside an exact K-best search over checksum residue and parity. The structural rules stay hard, so no source can produce a code the rules forbid, and each soft source is capped so it cannot outvote clearly read bars.
Agent. Deterministic rules decide stop or reshoot from the measurements. I also wired Claude on Bedrock to the same decision. In live use it added long, often pointless comments on codes already read, and every comment was a model call, so the rules are the default and the model is kept switchable for comparison. Bedrock (Claude Haiku 4.5) writes the two-sentence damage report from the measured facts. Model text that gives advice or runs long is thrown away and replaced by a template.
AWS. One Lambda container image in eu-west-1 serves both the phone page and the decoder behind a Function URL, so the phone gets https and a camera without a separate web host. The phone keeps the accumulated evidence (a few hundred numbers) and sends it back with each frame, so the function holds no state. Camera frames go to the function inline and are never stored. Full-size photos take a separate path through a private S3 bucket and are deleted after decoding, with a one-day lifecycle rule behind that. A DynamoDB counter and hard caps in code limit scans and model calls per day, and model calls per request. Deployed with CDK, dependencies pinned, tests in GitHub Actions.
Evaluation
The dataset was the first thing built. I printed 72 EAN-13 codes on six A4 sheets and damaged them: scratch, tear, crumple, fade, ink smear, tape, glare through a plastic sleeve, motion blur. Ground truth is exact, because I printed it. 23 photos from a Samsung phone, 4080 x 3060. Codes are matched to their truth by OpenCV's reads, the printed id and the sheet layout, never by ARPI's answer.
Two runs. Sheets: every code in the whole-sheet photos, 240 codes. Singles: each code cut out and run through exactly the path the phone runs. 225 crops, of which 15 caught more than one label and are not scored, so 210. "Asserted" means it committed to one code. Ranked candidates do not count.
| sheets, 240 codes | OpenCV | ARPI | ARPI + list | cascade, checked |
|---|---|---|---|---|
| right code first | 152 | 210 | 229 | 212 |
| asserted | 154 | 110 | 228 | 168 |
| asserted and wrong | 2 | 0 | 0 | 0 |
| singles, 210 crops | OpenCV | ARPI | ARPI + list | cascade | cascade + list |
|---|---|---|---|---|---|
| right code first | 160 | 188 | 196 | 186 | 194 |
| asserted | 161 | 103 | 194 | 162 | 190 |
| asserted and wrong | 1 | 0 | 0 | 0 | 0 |
OpenCV's wrong reads pass their checksums. The commonest is an EAN-8 read out of part of an EAN-13. A cascade that trusts OpenCV passes them straight through (2 on the sheets); checking OpenCV's read against the bars removes both.
The last row is the number I watch most. It is not zero everywhere I looked. In the 15 unscored crops ARPI once answered the code of the label above the one aimed at, with the list loaded. An earlier run had three wrong "only one code in your list fits" answers on steep angles and glare, which now come back as candidates. Answers of that kind always carry a Confirm button.
The sets are small (one phone, one day) and 72 random-prefix codes is a forgiving list. Synthetic, 224 rendered labels: OpenCV 144, ARPI 186 first, 85 asserted, 0 wrong. An earlier synthetic run had the harder list, 1000 codes from one company prefix where neighbours differ by one digit: on 48 tears and occlusions ARPI asserted 35 with 0 wrong.
The agent loop, measured on pairs of real photos of the same sheet (an angled or sleeved shot first, the straight one second), 80 labels: it stopped on frame 1 for 49, all right, and asked for another frame for 31, all settled right by frame 2. Never stopped on a wrong code. The second photo was not taken in answer to the request, and it is the easy one, so this tests when to stop, not which instruction helped. One trace: OpenCV read nothing, ARPI's top candidate was wrong in the right half, the agent said "Try the right half from another angle" because the right half's confidence was 0.55 against 1.00, and the next frame settled it. The diagram and full traces are in the report.
All numbers are in the repo under results/2026-10-05, with tool versions and hashes, and in the report.
Challenges
Sub-pixel module width. Everything downstream depends on it, and integer rounding is where naive readers fail.
The real photos found bugs the synthetic set never did. The paper edge is a closed loop of strong gradient and hid every code inside it. The torn-fragment merger once chained twelve codes and a keyboard into one region.
Tape. Translucent tape lowers contrast without lowering ARPI's confidence, so faint bars under it can read as confidently wrong. Those codes come back as ranked candidates, not answers, but the true code is not always first. It is on the list of things to fix.
What I learned
OpenCV's decoder is good and fast and I kept it as the first stage. But it is wrong sometimes, and wrong with a valid checksum. A plain cascade passes those straight through, so the cascade checks it.
An agent that talks is not automatically better. The rules made the same decisions as the model without the chatter.
What's next
Two real problems from my own use, both Code 128:
- GS1-128 labels on our 1 kg ink tins at work. The code is long and curves round the tin, and a static scanner sometimes cannot read it at all. With a streaming camera you can turn the tin, or your camera, and stitch the parts.
- Finnish invoice barcodes for payments, read from a pdf on a pc screen. The bank barcode has its own stack of checks (IBAN, reference number, due date), and a payment is exactly where a wrong code must never be asserted.
Built With
- amazon-bedrock
- amazon-cloudwatch
- amazon-dynamodb
- amazon-ecr
- amazon-web-services
- aws-cdk
- aws-lambda
- canvas
- claude
- docker
- getusermedia
- github-actions
- html5
- javascript
- numpy
- onnx
- opencv
- python
Log in or sign up for Devpost to join the conversation.