Inspiration
Plants are full of analog gauges that will never be replaced with connected sensors. They are read on rounds, by people with clipboards, and the numbers are typed up later, if at all. A bearing can warm for a week without ever leaving its normal band, and a filter's pressure drop lives in two rows nobody subtracts. We wanted the round to take a phone photo and get back a reading someone can trust, plus the arithmetic nobody has time to do.
What it does
- Takes a phone photo of a pressure gauge or dial thermometer and reads it with OpenCV 5, with a calibrated confidence and the photo checks behind it.
- Shows its work in nine stages: the dial it found, the straightened face, the tick ring, the true pivot, the needle line, the numbers it read and the fitted scale.
- On a round, an agent decides what happens next: log the reading, ask for a better photo and say exactly why, flag a photo of the wrong gauge, or hold a maintenance work order when a limit, a week's drift or a filter's pressure drop is breached.
- Only a person releases a work order, on the approvals desk, with the photo, the reading and the agent's full trace beside it.
- On a phone, OpenCV.js 5 coaches the shot live (tilt, distance, glare, focus, steadiness) before the shutter unlocks.
How we built it
- OpenCV 5.0 in Python on AWS Lambda (arm64): Canny and fitEllipse find the dial; one warpAffine removes the camera's tilt; a black-hat stroke map keeps needle, ticks and print; warpPolar finds the tick ring; tick lines meet at the true pivot; HoughLinesP and HoughCircles find the needle and its hub.
- Two OpenCV Model Zoo networks through cv.dnn: PP-OCRv3 finds the printed numbers and CRNN reads each box in its own orientation. OpenCV 5's new DNN engine runs the detector 2 to 4 times faster than the classic engine on CPU, while the classic engine runs the recogniser 3 to 4 times faster, so each model uses the engine that suits it.
- The scale is fitted per ring of numbers with RANSAC over every plausible reading of each number (lost decimal points, lost minus signs), then interpolated between neighbours.
- A logistic confidence model fitted on 600 rendered gauges kept apart from every test set; below 90% nothing is logged.
- The agent: a tool-calling model (Amazon Nova Micro on Amazon Bedrock, the cheapest Bedrock model with tool use, with Nova Lite as backup and a rule engine if neither answers) with six tools. Guards in code decide what each tool may do: a work order needs a measured breach, its priority can't be set below the breach, and any figure the model writes that isn't in the measurements is struck through.
- AWS (Sydney, ap-southeast-2): one CloudFormation stack with Lambda behind a Function URL, DynamoDB, S3, CloudFront, CloudWatch and Bedrock. The site is a static Next.js export. Spending guards are built in: at most 5 concurrent Lambda runs, 400 model-assisted captures a day, and 600 tokens per model reply.
Challenges we ran into
- Real photos fight back: needles with counterweight tails, two printed scales, numbers printed sideways, glare across the needle, and perspective that moves the pivot away from the ellipse centre. The pivot now comes from where the tick lines meet, and the pointer is the end whose unbroken run is longer.
- The text model has no decimal point or minus sign, so "05" might be 0.5 and "20" might be -20. Each is a hypothesis with a cost in the fit.
- Honest evaluation. Real photos don't come with a true value, so every one was read by eye. When the reader was tuned on its own photos, its numbers looked better than they were, so we set a blind batch aside and published that batch's single run as it came out.
Accomplishments that we're proud of
- On 300 rendered gauges never used for tuning, 219 readings were accepted and 99% of them were within 2% of the scale; none was off by more than 5%.
- Every real photo, every failure and every refused photo is on the Evidence page with its credit.
- The agent passes all 8 scenarios with the rule engine and with Amazon Nova Micro, and every decision is traceable from the photo to the work order.
What we learned
- Classical geometry still does the heavy lifting: tick lines and ellipses carry more truth about a dial than any single detector.
- A confidence that means what it says is worth more than a higher hit rate. The blind batch made that concrete: real photos are much harder than rendered ones, and the reader is right to refuse most of the hard ones.
What's next
- Per-gauge enrolment from one good photo, so later reads on a round skip the number reading entirely.
- Better handling of idle needles resting below the first printed number, and of compound vacuum and pressure gauges.
- Offline capture on the phone with sync when the round ends.
Built With
- amazon-bedrock
- amazon-cloudfront
- amazon-cloudwatch
- amazon-dynamodb
- amazon-nova
- amazon-web-services
- aws-cloudformation
- aws-lambda
- crnn
- next.js
- numpy
- onnx
- opencv
- opencv-5
- opencv-js
- pp-ocrv3
- python
- react
- typescript
Log in or sign up for Devpost to join the conversation.