Inspiration

Every circuit board coming off an assembly line gets inspected, and a lot of that inspection is still done by people squinting at boards or by rigid machines that only catch the exact defects they were programmed to find. People get tired and miss things. The machines miss anything they weren't told to look for. And neither one explains itself, so when a board gets rejected, nobody downstream knows why.

I wanted to see if an agent could do the whole job. Not just spot the defect, but decide how serious it is, explain what went wrong, and recommend what to do with the board. The actual quality decision, with a record you can audit afterward.

What it does

You give PCBPatrol a photo of a circuit board and it hands back a full inspection record.

It tells you which defects are on the board and where, with bounding boxes and confidence scores. It gives an anomaly score and a heatmap that flags odd-looking regions even when they don't match a known defect. It writes a plain-language explanation of the defect, what probably caused it, and how to stop it happening again. And it makes a call: a severity rating from 1 to 5, a recommended disposition (pass, rework, scrap, or human review), and how confident it is.

Here's a real run. I fed it a board with a soldering defect. The system found a short circuit at about 85% confidence, flagged the board as anomalous, rated it severity 5, and recommended scrap. The reasoning it wrote back: excessive solder forming a blob-like structure that can cause an electrical short. That's not a canned message. The vision model looked at the solder and described what it saw.

How I built it

The design splits the work in two. Fast specialized models handle detection, the what and where. A vision-language model handles reasoning, the why and what to do. I called it detect-then-reason because that's the order it runs in.

Three models sit behind one FastAPI service with a single endpoint, POST /inspect. Send a base64 image, get back structured JSON.

The first model is RF-DETR, a transformer object detector I trained on six PCB defect classes: shorts, open circuits, missing holes, mouse bites, spurs, and spurious copper. It draws the boxes. The second is PatchCore, an anomaly model from Anomalib. It's the safety net, because it wasn't trained on specific classes and catches the unexpected stuff the detector would miss. I exported it to ONNX so it runs on CPU without the full training stack. The third is Qwen3-VL-8B, an 8-billion-parameter vision-language model I fine-tuned with a LoRA adapter on defect-reasoning examples, running 4-bit quantized to fit on one GPU. It writes the description, probable cause, and corrective measure.

The disposition logic is deliberately not a model. It's a plain Python function, about twenty lines, and you can read every branch. Severity comes from the worst defect on the board. A short or open circuit is a 5 and gets scrapped. Spurious copper or a missing hole is a 3 and gets reworked. If the detector found nothing but the anomaly model is still suspicious, or if the detector isn't confident in its worst call, the board goes to a human. When a board gets scrapped, you can point at the exact rule that scrapped it.

The UiPath piece is a coded agent written in Python with the UiPath SDK. It takes an image by URL, base64, or local path, calls the vision service, and maps the result into a structured Orchestrator job output. I deployed it to UiPath Automation Cloud where it runs serverless. In the run above, the agent pulled the board image from a URL, called the service, and came back as a Successful job in about 40 seconds.

I built both pieces working with Claude Code the whole way. I'd describe what I wanted, it would implement, and I'd review the diff against the contract before anything got committed. That review gate caught real bugs, like the disposition function being handed the wrong data type before it shipped.

Challenges I ran into

Getting the anomaly model's preprocessing right took longer than the model itself. PatchCore was double-normalizing images, once with a divide-by-255 and again with ImageNet statistics, which smeared the heatmap into uniform mush. I ran four preprocessing variants side by side to find the one that gave a clean signal.

Deployment auth was its own adventure. My account lives on UiPath's staging environment, not production, so the CLI kept failing against the wrong server until I pointed it at staging. Then publishing kept returning a 403 because my access token was missing scopes the publish call needed. The fix was granting the token full Orchestrator access instead of guessing one scope at a time which one it wanted next.

Then there was the input size limit. UiPath caps input arguments at 10,000 characters, and a base64-encoded image blows past that instantly, mine was 96,100 characters. So I added a URL input to the agent and had it download the image itself, which made the limit irrelevant.

The anomaly score also saturates at 1.0 on the test boards. I dug into it and decided to accept it rather than fight it. The anomaly model was trained on a different image distribution than the test set, so it reads every test board as maximally anomalous. I leaned the disposition on the detector, which is calibrated, and kept the anomaly model as a binary flag for the unexpected-defect case.

Accomplishments that i am proud of

The thing actually runs end to end on real infrastructure. A coded agent on UiPath Automation Cloud, running serverless, fetching an image, calling a GPU service with three models, and returning a structured disposition into an Orchestrator job. That's not a mockup or a local demo, it's a Successful cloud run with real output.

Three models, three different jobs, cooperating behind one clean API. The detector localizes, the anomaly model backstops it, and the vision-language model explains in language a human can act on. And the final decision stays out of the black box, in readable Python, so it's auditable.

What I've Learned

Most of the hard-won lessons were the kind you only hit when something is actually running rather than just compiling. The Qwen model needed a specific argument popped out of its inputs before generation or it would error, which I only found by checking the current docs against the failure. The base64 size limit only showed up when a real image hit the agent. The auth scopes only failed at publish time, not at login.

I also learned that the boring part, the disposition rule, was worth keeping simple on purpose. It would have been easy to train another model to make the pass/rework/scrap call. Keeping it as plain logic means the quality decision is explainable, and in an inspection context that matters more than squeezing out another point of accuracy.

What's next for PCBPatrol

The natural next step is a UiPath Maestro case that wraps the agent with an Action Center approval step. Right now the agent recommends scrap. A Maestro case would hold that recommendation and route it to a human reviewer who signs off before any board actually gets rejected. The reviewer sees the image, the heatmap, the reasoning, and the recommendation, and either approves or overrides. The disposition logic is already built for this, since human_review is one of the calls it can return.

Beyond that, batch inspection so a whole tray of boards runs in one pass, and routing recurring defect patterns to supplier non-conformance workflows so the same soldering problem stops showing up.

Built With

  • anomalib
  • fastapi
  • lora
  • onnx
  • python
  • pytorch
  • qwen3-vl
  • rf-detr
  • uipath
  • uipath-automation-cloud
Share this project:

Updates

Submission history