Inspiration

I started out wanting to explain confusing government letters. Then I opened a real one: a 45-page EPA pesticide label. The application rate is on page 31, the re-entry interval is somewhere near the back, and getting either wrong is a fine or a lost crop. Four of the six labels I pulled were over 20 pages. That's when it stopped being a letter problem and became a finding problem — the answer is in the document, it's just buried under 44 other pages. And the people who need it most are the least likely to have a lawyer read it for them.

What it does

PigeonEye reads a government document on your own Mac and tells you the parts you're accountable for.

  • Reads every page locally. PDFs and scans, through Apple Vision. 45 pages in 23 seconds.
  • Pulls out the facts that carry obligations — dates, deadlines, application rates, registration numbers, addresses — each one carrying the exact sentence it came from and a confidence score.
  • Knows a form when it sees one. If the PDF declares fillable fields, it lists every one, exactly, from the file itself.
  • Explains it in plain language, and lets you ask questions about the page you're looking at.
  • Exports the lot — JSON, CSV, plain text and PDF, so the findings outlive the app.

The part I care about most: it cannot invent a deadline. Facts only enter through tools that carry a verbatim quote, enforced at one place in the code. There is no path where the model makes one up.

How we built it

Native macOS — Swift 6.2, SwiftUI, PDFKit, Apple Vision. Zero thir

Four layers, and imports only point downward: Tools (Vision, PDFKictors), Agent (orchestration, the step log), Gate (the single placeanything is allowed to touch the network), UI. A shell script greps for violations and runs before every commit — URLSession outside Gate fails the build check, not code review.

The model tier is one OpenAI-compatible client with a swappable baath runs against OpenAI or a local endpoint. (Privacy first for local only)

Separately, a Python evaluation harness in eval/ scores OCR engineh. Every engine decision in this project came out of that, not outof a blog post.

Challenges we ran into

Apple Vision cannot read a PDF. It returns invalidImage("Zero-dimensioned image (0.0 x 0.0)"). Every page has to be rasterised first — that's about 8 of the 23 seconds, and it's mandatory, not an optimisation. The coordinate origin was flipped, and it was almost invisible. Vigin. With it wrong, a letterhead reported y = 0.9375 instead of0.046, and cropping the last line of a page returned "january 11, 2017" — the mirrored line at the top. It looked like a real answer. That's the worst kind of bug in an app whose whole claim is "here's the sentence it came fr OCR confidence lies at the top end. On one page, four 口 glyphs — top 4 of 1,092 lines. Meanwhile a genuinely mangled seal read WEALPROTECTED at 0.08. So the low end is trustworthy and the high end is not, which is the opposite of how you'd naturally use it.
My own validator was 88% wrong, and said "checked" every time. I only found out by dumping every finding across the whole corpus — 12 PDFs, 18 scans, 8,421 findings — into a TSV and reading it. It was labelling 131 distincon numbers, including 42 on an IRS tax guide that contains no EPAanything. The fix took it to 7, which is how many the corpus actually contains, with the real one still found on all 30 of its pages.
Vision crashes under concurrency. Releasing a finished TextRecognition request from multiple threads takes the process down. Per-caller limits don't compose, sothe bound had to be process-wide. One damaged page used to cost you the other 50. A throwing TaskGroso a single unrenderable page unwound the entire read. It'snon-throwing now; each page reports its own outcome.
Local models didn't survive contact with the hardware. mere.run wants 16 GB of memory headroom on a 16 GB machine. RapidOCR scored 47.6% NSCER against Apple's 26.2% — a Chinese-trained recogniser can't read English government then ruled out, and the measurements are in the repo. Accomplishments that we're proud of Honesty is architecture here, not a prompt. A Finding can only be point that checks its quote against the transcript. There's nopublic initialiser. So "every fact is quoted" is a rule the compiler helps enforce, not a promise in a system message.
I measured my own output and it disproved me. The corpus dump was built to confirm the extractors worked. It showed the registration validator was wrong 88% of the time. Nothing about that was visible from reading the code — aluable thing I built.

The egress ledger is written after the wire, not before. It used tment the request was assembled — so an offline machine stillproduced a line claiming your text left. For an app selling privacy, that's the dangerous direction to be wrong in.

90 tests, zero dependencies, layer rules enforced by script. Test-first throughout.

What we learned

Confidence scores are asymmetric. Trust them when they're low, ignore them when they're high. Every gate in the app is built around that asymmetry now.

Reading your own output beats reading your own code. I could have shipped that 88%-wrong validator. It passed its unit tests. It took dumping 8,421 real findings to a file and actually looking at them.

Validators must whole-match, never contain. R G-2 26-O4871 is a res corpus. A containment check passes it the moment any fragmentlooks plausible.

Search patterns and validation patterns are different patterns. Using one for both silently dropped 10.5 0Z — the OCR of 10.5 oz — so the user was never told the rate existed at all. Find it and flag it as failed; discardingr the reader.

The right split for an agent is: deterministic tools decide what'shere to look. A 45-page label fits in no context window, so choosing which pages to read is genuine planning — and it's a decision the model can make without ever being in a position to fabricate a fact.

What's next for PigeonEye

  • Escalate with consent (F5). When a value comes back low-confidence, show the user the exact image crop and ask before sending it anywhere.
  • Fail honestly (F6). Every failure mode given a specific, non-gen
  • Full inspector mode (F7). The step log and egress ledger ship today; next is the whole agent run visible end to end.
  • Measure the confidence thresholds. They're named placeholders rifrom assets/golden/, not from my judgement.
  • A real page-selection strategy. Currently a page cap. A 45-page label deserves better than a constant.
  • Fill the form, don't just list it. The fields are read exactly fing back is the obvious next step.
  • Work more with

Built With

  • applevision
  • gemma
  • localmodel
  • openai
  • swift
Share this project:

Updates

Submission history