Inspiration
I once told a friend that a number was not in her document. It was on page 6, on a scanned page lying on its side. The text extraction had returned four pages of text and no warning, the model answered from what it was given, and the answer was confident and wrong.
Most "chat with your PDF" tools share that blind spot. A scanned page has no text layer, so it is silently skipped. A page OCR'd sideways becomes noise that still looks like text. The model never learns that part of the document was missing, so "it is not in the document" can really mean "it is not in the part I was shown". For a contract, a tender file or a court paper, that is the most expensive kind of wrong answer.
What it does
You pick a sample or upload a PDF, and before any question is asked EveryPage shows a table of what it actually read: each page, how it was read (text layer or OCR, and whether the page had to be rotated), and whether the result is legible.
Then you ask a question and get one of three verdicts:
- FOUND: the answer, with quotes. Each quote was checked in code to be literally on the page it names, and is shown next to a preview of that page.
- NOT FOUND: allowed only when every page was read and every page was examined. Then it is a real negative.
- CANNOT SAY: nothing was found, but some page could not be read or examined. The app names the pages and says plainly that this is not a "no".
The model is never the one that says "it is not in the document". Only the code says it, from the page count.
How I built it
- Reading: each page is read from its text layer, or by OCR (Tesseract, English and Romanian) when there is none. A legibility score, the share of everyday words on the page, tells real prose from OCR noise. A page that reads as noise is tried at other rotations, and marked unreadable if nothing works.
- Examining: every readable page is sent, in parallel, to NVIDIA Nemotron 3 Super (
nvidia/nemotron-3-super-120b-a12b) on Nebius Token Factory, with the question. There is no retrieval step that decides what the model may skip. The model must reply in JSON with verbatim quotes. - Verifying: the code checks each quote against the text of the page it names. A quote that is not there is discarded, whatever the model says about it.
- Answering: the answer is written by the same model, only from the quotes that survived, each with its page number.
- Deciding: the verdict is computed in code from the evidence and the page count. A page whose model call failed counts as "not examined", so a temporary outage produces CANNOT SAY, never a false NOT FOUND.
- The app: a FastAPI service with a single-page interface and no front-end framework, in one container with a read-only filesystem. Uploads live in memory-backed storage and are deleted after 30 minutes. Per-visitor limits and a daily model budget keep a public demo from running up a bill, and every answer shows how long it took, how many model calls it used and what it cost.
How Nebius Token Factory and NVIDIA Nemotron made the difference
The first version of this idea ran a small open model on a CPU: 77 to 185 seconds per question on a six-page document. On Token Factory the same question takes about 4 to 6 seconds, with all pages examined in parallel, and costs about 0.001 USD (seven model calls). That is the difference between a command-line experiment and something a person will actually use.
Nemotron 3 Super returned quotes exact enough to pass a literal, character-level check, in JSON, on every run I measured. The API is OpenAI-compatible, so the client is one short file of standard-library Python.
Challenges
- Being honest about failure. On the first live test, one of six parallel calls failed and the page that held the answer was not examined. The app answered CANNOT SAY instead of NOT FOUND, which is the design doing its job, but it showed that parallelism needed a limit, a proper wait and a retry.
- OCR noise looks like text. A vowel-based legibility test passed a page of garbage. Counting everyday words fixed it, and the same list approach extends to other languages.
- A page lying on its side. Rotating every page would triple the OCR cost, so rotation is tried only when the first reading is not legible.
Accomplishments that I'm proud of
The three-verdict rule is enforced in code and covered by tests that need no model: an invented quote cannot produce FOUND, a negative needs every page read, and a failed or unparsable model reply is never a "no".
What I learned
A retrieval step is a decision about what the model does not get to see. For questions where a wrong "no" is costly, it is worth examining every page, and fast, inexpensive inference is what makes that affordable.
What's next for EveryPage
More languages for the legibility check, larger documents with streamed progress during the examination, tables read as tables, and a self-hosted package for organisations that cannot send documents to a public demo.

Log in or sign up for Devpost to join the conversation.