Inspiration

Public bids are rejected on paperwork before anyone reads the offer. Not outbid — rejected. The price, the method, the team: never evaluated.

A French tender pack runs 150–200 pages and hides dozens of administrative obligations: certificates to attach, insurance minimums, turnover thresholds, forms to sign. A small firm can lose a contract because a professional indemnity attestation expired eleven days before the submission date.

Nobody catches that by reading. The obligations are scattered across four or five documents, they use different words for the same requirement, and the ones that matter most are dates — precisely what a human eye is worst at checking.

What it does

Give it a tender pack and your company's evidence library. It returns a compliance matrix:

┌──────────────┬──────────────────────────────────────────────────────────────────┐ │ Status │ Meaning │ ├──────────────┼──────────────────────────────────────────────────────────────────┤ │ Covered │ met, and here is the document that proves it │ ├──────────────┼──────────────────────────────────────────────────────────────────┤ │ Missing │ nothing in the library answers this │ ├──────────────┼──────────────────────────────────────────────────────────────────┤ │ Expired │ the document exists but will not be valid on the submission date │ ├──────────────┼──────────────────────────────────────────────────────────────────┤ │ Needs review │ a match was found, but not certainly enough to assert │ └──────────────┴──────────────────────────────────────────────────────────────────┘

Every row cites a page on both sides — the page of the tender that states the obligation, and the page of the evidence that answers it. A claim you cannot trace is a claim you cannot defend.

It also lists the documents still to obtain with the deadline each must beat, glosses every French requirement in English, writes a self-contained HTML report, and analyses a whole folder of open tenders in one command — nothing pooled, because four tenders have four deadlines.

It deliberately does not decide whether to bid, price anything, or submit anything.

How we built it

Python, the Strands Agents SDK, PyMuPDF. One rule shapes the architecture:

▎ The model observes, the code decides.

The model reads documents and proposes interpretations — it is good at that, and nothing else does it. Every consequence is computed deterministically.

Dates never reach the model. For a document with issue date $i$ and expiry $e$, against a submission deadline $d$ and a maximum age of $m$ months, the verdict is

$$\text{valid} \iff i \ge d - m\ \text{months} ;\wedge; e \ge d$$

and the slack a bid is running on is a plain subtraction, $\Delta = e - d$, reported in days. Arithmetic a model gets wrong looks exactly like arithmetic it gets right, so we never ask.

Thresholds are compared, not judged. Buyers write both $x > 3,124,998$ and $x \ge 138,000,000$, and the difference between strict and non-strict is somebody's bid. The model extracts the number and the operator; the comparison is code.

Counts are counted. "31 of 47 covered" is $\frac{|C|}{|O|}$, never an estimate.

Evidence is matched with a citation or not at all. A match that cannot name its page is downgraded to needs review — never to covered.

Even the language detection is deterministic: a line is glossed into English only if its English function words satisfy

$$n_{\text{en}} \ge 2 ;\wedge; n_{\text{en}} > 2,n_{\text{fr}}$$

decided per row, not per file, because these packs mix both languages on one page.

We tested against real public tender packs from BOAMP, TED and the EU portal, with a deliberately fabricated evidence library. Publishing which of a real company's certificates have lapsed is not something a demonstration gets to do.

Challenges we ran into

A file that lies about being readable. Page 13 of the ANTAI tender states a mandatory declaration on the honour of the bidder — legible to any human, invisible to pypdf, pdfplumber and PyMuPDF alike. Runs of text had been rasterised into image strips: ten on that page, 261 across the file. That is the worst failure a tool like this can have: not a wrong answer someone would question, but a confident and incomplete one. The tool now asks a different question — is any of this page a picture of words — names the pages, and refuses to conclude that anything is absent. "The tender does not ask for X" and "we could not read the part that asks for X" are different statements, and only one is safe to act on.

A detector that was right by coincidence. Our first version compared glyphs drawn against characters returned. It separated the two files cleanly, so it looked correct — and it was measuring section headings. It would have survived the demo and broken on the next file. We shipped the version whose mechanism was verified, not the one whose output looked right.

Real notices broke the design. Some obligations are answered by no document at all — an average turnover above a threshold is a number, and the evidence matcher would have reported missing on a company that meets it comfortably. Some obligations offer three satisfaction paths in one sentence. And a two-year-old company facing a three-year reference window has not failed; it falls on a different path. We return needs review there, never missing, because that error costs a winnable bid.

Windows and Linux disagree about everything. *.pdf expands under bash and arrives as a literal string under PowerShell; NTFS returns directory entries in name order and ext4 in hash order. Both were found by tests, not by users.

Accomplishments that we're proud of

  • 438 tests, all passing with no API key and no network call.
  • Every number in the report can be recomputed by hand; every assertion points at its page.
  • We mutation-tested the suite — a test that survives the reintroduction of its own bug proves nothing. It caught 21 real gaps we would otherwise have shipped believing they were covered.
  • One document's API failure costs that document and no other; the batch names what failed and exits non-zero.
  • The honest failure mode — refusing to answer where we could not read — turned out to be the hard part to build, and it is the part that makes the report usable.

What we learned

That the interesting engineering in an agent is not the prompt: it is the boundary. Deciding what the model is permitted to conclude, and computing the rest, is the entire design.

That a passing demo is not evidence a system works. We caught two components producing correct output through incorrect mechanisms, and only one of them by accident.

That documents written by real buyers break assumptions no specification would ever expose. We wrote the design first, then read four published tenders, and every one of them cost us a rewrite.

What's next for Tender Compliance Agent

  • OCR the rasterised strips, so pages we currently flag as unreadable become readable rather than merely honest.
  • Candidature vs offre — the same missing paper is fatal in one pile and regularisable in the other.
  • Conditional obligations ("le cas échéant le DC4"), so the report stops flagging requirements that do not concern the bidder. Noise is how a report stops being read.
  • Groupements — every member supplies the full document set, while capacity is assessed on the group as a whole.
  • Scoring grids, not only admissibility: a bidder can be admissible and still lose on points.

Built With

Share this project:

Updates

Submission history