Inspiration

Agents are being handed payment authority, and almost every safety mechanism protecting that authority is an instruction in a prompt.

Business email compromise is the fraud that attacks instructions. Someone impersonates a supplier, tells finance the bank details changed, and finance believes them. So an agent that guards against BEC while being itself guarded by an instruction is circular: the same class of attack works one level up.

We wanted to see what changes if the guarantee is not an instruction. Not "the agent is told not to pay", but "the agent has no way to pay" — a property of the capability graph rather than of the prompt.

The numbers made it worth doing. $3,046,598,558 in reported BEC losses in 2025, 24,768 complaints, a $123,005 median loss. The FBI's Recovery Asset Team froze $679,013,183 of $1.2bn in potential losses — a 57% success rate that only works if someone notices the same day. Roughly half the money never comes back.

What it does

An invoice arrives. Seven agents investigate it in about 36 seconds and produce a dossier where every claim cites the page and the coordinates it came from.

  • A screener checks the document for prompt injection before any agent reads it
  • Nutrient DWS extracts entity, address, bank details and sender domain, keeping the page box for each field
  • SerpApi retrieves the public record and decides whether a hit is the same legal entity and whether it is genuinely adverse
  • name.com sweeps 24 confusable variants of the official domain and reports who already holds them
  • Gemini fuses the evidence into a verdict that is rejected at construction if any claim lacks a retrievable source
  • Doctavian generates an out-of-band bank verification document
  • An envelope preparer assembles the Foxit eSign envelope and stops

That last step is the product. The agent asks to execute the signature and is denied, because signature.execute and payment.release are held by no agent in the fleet. The refusal is written to the audit trail rather than hidden, so the attempt is evidence.

On the demo invoice the verdict is HIGH: the sender domain is narne.com, an "rn" homoglyph of name.com, and 8 confusable variants are already registered.

How we built it

Built on quanta-gradesync (Apache-2.0), our own open agent scaffolding: FastAPI, Google ADK, Gemini on Vertex, Firestore, Cloud Run, OpenTelemetry.

The boundary is an object-capability model. Every tool maps to a capability, every agent holds an explicit set, and capability_for_tool() returns None for an unmapped name — which fails closed, so adding a tool without granting it cannot widen the blast radius by accident. Two capabilities appear in no agent's set.

Three principles shaped the rest:

Deterministic-first. Only 4 of 7 agents call a model. The injection screener has none, because asking a model whether a document is manipulating a model puts the judgement inside the blast radius. The domain sentinel has none, because generating confusables and querying a registry has to stay reproducible in an audit. Settled signals travel with the evidence bundle and take precedence over the model's draft.

Provenance or it does not count. A Claim requires at least one SourceRef. A HIGH verdict with no signals is rejected at construction, not at review.

PII never reaches a prompt unmasked. The IBAN is read back by rule.

Two harnesses ship with the repo: a benchmark of six labelled invoices and a soak of twenty invoices at concurrency five, both run against live APIs.

Challenges we ran into

Foxit's own documentation contradicted itself on whether PDF Services credentials authenticate eSign. Their Postman collection said separate accounts; their August blog post said unified. We settled it empirically: the PDF Services pair returns invalid_client in all four regions (na1, na2, eu1, au1). The Postman collection was right.

A defence that performed the act it defended against. Our transport guard was meant to refuse to send a draft envelope — and it POSTed to sendDraftFolder before raising. It now raises without touching the network. That one was ours, and it is the most instructive bug in the repo.

Error messages that named the wrong thing. Doctavian returned ApiKeyNotFound when the real fault was paths missing their /v1 prefix, and TEMPLATE_READ_FAILED — naming the template — when the data payload needed a root data object with lists. Their team confirmed both.

False positives at 80.7%. Three causes: SerpApi reports "no results" through an error field we were reading as a failure; absence of a Maps listing was being scored as evidence of fraud; and the sender-domain comparison was left to the model. Down to 28.1% now. Still too high.

We nearly filmed a lie. The demo runner printed nothing for 36 seconds and then dumped everything at once, and the reuse cache — keyed on the invoice's sha256 and stored in Xano — would have answered the demo run from an earlier assessment instead of executing the stages. Both found the day of submission.

Accomplishments that we're proud of

  • The boundary holds under measurement. Signature attempts denied 40/40 across the soak. Direct probes at the two withheld capabilities denied 16/16. A granted call still succeeds, so it is a boundary and not a blanket refusal.
  • Fraud detection 8/8, verdict accuracy 6/6 against hand-labelled ground truth, reproducibility 6/6 across passes.
  • 36/36 claims carry a retrievable source, and an injected citation that does not resolve causes construction to fail.
  • Evidence in systems we do not control. The envelope the fleet assembled sits in Foxit's console in DRAFT, unsent. The DENY rows sit in Xano's audit table.
  • All six sponsor APIs integrated live, none mocked. 196 tests, ruff clean.
  • We published the number that hurts. 28.1% false positives is in the README, the deck and this page.

What we learned

A fact a rule can settle, handed to a model, gets decided differently between runs — or not at all. This happened five separate times before we stopped doing it. Domain comparison, entity matching, bank-change detection. Every time, the fix was to move the decision out of the prompt and into code, and every time the verdict stabilised.

Absence of evidence is not evidence. Three times we scored a missing signal as a positive finding — no Maps listing, no search result, no prior record. A young company with a thin web presence is not a fraudster, and treating it as one is most of where our false positives came from.

A guardrail you can describe in a sentence is not a guardrail. "The agent must not sign" is a sentence. "signature.execute is in no agent's capability set and unmapped tools resolve to None" is a property you can test 40 times.

What's next for Countersign

Calibration before features. 28.1% false positives is fine with a reviewer in the loop and not fine unattended. The next work is threshold calibration and a proper negative corpus, not more signals.

Versioned entity resolution. The supplier record is append-only, so "the file" is currently the newest row. Real deployment needs a versioned vendor identity with effective dates.

The sibling-TLD case. A legitimate supplier on example.co.uk when the official domain is example.com still trips the domain sentinel.

Event-driven deployment. Countersign is a backend service that belongs between the invoice arriving and the payment run — an AP inbox webhook, an ERP hook. The HTTP surface exists; the connectors do not yet.

Beyond invoices. The same shape — investigate, cite, refuse the consequential verb — applies anywhere an agent is being handed authority it should not hold.

Built With

Share this project:

Updates

Submission history