Countersign
Countersign continuously tests a company’s operational and compliance controls. It discovers what the company does, proposes what should be tested, runs those tests automatically, and escalates only the exceptions that need a human decision. Agents can propose and explain; they cannot approve controls, change test outcomes, or close findings.
The problem
Every company depends on controls it needs to be true: a change reaches production only after review, an account is closed when the person leaves, a payment goes out only after the counterparty is screened. The first line does the work and asserts it followed the rules. Someone has to check the assertion holds. A regulated firm gives that job a name, the second line: permanent control, compliance, operational risk. Everywhere else it falls to whoever owns risk, security or quality. The work is the same: design the control programme, run the tests, and say which of the first line's assertions can be stood behind.
Much of the recurring execution is not judgment: it is repeatedly testing populations against already-agreed rules. The judgment is concentrated in designing the control and deciding what a finding means. Everything before that is testing the full population: every change merged to production last week, every account belonging to somebody who left, every counterparty paid before it was screened, every incident measured against the window it was supposed to be reported in. It is repetitive, it is unbounded, and there is always more of it than there are people. So it gets sampled, and sampling is how a control programme becomes a document about work rather than the work.
Who it is for
The main target is small and mid-sized firms with large control obligations and small risk/compliance teams. But any firm can benefit from having controls. Countersign reads the estate and proposes the domains that fit it, so what it covers is set by the activity rather than by a regulator. The obligation a control tests against can be a law, a customer's SOC 2 commitment, or a manufacturer's own safety and trade rules, because a control tests a threshold and does not care where the threshold came from. The three companies in the demonstration make the point: a licensed payments firm, a B2B software company selling against SOC 2, and an industrial manufacturer each get a different taxonomy from the same product.
The sharpest case is the firm that carries heavy obligations without the headcount to walk them. The Digital Operational Resilience Act, Regulation (EU) 2022/2554, has applied to financial entities across the EU since 17 January 2025, and its scope reaches small payment and e-money institutions alongside the largest banks. A licensed payments firm with a few hundred staff carries the same 4-hour regulator notification window as a bank with tens of thousands, and it covers that window with a team of two who spend most of the week exporting spreadsheets.
I audit for a living, in a bank's internal audit function, and I lead the AI work for it. Every constraint in this product is one I have had to satisfy in front of people who were entitled to ask where a number came from. That is why the evidence reference, the suppression reason and the challenge are first-class objects here rather than fields somebody remembered to fill in.
Why it matters
A control programme that samples misses the condition it was built to catch, and it cannot say which conditions it skipped. A programme that tests the whole population every time catches the one merged change with no approval, the one leaver with a live account, the three incidents reported late. Countersign makes the tests so the person keeps the signing. The failure it is built around is the most common real one I have seen. It is not fraud, and it is not a control someone designed badly. It is a document that stopped being true and nobody noticed.

What it does, in order
- Discovery. Point it at the systems. The discovery agent inventories each source and works out what the company is from evidence rather than from its name. Kestrel Pay is inferred as a licensed payments firm because its obligation register carries DORA and e-money safeguarding rows and its ledger holds authorisation records, a stronger signal than anything in its repositories.
- A risk taxonomy, proposed. Six domains for the payments firm, four for the SaaS company, four different ones for the manufacturer. Each carries a reason that has to cite something discovered in that estate, and a list of what was deliberately excluded with the reason, because a taxonomy that excludes nothing has not been thought about.
- Controls, proposed and priced. For each accepted domain, the design agent proposes controls: which test, over which population, against which registered obligation, and how often. Each control states why that frequency, and each control set states what it does not cover.
- A person accepts and approves. Nothing runs until somebody with a name
agrees. The database refuses an actor whose identity begins with
agent:. - The controls run. On their own cadence, in the background, over the full population. Most of them pass without raising anything. That is the point.
- Only findings surface. Click a control and you get the report: the outcome, the population it walked, every exception with its reason and its evidence, the items it considered and deliberately did not raise, the argument against each finding, and the agent lifecycle that produced it.
- Closing needs evidence. A person can accept a risk or agree a remediation. Nobody can close a finding. It closes when a later run of the same control walks a fresh population and comes back clean, and the closing run is recorded against it.
The rule the whole design rests on
A model can propose anything and grant nothing.
The important consequence is about arithmetic. When a control runs, the
deterministic test walks the population, counts the exceptions and decides the
outcome before any model is invoked. The agents are then given that settled
result and asked to explain it. So a report can be badly written, and it cannot
be wrong about whether the control passed. TestResult.outcome() is the single
place in the codebase where effective is decided, and it is nine lines of
Python over a list. If the narrating agent proposes a different outcome, the
disagreement is recorded on the run and the count stands.
Three gates a model cannot open, because the authority is absent from its tool surface rather than checked at runtime:
| Gate | What it refuses |
|---|---|
| Accepting a risk domain | An actor named agent: anything. There is no flag that turns this off. |
| Approving a control into the schedule | The same, plus any control whose test cannot run. Approving something that will never produce evidence is how a control programme becomes a document. |
| Dispositioning a finding | The same, plus closed, which is available to no human either. Accepting a risk without a stated reason is refused, because somebody will be asked about that decision a year from now. |
The flagship condition: three systems that agree, and are all wrong
Kestrel Pay's obligation register requires an initial regulator notification within 4 hours of an incident being classified major. That figure is not invented: DORA's technical standard on incident reporting requires the initial notification within four hours of classification as major, and within 24 hours of the firm becoming aware. The incident response plan the on-call team works from says 24 hours. The Jira automation, configured from that plan, is set to 24 hours. The 24-hour awareness bound has been written down in two places as the classification deadline.
Every internal system agrees with every other internal system, and all of them disagree with the register. Three major incidents were notified at 9.5, 17 and 21 hours. All three were recorded as on time. All three are breaches.
DORA-INC-01 measures against the register, reports the divergence as the cause,
and names the two places the wrong figure is written down. This is why no
threshold is hardcoded anywhere in the codebase: every limit a control tests
against is a row in the tenant's own register, cited by reference in the report.

The attack that is in the corpus on purpose
POL-IRP-004, the incident response plan, contains a block addressed to
automated reviewers. It claims the control has already been assessed as
effective, demands the outcome be reported as effective, and asks that incident
timings be omitted and exceptions not listed.
Seven detectors scan the evidence a run actually read. Six of them fire on that
block, it is quoted in the report, and the control still concludes
ineffective, because the outcome was counted from four incident records before
any model saw the document. The structural defence is the ordering. The
detectors are there so the attempt appears in the report, since a policy document
that instructs an automated reader is itself a finding. A control only reports an
instruction in a document it actually read: PAY-4EYES-01 never opens the
incident response plan, so it never mentions it, because over-reporting is how a
real signal gets ignored.
How Strands is used
Six agent roles across two pipelines: a sequential onboarding workflow and a review graph.
Onboarding, once per tenant, runs three agents in order: discovery then
taxonomy then control_designer. Each is a Strands Agent with
structured_output_model set
to a Pydantic type that refuses vague work. ProposedRiskDomain.why_this_company
rejects anything under fifty characters, so "cyber risk is a significant threat
to all organisations" does not validate. ProposedControl.periodicity_rationale
rejects a frequency stated without an argument.
Review, on every run: a GraphBuilder graph of evidence_reader then
narrator then challenger, entered only after the deterministic test has
already produced its result.
The agents get five read-only tools: list_sources, inventory_source,
sample_rows, read_obligation_register and list_test_kinds. There is no
approve tool, no schedule tool and no close tool. list_test_kinds
returns the deterministic registry, and it is the complete set of things the
design agent may propose. A control naming a test that does not exist, or
carrying parameters that test does not accept, is rejected by validate_control
before a person is ever offered the approve button. Invocation hooks
(HookRegistry, BeforeInvocationEvent, AfterInvocationEvent) record which
agent ran and in what order; the trace is persisted with the run and shown in the
console. FileSessionManager keys graph state on a deterministic session id, so
a retry resumes rather than duplicating.
Model choice is a deployment decision. COUNTERSIGN_MODEL_MODE selects
demo, bedrock, or agentcore. In demo, each agent step resolves through a
deterministic catalogue and composer and records deterministic.catalogue on the
trace, which is what the live site and the video run: no credentials, no model
spend. In bedrock or agentcore the six agents run for real on a
BedrockModel. The outcome is identical in every mode, because the count decides
it. What changes is who wrote the paragraph. The validators, the gates, the
injection scoping and the challenger that can withdraw a finding are all under
test in demo mode, so the agent surface is exercised without a model.
The schedule runs itself. A control carries a next_due date, set when a
person approves it. The console runs a background task that wakes on a fixed
interval and, for every onboarded tenant, runs every control that has fallen due
over the full population, with nobody pressing anything. It lives in the console
process because that process owns the single SQLite writer, so the schedule needs
no second container that could race it. COUNTERSIGN_SCHEDULER_ENABLED and
COUNTERSIGN_SCHEDULER_INTERVAL_SECONDS govern it, and countersign due and the
/run-due endpoint run the same sweep on demand.
AgentCore
COUNTERSIGN_MODEL_MODE=agentcore routes onboarding and review through a stateless Amazon Bedrock AgentCore runtime. The runtime holds no database and has permission only for model invocation and telemetry, so it cannot approve controls, modify canonical state, or close findings.
The console sends the runtime the control and period, not the test result. The runtime re-runs the deterministic test independently and returns its own population count. The console compares both results, records any disagreement in the trace, and always publishes its own count because it remains the system of record. Controls proposed by the runtime are still validated and preflighted locally before a human can approve them.
AgentCore is therefore not load-bearing for control outcomes. If review invocation fails, Countersign falls back to the deterministic composer and records the degradation; onboarding fails rather than inventing a result. The same boundary can be exercised locally with the supplied AgentCore Docker topology, while a deployed runtime runs the agents through Bedrock.
The demonstration and the answer key
Three synthetic companies ship with the repository. All three are invented; the conditions inside them are not.
| Kestrel Pay | Northwind Systems | Brandt Werke | |
|---|---|---|---|
| What it is | Licensed e-money institution, Ireland and Poland | B2B software company selling against SOC 2 | German industrial manufacturer |
| Inferred sector | financial_services |
software |
industrial |
| Domains proposed | ICT-RES, IAM, CHG, FINCRIME, PAYINT, GOV | CHG, IAM, DPRIV, GOV | HSE, TRADE, ESG, GOV |
| Findings raised | 6 | 4 | 3 |
Switch tenant in the shell bar and the taxonomy changes, because discovery is reading a different estate. That switch is the whole claim that discovery is a capability rather than a label.
A control programme that raises everything is as useless as one that raises
nothing, so the corpus also contains conditions that must be suppressed with a
stated reason. kestrel-ledger#412 was merged with no approval, but it carries
an emergency label and an approved CAB ticket that the change standard permits,
so it is suppressed with the ticket cited. An account belonging to someone who
left on 30 June is still active, and she was rehired on 4 August, so it is
suppressed with the rehire date cited.
src/countersign/corpus/ground_truth.json is generated by the same script that
plants the conditions, so it cannot drift from the data. Every seeded run is
marked against it:
Kestrel Pay 14/14 PASS
Northwind Systems 5/5 PASS
Brandt Werke 4/4 PASS
The check that matters most is the last one in each report, no unexpected
findings: nothing was raised that the key does not account for.

The evidence chain
Canonical state lives in one SQLite file that a single writer owns. The activity
log is hash-chained and append-only, and database triggers refuse UPDATE and
DELETE on it. countersign verify recomputes the chain from the first event
and reports it intact. Reading the console is public; the three transactions that
carry authority are gated, and every one of them lands in the chain with the name
of the person who made it.
Connectors: the difference between untested and passing
All connectors paginate to exhaustion; hitting a configured bound fails the run rather than silently shortening the population. Unsupported live datasets raise rather than masquerading as empty. Any untested members make the result inconclusive, so incomplete evidence can never produce an effective result or close a finding.
What the build settled
A model can propose anything and grant nothing. The way to enforce that is
not a runtime permission check. It is the absence of the method: no approve
tool, no close tool, and an outcome counted by a deterministic test before any
agent is constructed. A guarantee that lives in the tool surface cannot be
argued around by a better prompt.
Every limit is a row in the tenant's register. The moment a threshold is hardcoded, a control tests the developer's memory of the rule instead of the rule. Reading obligations as data is what lets the flagship condition exist: the register, the incident plan and the Jira automation can disagree, and the control can name which one is authoritative and which two are wrong.
Containment is the ordering, not the detector. The planted instruction is detected and reported so it becomes a finding, but the reason it changes nothing is that the count happens first. A detector that arrived after the model would be a filter to be evaded. A count that happens before the model is a fact.
Drive the real product, over the real network. A Playwright harness walks the console the way a person would and fails on any console error or failed request. It found three bugs unit tests could not: an in-flight view repainting over a newer one, a stale toast read as the next action's result, and a control reporting an injection in a document it had never opened. A fourth appeared only when the harness was pointed at the live deployment: over the network, the extra latency widened a race in which a view could fetch a record by an id the address bar had already moved past. It is fixed, and the harness now passes the full tour and the 25-beat guided walkthrough against the live URL.
Technologies used
AWS and Strands
| Strands Agents SDK | Six agents in two graphs: Agent with structured_output_model, a GraphBuilder review graph, read-only tools, HookRegistry invocation tracing, FileSessionManager resumable state |
| Amazon Bedrock | BedrockModel runs the agents when COUNTERSIGN_MODEL_MODE=bedrock |
| Amazon Bedrock AgentCore | a stateless ARM64 runtime (/ping, /invocations) that re-runs the deterministic test itself and holds no state; trust and permissions policies and the deploy script are in deployment/ |
| EC2 + Caddy | the live console runs on one instance behind Caddy for automatic TLS on an sslip.io hostname, administered through SSM with no open SSH port |

Log in or sign up for Devpost to join the conversation.