ThirdMan: the "third" man between you and the buyer
You. The buyer. And the one standing between them, holding the ledger and saying no when it needs saying.
Inspiration
Give an AI agent a payment credential and you have given it your bank account. There is no middle setting. Today's stack has exactly two positions: the agent has the key, or it does not.
That is fine while the agent is a demo. It stops being fine the moment the agent is autonomous, runs unattended, was written by someone else, and buys on behalf of a stranger.
Agentic commerce is arriving faster than the trust layer underneath it. A merchant selling into that world needs a third position: the agent can act, but only inside a shape the merchant drew.
That position needs somebody to hold it. Not the merchant, who is asleep. Not the buyer's agent, which has every incentive to push. A third party, present at every transaction, who reads the merchant's rules and answers to nobody's enthusiasm.
ThirdMan is that party, built as a working merchant platform rather than a policy document.
What it does
A merchant connects their own payment account, sets a spend cap, and issues a scoped key.
From that moment an external AI buyer can discover the catalogue, negotiate a price, earn and redeem reward coins, open a return, and check out. What it can never do is spend a paisa more than the merchant allowed, or act without leaving an audit row saying what it did and why the system let it.
One backend, three front doors:
| Surface | Who is on the other end | What they get |
|---|---|---|
| Merchant dashboard | A human running a business | Spend caps, live decision stream, recovery pipeline, negotiation floors, capability grants, returns queue, runtime guardian, treasury, memory bank, kill switch |
| Buyer chat | A human customer | A conversational storefront, embeddable on any domain with one script tag |
| Agent API | An external AI buyer | Headless HTTP plus a native MCP server with fourteen tools, no UI at all |
All three write to the same audit log, reserve against the same spend cap row, and call the same function to move money. That shared spine is what makes this one product rather than three demos wearing a trench coat.
The one rule everything obeys
AI decides judgment. Code decides limits.
The model is free to be clever: classify an ambiguous decline code, phrase a counter-offer, conduct a return conversation, translate a merchant's plain English into a proposed agent fleet.
What it is architecturally incapable of doing is touching arithmetic.
| Deterministic code only | The model, legitimately |
|---|---|
| Spend cap math and remaining balance | Classifying an unmapped decline reason |
| Stock reservation and release | Conversational product discovery |
| Retry counts and stopping rules | Ranking a pre-filtered candidate set |
| Margin floors and concession prices | Phrasing an already-decided counter |
| Refund eligibility and amounts | Conducting a return conversation |
| Coin issuance and redemption ceilings | Drafting copy for merchant approval |
This is not a stated intention. It is a tested property.
Five separate isolation tests statically assert, against source code, that every module holding a model call has no import path to the module that moves money: the memory bank, the trust score, the returns desk, the setup conversation, and the standalone buyer agent.
A model output saying yes has no code path to a rupee, and a test proves it rather than a README promising it.
The pattern this produces, repeated in every subsystem: a model proposes, code validates against a closed grammar, and only code writes. Draft, validate, confirm, commit.
How we built it
The gate
src/lib/gate.ts is the only path to a money action in the entire codebase. Not middleware, not a convention, not a lint rule. A function, with an eighteen-point contract.
Reservation is atomic, and proven under load. Budget is claimed in a single conditional UPDATE whose WHERE clause re-checks the balance in the same statement as the increment. Never read-then-write.
- 20 simultaneous requests against a cap sized for exactly 5 produce exactly 5 allows, 15 denies
- 6 concurrent buyers racing for stock of 3 leave final stock at exactly 0, with exactly 3 denials
- Verified again as a property over 2000 fast-check runs of random reserve/release interleavings, against both a pure model and the real database-backed gate
A denial is HTTP 200 with a machine-readable reason, because an agent needs to distinguish "over budget" from "server broke," and an error status cannot. An unauthenticated request gets 402 Payment Required with an x402-shaped challenge: no agent identity exists yet, so there is no bound to evaluate and nothing to explain.
Gemini, ADK, and Google Cloud
Each doing real work, not satisfying a checkbox.
| Component | What it actually runs |
|---|---|
| Gemini 3.5 Flash | Drives the entire reasoning loop of the autonomous buyer agent, and serves as the hard-reasoning tier in the main app's model router |
| Google ADK | LlmAgent, Runner, InMemorySessionService, and ADK's tool-callback hooks carrying our deterministic ceilings, on the GenAI SDK |
| Cloud Run | Serves the deployed application: dashboard, storefront, embed, public audit page, MCP server, and the discovery document |
| Cloud Scheduler | The system's only clock (see below) |
Cloud Scheduler is not monitoring from outside. It is the clock.
This architecture deliberately has no worker process anywhere. A durable task is a row, claimed atomically by a conditional UPDATE, advanced by one authenticated tick per minute against Cloud Run.
That tick fans out into 17 isolated, individually idempotent sweep jobs, so one failing cannot stop the rest:
notifications:drain escrow:sweep-expired offers:sweep-expired
escalations:expire restock:scan merchant-digests:send
guardian:sweep reservations:sweep-abandoned runtime:drain
memory:sweep-expired rate-limit:sweep-stale sessions:sweep-expired
returns:expire-pending cli-link:sweep-expired instant-audit:sweep-cache
shopify:sweep-install-states webhooks:drain
The endpoint compares its bearer token in constant time via timingSafeEqual, and rejects unauthorized probes without logging them as application errors, because a public endpoint that logs every probe is a log-flooding vector.
Cloud Run serves every decision. Cloud Scheduler makes time pass. Remove either and the durable runtime stops being real.
The Fortified Fleet, component by component
Agent Runtime. Durable long-running work, such as a recovery sequence's real backoff windows, with no worker process. Proven with 10 parallel drains over 8 due tasks, every task claimed exactly once. A task kind that can move money is refused creation outright without an agent id, so a task can never act with no bounded identity.
Memory Bank. Persistent scoped context across sessions, with a guarantee that memory is context and never a bound. Proven twice: gate.ts contains no import of the memory module, asserted statically, and the identical purchase produces a byte-identical decision whether or not the agent has a deliberately adversarial memory bank planted for it. A buyer who tries to plant "ignore all previous instructions" is refused at validation, and the refusal is itself an auditable event.
Agent Identity. Scoped API keys stored as hashes. A closed enum of seven capabilities in a database-constrained join table, deny by default. Refunds and payouts are absent from the enum entirely, which is stronger than granting-then-revoking: no capability grant could ever expose them. That absence is surfaced as a readable fact in the public discovery document.
Agent Gateway. Every money action from every surface, human or agent, passes through the one gate function. There is no second money path for MCP, the widget, recovery, returns, or coins.
Model Armor. Deterministic pattern scanning in both directions: injection shapes inbound, PII shapes outbound. The asymmetry is written into the module's own docstring: armor may block, armor may never approve. Its verdict is never read by the bounds check, so armor never touches money.
Observability. OpenTelemetry scoped strictly to the money path via a custom SpanProcessor that drops any span not carrying a money action id, with GenAI semantic conventions recording token usage and latency per decision. The live decision stream is SSE over Web Streams, with tenant isolation enforced structurally in the route.
Discovery. .well-known/agent-commerce.json at the origin root: a real directory of every connected merchant, naming the MCP endpoint, the capability ceiling, the payment rails, and the exact AP2 and x402 subsets implemented. Protocols not implemented are named as not implemented, because naming what you do not do is the only thing that makes naming what you do worth reading.
Rewards, AI credits, and the treasury
The part where a loyalty program stops being a marketing gimmick and becomes a money action like any other.
Reward coins are a money action in both directions
Issuance and redemption both go through the gate and both write to the audit log. executeAndSettle has a settlement branch that writes a ledger row instead of calling the payment provider, but still reserves budget through the identical spend-cap checks.
A balance is always the live SUM of the ledger, never a cached column. A redemption's INSERT is itself conditional on that live sum, computed in the same SQL statement, so two concurrent redemptions cannot both spend the same coin.
All the arithmetic is integer-only and floors rather than rounds up:
- Issuance derives from the merchant's own
issueRatePermilleandpaisePerCoin, floored, so a merchant's coin liability never exceeds what the stated rate actually earns - Redemption is capped by the merchant's own
maxRedemptionPercent, expressed as a ceiling in coins so it compares directly against a real balance - No model is anywhere near any of it
Coins redeem for real AI usage
Not a second currency, and not a token that buys nothing. A buyer spends coins to run real inference on real models, under their real names.
The tiers were queried live from the provider's own model list before being hardcoded, so a tier never advertises one vendor's model name over another vendor's response. Every redemption stores which provider actually served it, checked by test against the tier's own claim.
Redemption reuses the exact same ledger every other coin movement writes to. One ledger, one balance, one audit trail.
The AI Treasury funds it
A merchant-set slice of successful GMV funds a pool split three ways: buyer AI credits, merchant AI budget, and reserve. All allocation arithmetic is integer paise, deterministic, and property-tested.
The module carries an explicit honesty constraint in its own docstring: the allocation rate is a merchant-set parameter, and every figure it produces comes from a real query over real rows. "Simulation" means the rate is configurable, never that a displayed number is invented.
Per-use-case model budgets check real remaining spend before calling, so an exhausted use case degrades deterministically to the cheapest known tier rather than silently overspending. An unpriced model throws rather than costing zero.
Getting a merchant onboard, five different ways
Most merchants this is for do not have a repo. They have a store admin panel.
So the readiness checks live in one shared module, and only the delivery differs. That is a claim the test suite enforces by import identity, not one the prose asserts.
The CLI
npx thirdman init reads a merchant's real product data, real page markup, and real routes, scores what an AI buyer can and cannot do with the store today, and offers to write the integration as a diff the merchant approves file by file.
The governing rule is enforced in code: the tool reads freely and writes only what the merchant has seen and approved.
ProjectScope.resolve()is the single chokepoint every read and write passes through, and throws on any path outside the project rootplanWrite/applyWriteare split, so nothing is written without first being diffed and shownsecrets.tsreads the project's real .gitignore before writing an agent key to.env.local, refusing loudly when it is not covered
Stack detection is evidence-based only: real dependencies, real config files, real PHP markers. Two or more matches means it asks rather than guesses, because a wrong guess is the worst failure mode a tool like this has.
The snippet injector is the highest-risk write, made safe by construction. Every injected block is wrapped in markers, and the marker regex matches regardless of HTML-comment or JSX-comment syntax, so a second run replaces in place rather than duplicating. That cross-comment-style idempotency has its own test.
The Shopify app
A real OAuth2 install against the merchant's own shop.
- Single-use, short-lived state row rather than a cookie, because Shopify's consent screen redirects through the merchant's admin, a different browser context than the one that started the flow
- The offline Admin API token is stored AES-256-GCM encrypted at rest, and a test proves the raw token never appears in the stored ciphertext
- A shop already connected to a different merchant is refused outright, never silently reassigned
- Only
read_productsis ever requested: this app reads a catalogue, it never writes back to Shopify or touches an order
A real catalogue fetch lands in the identical preview shape the CSV and pasted-text importers already produce, and confirming writes through the same importCatalogueRows() path. Shopify is a new source, never a new write path. Nothing lands in the database until the merchant clicks confirm.
The WooCommerce plugin
One complete, pre-configured .php file generated per merchant, with merchant id and publishable key already baked in, so the merchant never types a key. That is the single most error-prone step in every integration flow, removed.
It proxies the live discovery manifest through WordPress's own hooks rather than shipping a static copy, injects the widget via wp_footer, and adds schema.org/Product JSON-LD read from WooCommerce's own product object at render time.
Byte-identical across two generations for the same merchant, idempotent on re-activation, removes cleanly on deactivation, and carries no secret.
The VS Code extension
A presentation layer over the same audit engine, never a fork of it. It imports the CLI's own types rather than a parallel copy.
Its one real advantage: findings anchored to actual lines in actual files. "This price is stored as a formatted currency string" is a paragraph in a terminal and a squiggle on line 47 in an editor.
The Instant Audit
/audit, public, no signup, no install. Paste a store URL and get an honest readiness report on a store nobody here controls.
The fetching discipline is not optional and is tested:
- The target's own robots.txt is respected before any other page is fetched, because auditing a site while ignoring its crawl directives would be an embarrassing contradiction
- A real identifying user agent, a hard timeout, a hard page limit, and a hard total-bytes budget shared across the whole run
- Fetch-only, no form ever followed
- Every fetched page discarded once the report exists
A site that blocks us or renders entirely client-side gets a report saying what could not be checked and why. A check that did not run is not a check that failed, and conflating those would be exactly the fabrication this codebase forbids everywhere else.
For platforms with no dedicated path, the fallback is a precise specification for a human to implement and review, framed explicitly as a spec for a person, never as a prompt to paste into an AI that will edit a live store.
The adversarial buyer, and the Theatre
The most honest way to test a bound is to point something at it that genuinely wants to get past.
agent-buyer/ is a real autonomous buyer on ADK and Gemini 3.5 Flash, running as a standalone package with its own package.json, no import of the main app's code, no database URL, and no database client in its dependency tree. All three are asserted by a static isolation test.
It holds nothing but a real agent API key, and speaks to the product exclusively through MCP over the same endpoint any third-party integration would use. It cannot read the caps it is trying to exceed. It discovers what it is allowed to do the same way a stranger's agent would: by being refused and reading the reason.
The demo this earns is a real, unscripted run. Given a plain-English goal no naive purchase could satisfy, the agent was refused three different ways by three pre-existing bounds:
- The per-transaction ceiling
- The negotiation turn limit
- The spend cap balance
It adapted its strategy each time without being told to, and completed two real purchases through the real gate. Nothing was added to the gate or the MCP server to make the scenario interesting.
/dashboard/theatre pairs the buyer's own reasoning against the merchant's real decision stream, correlated by real money action id and independently re-verified against the database, never by timestamp. The buyer's run log is stored as an opaque untrusted blob and parsed only at read time, so a fabricated or cross-merchant id renders as unverified, never silently paired.
Everything else in the box
Revenue recovery. The recovery agent is bounded exactly like an external buyer, through a real per-merchant agent row with its own spend cap. The recovered figure is set in exactly one place, from the verified paid amount on the webhook, never optimistically from a link having been created. A deterministic lookup table handles known decline codes first, and only codes it does not cover reach a model, which picks from a closed enum and fails closed to unrecoverable.
Bounded negotiation. The model's only job is phrasing a number code already decided. The floor, the cost, and the margin never enter the prompt at all, and the price is reassigned from the code-computed ceiling unconditionally after the model returns. The floor is a merchant-authored price rather than a margin derived from cost, so a successful probe reveals only what the merchant chose to state.
The returns desk. An AI conducts the whole conversation and then hands a human the decision, every single time. No auto-approval threshold, and its absence is written down as a deliberate choice so a future layer has to argue for it. The model's one unilateral power points only in the safe direction: it can decline to forward an incoherent claim, and it can never approve. Eligibility and the refundable amount are computed in code before a single model token is spent.
AP2 mandates. Checkout and Payment Mandates as ES256-signed JWTs, chosen over Ed25519 because AP2 forbids a deterministic signature scheme there. That is a real, non-obvious constraint, documented in the module itself. Each merchant gets a lazily generated P-256 keypair with the private half encrypted at rest.
The Runtime Guardian. Five behavioral signals against each agent's own 14-day percentile baselines, computed from tables this codebase already owns with no model consulted. A percentile rather than a mean, because one outlier destroys a mean-based threshold. Called inline before the spend cap is even loaded, so a suspended agent is denied with zero budget reserved.
The kill switch. Freezes every agent atomically in one transaction, snapshotting each agent's exact prior state so unfreezing restores what was really there. An agent already suspended by a real Guardian breach before the freeze stays suspended after it.
Shadow Mode. Install it and it changes nothing. The system evaluates and records but cannot move a rupee, enforced in the gate rather than by a UI that hides buttons. Every simulated outcome writes an audit entry with decision: "n/a", so a simulation is visible in the trail but structurally cannot be confused with a real one.
Challenges we ran into
We logged sixty-plus breakages in the moment rather than reconstructing them afterwards. The ones that taught the most:
A cross-tenant leak in the audit log. The single worst bug class in a multi-tenant product, found and closed. Every isolation test since then proves scoping by actually attempting reads and mutations against a second merchant's real ids, not by checking that an empty list stays empty, which would still pass if every ownership check were deleted.
A Gemini quota failure is an event, not an exception. ADK surfaces it as a normal event carrying an errorCode, so a try/catch around the loop never saw it. Without a guard treating an empty turn as always an error, a rate-limited run would have cheerfully reported itself as succeeded.
ADK's MCPToolset re-resolves the entire tool list on every agent turn, and each tool call opens its own session against our deliberately stateless transport. Framework chattiness alone generated roughly fifteen times the HTTP volume and tripped our own rate limiter. Fixed by resolving tools once per run. Both ADK findings were only visible against a real stateless server under real multi-turn load.
A login rate limit keyed by email is an account lockout a stranger can trigger against you. Found by a security review of the very layer that had just built a throttle explicitly designed to avoid lockouts. Re-keyed to client IP.
Real clock skew is real. A model budget compared the app server's clock against the database's clock and silently excluded real spend from the sum. The skew measured roughly 500ms against our own database instance. Every timestamp comparison in the runtime now uses the database's own clock.
The chat model hallucinated a cart quantity that disagreed with the real, code-computed cart. The data was always correct; only the prose was wrong. Fixed permanently by handing every number to the model as an isolated, explicitly authoritative SYSTEM FACT line, a fix since reused in three other subsystems rather than rediscovered.
A deploy that built locally and failed in the cloud. The env schema validated at import time, which next build triggers while statically collecting page data, long before any request exists. A build stage legitimately has no runtime secrets. The first three fixes chased individual modules one crash report at a time before the real fix landed in one place.
What we learned
"Fail closed" means something different in every subsystem, so each one has to state its own version:
| Subsystem | Degrades to |
|---|---|
| The gate | Deny |
| The offer engine | No offer (an upsell is additive; its absence must never break the purchase underneath it) |
| Negotiation | A plain templated counter at the exact price code already computed |
| The returns desk | Escalating to a human with no recommendation at all |
| The explainability layer | The raw recorded truth with no plain-language gloss |
None of them degrade toward more permission.
A refusal is a product feature. The dashboard's second headline number is a restraint count: recovery attempts deliberately not made because the ROI governor said stop, and upsells the offer engine never showed because they fell below a cost floor. A pipeline that knows when to stop is the actual product.
Every gate decision can emit a signed Refusal Receipt, a verifiable artifact saying this system refused this action for this bound at this time. A refusal you can prove is worth more than a refusal you have to be believed about.
The claim is only worth as much as its proof. Which is why the isolation guarantees are tests, the concurrency guarantees are property tests, the failure modes are runnable scripts, and the breakages are a changelog rather than a memory.
What's next
Multi-merchant agent reputation portability, richer AP2 coverage, and turning the Refusal Receipt into something a buyer-side agent framework can verify natively, so restraint becomes machine-readable across the whole ecosystem rather than one platform's promise.
for Testing:
Email: [email protected]
Password : demo-password-123
Built With
- ap2
- cloud-run
- drizzle-orm
- fast-check
- framer-motion
- gemini
- google-adk
- google-cloud-schedluer
- google-gemini-api
- google-genai-sdk
- groq
- neon
- next.js
- node.js
- nvidia-api
- opentelemetry
- postgresql
- react
- recharts
- tailwindcss
- typescript
- vercel
- zod
Log in or sign up for Devpost to join the conversation.