Inspiration

Two problems collided for us.

The first is the shopper's: buying something online means browsing in one tab, comparing in three more, hunting a discount code, and finally being redirected to a checkout that asks for everything again. The decision is a conversation, but the software isn't.

The second is the merchant's. A large retailer can afford to build an AI shopping experience. The skincare shop with six products cannot — and they're the ones who'd benefit most.

But the thing that actually shaped the build was a third realisation, when we read Visa's problem statement closely. Visa is a network and a trust broker, not a shop. Their agentic push — Intelligent Commerce, the Trusted Agent Protocol — isn't about better chatbots. It's about answering one question: is this agent legitimate, and did a human actually authorise this?

Anyone can ship a shopping chatbot in 24 hours. Almost nobody ships one that can prove it was authorised. That became our wedge.

What it does

Any merchant signs up themselves — no admin gatekeeping an API key — brands their store, uploads or edits a catalog, and publishes the same grounded agent as a hosted storefront or a one-line embeddable widget. They then work out of a CRM dashboard built from their own live sales.

Their shopper discovers by chat or by browsing, compares products in a deterministic table, asks general skincare questions, sets a spending limit, previews a server-priced cart, consents explicitly, clears a bank OTP challenge, and gets a receipt — all without leaving the conversation. A live trust rail shows every verification step as it happens.

How we built it

Python 3.11 / FastAPI / SQLite on the back end, React 19 / Vite / TypeScript on the front. One process, several routers — because on demo day, restartability beats microservice purity.

Every architectural decision fell out of a single rule we adopted in hour one and never broke:

Facts travel through deterministic code. Only phrasing travels through a model. No model runs from cart creation downward.

Concretely, the model never sees a price it can change, and never authorises anything:

  • The cart is priced by the server, from catalog rows. The model can propose a basket; it cannot compute a total.
  • Every purchase carries a mandate chain — intent → cart → payment — each link Ed25519-signed, in the shape AP2 describes.
  • The cart hash commits to the shipping address, so an agent cannot redirect goods after you consent:

$$H_{\text{cart}} = \text{SHA-256}\big(\text{canonical_json}({\text{items}, \text{total}, \text{currency}, \text{merchant}, \text{address_fp}})\big)$$

Change any field after approval and authorisation fails with SHIPPING_ADDRESS_MISMATCH.

  • The payment request is a TAP-shaped HTTP Message Signature, verified with a nonce, so a replayed request is rejected rather than charged twice.
  • The issuer can refuse us. The mock ACS runs on its own database and mints a single-use token bound to cart hash, amount, and merchant. If refusals came from our own code, they'd prove nothing.

Catalog ingestion is deterministic too: a canonical Excel template, versioned column aliases, and a staged cleaning run the merchant reviews before anything goes live. A model may rewrite the explanation of what went wrong — never the row counts.

The merchant dashboard follows the same discipline. Every figure is computed from that merchant's own rows on every request, with no metrics table to drift out of sync, and each derived number states its own denominator:

$$\text{conversion} = \frac{\text{paid orders}}{\text{conversations}}, \qquad \text{repeat rate} = \frac{|{c : n_c \ge 2}|}{|{c : n_c \ge 1}|}$$

The revenue forecast is deliberately the dullest method that answers the question — a trailing seven-day mean carried forward, labelled as exactly that on screen:

$$\hat{r}{t+k} = \frac{1}{7}\sum{i=t-6}^{t} r_i, \qquad k = 1, \dots, 7$$

And when the assistant summarises the business in plain English, the guardrail is set membership: let $N(s)$ be the numbers appearing in a generated summary $s$, and $F$ the figures we handed it. The rewrite is accepted only if

$$N(s) \subseteq F$$

Fail that check and the merchant reads the deterministic sentence instead. A dashboard that invents a revenue figure is worse than no dashboard.

Challenges we ran into

Trusting the AI exactly as far as we should. Our first catalog importer used a model to map spreadsheet columns. It worked — until it didn't, silently, on a column named unit_price. We tore it out and replaced it with a canonical template plus deterministic aliases.

Four people and five AI windows on one repo. Mid-build, five commits landed on main while a feature branch was being written — including one that made the platform genuinely multi-tenant. A clean merge would have left the new dashboard silently querying the wrong merchant's data, with no conflict marker pointing at it. We caught it by diffing intent, not just text.

Assumptions that only break later. Three tests asserted COUNT(*) FROM orders == 0 as shorthand for "this flow created no order." The moment we seeded demo history, they failed — correctly. The tests weren't wrong about the behaviour; they were wrong about the world.

A hostile dev environment. One machine had no Node at all, and a parent directory containing a ? that esbuild refuses to parse. We type-checked the entire front end by driving the bundled TypeScript compiler through macOS's built-in JavaScript engine — then deliberately broke a file to confirm the harness could actually catch errors before trusting it.

Accomplishments that we're proud of

We shipped the trust layer, not just the chatbot. Three refusals are live and rehearsable on demand: an over-budget purchase declined before the bank is ever contacted, a replayed bank token rejected, and a cart edited after approval refused. Those are more convincing than any happy path.

We deleted our own AI feature and the product got better. Replacing model-based column mapping with a deterministic template was a real capability loss on paper. It made ingestion trustworthy. That trade is the whole thesis of the project in miniature.

210 passing tests, plus a route-authorisation test that fails the build if any new endpoint ships without a credential check — because "we'll remember to add auth" does not survive hour nineteen.

The single-merchant assumption is genuinely dead. Catalog rows, receipts, dashboard figures, and the agent's own answers all name the actual shop a shopper is standing in. Nothing is hard-coded to our demo store.

What we learned

  • Constraints are a feature. "The model never touches money" sounded limiting on a whiteboard. It turned out to be the whole answer to what if the AI hallucinates a price? — it never sees one.
  • A refusal is a better demo than a success. Showing what the system won't do communicates trust faster than anything it will.
  • Verify, don't assume. Every claim we made about the build, we made after running it — including the ones that came back red.
  • Write down what's true, not what you hoped. Our handoff docs record failures and unverified claims as loudly as wins. That's what let five parallel AI windows and four people work without stepping on each other.

What's next for Sway

More category packs beyond skincare — the pack mechanism already exists, it just needs populating. Real capture and refund flows once acquirer approval allows. Passkey authentication to replace the OTP simulator. And merchant-side agent analytics: not just what sold, but which conversations nearly converted and why.

Share this project:

Updates