Inspiration
Razorpay's Track 1 brief was direct: make a merchant transactable by an AI buyer, end to end. Most "AI shopping agent" demos stop at chat that describes products. We wanted the harder version: an agent that actually holds a cart, actually pays through Razorpay, and can't be talked into moving money it shouldn't. And we wanted the same architecture to serve the merchant side too, not a bolted-on dashboard with no agent of its own.
What it does
Vendra is two agents sharing one backend for a demo apparel store, ACME Clothes:
- A shopper agent embedded in the storefront — searches the catalog, manages the cart, walks the customer through checkout, and completes payment via Razorpay in test mode, all through conversation.
- A merchant agent on the dashboard — sales performance over a period, low-stock visibility, and staged price updates, scoped to tools a shopper session can never touch.
Every write that matters — placing an order, changing a price — is staged behind an explicit confirmation and re-verified server-side before it executes. The model proposes; it never directly executes.
How we built it
FastAPI backend, a single LangGraph loop shared by both agent roles, Postgres for catalog/cart/orders, MongoDB for conversation history and skills, Redis for cart locking and session state, Next.js on the frontend.
Key design decisions:
- One agent, not a subagent swarm. A shopping session is one long-running, tightly coupled thread across many intents. Subagent handoffs lose state and add a round trip for no benefit — this follows Anthropic's own commerce-agent design guidance directly.
- Role-scoped tool binding as the actual security boundary. Shopper and merchant sessions share the graph but not the tool set.
add_to_cart/create_orderare never bound in a merchant session;update_product_priceis never bound in a shopper session. Enforced at bind time, not by a prompt telling the model what not to do. - A harness boundary around money and catalog writes.
create_ordertakes only an address_id — the server re-reads the cart and computes the real total itself every time. Orders flip to paid only behind a verified Razorpay signature, never a client redirect.update_product_pricefollows the identical staged-approval pattern on the merchant side. - A per-session server-issued id registry. Every product/cart/order id the model uses has to be one the server actually returned this session. A hallucinated or guessed id is rejected before it reaches Postgres, with a message the model can recover from instead of a raw DB crash.
- A three-segment prompt, assembled fresh every call. Global (static rules), session (skills + memory, append-only), volatile (current page, cart count, timestamp) — the first two byte-identical across calls in a session so they're cache-friendly, the third always last since it always changes.
Challenges we ran into
- Latency was a measurement problem, not a vibes problem. Instrumented every LLM and tool call. Tool calls: single-digit milliseconds. LLM calls: 30 to 45 seconds on one provider, because its reasoning mode was silently burning that time on invisible tokens. Went through four providers in two days — local vLLM, Groq (hit an 8K token/request cap on system prompt plus ~17 tool schemas before any conversation history), a detour through Gemini, back to Groq, landed on Sarvam for its 128K context window. Every swap was a decision driven by a specific measured failure.
- A real degenerate-output bug. Asked a smaller model to regenerate full product objects inside a presentation tool call — data the server already had and was going to overwrite anyway. Under load it degenerated into repeated garbage tokens and hard-failed the call. Fix: the tool takes ids only now, server re-fetches authoritative data.
- Cross-turn state loss in the confirmation flow. A staged action (present approval, wait, then write) spans two separate HTTP turns. Tool results don't get reloaded between turns, only plain text does — so the model lost a real id it had found by searching and reinvented one on confirm. Fixed by echoing the real id back into the confirmation message itself, the same pattern already used for address_id on checkout.
- An IDOR on order lookup. Any authenticated user could fetch any order by id. Fixed with an ownership-scoped query.
- No public webhook for local dev. Added a second, independently signature-verified confirmation path using the signature Checkout.js already hands back on success — cryptographically equivalent to the webhook, just Razorpay's other signed channel.
Accomplishments that we're proud of
Getting the harness boundary right: the model never computes a price, never restates a charged amount from memory, and never touches a write path without staged, re-verified confirmation — verified live by catching a hallucinated id get rejected in production logs, not just in theory. Also standing up a second, role-scoped agent on the same architecture in the time we had, instead of hardcoding a separate merchant flow.
What we learned
Where the actual risk lives in an LLM-driven agent isn't the model being wrong in conversation — it's the model being trusted with a side effect. Every bug worth fixing this build was at that boundary: a hallucinated id, a restated total, a lost id across a turn. The fix was never "prompt it harder," it was moving the guarantee into the harness so the model's mistake has nowhere to land.
What's next for Vendra
Turn on real prefix caching at the provider level instead of just keeping the prompt structure cache-friendly by discipline. Add Langfuse for trace-level visibility into tool calls, token spend, and per-turn latency instead of reading raw logs. Move from a single shared demo catalog to real multi-tenant merchant ownership.
Built With
- ai-gateway
- fastapi
- lang-fuse
- nextjs
- sarvam
Log in or sign up for Devpost to join the conversation.