Inspiration
Quick-commerce apps made buying groceries fast. They didn't make deciding what to buy any easier - you still open the app, remember what's running low, search for each item one at a time, swap out whatever's gone out of stock, and pull out a card just to pay for milk and bread. Every week.
Agentry isn't a from-scratch idea, and we want to be upfront about that: the product concept and the Playwright browser-automation approach come from a pre-existing project of ours (it's gone through a few names - Zepto402, then Pantry, now Agentry). What is new for this hackathon is the agent itself: we tore out the old hand-rolled orchestration loop and rebuilt it from scratch on the Strands Agents SDK, for the Agents for Humans Hackathon's Everyday Agents track. The inspiration for entering wasn't "let's automate shopping" - we'd already done that. It was "what does this look like when a real agent framework, not a bespoke state machine, is doing the planning and tool-calling."
What it does
You give Agentry a goal in plain language - "restock the pantry," or something specific like "get me a wheat bread and some ice cream." It:
- Plans - checks a household's own purchase history for what's likely due, and turns the goal into a concrete cart: items, quantities, substitutes, a budget it won't exceed.
- Shops - drives an actual live storefront (Zepto or Blinkit) through stealth browser automation: search, compare, add to cart, ride out whatever the DOM throws at it.
- Pays - checks out from the storefront's own platform wallet balance. No card entry, nothing on-chain - just the money that's already sitting there.
- Reports - sends exactly one Telegram message when it's done, and only interrupts you mid-run if something genuinely needs a person's judgment (out of stock, over budget, a price that looks wrong).
There's a web console (chat + a live purchase-history knowledge graph you can watch grow), and a real-time voice mode - you can talk to it, and it shops out loud, using the exact same tools as the text agent.
How we built it
The core is a single Strands Agent, given a system prompt and a set of @tool-decorated Python functions - one file per capability (search_products, add_to_cart, checkout, check_wallet_balance, manage_address, and so on), each with a docstring written for the model, not just for humans. Model provider is swappable through Strands' own abstraction: Gemini by default, with AWS Bedrock and Groq as alternates.
Every tool drives a persistent, already-logged-in Playwright browser session - one profile per storefront, reused across a whole run so cart state survives from search_products through checkout, the way a real shopping session would. A small NetworkX knowledge graph tracks restock timing and co-purchase patterns, read and written on every completed order.
For voice, we went further than a record-a-clip-and-transcribe loop: real-time audio streams both ways over a WebSocket through Amazon Nova 2 Sonic, using Strands' experimental BidiAgent - and it's handed the same tool list the text agent uses, so a spoken "add that to my cart" does exactly what typing it would.
The web console is Next.js talking to a FastAPI bridge that keeps one live Agent instance per session. The landing page and a six-slide pitch deck share a handful of custom WebGL background components (flowing "strands," ghost fibers, floating lines) built for this submission.
Challenges we ran into
Most of the hard problems here weren't design decisions - they were bugs we had to actually chase down against live systems, and being honest about them is more useful than pretending it was smooth:
- A real concurrency bug, not a theoretical one. Strands dispatches each tool call on its own thread via
asyncio.to_thread. Playwright's sync API is permanently bound to whichever single thread created it. Two tool calls landing on different threads crashed outright withcannot switch to a different thread. The fix was routing every bit of browser work through one dedicated worker thread - anything less was a race waiting to happen. - Live DOM drift, mid-project. Blinkit's quantity stepper used to render as plain +/- text; at some point it silently switched to an icon-font glyph with no text content at all, breaking cart-removal in a way that looked like a "not in cart" false negative rather than a selector miss. Caught it by literally screenshotting the live page and diffing what changed.
- An SDK gap, discovered the hard way. Voice mode's cart display initially showed nothing. The obvious fix - hooking
BidiAfterToolCallEvent, an event that exists specifically for "a tool just finished" - turned out to be defined in the installed Strands package but never actually invoked anywhere in the bidirectional agent loop. Found the one event that genuinely does fire (BidiMessageAddedEvent) by reading the SDK's own source, not its docs. - A dependency landmine. Installing
strands-agents[bidi]for the voice work quietly pulled in a newermcpthat needed a starlette version incompatible with our already-pinned FastAPI - the kind of conflict that breaks a running server on the nextpip installwith no warning until it does. - A one-line CSS property with a two-hour blast radius. The architecture diagram overflowed on every screen size. Root cause:
transform-origin: top centerfighting a flex container alignedflex-start- the scale was shrinking the box toward its own middle instead of the edge the layout was expecting.
Accomplishments that we're proud of
Everything above got verified live, not assumed. Every tool was run against a real logged-in storefront session before being trusted; the voice pipeline was confirmed against a real Bedrock connection with real credentials before a single line of frontend audio code was written. We're proud that when something turned out not to work - a stale selector, an unfired SDK event - we changed course and said so, instead of shipping a demo that only works in a screen recording.
We're also proud of the shape of the final thing: one tool set, one set of shopping behaviors, used identically by a chat window and a live voice call - not two parallel implementations that could drift apart.
What we learned
Experimental SDK features can have real gaps between what's documented and what's actually wired up - reading the installed source paid off more than once. Thread-safety assumptions in an async agent framework wrapped around a synchronous automation library bite quietly, and only under real concurrency, not in a quick manual test.
What's next for Agentry
- More storefronts, and eventually price comparison across them using the same knowledge graph that already tags every purchase by platform.
- Deploying to Bedrock AgentCore Runtime - a deliberate stretch goal we parked in favor of shipping a solid core submission first.
- Hardening voice mode further: cleaner reconnection UX around Nova Sonic's 8-minute session cap, and better handling of overlapping interruptions.
Built With
- agentcore
- amazon-web-services
- knowledge-graph
- networkx
- nextjs
- pnpm
- python
- strands
Log in or sign up for Devpost to join the conversation.