Inspiration

Picture a small bakery owner. We'll call her Sarah, though she's a character we made up to tell the story. She opens at four in the morning, and at nine at night she's still deciding how many croissants to bake tomorrow and what to order for the week. She does it from memory, tired, after a full day on her feet.

The problem behind her is real. Small bakeries rarely have anyone whose job is operations or planning, so the owner does it, usually last thing at night.

What really got us was a flaw hiding in the one piece of data every shop already has: the till. When a tray sells out at noon, the till records a good day: sixty sold, nothing left over. What it can't record is the people who came in at two and left with nothing. As far as the data knows, demand was sixty.

That matters because every forecast is built on that data. Plan from raw sales and you plan for sixty again. You sell out again, the till records sixty again, and the forecast learns that sixty was right. The more often a product sells out, the more the numbers understate how many people wanted it. Your most popular items are the ones the data gets most wrong.

So we wanted an agent that doesn't just read what sold, but works out what people actually wanted, and plans from that.

What it does

Sonnabon is an operations and planning manager for a small bakery.

  • Reads the till all day. Receipts land as they happen, and the page follows along.
  • Works out what actually ran out. A product that stops selling dead while everything else keeps selling has sold out. Sonnabon estimates how many people actually wanted it.
  • Plans tomorrow's bake product by product, from corrected demand instead of raw sales.
  • Holds the calendar. It sees an occasion like Halloween coming, measures what that occasion did in the shop's own past sales, and says what to order and when.
  • Hands out the work. Each person gets their list, and it reads back what the team ticked off.
  • Learns the shop. Things the receipts can't say, like a supplier needing three weeks' notice, go into a notes file the owner can read and correct.
  • Stays quiet. It emails the owner only when a decision needs a person.

You brief it once, like a new manager, and it runs from there.

How we built it

One agent on the Strands Agents SDK, with 18 tools grouped by job: reading the shop, doing the math, looking ahead, remembering, and one tool that can reach the outside world.

The core idea is correcting censored demand. For each product we learn its normal sales curve through the day from days it didn't sell out. On a sell-out day we see how far along that curve it got before stopping, and scale up:

$$\hat{D} = \frac{\text{sold}}{F(t_{\text{out}})}$$

where $F(t)$ is the share of a normal day's sales usually done by time $t$.

Tomorrow's quantity comes from the newsvendor model. Running out costs the margin $C_u$, and a leftover costs what it took to make $C_o$, so each product gets its own service level:

$$\text{CR} = \frac{C_u}{C_u + C_o}, \qquad q^* = F_D^{-1}(\text{CR})$$

That's why cheap items get baked past the point of certainty and expensive ones don't.

The agent's limits live in code, not in the prompt. A HookProvider sits inside the Strands loop and cancels the run if it overspends, loops, or tries a second irreversible action. So whether it emailed the owner is answered by the tool call itself, not by what the model says it did.

For AWS we wrote the deployment path: an Amazon Bedrock AgentCore Runtime entrypoint that wraps the same agent, EventBridge Scheduler rules that wake it after closing time, the till export read from Amazon S3, and the owner's email sent through Amazon SES. Each piece switches on once its AWS setting exists, so the same code runs locally.

Around the agent: a Python backend, a web app in HTML, CSS and JavaScript, and a year of simulated trade to test on.

Challenges we ran into

Bedrock. We built the Bedrock provider, but neither of our AWS accounts could get model access. The console kept sending us back to a plan upgrade and registration step we couldn't finish, and the server never responded. Since AgentCore runs the agent on Bedrock, that also blocked the deployment. So we ran everything on a local model through Ollama. It's the same Strands agent with the same tools and hooks; switching provider is one flag.

Bugs hiding in the agent's own loop. Our email safety gate was quietly blocking every email, including the first one. And during the live replay, the shift froze because it rebuilt a whole year of history on every new receipt. Testing the tools one by one didn't catch either; testing the loop around them did.

No real data. We had no bakery's till data, and no public dataset records people who wanted something and left. So we generated a year of simulated trade where the true demand is known, which let us check our estimates against it.

Accomplishments that we're proud of

  • A complete product, not a script: onboarding, a live day view, a team board, a diary of everything the agent did, and a notes file that makes it smarter about one shop.
  • Restraint as a feature. On the simulated data, over thirty nights it reached the owner eight times and handled the rest on its own.
  • On the same simulated year, replaying 120 days, lost sales fell by about half compared with planning from raw sales.
  • 90 tests, including the AWS pieces, and limits enforced by the SDK instead of by asking the model nicely.

What we learned

A sold-out day isn't a good day; it's a day you guessed low. Agent safety belongs in code, because a prompt is a request and a model can talk itself out of it. And testing the tools isn't testing the agent: you have to exercise the loop.

What's next for Sonnabon

  • Test it with real bakeries and real till exports.
  • Deploy on AgentCore with Bedrock as soon as model access works; the code is ready.
  • Connect tills through POS webhooks like Square.

Built With

Share this project:

Updates

Submission history