Inspiration

I spent nearly 30 years in food industry.

Stock never runs out at a convenient time. It runs out at the exact moment you cannot do anything about it — mid-rush, both hands busy, twenty people in line. And when the rush finally breaks, you still can't stop: that gap is for reloading. Skewering meat, cutting vegetables, refilling sauce bottles for the next wave.

So the honest answer to "why don't food trucks track inventory properly" is not laziness. There is no moment in the day to do it.

The other half of the problem is that recipes lie. Before the truck, I spent four years making sushi. A piece of nigiri is supposed to be 28 grams. It takes about twenty sessions before your hand reliably produces 28 grams — until then you're somewhere between 20 and 30, every single time. Shaving meat off a kebab cone is exactly the same. The recipe says 75g. The hand says something else.

Every inventory system I've seen assumes the recipe is true. On a food truck, it isn't.

What it does

teruo is an inventory agent for food trucks. The name comes from the English word tell — it calculates, but it never acts. It only tells you.

It learns the gap between the recipe and the plate. Each ingredient carries a coefficient that starts at 1.0. When you weigh what's actually left, teruo compares it against what should be left, and adjusts. After a handful of counts, it knows that your hand serves 6% more chicken than the recipe claims — and every forecast from then on uses your number, not the recipe's.

It knows when to speak.

Before service: today's outlook, what will run short. This is the only moment in the day when you have both hands free and can still buy more. During service: silent. If something is genuinely about to run out, it says "sauce, about 10 servings left" — a fact, nothing else. It never asks "should I reorder?" while you're holding a knife. After close: results, what to count tonight, what to order.

It admits it doesn't know yet. For the first two weeks it says so, with a progress count: learning (count 2 of 5). And it puts a deadline on that — after five counts, it stops saying "still learning" whether or not it got better, and instead reports why the numbers won't settle. "Still learning" is an excuse; "the person shaving changes day to day" is information.

It asks you to count as little as possible. Twenty ingredients on the shelf, two or three requested tonight. Choosing what's worth counting is the agent's job. Make someone count everything daily and they'll quit in three days.

How we built it

Built on the Strands Agents SDK, running Claude on Amazon Bedrock.

The core architectural rule: Python calculates, the AI decides. No arithmetic is ever delegated to the model. The agent decides what to ask, which items are worth counting tonight, and whether a number looks wrong — then calls a Python tool that does the actual math against a single JSON state file.

This separation is why the numbers can be trusted. An LLM that does mental arithmetic on your inventory is a liability.

The tools are deliberately few:

record_sales — sales in, stock down record_count — a physical count, which updates the coefficient record_purchase — free-text delivery notes ("chicken, 1 cone, 3850g") get_stock_status — what's left plus registration tools so the owner configures their own shop by talking

Setup happens through conversation, not a form. The agent asks about the menu, then asks the one question that determines everything:

"When you plate it, is it the same amount every time? Or does it change depending on who's working?"

That answer — not the cuisine type — decides whether an ingredient's coefficient is allowed to move. A shaved-ice stand and a popsicle cart are both "ice cream shops" and behave nothing alike. So teruo never asks what kind of shop you run. It asks how you make the food.

Challenges we ran into

A timezone bug that would have corrupted the learning, silently.** Sales and counts were compared as ISO strings. The moment records moved to Asia/Tokyo, a sale stamped +00:00 and a count stamped +09:00 would sort by text instead of by time — a sale that happened after a count would be read as before it, drop out of the consumption calculation, and push the coefficient the wrong way. Nothing would have raised an error. The numbers would just have been quietly wrong, and the system would have learned from them. The fix has a required order: convert every comparison to datetime.fromisoformat(), migrate or reset the existing records, then change the timezone. Any other order breaks the data.

Accomplishments that we're proud of

The coefficient moves, and moves to the right place. Count, sell, count — and the number leaves 1.0 on its own. Watching a machine discover that a hand serves more than the recipe says is the whole idea, working.

The items that are supposed to not learn, don't. One napkin per serving can't drift. Their coefficient stays at 1.0 by design, which makes them a control line: if the paper ever started learning, the math would be wrong. Having something that must stay still is how we know the rest is moving for real reasons.

No arithmetic reaches the model. The calculation layer is plain Python against a single JSON file, and it can be tested with no model involved at all. That separation is what makes the numbers defensible.

It stays quiet during service. Most assistants are judged on how much they say. This one is judged on when it doesn't.

"Still learning" has an expiry date. After five counts it stops saying it, improved or not, and reports why the numbers won't settle instead.

Scope held. No database, no accounts, no weather API, no automatic ordering. Every one of those was considered and written down as a refusal.## What we learned A guess in a design document is fine — if you go back and replace it with a measurement. "Five counts" was invented at the whiteboard. It survived only because we later simulated it and it happened to be close. Two other numbers didn't survive.

Domain knowledge is a constraint, not decoration. Silence during service isn't a UX preference. It comes from the fact that when the rush breaks, the hands still aren't free — that gap is reloading time. Nothing in the interface makes sense without that.

Never let code infer meaning from a name. Whether an ingredient's amount can vary is stored as metadata, set by the owner's answer to one question. An ingredient called "sauce" tells the code nothing. The moment you pattern-match on names, the system breaks in the first shop that phrases things differently.

Writing the non-goals is a feature. The list of things teruo will never do — and the reasons — turned out to be the part of the design document we referred back to most.

Shipping alone against a deadline means choosing which failures you can survive. A missing feature is survivable. A number that's quietly wrong is not.## What's next for teruo Split the agent by role. Recorder, observer, reporter — the seam is already in the design: one takes input and calls tools, one decides what looks wrong and what's worth counting tonight, one decides when to speak. It ships as a single agent for now, which is honest about where it is.

Tier 2 — ingredients the customer changes. Sauces, toppings, anything poured or piled on request. The variance has a different source, so it needs a different model.

Oil. Fryer oil doesn't deplete, it degrades. Topped up by the can, replaced by the cycle. Judged by how long since the change and how much has gone through it, not by how much is in the tank. Wrong model entirely — it gets its own.

Dry goods and seasonings. Shelf life, not stock level. Blends degrade faster than whole spices.

Concurrency. Two people entering data at once currently means a read-before-write guard that refuses to overwrite and asks for a re-entry. That is enough to stop numbers vanishing, and not enough for real service. It moves to a database when it leaves the hackathon.

Unit systems. unit_type is already in the data model, so grams to ounces is additive rather than a rewrite.

One we haven't built. When a new person joins the crew, the coefficient should probably reset. The system learned a hand, not a shop.

Built With

  • aws-bedrock
  • claude
  • python
  • replit
  • strands-agents
Share this project:

Updates

Submission history