Inspiration
In retail, one bad decision can mean lost sales, wasted product, and misused resources. And most of the time, the problem isn't a lack of data. It's knowing how to read it and what to do about it.
We saw this firsthand in the data from Glazed Co., a coffee shop chain with locations in airports, train stations, and resorts (38 stores, 8 suppliers, one year of operations). They were losing money in two directions at once:
- Waste: 21.9% of everything received ends up in the trash (≈ $1.19M at cost).
- Stockouts: 20.3% of store-product days end with the shelf empty (somewhere between $1.4M and $3.5M in lost sales, depending on assumptions).
Kiosks were throwing product away while high-traffic stores were running out of it. The suppliers weren't the problem. The reorder logic just wasn't tuned to each store.
What it does
Glazed Copilot turns daily operational data into clear, well-reasoned, results-driven recommendations:
- Spots problems: upcoming stockouts, over-ordering, product about to expire, and late or incomplete deliveries.
- Suggests options: each one comes with its benefit, cost, urgency, risk, and confidence level, including the option of doing nothing.
- Shows its work: what happens if you act, and what happens if you don't.
- Lets people decide: the person in charge can approve, reject with a reason, or ask for another option.
- Measures and learns: it compares expected vs. actual results and adjusts future recommendations.
It's not about replacing people. It's about making sure that even someone with little experience or training can understand the problem, see their options, and have a real say in decisions.
Everyone sees what they need
| Role | What they can do |
|---|---|
| Floor employee | Check the daily summary and drill into each issue |
| Store manager | Review options, make decisions, and track results and ROI |
| Area manager | Oversee stores in their region and propose product transfers between them |
| Admin | Set goals, traffic-light indicators, and thresholds |
How we built it
Our core rule: code does the math, the LLM does the talking. Forecasts, quantities, deadlines, losses, and confidence levels all come from deterministic, tested functions. The AI looks things up, chats, explains, and suggests, but it never makes up a number.
- Reorder engine. It forecasts demand by store, product, and day of the week, then projects inventory day by day (what's on hand, what's on the way, and what's expiring) to figure out when each product would run out. From there it decides when to order, based on each supplier's delivery days, and how much, balancing waste against stockouts (a newsvendor model for perishables and a reorder-point model for everything else).
- AI agents with a clear chain of command. An Orchestrator coordinates four specialists (Present, Past, Supply, and Strategist). The specialists don't talk to each other, so there's a single point for synthesis and auditing. On top of that, the Sentinel, an independent agent, checks that whatever got approved was actually carried out and actually worked. Whoever proposes doesn't get to audit.
- Tiered autonomy (L0–L3). The system only acts on its own when a decision is safe, reversible, and low-impact. Otherwise it makes a suggestion or opens a conversation with the manager. The level is calculated by code, not by the LLM.
- Transfers between stores. A Network Mediator connects stores without exposing their data: they only share ranges in boxes and stay anonymous until both managers approve.
- Per-store isolation with Row-Level Security in PostgreSQL. Privacy is enforced by the database, not by the prompt.
- Experience memory. Every decision and its outcome get logged, and before citing a past experience, the system double-checks it against real data.
Challenges we ran into
- Censored demand. When a product sells out, sales underestimate real demand. We had to correct for that so we weren't forecasting from biased data.
- Data leaking from the future. Actual delivery dates and quantities aren't known when the order is placed. We defined in-transit orders as of the cutoff date and projected their arrival using each supplier's historical profile.
- Business rules nobody follows. The minimum order quantity is almost never respected, so we treat it as a warning instead of a hard constraint.
- Daily data, no hourly detail. We can't answer hour-by-hour questions yet, so we left an hourly profile ready for when ticket-level data is available.
- Being honest about impact. Historical data doesn't tell us how things would have played out under our decisions, so we measure savings with a backtest and don't promise numbers before running it.
- Scope. There was way more to build than hours in the hackathon, so we prioritized the reorder engine and the manager experience.
Accomplishments that we're proud of
- An AI that never makes up a number. Every figure the copilot shows comes from tested, deterministic code, so managers can trust what they see.
- Finding the real root cause. The data showed the issue wasn't suppliers but reorder logic that wasn't calibrated per store, and we built the engine to fix exactly that.
- Dealing with messy, real-world data. We corrected for censored demand and avoided data leaking from the future, so our forecasts reflect what a manager would actually know on the day they decide.
- Autonomy with guardrails. The system knows when it can act on its own and when it needs a human, and that call is made by code, not by the LLM.
- Privacy by design. Stores can collaborate on transfers without exposing their data, and isolation is enforced at the database level.
- Built for everyone on the team. From floor employees to admins, each role gets a view that helps them understand the problem and take part in the decision.
What we learned
- Statistical confidence never hits 100%. That's why autonomy depends on three things: calibrated confidence, financial impact, and whether the action can be undone.
- For perishables, covering 100% of demand isn't optimal. The extra waste costs more than the sales you recover.
- AI adds more value when it explains and a human decides than when it tries to decide on its own.
- Defining who can see, say, and approve what, both among agents and among people, matters just as much as the model itself.
What's next for Glazed Copilot
Real integration with point-of-sale, warehouse, and shift-scheduling systems; monthly model recalibration; a message queue with guaranteed delivery between stores; and deeper support for promotions and staffing.
We want every decision to have a reason, every action to be measurable, and every result to make the next one better. A smart store isn't just one that uses AI. It's one that keeps learning and makes better decisions with people at the center.
Log in or sign up for Devpost to join the conversation.