Inspiration
I ordered 30 eggs because I wanted super, super fresh eggs.
But 60 arrived.
So what do you do with the extra 30?
I asked an AI. I asked customer support. I waited for a human agent who never picked up. And while I waited, my super-fresh eggs became less and less fresh.
Here's the part that actually stung: if I had simply known I was allowed to keep them, I could have given them to my neighbors. Instead they sat there. And in Korea, even throwing away food waste costs money — you pay by weight. I was being billed for the privilege of disposing of eggs I never ordered.
The platform wasn't malicious. It was just contradictory. Misdelivered items must be returned. Also: fresh food cannot be returned. Two rules, both reasonable, pointing in opposite directions. Humans survive contradictions like this with heuristics and discretion. A support agent shrugs and makes a call. I shrug and eat the egg.
An agent has no shrug.
So it does one of two bad things: it reinterprets the contradiction on its own and invents a loophole, or it stalls and does nothing at all. Neither of those is what I wanted. What I wanted was something that knew it wasn't allowed to decide — and went to find someone who was.
Most agents are designed to act. This one is designed to know when not to.
What it does
Let Me Eat The Eggs is a WebMCP-enabled site where the site, not the model, owns the guardrail.
A split screen runs the same problem twice. On the left, a human works it through customer support — and gets shredded by a work call, a lunch break that ends in three minutes, a new agent who asks for the order number again, and finally the words support hours have ended. On the right, an agent works the same problem through five WebMCP tools.
The agent's day goes like this. It reads the user's standing policy — resolve perishable exceptions within 24 hours, never decide edibility without evidence — and then does something small but important: it narrows its own deadline to 18:00 today, because seller phone lines don't answer at 3am and eggs don't get fresher overnight.
It tries the AI support chat, which offers refund or return and nothing else, and closes the session as low utility. It reads the policy and finds the contradiction. And here it stops, because the site tells it plainly: agentMayReinterpret: false. You do not have the authority to resolve this. Here are the people who do.
So it waits. Nine hours of virtual time, while a human support agent says let me check and get back to you on the other side of the screen. The wait is not a bug in the demo. Human discretion was genuinely in motion; the agent was right to let it work. It just never arrived.
When it becomes clear discretion won't reach in time, the agent doesn't grant itself more power. It shortens the distance to whoever has it. Least intrusive channel first: a text to the seller. An hour of silence later, with two hours left, it escalates to a call — and only then, because the site checks that a text actually went unanswered first.
The seller says: yeah, we packed an extra one, our mistake, just eat it. Now there's evidence. Only now does resolve_case return resolved.
Try to skip any of that and the site says no. Ask it to conclude immediately — blocked. Hand resolve_case four real pieces of evidence that happen to be the wrong ones — blocked, because it verifies content, not count. Try to call the seller thirty minutes after texting — blocked. Two minutes before the deadline with no seller confirmation — still blocked.
The agent cannot mint evidence, cannot mint approval, and cannot mint its own authorization to escalate.
The demo uses one egg instead of thirty. Sixty egg emojis did not fit on a phone screen.
How we built it
Vite + React + TypeScript, a static SPA on Cloudflare Pages, in-memory state, no backend, no database, no auth.
Five tools registered through document.modelContext.registerTool(). The part that mattered architecturally: every execute() does nothing but call a shared action layer. The 35-second scripted playback calls the same functions. The manual buttons in the footer call the same functions. There is no separate demo code path, which means the animation can't lie about what the tools do — if the gating works in the demo, it works for a real agent, because it's the same code.
The gating itself lives in the tools, not in prompts:
check_policyreturns the contradiction and refuses to resolve it, naming the authority holders insteadcontact_supportrefuses a more intrusive channel without a recorded failure on a less intrusive oneresolve_casetreatsevidenceIdsas lookup keys only, readingauthoritativeandsupportsoff the stored record and ignoring anything the agent claims alongsidepay_sellerignores the agent'sapproved: trueentirely and readspayment.humanConfirmedAt, which only a human UI click writes
Thirteen invariant checks run headlessly against all of it.
External systems — support center, seller, telephony, payment — are simulated. Tool registration, tool invocation, evidence gating, escalation gating, deadline-dependent routing, and state transitions are real.
Challenges we ran into
We built a project about deadlines under a deadline, which was funnier in concept than in practice. Three hours, self-imposed. The virtual clock in the demo moves one hour every four seconds. Ours moved about the same.
We gated the conclusion and forgot to gate the escalation. This was the real one. The first version locked down resolve_case beautifully — no evidence, no answer — and then let contact_support escalate to a phone call based purely on what time it was. Which meant an agent could skip the text entirely and cold-call a stranger's personal number at 16:01, and the site would happily connect it. The whole thesis is deadline changes the route, not the authority, and the route change had no requirements at all. We caught it with about an hour left and fixed it by requiring a recorded failure of the less intrusive channel. It turned into the second-best scene in the demo.
An agent that waits looks identical to an agent that's broken. The wait is the intellectual center of the project and it is, visually, nothing happening. We had to promote it from a passing beat to a persistent badge holding for nine real seconds, anchored to the human support scene on the other screen, before it read as a decision instead of a hang.
Restraint is hard to film. Everything interesting here is a refusal, and refusals don't animate. We ended up building the demo around three blocked calls instead of one successful one.
Accomplishments that we're proud of
The tool table in our README has a column header that reads "what it refuses." That was the moment the project clicked.
Thirteen invariants that hold under adversarial input, including the one we're fondest of: hand resolve_case four genuine, real, correctly-formed pieces of evidence that don't happen to support the conclusion, two minutes before the deadline, and it says no. Not because a prompt asked it to be careful. Because the site checked.
One action layer, three entry points, zero divergence. The demo cannot show something the tools don't actually do.
And we shipped a working, deployed, agent-testable site in three hours without cutting the part that made it worth building.
What we learned
Guardrails belong to whoever owns the state, and that isn't the model. Every attempt to enforce this in the agent's reasoning is a suggestion. Enforced in a tool's return value, it's a fact. WebMCP turns out to be a surprisingly good place to put policy, because the site already knows what actually happened.
Partial gating is worse than no gating, because it looks finished. We locked the front door and left the escalation path wide open, and the demo ran perfectly the whole time.
Contradictory policy isn't an edge case, it's the steady state of every legacy system. And it gets more visible as agents take over, not less, because humans have been silently absorbing these contradictions for decades with judgment calls nobody wrote down. Every one of those is a decision an agent will hit and have no idea what to do with.
The goal is not maximum autonomy. It is the minimum necessary agency, granted at the moment it becomes justified. Waiting is not failure. Sometimes waiting is the correct action — but waiting has a cost too, so the system keeps asking a different question: has enough evidence accumulated to justify doing more?
What's next for Let Me Eat The Eggs
An authority-holder registry. Right now check_policy names who can resolve a contradiction because we hardcoded it. A real version lets a merchant declare which contradictions exist in their policy set and who holds discretion over each — turning "our policies conflict" from an embarrassment into structured, agent-readable data.
Real platforms. The Shopify and Cloudflare surfaces are right there. A merchant-side WebMCP layer that publishes both its rules and its exceptions would let agents resolve the long tail of small disputes without a human ever opening a ticket.
More contradictions. Damaged-but-usable. Wrong size, opened packaging. Perishables past a date that are demonstrably fine. Every one is an item currently being thrown away because nobody was authorized to say "it's fine."
A less silly demo of a serious idea. Or possibly not. The eggs have been good to us.
This project is dedicated to my once-super-fresh eggs.
Log in or sign up for Devpost to join the conversation.