Who this is for
A household where one person reacts to something. The parent who reads every label twice, the flatmate who keeps a separate shelf, the adult child doing the shopping for a father who no longer trusts his own memory for it. Not a compliance team, not a grocer, not a food safety professional. Somebody standing at a freezer at nine in the evening with a tub in one hand.
That person already knows recalls exist. What they do not have is a way to ask about the tub in their hand without putting it down, opening a phone, finding the right agency, and reading a notice written for someone else. Two agencies publish separately, neither covers the other, and the notice that matters is mixed in with three hundred that do not.
The reason this is worth running after the hackathon is that the question comes back. Recalls are published every week, allergens repeat within a household, and nothing about that ends. On 15 April 2026 one creamery filed forty eight recall records, forty three of them flavours of the same ice cream, carrying thirty one different recall reasons between them, and the difference between the tub that is safe and the tub that is not was one word on the front of the pack. That is a question a person cannot answer alone at a freezer, and it is exactly the shape of question a voice assistant is good at, provided it is allowed to ask something back.
It is self hosted on purpose. A household that sets its own allergens is telling a machine something about a member of the family, and that record has no reason to leave the house.
What it does
pulled is a self hosted MCP server that answers one question in a kitchen. Has this been recalled, and does it matter to me.
Three tools, over Streamable HTTP, protocol 2025-11-25.
check_item(description) takes what the person actually said, in their words,
and answers recalled, unclear or clear. recent_recalls(days, only_mine)
lists what has been pulled lately whose stated reason names an allergen this
household reacts to. set_allergens(allergens) records what that is. There is
also a prompt, kitchen_check, that tells a client how to read the answers out
loud without turning a clear result into a promise.
Why the middle outcome exists
A voice assistant that says yes that is recalled about the wrong jar sends
somebody to bin their dinner. One that says no about the right jar is worse.
So check_item has three outcomes rather than two, and the middle one is a
question.
On 15 April 2026 the FDA published forty eight recall records from one creamery, forty three of them flavours of the same ice cream, carrying thirty one different recall reasons between them. Peanut Butter Fudge is undeclared milk and peanuts. Pistachio is milk and pistachios. Rocky Road is milk, walnuts and eggs. Somebody standing at a freezer saying my Loard's ice cream has told you nothing that can decide the question, and naming the top hit would be a one in forty three guess dressed up as an answer. So the server names the part of the label it matched back to them, says how many open records fit it, and asks what else the front of the pack says.
How the asking works
The question is not a field in a payload that the client may or may not read. When the server is unsure it raises an elicitation, folds the answer into the description, and decides. One tool call, one round trip to the person, one verdict.
A client that does not implement elicitation, and a person who declines to
answer, both lose nothing. The outcome stays unclear and the same question
travels back in the payload for the client to ask however it likes. Both paths
are covered by tests.
How the matching works
Every candidate is scored on how much of what the caller said appears in the record, weighted by how rare each word is inside the candidate pool. Two rules then stop a confident answer. If several records fit equally well we have found a family rather than a product, so the shared part of the name goes back to the caller as a question. If the winning record hinges on a word at least as distinctive as anything the caller said, and they did not say it, that word is probably the difference between their tub and this one, so we ask about it.
Sizes, packaging and supply chain boilerplate are stripped first, because
3.17oz, pouch and per look rare and mean nothing. That list did not come
from imagination. It came from running the matcher against the live feed and
watching it ask a caller whether their pack said per.
Both agencies, and neither able to sink the other
US food recalls are split. The FDA covers most of the shelf through openFDA. Meat, poultry and eggs belong to the USDA through a separate feed with a different shape. A checker wired to one of them is silent on chicken, deli meat and ground beef, which is where the Class I recalls live. pulled reads both. A source that fails is skipped rather than raised, because an answer from one agency beats an error naming two.
What running it against the real feed changed
Twelve handwritten tests passed while the server, on live data, asked every
caller Does yours say per on the pack and never gave an answer. Everything
below was found the same way, by pointing the code at the real thing and
reading the output rather than counting green ticks.
- The USDA feed publishes every recall twice, once in English and once in Spanish, under the same recall number. 789 numbers of 2 023 records. Left alone, every single query looks like a tie and the server never answers. Deduplicated, the feed holds 1 234 recalls.
- 361 of those carry an allergen reason and not one names the allergen in it. The press release names it in 354 of them, so that is where it is read from.
- The FDA puts the product name first and the packaging after it. The USDA does the opposite and puts the name in quotes. Reading the first words, correct for one agency, made the server ask somebody whether their pack said vacuum.
- 169 USDA records are public health alerts rather than recalls. Announcing one as a recall tells a person their dinner has been pulled from sale when it has not, which is exactly the wrong alarm this server exists to avoid.
Heard, not typed
The video plays a real spoken exchange on a simulated Alexa+ kitchen display,
the alternate path the Alexa+ track allows. A synthetic voice plays the person,
Whisper transcribes it, a language model client with nothing but the three
pulled tools decides what to call over MCP, and a second voice reads the answer.
Nothing on the assistant side is scripted. The display is sim/device.html and
the script that plays the take on it is sim/film.py, both in the repository.
The caption, the reply, the result card and the tool calls on screen are read
from the take, and only the model's thinking time is cut.
The first take failed. Whisper wrote Lord's for Loard's, the FDA search came back empty, and the server asked whether the tub came from a maker of Brazilian pastries. On 200 live FDA records with one letter lost from a word, the matcher as it was answered no recall matching that 34.5 % of the time and named a different record with confidence 3.5 % of the time. After three small changes, both are at zero on the same 200 records, on that test and for that kind of mishearing only. Most of those descriptions now end in a question, which is what the middle outcome is for. The method, its limits and the before and after runs are in the README.
How to run it
pip install -e ".[dev]"
python -m pulled.server # http://127.0.0.1:8931/mcp
python demo.py --offline # the whole kitchen conversation
python talk.py # the spoken take, voices and all
python sim/film.py # that take on a simulated Alexa+ display
pytest # 48 tests, listens on nothing
python no_network.py # the same 48, with the network refused
The suite drives the server through an in memory client and server pair, so it listens on nothing and a failing run cannot leave a listener behind.
What is not claimed
A clear result is not a safety certificate. It means nothing matched in two
public feeds, and the sentence the server hands back says so in those words,
every time. Nothing here measures precision or recall against a labelled set,
because no such set exists for spoken kitchen descriptions. The numbers about
the feeds and the matcher sit in AFFIRMATIONS.md next to the command that
disproves them, and the speech measurements sit in the README with the command
that reproduces them.
Log in or sign up for Devpost to join the conversation.