Inspiration

U.S. agencies announce thousands of recalls a year. By early October 2026 the Consumer Product Safety Commission alone had announced 475, and the FDA has more than 87,000 enforcement reports on record. Each recall lives in its own database, in its own words. The detail that decides whether yours is covered is a model number, a lot or a date window, and it's usually buried in a paragraph. CPSC's own notice for a heater recall of 255,000 units says the model "TYPE SRTH" is printed on the silver rating label on the bottom of the product. Nobody reads that. We wanted to turn "is any of my stuff recalled?" into a two-minute task, with the proof shown, and without an AI guessing.

What it does

  • Reads what you own. You type it, one thing per line, or photograph labels.
    • Nemotron turns each line into an item: kind, brand, and the codes a recall turns on (model, lot, VIN, UPC, NDC).
    • A code the model returns that isn't in what you typed is dropped.
    • VINs are decoded with NHTSA's decoder.
  • Searches like an investigator. For each item, an agent on NVIDIA Nemotron 3 Ultra calls six tools: NHTSA vehicle recalls, NHTSA child-seat recalls, CPSC, openFDA, Tavily search and Tavily extract.
    • It reads the results and searches again when the first try misses. CPSC calls a space heater a "tower heater"; the agent finds that out from the results.
    • It always runs a Tavily search too, because the databases trail the agencies' newsrooms by one to six weeks.
  • Reads every notice's list. Nemotron reads the affected models, lots, UPCs and manufacture windows out of each notice's prose. Every code is then looked up in the notice's own text, character for character, and struck if it isn't there.
  • Decides by rules, not by a model. A deterministic matcher compares your codes with each list: exact codes, printed ranges, family codes and date windows. It hangs one of three tags:
    • Do not use: a code on your product is listed. You get the remedy, the firm's contact and the notice.
    • Check the label: the product line is recalled but yours isn't confirmed either way. You're told exactly where the code is printed. Cars are always this, because only NHTSA's VIN lookup can confirm one car.
    • Checked: nothing lists it today. The tag says which sources were searched and when.
  • Shows its work. Every check streams its searches live and keeps a shareable link. The proof sits next to every red tag: your code beside the recall's line.
  • Keeps you current. The home page has a live wall of every CPSC recall this year, and "Recalls today" lists the newest from CPSC and the FDA with a one-click "check if yours is affected".

How we built it

  • NVIDIA Nemotron on Nebius Token Factory (the OpenAI-compatible API at api.tokenfactory.nebius.com/v1):
    • nvidia/Nemotron-3-Ultra-550b-a55b runs three jobs:
    • the list reader (strict json_schema);
    • the search agent (OpenAI-style tools, parallel tool calls, up to six rounds, finish forced on the last);
    • the notice reader (strict json_schema).
    • nvidia/nemotron-3-super-120b-a12b is the fallback for all three.
    • Qwen/Qwen3.8-27B reads label photos, because Token Factory has no NVIDIA vision model today. GLM-5.3-Flash is its fallback.
    • On day one we measured all four Nemotron models with the same tool call and the same strict-JSON task. Ultra was both the fastest (0.5 s for a tool call, 0.9 s for structured output) and the cleanest, so it does the serious reasoning, as the track suggests.
  • Tavily:
    • search, limited to official recall domains (cpsc.gov, nhtsa.gov, fda.gov, fsis.usda.gov) plus a maker's domain when the agent asks for one;
    • extract, to read a recall page in full when its snippet doesn't show the affected models.
    • When cpsc.gov turns away our server, Tavily also finds the CPSC newsroom's newest recalls. Those run about a week ahead of CPSC's database.
  • Public data: NHTSA vPIC (VIN decoding), NHTSA recalls by make, model and year, NHTSA child-seat search, CPSC's SaferProducts.gov recall API and newsroom feed, and openFDA enforcement reports for drugs, food and devices.
  • Guardrails we added after testing:
    • The model can't drop recalls the rules say could cover the item, such as one containing your code or a brand's car-seat campaigns.
    • The agency database for the item's kind is always queried, even if the agent forgets.
    • Food and medicine recalls too old to cover what's in the cupboard are listed but not used.
    • Per-visitor and daily caps protect the demo budget; past them, a fixed search plan runs with the same matcher, and the page says so.
  • App:
    • Next.js 15 and React 19, streaming each check as NDJSON.
    • A private Vercel Blob store for checks.
    • Deployed on Vercel, with CI on GitHub Actions.
    • The matcher has unit tests.
  • Design: each item you own is drawn as its own identification label: an aluminium rating plate, a car's door-jamb sticker, a car-seat label or a medicine carton. The verdict is an industrial lockout tag, colour-coded as safety tags are and swung onto the label.

Challenges we ran into

  • CPSC's database never fills in its model-number field. It was empty in all 2,276 recalls since 2020. Model numbers live only in prose, so a model has to read them, and that's why every code it reads is verified against the text.
  • CPSC's API returns a fake error record and caches it per URL; every call carries a cache-buster and a retry. When several fields are searched together, it ANDs them, so the tool retries each field alone.
  • There's no public recall-by-VIN API. We made cars orange at most and link to NHTSA's VIN lookup, rather than claim a match we can't prove.
  • NHTSA spells models its own way. "F-150" is listed only as "F-150 (SUPER CREW) GAS" and similar. The tool tries the name as written, then NHTSA's spellings.
  • Honest green tags. An early version said "none lists your lot code" when the person hadn't given a lot. Green tags now say only what was actually compared.

Accomplishments that we're proud of

  • Every red tag is provable by the person reading it: their code next to the agency's line, with the notice one click away.
  • The model does what models are good at (reading prose, planning searches, recovering from a miss), and rules do the part that must be exact.
  • It finds real recalls on real household items in well under a minute per list, including ones newer than the agency databases.

What we learned

  • An agent is only as honest as its last step; the verdict should come from a check you can explain.
  • Tool calling on Nemotron 3 Ultra was fast enough to let the agent search iteratively in front of a user, instead of planning everything up front.
  • Agency data is messier than its docs. Measure every API on day one.

What's next

  • Saved households that are re-checked when new recalls are published, with an email when a tag turns red.
  • More countries: Health Canada, EU Safety Gate and UK product recalls are already probed and have usable data.
  • Barcode scanning and receipt import.

Built With

  • cpsc-api
  • github-jobs
  • nebius-token-factory
  • nemotron-3-ultra
  • next.js
  • nhtsa-api
  • nvidia-nemotron
  • openfda
  • qwen
  • react
  • tavily
  • typescript
  • vercel
  • vercel-blob
Share this project:

Updates

Submission history