We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Everyone is shipping agents; nobody can explain their failures. We watched a browser agent get trapped inside a cookie-consent modal loop for 11 LLM calls — and realized the only evidence was a wall of stdout. Planes carry black boxes for exactly this reason. Agents should too. What it does

Mayday is a local proxy + forensic cockpit:

  • mayday record starts a tap on localhost. Point any agent's OPENAI_BASE_URL at it — zero code changes, works on agents you didn't write — and every LLM request/response (including streamed SSE, reassembled) is captured with tokens, cost, and latency.
  • A POST /events ingest API lets agents attach what observability tools can't see: screen frames, world actions, errors.
  • Everything seals into a single .mayday file (SQLite) with secrets redacted at write time — the tape is the repro and the evidence.
  • mayday open launches the cockpit: a timeline scrubber where the agent's screen replay stays in sync with its reasoning, a film-strip of thumbnails for hover-preview scrubbing, a live cost ticker, a run comparison that pins behavioral divergence to an exact step (with an improvement/regression verdict pill), and a BLAME button that writes an NTSB-style postmortem — root cause, detected loops with step numbers, per-tool fail-rate ledger, and the cost of confusion.
  • Every event is hash-chained (sha256(prev_hash ‖ payload)), so a shared tape is tamper-evident — mayday verify re-checks it and the cockpit shows a ⛓ badge. LLM events carry OTel gen_ai.* semantic-convention fields. How we built it

  • Core is dependency-free Node (node:sqlite, node:http) — git clone, run it, done.

  • The tap tees SSE streams while reassembling them into response objects; replay mode re-serves recorded calls so you can re-run the agent for $0.

  • The cockpit is a build-free web app that reads tapes via API locally, or entirely in-browser via sql.js (wasm SQLite) — drop a .mayday on the page. The live demo link is a static GitHub Pages site hosting a real crash tape.

  • The demo tape is real: a Playwright agent that genuinely clicked through a purpose-built adversarial site (decoy buttons, modal loops) while the tap recorded. 58 events, 113k tokens, $0.30 — doomed by one click at step 4. One policy line in the system prompt, and the same agent finished in 4 decisions for $0.03. mayday diff shows the fork. Challenges we ran into

  • SSE is a stream in production but a single object in forensics — we reassemble OpenAI chat/Responses and Anthropic message streams chunk-by-chunk while passing them through untouched.

  • Run attribution: agent events and tapped LLM calls must land on one timeline — solved with an active-run model + x-mayday-run override.

  • Playwright's stability checks timeout on CSS-pulsing buttons — which is exactly the class of failure Mayday exists to explain. (Recorded, kept.) Accomplishments that we're proud of

  • The whole loop — record → crash → investigate → blame → fix → diff-verify — runs end-to-end on real artifacts, not mockups.

  • The cockpit is fully client-side: the live demo serves the bundled crash tape through in-browser wasm SQLite.

  • Redaction is built into the write path, so a shared tape can't leak keys — and the hash chain means a shared tape can't be quietly edited either. What we learned

Agent failures are causal chains, not single bad outputs — the transcript alone can't explain them. The frame the agent saw and the action it chose have to live on one timeline or you're guessing. What's next for Mayday

  • MITM tap: record agents you can't reconfigure (Cursor, Claude Code).
  • Counterfactual fork: replay to step N, inject a different answer, watch the run diverge.
  • CI gate: mayday diff as a regression test for agent behavior.

Built With

  • ai-agents
  • debugging
  • developer-tools
  • llm
  • node.js
  • observability
  • opentelemetry
  • playwright
  • sqlite
Share this project:

Updates

Submission history