Inspiration

I run code I cannot fully trust for a living. When something falls over at 2am, I am the one who looks. So I pointed a model at my logs. It gave me a fluent account of the outage, quoting a memory figure that appeared in none of the readings, and sounded exactly as certain when it was wrong as when it was right. Explanation was never the problem. I can write my own theory of what broke. What I want at 2am is what a good teammate gives you: the evidence, the failure reproduced somewhere safe, the fix passing three times.

What it does

Night Call investigates a production incident and hands back a pull request with the evidence attached.

  • Reads production itself: failure rates, crashes, memory, CPU, logs, traces, deploys, live flags. Every number links to the query that produced it. Missing data says "not measured".
  • Asks one question, whether the customer impact is tolerable, and treats silence as urgent.
  • Proposes causes, but marks one supported only when two independent readings back it.
  • Rebuilds the failure in a sealed copy of the app, replaying the real traffic it captured from traces.
  • Verifies the fix 3 of 3 on fresh copies, hashing production before and after each round.
  • Opens a one line pull request. A human merges it.

How we built it

Agents are TypeScript on the Strands Agents SDK, running on Amazon Bedrock AgentCore Runtime.

Models run on Featherless. GLM-5.3 leads. Kimi-K3, deliberately a different model family, reviews the reproduction and the verification. The reviewer writes its reasons, but code refuses any approval of a failed check, so it can reject and never rubber stamp. Each role's model is a setting.

The service is NestJS. Every change is an event in an append only log with zod checked payloads, and a reducer rebuilds the snapshot and streams it to the page. That log is the only source of truth, which is why the pull request body is generated from it rather than written by a model. Evidence comes from server routes that read Prometheus, Jaeger and Docker and record each reading with its source link, so the model never fetches anything. It sees only what code went and got. The sandbox is a sealed Compose copy on an internal network with no published ports.

Challenges we ran into

Bedrock access was blocked with Error 002 and the AgentCore quota was zero in most regions. I lost most of a day. The fix was to stop treating the model provider as fixed.

The crash never triggers a memory alarm. Memory peaked below half the container limit and Docker killed the process anyway, so the alarms watch restarts and failed requests instead.

Models over-claim, constantly. Every time I caught one quoting a number that existed in no reading, I moved the judgement into code rather than tightening the prompt. That migration is most of what the project is.

Jaeger keeps traces in memory and restarts if you ask for too much at once, so the traffic recipe reads in bounded pages with two fallbacks. One model's stream sometimes dropped the assistant role and broke tool calls inside the SDK, so I patch the stream before Strands reads it.

Accomplishments that we're proud of

A real run from button press to reviewable pull request in under seventeen minutes, with a reproduction and 3 of 3 verification on fresh copies.

A report where every number is a link to the query behind it. If you do not believe a figure, click it.

A reviewer that can reject but cannot approve. That asymmetry took a while to see and I think it is the most reusable idea here.

What we learned

The most useful thing an agent can do on call is gather and cite, not conclude. I started out trying to make the model a better diagnostician, and it turned out to be more valuable as a fast, literal assistant that shows its work.

Every rule I wrote as a sentence in a prompt was eventually broken. Every rule I wrote as code was not.

A sandbox only convinces anyone if you can show it used real traffic and didn't touch production.

What's next for Night Call

Night Call knows one incident type, and its mitigation is turning a flag off, which is a mitigation and not a fix. The leak in the cache path is still there. Next are code-level mitigations, put through the same proof loop so a real fix has to earn 3 of 3 the same way.

  • Alarms starting investigations instead of me pressing the button.
  • Bedrock models once account access is enabled. Already a setting, so this is configuration.
  • Slack and Discord for the impact question.
  • Onboarding any Docker Compose application, not only the demo shop.

Built With

Share this project:

Updates

Submission history