Inspiration
Production incidents rarely fail because engineers do not care. They fail because critical context is fragmented across logs, GitHub commits, Slack messages, dashboards, runbooks, and the memories of exhausted responders.
The engineer handling an incident has the most context while the system is failing, but the least time to document it. Later, when the team writes the postmortem, important details have already disappeared. The same failure pattern is then repeated because the organization never converted the incident into reliable institutional memory.
We built Nightshift to close that gap.
Nightshift is an AI incident-response teammate that stays with an issue from the first alert through diagnosis, repair, documentation, and future prevention. Instead of acting as another chatbot, it operates inside real sandboxes, tests competing hypotheses, records evidence, updates shared team memory, and protects that memory from incorrect or poisoned conclusions.
We deliberately designed the experience without depending on phone calls. Engineers interact with Nightshift through a persistent desktop notch and Slack, allowing them to see the agent’s reasoning, inspect its evidence, approve sensitive actions, and continue the same incident session without re-explaining context.
What it does
Nightshift manages the complete incident lifecycle:
1. Detect and understand the incident
When an incident begins, Nightshift opens a persistent incident session in the desktop notch and posts a structured status card in Slack.
It collects:
- Application logs
- Recent GitHub commits and pull requests
- Service ownership information
- Existing runbooks
- Previous related incidents
- Screenshots or browser-state evidence
- Relevant Slack discussion
The notch acts as the engineer’s live command center. It shows the current hypothesis, evidence collected, active sandbox experiments, candidate fixes, and actions requiring approval.
Slack serves as the team-wide incident timeline. Nightshift posts concise updates, evidence receipts, proposed fixes, and final artifacts so everyone can follow the incident without interrupting the responder.
2. Reproduce the failure in Daytona sandboxes
Nightshift does not immediately guess at a root cause or modify production.
It first creates an isolated Daytona sandbox containing a reproducible version of the affected service. The agent replays the failing request, runs tests, inspects logs, and can use a real browser inside the sandbox to reproduce user-facing failures.
It then launches multiple controlled experiments in parallel:
- One sandbox tests whether a recent database migration caused the failure.
- Another tests a configuration change.
- Another checks a dependency update.
- Several commit snapshots replay the same request to identify the exact breaking commit.
- Candidate fixes are tested independently before one is recommended.
This turns incident response from sequential guessing into a parallel hypothesis race.
The winning diagnosis must include evidence: logs, test output, screenshots, the breaking commit, and a reproducible successful fix. Nightshift therefore produces a receipt rather than simply claiming that a patch works.
3. Recommend and execute a fix safely
After identifying the likely root cause, Nightshift presents the engineer with:
- The suspected cause
- Confidence and supporting evidence
- The safest rollback point
- A proposed patch
- The results of the sandbox reproduction
- Potential blast-radius risks
The engineer can approve, reject, or modify the action directly from the notch or the CopilotKit interface.
WorkOS protects high-risk actions through authentication, role-based access control, organizational permissions, and step-up approval. Nightshift may investigate independently, but destructive production actions remain human-controlled.
4. Convert the incident into durable team memory
Once the issue is resolved, Nightshift combines the complete timeline with the engineer’s comments and automatically produces:
- A structured postmortem
- A reusable runbook
- Action items
- A concise Slack summary
- An optional ElevenLabs-generated audio briefing
- A record of the root cause, fix, affected service, and evidence
The system stores both structured incident data and searchable semantic memory.
Each engineer has a personal memory layer that learns how they investigate incidents, which dashboards they prefer, and how they communicate. The team also has a shared memory layer containing verified incidents, runbooks, service ownership, fixes, and historical failure patterns.
Months later, another engineer’s Nightshift agent can recognize that a new incident resembles an earlier outage and immediately retrieve the verified fix.
5. Prevent the same failure from returning
When a new pull request touches code associated with a previous incident, Nightshift can surface a pre-merge warning:
This change modifies the same connection-pool configuration involved in a previous checkout outage. Recommend replaying the historical workload before merge.
Nightshift can then launch a new Daytona sandbox, apply the proposed branch, and replay the original failure against it.
The incident therefore becomes a permanent regression test rather than a forgotten document.
How we built it
Nightshift combines a personalized agent interface, shared organizational memory, real execution environments, evaluation infrastructure, and human-controlled security.
Daytona — isolated incident investigation
Daytona provides Nightshift’s execution layer.
We use Daytona to:
- Create isolated forensic sandboxes
- Restore known service snapshots
- Replay failing requests
- Run browser-based reproductions
- Compare multiple commit snapshots
- Test several hypotheses simultaneously
- Validate candidate patches
- Generate evidence receipts
- Replay previous failures against new pull requests
Daytona is what separates Nightshift from an incident chatbot. The agent can prove its diagnosis by running the system instead of reasoning only from text.
Fireworks AI — the reasoning and orchestration layer
Fireworks powers the agent responsible for:
- Selecting diagnostic tools
- Forming root-cause hypotheses
- Coordinating sandbox experiments
- Ranking competing explanations
- Producing structured incident output
- Generating postmortems and runbooks
- Reasoning over memory provenance
We use a capable function-calling model for the agent loop and a structured-output fallback for schema-sensitive artifacts.
CopilotKit — the live incident interface
CopilotKit powers the interactive incident dashboard and generative UI components.
It streams:
- Sandbox status
- Agent hypotheses
- Logs and screenshots
- Candidate patches
- Approval requests
- Root-cause evidence
- Postmortem progress
Instead of forcing engineers to communicate through a basic chat window, CopilotKit presents the information as actionable incident cards.
The notch — the persistent local presence
The desktop notch is the engineer’s personal Nightshift interface.
It remains attached to the active incident and provides:
- Current incident state
- The leading diagnosis
- Running sandbox experiments
- Pending approvals
- Evidence collected
- A compact command surface
Because the notch and Slack reference the same incident session and memory, engineers can move between private investigation and team coordination without losing context.
Slack — the shared coordination layer
Slack is Nightshift’s team-facing timeline.
Nightshift posts:
- Initial incident context
- Investigation progress
- Evidence-backed hypotheses
- Fix recommendations
- Approval outcomes
- Resolution summaries
- Postmortems and runbooks
- Optional generated audio briefings
The agent does not flood the channel with every internal thought. It posts only meaningful state changes and evidence that other engineers need.
Braintrust — ground-truth evaluation
We did not want to evaluate Nightshift based only on whether its postmortems sounded convincing.
We created seeded incidents with known:
- Root causes
- Breaking commits
- Correct rollback points
- Reproduction steps
- Expected fixes
Braintrust scores Nightshift against this ground truth using metrics such as:
- Root-cause correctness
- Rollback-point correctness
- Negative-control specificity
- Action-item specificity
- Postmortem quality
We also include benign negative-control incidents where the correct conclusion is that nothing is broken. This measures whether Nightshift can avoid inventing problems.
The evaluation set is divided into development incident families and held-out incident families, reducing the risk that we simply tuned the agent to the exact examples being scored.
CodeRabbit — pull-request forensics
CodeRabbit analyzes the pull request that introduced the failure and surrounding high-risk changes.
Nightshift uses this information to answer:
- What should code review have caught?
- Was a relevant warning raised?
- Was the warning dismissed?
- Which nearby pull requests could have the same problem?
- What review rule should be added for the future?
This gives the postmortem a forward-looking engineering review section rather than only a timeline of what already happened.
WorkOS — permissions and human control
An agent capable of interacting with production systems must be constrained by real organizational identity and authorization.
WorkOS provides:
- Single sign-on
- Role-based access control
- Organization membership
- Service-specific permissions
- Step-up authentication
- Audit trails
Nightshift can gather evidence and prepare an action independently, but a rollback, deployment, or destructive operation requires permission from an authorized engineer.
ElevenLabs — optional voice inside the product
We intentionally removed dependency on cellular calling.
Instead, ElevenLabs supports optional voice interaction directly inside the notch and produces concise audio incident briefings for Slack. Engineers can speak to the agent while working, but the product remains completely usable through text and visual controls.
This preserves voice as a useful interface without making telephony a failure point.
Challenges we ran into
Making the agent prove its diagnosis
Language models can produce explanations that sound technically credible even when they are wrong.
We addressed this by requiring Nightshift to reproduce incidents in Daytona, test competing hypotheses, and attach evidence to every recommendation. The system distinguishes between an untested hypothesis and a sandbox-verified result.
Maintaining one continuous incident state
The notch, Slack, sandboxes, memory system, and evaluation layer all need to refer to the same incident.
We created a shared incident state machine that tracks the lifecycle from detection to investigation, resolution, documentation, and prevention. Every tool call and memory write is attached to the same incident identifier.
Preventing unsafe autonomous actions
An incident-response agent may eventually have access to sensitive systems.
We separated investigation from execution. The agent can inspect, reproduce, compare, and recommend automatically. High-impact actions are blocked behind WorkOS permissions and explicit human approval.
Preventing memory poisoning
Shared memory creates enormous value, but it also creates a dangerous failure mode.
If one agent stores an incorrect root cause, every future agent could retrieve it and make the same mistake. A single hallucinated conclusion could become false institutional knowledge.
To address this, we built a Guardian agent dedicated to protecting shared memory.
Accomplishments that we're proud of
We are proud that Nightshift is not only a wrapper around an LLM.
It combines:
- A persistent per-engineer agent through the notch
- A shared Slack incident timeline
- Real execution inside Daytona sandboxes
- Parallel root-cause testing
- Evidence-backed fixes
- Ground-truth evaluation through Braintrust
- Pull-request forensics through CodeRabbit
- Human-controlled authorization through WorkOS
- Structured reasoning through Fireworks
- Optional ElevenLabs voice and audio summaries
- Personal and shared organizational memory
- A dedicated Guardian agent for memory integrity
The most important accomplishment is that every part contributes to one continuous product rather than operating as an unrelated integration.
An incident begins in the notch, is investigated in Daytona, coordinated through Slack, scored in Braintrust, documented into memory, protected by the Guardian, and later reused to prevent another failure.
What we learned
We learned that the most valuable part of an incident agent is not generating an answer quickly. It is creating a trustworthy chain from hypothesis to evidence to action.
We also learned that organizational memory must be treated as a security boundary.
Traditional security systems protect databases from unauthorized access. Agent systems must also protect knowledge from authorized but incorrect writes. An agent may have permission to update memory while still storing a hallucination, adversarial instruction, or mistaken diagnosis.
This means agent memory requires:
- Provenance
- Confidence
- Evidence
- Validation
- Quarantine
- Human review
- Recovery mechanisms
We also learned that personalization does not require continuously retraining a model. Nightshift becomes more useful through structured memory, behavioral history, and verified incident data rather than modifying model weights.
Finally, we learned that interfaces matter. Engineers do not want another giant chatbot during an outage. They need concise status, visible evidence, clear actions, and control. The notch, Slack, and generative incident cards make the agent easier to trust.
What's next for Nightshift
A stronger Guardian for memory-poisoning prevention
Every shared-memory write includes provenance:
- Which agent created it
- Which incident produced it
- Which sources support it
- Which sandbox experiment verified it
- Whether a human approved it
- Which later memories depend on it
The Guardian continuously checks new memory against ground truth, existing evidence, and later incident outcomes.
When it detects a suspicious write, it:
- Identifies the conflicting claim.
- Traces it to the exact agent, incident, source, and memory operation.
- Determines which derived memories or recommendations depend on it.
- Quarantines the questionable knowledge.
- Regenerates affected summaries from clean evidence.
- Requests human approval before permanently replacing shared state.
The Guardian is evaluated using both poisoned-memory tests and negative controls. It must catch known false memories without deleting correct ones.
Broader production integrations
We plan to connect Nightshift to additional observability and deployment systems so it can reason across traces, metrics, feature flags, infrastructure changes, and service dependencies.
Automatic regression generation
Every verified incident can become:
- A regression test
- A monitoring rule
- A code-review check
- A runbook trigger
- A pre-merge sandbox simulation
This would allow Nightshift to turn operational history into continuously improving engineering infrastructure.
Cross-team incident intelligence
As the memory layer grows, Nightshift can identify patterns across services:
- Repeated migration failures
- Common deployment risks
- Fragile dependencies
- Services with recurring ownership gaps
- Review patterns correlated with incidents
A trusted operating layer for engineering agents
Our long-term goal is for Nightshift to become the trusted operating layer between AI agents and production systems.
Agents should be able to investigate quickly, execute safely, remember accurately, and prove every important conclusion.
Nightshift is the on-call engineer that never sleeps, never forgets, and—most importantly—can show its work.
Log in or sign up for Devpost to join the conversation.