Inspiration
Autonomous coding agents are already in production — Uber's FlakyGuard has processed over 1,100 flaky tests and shipped 197 merged fixes without an engineer ever being assigned the work. But practitioners keep hitting the same failure mode: agents that confidently patch the symptom of a bug, not the cause. Nightwatch is built specifically against that problem.
What it does
Nightwatch watches a codebase for failing tests and fixes them end to end, autonomously, with no human in the loop. A GitHub webhook fires straight to Cloud Run, and three independent Gemini agents take over:
- Triage decides if the failure is even worth investigating, or should escalate immediately.
- Diagnosis generates 2-3 ranked hypotheses and checks each against the source code, informed by a growing memory of past confirmed root causes.
- Verifier — a completely separate agent, blind to Diagnosis's reasoning — independently reviews the proposed fix before anything ships. If it rejects the fix, the system tries again with fresh hypotheses.
Once approved, Nightwatch generates a visual illustration of the fix with Gemini's image model, patches the code, reruns the full test suite, and opens a real pull request. Slack and email notifications fire the moment it's done. If it can't reach a confident, verified answer, it synthesizes a handoff report instead of guessing — what it ruled out, its best lead, and an honest confidence note for a human to pick up from.
A live dashboard streams every hypothesis, rejection, and verdict from Firestore in real time, so the investigation itself is visible, not just the outcome.
How we built it
Built solo using Antigravity as the development environment. The backend runs on Cloud Run, triggered by a GitHub webhook (HMAC-verified), using Gemini 3.5 via Vertex AI and the GenAI SDK for all reasoning, with Firestore holding live pipeline state. The frontend is React + Vite + Framer Motion, reading Firestore in real time. Every feature — the Triage/Diagnosis/Verifier split, case memory, dual-path handoff reports, notifications — was built, then verified on both its success and failure paths before being accepted.
Challenges we ran into
Region-specific model availability on Vertex AI (Gemini 3.5 and Gemini 2.5 Flash Image are both served from the global location, not us-central1), an IAM/env-var reset that silently wiped notification config after an unrelated fix, and — memorably — the agent's own fix-branch pushes re-triggering its own webhook, causing a real infinite loop that we caught, root-caused, and fixed by filtering on refs/heads/main before any pipeline work begins.
Accomplishments that we're proud of
A genuine three-agent architecture with real independence — the Verifier never sees Diagnosis's reasoning, only the proposed fix. Both the hypothesis-rejection and full-retry-exhaustion paths are tested and proven, not just theoretical. Case memory correctly avoids false-matching unrelated bugs (verified against a genuinely different failure class). And a live dashboard that shows the actual reasoning process, not just a status bar.
What we learned
Multi-agent independence is a real, cheap-to-add lever for trustworthiness — a second agent whose only job is to disagree with the first one meaningfully changes what "autonomous" means. Also: production infrastructure (billing, IAM, container churn) is where real engineering discipline gets tested, just as much as the AI reasoning itself.
What's next for Nightwatch
Multi-tool coordination beyond GitHub (ticketing systems, CI dashboards), broader bug-class coverage beyond single-repo Python test failures, and opening the Case Memory index as a shared, cross-repo knowledge base.
Built with
Gemini 3.5 (Vertex AI), Google GenAI SDK, Gemini 2.5 Flash Image, Google Cloud Run, Cloud Firestore, Cloud Storage, Secret Manager, Python, Flask, GitHub API, React, Vite, Framer Motion, Firebase
Log in or sign up for Devpost to join the conversation.