Inspiration AI agents are incredibly powerful at writing code, but they have a fatal flaw: they can't prove their code is safe to merge. We realized that as autonomous coding scales, the bottleneck won't be generating code, but verifying it. We wanted to build a definitive trust layer—a single, quantifiable gatekeeper that AI agents must call before they commit.
What it does Aegis is an automated trust layer and PR gatekeeper exposed over the Model Context Protocol (MCP). It intercepts pull requests and runs them through a rigorous, isolated verification pipeline before merging.
When a PR is opened, Aegis kicks off a sequence of deep checks: it generates candidate patches, spins up ephemeral sandboxes to test the code dynamically, runs security reviews, and maps out a call-graph logic diff. It then fuses all these metrics into a single Trust Score.
If a PR fails the trust gate, Aegis doesn't just block it—it autonomously generates a verified fix, opens a corrective counter-PR, and pings the team in a Discord war-room with a full debrief. It’s a fully closed, self-healing trust loop.
How we built it We architected Aegis using a powerful stack of sponsor technologies, ensuring every tool wasn't just decorative, but load-bearing:
Fireworks AI: Powers the heavy lifting. We use deepseek-v4-pro to generate candidate patches, act as our LLM-as-judge scorer, and drive the CopilotKit model.
Daytona: Handles the dynamic verification. It spins up ephemeral, isolated sandboxes to install dependencies, apply the patch, and run tests before a commit even exists.
Braintrust: The brain of our scoring engine. It fuses the Daytona tests, security reviews, LLM-judge feedback, and minimality into a single composite Trust Score and logs the entire trace.
CodeRabbit: Runs our in-sandbox CLI security review. It also exposes the get_trust_report over MCP and operates as a Discord agent for enterprise-style debriefs.
CopilotKit: Powers our human-facing interactive dashboard, ensuring human reviewers and AI agents are driving the exact same tools over the same MCP.
Discord: Serves as our developer war-room via the "Captain Aegon" webhook. If a PR is blocked, Aegis fires an [AEGIS_GATE:BLOCKED] trigger, waking up the CodeRabbit Discord agent to debrief the team.
Challenges we ran into One of the hardest parts of building Aegis was orchestrating a seamless, synchronous pipeline across so many distinct platforms. Spinning up a Daytona sandbox, injecting the CodeRabbit CLI, running tests, and streaming those logs back out to a Braintrust scoring matrix required incredibly precise timing and state management. Additionally, designing a unified "Trust Score" that accurately weights static security analysis against dynamic sandbox testing took significant trial and error to get right.
Accomplishments that we're proud of We are incredibly proud of building a fully closed, self-healing loop. Getting an AI to review code is one thing; getting it to spin up an ephemeral sandbox, realize the code breaks, dynamically generate a fix, verify its own fix, and open a corrective counter-PR without human intervention is a massive leap forward. We're also proud of how we utilized MCP—proving that an AI agent (CodeRabbit) and a human (via the CopilotKit dashboard) can share the exact same explicit, measurable trust metrics.
What we learned We learned that static analysis is no longer enough for AI-generated code; dynamic, in-sandbox testing (via Daytona) is absolutely mandatory to establish real trust. We also learned how powerful the Model Context Protocol (MCP) is for standardizing agentic tooling. By exposing our trust gate over MCP, it became incredibly easy to route data between GitHub, our dashboard, and Discord.
What's next for RabbitWall (Aegis) Moving forward, we want to expand RabbitWall's capabilities by adding broader language and framework support for the Daytona sandbox environments. We also plan to introduce more granular customization for the Braintrust scoring weights, allowing enterprise teams to set custom strictness thresholds for different repositories, and expanding our Discord debrief agent to handle interactive, natural-language rollbacks directly from the chat UI.
Built With
- braintrust
- coderabbit
- daytona
- fireworks
Log in or sign up for Devpost to join the conversation.