Inspiration
Modern engineering teams lose hundreds of hours triaging midnight production alerts, isolating broken commits, and writing boilerplate bug fixes. While existing LLM developer tools act like passive chatbots waiting for prompts, we asked: What if an AI engineer could respond to an alert, diagnose the codebase, run tests in an isolated sandbox, fix the bug, and open a verified pull request entirely on its own? This led us to build DevSentry, an autonomous incident remediation agent that acts rather than just answering questions.
What it does
DevSentry automates the Site Reliability Engineer (SRE) and bug triage workflow:
- Incident Intake: Ingests crash reports and error logs directly from alerting systems.
- Environment & Diagnosis: Uses Git tools to clone the target repository, checks out the suspect commit, and inspects the stack trace.
- Sandbox Execution: Runs the existing test suite in an isolated sandbox environment to reproduce the bug.
- Autonomous Self-Correction Loop: If a proposed patch fails or throws unexpected errors, DevSentry parses the compiler/test output, reformulates its hypothesis, and refines the patch until all assertions pass.
- Verified Resolution: Pushes a dedicated bugfix branch, opens a GitHub Pull Request with a root-cause analysis summary, and notifies the team.
Why it is Agentic
DevSentry strictly follows the autonomous agent lifecycle: Goal ➔ Decide ➔ Tool ➔ Observe ➔ Finish.
- Planning & Action Sequences: Instead of completing a task in a single prompt, it breaks down the incident into discovery, reproduction, code patching, and verification stages.
- Real Tool Execution: It interacts directly with external systems including the local shell, isolated sandboxes, Git CLI, and the GitHub REST API.
- Dynamic Self-Correction: When a patch fails unit tests, DevSentry does not give up or ask the user for advice. It consumes the runtime failure logs as observation feedback, recalibrates its plan, and applies updated fixes iteratively.
How we built it
- Orchestration: Built with LangGraph to model stateful, cyclical agent graphs, fallback branches, and conditional decision points.
- LLM Engine: Powered by state-of-the-art models via function calling to guarantee structured tool invocations.
- Execution Sandbox: Containerized shell/Docker runner to safely clone code, execute unit tests, and capture stdout/stderr exit codes.
- Tool Integrations: GitHub REST API for branch manipulation and PR generation.
- Dashboard: Interactive UI showing real-time agent thoughts, tool executions, and terminal logs.
Challenges we faced
- Preventing Infinite Loops: LLMs can get stuck cycling through identical incorrect fixes. We engineered strict loop detection and state history checks that inject past failure context into subsequent planning steps.
- Parsing Dynamic Test Outputs: Different test runners output errors in diverse formats. We standardized error-capture logic so the agent consistently isolates relevant stack traces from noisy logs.
- Sandboxed Tool Safety: Ensuring autonomous terminal commands could safely edit files and execute tests without corrupting parent directory environments.
What we learned
Building true agentic autonomy requires treating the LLM as an orchestrator and reasoning engine, not a monolithic problem solver. Deterministic feedback loops (like test exit codes and compiler errors) provide the exact grounding LLMs need to reliably self-correct.
What's next for DevSentry
- Integration with live Datadog, Sentry, and AWS CloudWatch incident streams.
- Multi-agent collaboration with a dedicated Security Auditor agent reviewing code diffs before PR creation.
- Support for complex multi-service microservice sandboxes via Docker Compose.
Log in or sign up for Devpost to join the conversation.