Inspiration
Engineering teams spend too much time running post-mortems after a system breaks, rather than pre-mortems to catch the break before it happens. Most AI repository reviewers attempt to solve this, but they are just single-prompt LLM wrappers. They guess line numbers, hallucinate non-existent files, and generate alert fatigue through speculative scoring. We were inspired to build a tool that didn't just use AI to guess where code would fail, but used deterministic engineering tools to prove it.
What it does
FaultLine is an agentic software risk and failure prevention platform. When you provide a GitHub repository, FaultLine clones it into a secure, isolated Docker sandbox and deploys 7 specialized parallel agents. These agents analyze code complexity, test suite coverage, git churn velocity, architecture, and dependencies.
Instead of dumping isolated alerts into a document, our Risk Correlator Agent merges multi-modal signals to flag compounded hotspots (e.g., McCabe complexity > 20 + 0% test coverage + 14 recent bugfix commits). Finally, our Verification Agent tests every single claim against the isolated filesystem. Any unverified claim is given a 0.0 penalty multiplier, ensuring the final health score is based strictly on verified ground truth.
How we built it
- Frontend: We built a high-fidelity dashboard using Next.js (App Router), Tailwind CSS, shadcn/ui, and Recharts to visualize the deterministic risk breakdowns and verified evidence cards.
- Backend Engine: Powered by Python and FastAPI, handling concurrent agent execution and SQLite for lightweight, reproducible persistence.
- Agent Orchestration: We utilized the Google Gemini API via a custom
LLMProviderabstraction, running the agents in parallel usingasyncio. - Deterministic Tooling: We wrapped static analysis tools—including Ruff (for AST complexity), Pytest, and GitPython—to ensure the LLM was fed hard data, not just raw text.
- Security & Verification: A local Docker integration handles the repository cloning and provides the secure filesystem sandbox for our Verification Agent to test claims safely.
Challenges we ran into
Our biggest challenge was managing agent hallucinations at scale. During our third iteration, we experimented with using 10 highly granular micro-agents. We hypothesized this would give us deeper coverage. Instead, we found that unrestrained micro-agents just multiplied hallucinations, and our pipeline latency spiked by 2.5x with zero gain in precision.
We had to pivot. We consolidated down to 7 core agents and engineered the Verification Agent Sandbox. By forcing the LLM to cite exact file paths and line bounds, and programmatically checking those bounds in the sandbox before rendering the report, we successfully filtered out the hallucinations.
Accomplishments that we're proud of
- Crushing the Benchmark: We built an automated evaluation script to test FaultLine against a standard LLM baseline. FaultLine achieved 100% evidence accuracy and 91.1% finding precision, completely outperforming the baseline's 61.1% accuracy.
- The Verification Engine: Successfully executing a multi-agent loop that found 184 distinct risks in a production-grade repository (like
fastapi/fastapi) and verifying every single one of them against a ground-truth filesystem without a single test suite failure. - End-to-End Polish: Delivering a fully functioning backend pipeline and a beautiful, production-ready React dashboard within the strict 3-day hackathon sprint.
What we learned
More agents do not make software analysis more reliable. LLMs are incredible at generating hypotheses, but they cannot be trusted as the final source of truth for code quality. We learned that reliability in agentic software engineering is only achieved when agents are forced to support their claims with deterministic tool output and pass an automated verification filter.
What's next for FaultLine
We are transitioning FaultLine from an open-core MVP into a Developer-First SaaS platform.
Phase 2 (Cloud Workspaces): Introducing GitHub OAuth, Role-Based Access Control, and moving our data layer to PostgreSQL for team-wide risk tracking. Phase 3 (Active CI/CD Gatekeeping): Releasing a FaultLine GitHub Action that automatically blocks PR merges if a commit introduces compounded hotspots that exceed a team's risk threshold. Phase 4 (Autonomous Remediation): Upgrading the Verification Agent so it doesn't just check for missing regression tests, but autonomously writes them and opens a draft PR to fix the risks it finds.
Built With
- agentic
- developer
- devops
- learning
- llm
- machine
- next.js
- python
- security
- testing
- tools
Log in or sign up for Devpost to join the conversation.