AETHRIX — Think. Engineer. Verify.
Inspiration
AI coding assistants have become remarkably good at generating code. But generating code is only one part of software engineering.
The harder question is:
How do we know the code actually works?
A generated patch can look correct while introducing regressions, failing edge cases, or fixing the symptom instead of the root cause. Most coding assistants stop after producing an answer. AETHRIX was inspired by the idea that an AI software engineer should not simply write code — it should investigate, test, repair, verify, and provide evidence for its conclusions.
We wanted to build something closer to an autonomous engineering team than a traditional chatbot.
What is AETHRIX?
AETHRIX is an autonomous, evidence-driven software engineering platform that investigates, repairs, and verifies issues inside unfamiliar software repositories.
Instead of:
Prompt → LLM → Code
AETHRIX follows:
Problem
↓
Repository Investigation
↓
Test & Failure Detection
↓
Root Cause Analysis
↓
Patch Generation
↓
Sandbox Execution
↓
Regression Testing
↓
Verification
↓
Human Approval
↓
Engineering Evidence
The goal is simple:
AI can write code. AETHRIX proves whether it actually works.
How We Built It
AETHRIX is designed as a coordinated multi-agent engineering system.
Different agents have different responsibilities:
- Archaeologist Agent — understands an unfamiliar repository, discovers relevant files, entry points, dependencies, and test commands.
- Tester Agent — executes tests and converts raw failures into structured information.
- Root Cause Analyst — traces failures and forms hypotheses about the underlying defect.
- Fix Engineer — generates targeted source patches rather than rewriting entire files.
- Verifier Agent — compares before/after results and checks for regressions.
- Recovery Agent — analyzes failed attempts, revises the strategy, and retries.
- Evidence Agent — compiles the complete engineering trajectory into a structured report.
- Orchestrator — coordinates the agents through a controlled state machine.
This architecture allows AETHRIX to behave more like an engineering pod than a single model call.
Evidence-Driven Engineering
One of the core ideas behind AETHRIX is traceability.
Every important conclusion can be connected through an evidence chain:
Claim
↓
File Path
↓
Code Location
↓
Test Command
↓
Observed Failure
↓
Root Cause
↓
Patch
↓
Sandbox Execution
↓
Regression Test
↓
Verification
Instead of simply saying:
"I fixed the bug."
AETHRIX can show why it believes the bug was fixed.
The platform visualizes this process through an interactive Evidence Graph and also generates structured engineering reports.
Safe Autonomous Execution
Giving an AI permission to modify and execute code introduces another problem: safety.
AETHRIX therefore uses isolated sandbox environments for consequential code execution.
The repository is copied into a temporary execution workspace, where tests and generated changes can be evaluated without directly modifying the user's primary workspace.
A human approval gateway is also built into the workflow.
When verification succeeds, AETHRIX pauses for review so the developer can inspect:
- the generated patch
- test results
- execution logs
- verification status
- evidence collected during the run
Only after approval should the change be applied to the main workspace.
Self-Correction
Real software debugging rarely succeeds on the first attempt.
AETHRIX therefore includes a recovery loop.
When a generated patch fails verification:
Patch Failed
↓
Failure Analysis
↓
Wrong Assumption Identified
↓
New Strategy
↓
New Patch
↓
Re-test
This allows the system to learn from its own failed engineering attempts instead of repeatedly producing the same type of solution.
What We Learned
Building AETHRIX taught us that autonomous software engineering is not primarily about generating more code.
It is about building feedback loops.
The most important components became:
- Understanding the repository before changing it
- Testing before and after modifications
- Separating diagnosis from implementation
- Using isolated execution environments
- Recovering from failed attempts
- Making AI decisions traceable and auditable
- Keeping humans in control of consequential changes
We also learned an important lesson about evaluating AI systems: a benchmark result is only meaningful when its evaluation methodology is transparent.
Our repository therefore distinguishes deterministic development benchmarks from live model validation rather than presenting simulated results as real-world model performance.
Challenges
The hardest part was not connecting an LLM to a code editor.
The difficult engineering problems were:
1. Coordinating multiple agents
Each agent needs a clear responsibility and structured handoff. We implemented an orchestrated state-machine workflow rather than allowing agents to operate without control.
2. Understanding unfamiliar repositories
An AI cannot reliably fix code if it does not understand the surrounding project. The Archaeologist Agent therefore maps repository structure, dependencies, entry points, and test execution paths before the repair process begins.
3. Verifying generated patches
A syntactically valid patch does not necessarily mean a correct patch. AETHRIX runs tests before and after modifications and uses the Verifier Agent to determine whether the change actually resolves the failure without introducing regressions.
4. Recovering from failure
AI-generated fixes can fail. Instead of treating failure as the end of the workflow, AETHRIX feeds failure information back into a Recovery Agent that can revise the strategy and retry.
5. Building trustworthy evaluation
We initially explored broader benchmark claims, but during development we discovered the importance of clearly separating simulated benchmark execution from live API validation. The final project documentation explicitly records this distinction.
Evaluation
AETHRIX includes a synthetic benchmark repository containing intentional software bugs covering different failure patterns.
The repository also includes a baseline single-prompt solver for comparison.
The deterministic development benchmark demonstrated the complete multi-agent workflow across the benchmark cases.
For live model validation, we successfully validated the simple_bug case using Gemini 2.5 Flash:
Repository Investigation
↓
Failing Test Detected
↓
Root Cause Identified
↓
Patch Generated
↓
Sandbox Test
↓
Verification PASSED
We intentionally document the evaluation methodology so that benchmark results are not confused with live model performance.
Why AETHRIX?
Traditional coding assistants optimize for:
"Generate code quickly."
AETHRIX optimizes for:
"Reach a verified engineering result."
That difference changes the role of AI from a code generator into an engineering system with:
- investigation
- execution
- verification
- recovery
- evidence
- human oversight
What's Next?
Our long-term vision is to evolve AETHRIX from an autonomous debugging system into a broader AI Software Reliability Engineer capable of assisting with the complete software change lifecycle.
Future capabilities could include stronger adversarial testing, deeper security analysis, larger real-world repository evaluations, CI/CD integration, and continuous reliability monitoring.
The Vision
Software engineering is moving from AI-assisted coding toward AI-assisted decision making and execution.
As AI becomes capable of changing real software, verification becomes just as important as generation.
AETHRIX is built around one principle:
Don't just let AI write the fix. Make AI prove the fix.
Log in or sign up for Devpost to join the conversation.