Inspiration

Modern software engineering is still dominated by repetitive workflows: converting product requirements into tasks, writing boilerplate code, creating tests, debugging failures, and managing Git branches and Pull Requests. While AI coding assistants can generate code, they often stop at code generation and leave developers responsible for running the code, interpreting failures, fixing bugs, and completing the delivery pipeline.

We asked ourselves: What if an AI engineering platform didn't just suggest code, but could own the complete lifecycle from a raw software requirement to a tested, self-healed, and delivered GitHub Pull Request?

Inspired by autonomous control loops, compiler pipelines, and resilient distributed systems, we built DaedalusAi — an autonomous multi-agent software engineering platform that plans, researches, implements, tests, debugs, reviews, and delivers software with minimal human intervention.

At its core is an autonomous feedback loop:

while not test_suite.passed and iterations < MAX_HEALING_CYCLES:
    traceback = reviewer.extract_root_cause(sandbox.run_tests())
    code_patch = healer.remediate(code_patch, traceback)
    iterations += 1

The system is deliberately bounded by a safety condition:

$$ \text{Stop Condition} = (\text{Tests}(\text{Patch}_i)=\text{PASS}) \lor (i\ge3) $$

This allows DaedalusAi to autonomously recover from failures without becoming trapped in an infinite agent loop.


What it does

DaedalusAi accepts natural-language software requirements and transforms them into tested and reviewable GitHub Pull Requests.

  • Autonomous Requirement Decomposition: Converts natural-language requirements into dependency-aware tasks and a Directed Acyclic Graph (DAG), ensuring foundational components are implemented before dependent functionality.

  • AI Architectural Planning: Designs APIs, database schemas, component boundaries, dependencies, and implementation strategies before code generation begins.

  • Semantic Context Retrieval: Uses Retrieval-Augmented Generation (RAG) to retrieve relevant project documentation, coding standards, architectural decisions, API conventions, and testing patterns.

  • Paired Code & Test Generation: Generates implementation code together with pytest test suites covering normal execution paths, edge cases, and boundary conditions.

  • Deterministic Execution: Runs generated code and tests in an isolated workspace, using actual runtime results rather than allowing an LLM to decide whether its own code works.

  • Autonomous Self-Healing: When tests fail, the system extracts the high-signal traceback, identifies the likely root cause, generates a corrective patch, and automatically re-runs the tests.

  • Bounded Recovery: Self-healing is limited to three iterations. Persistent failures are escalated instead of allowing uncontrolled recursive agent execution.

  • Automated GitHub Delivery: Creates feature branches, commits validated changes, and opens structured GitHub Pull Requests containing implementation summaries and changes.

  • Real-Time Engineering Cockpit: Streams agent states, execution events, test results, and workflow transitions through SSE into a live mission-control dashboard with a visual DAG.

The resulting workflow is:

Requirement → Plan → Research → Architecture → Code → Test → Failure → Diagnose → Heal → Retest → Review → GitHub PR


How we built it

DaedalusAi follows a Producer–Verifier architecture with strict contracts between components, allowing different members of the team to develop and test their components independently.

AI Orchestration

We use LangGraph as the orchestration layer. Its cyclic StateGraph manages shared agent state, conditional routing, retries, and iteration limits.

The core agents include:

  • Planner Agent — decomposes requirements and creates dependency-aware tasks.
  • Research Agent — retrieves relevant engineering knowledge using RAG.
  • Architect Agent — produces system architecture and implementation plans.
  • Developer Agent — generates and modifies repository code.
  • Tester Agent — creates and executes automated tests.
  • Reviewer/Healer Agent — analyzes failures and produces corrective patches.
  • Documentation Agent — generates implementation documentation and changelogs.

LLM Gateway

Instead of coupling the system to a single model, DaedalusAi uses a resilient multi-LLM gateway.

The gateway handles:

  • Model selection
  • Structured output
  • Retry and timeout handling
  • Rate-limit recovery
  • Provider fallback
  • Token tracking
  • Latency monitoring

Our current routing prioritizes Gemini 2.0 Flash, with Groq/Llama as a fallback and a deterministic mock provider for development and offline testing.

RAG & Memory

DaedalusAi uses Qdrant as its semantic memory layer.

Engineering documents are:

Parsed → Chunked → Embedded → Stored → Retrieved → Injected into Agent Context

This allows agents to ground their decisions in project-specific documentation instead of relying entirely on model knowledge.

Execution & Verification

Generated code is executed through controlled subprocess workspaces. Tests are run using:

pytest -v --tb=short

Instead of feeding an entire terminal output back to the LLM, a dedicated traceback parser extracts the highest-signal failure information. This reduces context pollution and gives the healing agent focused evidence.

GitHub Delivery

The delivery engine interacts with GitHub programmatically to:

  1. Create feature branches.
  2. Apply validated changes.
  3. Commit changes.
  4. Push the branch.
  5. Generate a Pull Request.
  6. Attach architectural summaries and changelogs.

Backend & Real-Time Telemetry

The backend is built with FastAPI and exposes asynchronous APIs and Server-Sent Events (SSE).

The frontend receives events such as:

agent.started
agent.completed
task.created
code.generated
tests.started
tests.failed
healing.started
tests.passed
pr.created
workflow.completed

These events power a live engineering cockpit.

Observability

The platform exposes Prometheus-compatible metrics for monitoring agent execution, test cycles, latency, failures, and system health.

This allows DaedalusAi to evolve from simply being an AI application into an observable autonomous engineering system.


Challenges we ran into

The Hallucination Loop

Early versions of the system sometimes generated incorrect code and then attempted to "fix" it with another incorrect solution, repeatedly producing similar failures.

We addressed this by separating generation from verification and using deterministic runtime results as the source of truth.

Context Pollution

Raw test output contained warnings, ANSI escape sequences, environment information, and irrelevant terminal output.

We built a traceback extraction layer that filters the output down to the most relevant failure information before sending it to the healing agent.

Infinite Agent Cycles

A self-healing architecture naturally introduces cycles:

Developer → Tester → Reviewer → Developer

Without safeguards, an unresolvable bug could trigger an infinite loop.

We solved this with explicit iteration budgets and conditional LangGraph routing:

$$ \text{max_iterations}=3 $$

After three unsuccessful attempts, the workflow stops and surfaces the failure instead of continuing autonomously.

LLM Rate Limits

Multi-agent execution can generate many model requests in a short period.

We therefore built the LLM gateway with provider fallback so that temporary rate limits or provider failures do not necessarily terminate the entire workflow.

Concurrent Team Development

Multiple team members were developing orchestration, verification, backend, and frontend components simultaneously.

Strict Pydantic contracts and shared schemas allowed each component to be mocked and developed independently, significantly reducing integration conflicts.


Accomplishments that we're proud of

True Autonomous Self-Healing

Our biggest achievement is the ability to observe a real test failure, extract its root cause, generate a corrective patch, and verify that the fix works — without human intervention.

For example:

Generated Code → pytest failure → traceback analysis → healing patch → pytest → PASS

This transforms the LLM from a code generator into an active participant in a deterministic engineering feedback loop.

End-to-End Autonomous Delivery

DaedalusAi is not limited to generating snippets.

It can progress through:

Requirement → Architecture → Implementation → Testing → Healing → Review → GitHub Pull Request

This gives the system ownership of an entire development workflow rather than an isolated coding step.

Resilient AI Architecture

The multi-provider LLM gateway allows the platform to gracefully handle provider failures and rate limits while maintaining the workflow state.

Real-Time Engineering Visibility

The mission-control interface provides visibility into what the autonomous system is doing at every stage, including agent transitions, task execution, test failures, healing cycles, and workflow completion.

Deterministic Verification

Rather than trusting an LLM to evaluate its own output, DaedalusAi uses the actual execution environment and automated tests as the final arbiter.

This principle is central to the platform:

The model proposes. The runtime verifies.


What we learned

Deterministic Verification Beats Self-Evaluation

LLMs are powerful generators but unreliable judges of their own code. Actual execution and automated tests provide objective feedback that can drive autonomous correction.

Autonomous Systems Need Boundaries

An agent that can continuously retry is not necessarily autonomous — it can simply be uncontrolled.

Explicit iteration limits, conditional routing, failure escalation, and structured state transitions are essential for building reliable autonomous systems.

State Machines Are Better Than Linear Chains

Traditional sequential agent pipelines work well when every step succeeds.

Software engineering does not work that way.

Tests fail. Reviews reject code. Dependencies break. Agents need to revisit earlier decisions.

LangGraph's cyclic state-machine model allows DaedalusAi to react to these failures dynamically instead of following a rigid linear pipeline.

Contracts Make Multi-Agent Development Scalable

Strict schemas between agents allow the system to remain modular. Each agent can evolve independently as long as it respects the shared contract.

Observability Is Essential for Agentic Systems

Traditional application monitoring is not enough for autonomous AI systems. We need to understand not only whether the system failed, but which agent made the decision, which model was used, what tools were called, how long the operation took, and why the workflow entered a recovery loop.


What's next for DaedalusAi

Secure Ephemeral Sandboxes

Move from local subprocess execution to isolated ephemeral Docker or WebAssembly sandboxes for safer execution of generated and third-party code.

Multi-Language Engineering

Expand beyond Python/FastAPI to support:

  • TypeScript / Node.js
  • Go
  • Rust

with framework-aware generation and native test runners such as Jest and Cargo Test.

Human-in-the-Loop Engineering Gates

Introduce approval checkpoints where senior engineers can review architectural changes, security-sensitive modifications, or production deployments before execution.

Enterprise-Scale Codebase Understanding

Integrate AST-based analysis using technologies such as Tree-sitter to understand large multi-file repositories without relying entirely on raw context windows.

Intelligent Engineering Optimization

Introduce ML-driven predictions for:

  • Task completion time
  • Bug probability
  • Task complexity
  • Agent performance
  • Model selection
  • Failure likelihood

These predictions can eventually allow DaedalusAi to dynamically select the best agent/model strategy for each engineering task.

Full Engineering Observability

Expand the telemetry system into a complete AI engineering observability layer measuring:

Agent Success → LLM Cost → Token Usage → Latency → Test Failures → Healing Cycles → PR Quality → Human Intervention

Our long-term goal is to make DaedalusAi not just an AI coding system, but an autonomous, measurable, self-correcting software engineering organization.

Built With

  • autonomous-agents
  • cytoscape-js
  • devops
  • docker
  • fastapi
  • gemini-api
  • github-api
  • grafana
  • groq
  • html5
  • javascript
  • langgraph
  • multi-agent-systems
  • prometheus
  • pydantic
  • pygithub
  • pytest
  • python
  • qdrant
  • rag
  • self-healing-code
  • server-sent-events
  • tailwind-css
  • vector-database
Share this project:

Updates

Submission history