Inspiration

AI agents are becoming easier to build, but much harder to trust in production.

A tool-using agent can look impressive in a happy-path demo and still make dangerous decisions when it encounters incomplete evidence, ambiguous failures, or unfamiliar context. We wanted to build something that does more than evaluate an agent after the fact.

WINDTUNNEL asks a harder question:

Can an agent learn from its failures, change how it operates, and still be blocked from shipping if it remains unsafe?

What it does

WINDTUNNEL is an autonomous engineering and reliability system for AI agents.

It runs a tool-using agent through realistic simulated incidents, records the full execution trajectory, evaluates the result deterministically, reflects on failures, stores structured operational memory, and reuses that memory in later runs.

If memory alone is not enough, WINDTUNNEL treats the failure as structural and mutates the AgentSpec itself.

The loop is:

Observe → Learn → Repair → Compare → Freeze → Certify

The main demo uses an incident-response agent with tools such as:

  • query_metrics
  • inspect_logs
  • get_recent_deployments
  • restart_service
  • rollback_deployment
  • escalate

WINDTUNNEL evaluates not only the final answer, but also:

  • which tools were used
  • tool ordering
  • required checks
  • unsafe remediation
  • escalation correctness
  • completion
  • latency
  • regressions

Learning across runs

After a failed execution, WINDTUNNEL creates structured memory from the actual trace and evaluator result.

A learned rule includes:

  • trigger/context
  • recommended action
  • action to avoid
  • rationale
  • confidence
  • evidence count
  • source runs

That memory is retrieved in later, distinct runs and can change the agent's behavior.

We also implemented memory revision so new evidence can narrow or correct an earlier learned rule instead of simply appending more text.

Structural agent engineering

WINDTUNNEL is not only a prompt optimizer.

When a serious failure repeats despite relevant memory being retrieved, WINDTUNNEL can classify the issue as structural and mutate the AgentSpec.

Supported mutation families include:

  • ADD_REQUIRED_CHECK
  • ADD_SAFETY_GATE
  • REORDER_WORKFLOW_STEP
  • CHANGE_ESCALATION_POLICY

For example, one repeated unsafe pattern caused WINDTUNNEL to make inspect_logs a permanent required diagnostic before remediation.

The mutation is driven by failure categories, RunTraces, evaluator evidence, and the current AgentSpec — not by hardcoded scenario IDs.

Candidate comparison and Regression Guard

WINDTUNNEL evaluates multiple candidate AgentSpecs on the same fixed regression suite.

In the committed benchmark:

  • Baseline V1: 2/6 success, 12 unsafe actions
  • V2: 2/6 success, 6 unsafe-action records
  • V3: 3/6 success, 4 unsafe-action records
  • Regressions: 0
  • Critical safety regressions: 0

V3 was selected because it reduced unsafe behavior and improved success while preserving existing successes.

But selection is not the same as certification.

Any critical safety regression blocks promotion.

Freeze and sealed certification

Before the final holdout test, WINDTUNNEL freezes the selected AgentSpec and computes a stable SHA-256 hash.

The sealed scenarios are not used during learning, reflection, mutation, or candidate selection.

Final sealed result:

  • 2/3 successful scenarios
  • 1 unsafe action
  • 1 policy violation
  • PROMOTION BLOCKED

This failure is intentional and important.

WINDTUNNEL does not force a green result just because the agent improved. If an unseen test still exposes unsafe behavior, the agent does not ship.

Live browser run

Judges can also run a fresh AI-agent execution directly from the browser.

The live path uses TensorMux with glm-4-7-flash to choose tools, while WINDTUNNEL's simulator and evaluator remain deterministic.

The browser shows:

  • run ID
  • incident
  • real model-driven tool calls
  • actual simulator tool results
  • deterministic evaluator checks
  • unsafe actions
  • reflection and learned lessons
  • final promotion status

There is no heuristic fallback that fabricates a successful live result.

Generalization proof

To verify that the learning abstractions were not specific to incident response, we reused the same memory/retrieval/spec-application abstractions in a tiny refund workflow.

Results:

  • Order A before learning: FAIL
  • Order B with recalled memory: PASS
  • Order B without memory: FAIL

The same learned-memory mechanism prevented an invalid refund in a different domain.

How we built it

WINDTUNNEL was built from scratch during the hackathon using AO (Agent Orchestrator) throughout the development process.

We used separate AO sessions for:

  • deterministic simulator and baseline agent
  • evaluation harness
  • learning and persistent memory
  • memory retrieval and revision
  • structural mutation
  • regression protection
  • candidate comparison and certification
  • final live browser execution and productization

Claude Code and OpenCode were used as coding agents inside AO.

This gave us an auditable development history with isolated tasks, iterative testing, and clear engineering milestones.

Challenges we ran into

The biggest challenge was avoiding a fake-looking optimization loop.

It would have been easy to hardcode a scenario-specific fix or report only improving aggregate scores.

Instead, we focused on:

  • deterministic evaluation
  • real execution traces
  • persistent memory reuse
  • structural mutations derived from evidence
  • regression protection
  • sealed holdout testing
  • honest failure reporting

Another challenge was ensuring the live browser run remained real while keeping the committed benchmark reproducible.

We solved this by separating:

LIVE RUN — a fresh model-driven execution

from

EXECUTED PROOF — a committed, reproducible benchmark artifact

What we learned

The biggest lesson was that making an agent “better” is not the same as making it safe enough to deploy.

A system that improves average performance but hides regressions is not reliable engineering.

The most valuable part of WINDTUNNEL is therefore not the score increase itself — it is the decision boundary around what is allowed to ship.

What's next

Future work could extend WINDTUNNEL with:

  • larger multi-domain evaluation suites
  • richer candidate search
  • broader tool/MCP integration
  • learned mutation policies
  • production trace ingestion
  • team approval workflows
  • more advanced cost/latency optimization

The long-term vision is: CI/CD for learning agent architectures — build, break, learn, repair, and certify before production.

Built With

  • agent-orchestrator
  • ai-agents
  • ao
  • claude-code
  • glm-4.7-flash
  • llm
  • next.js
  • node.js
  • opencode
  • react
  • tensormux
  • tool-calling
  • typescript
  • vercel
Share this project:

Updates

Submission history