Inspiration
AI agents are becoming easier to build, but much harder to trust in production.
A tool-using agent can look impressive in a happy-path demo and still make dangerous decisions when it encounters incomplete evidence, ambiguous failures, or unfamiliar context. We wanted to build something that does more than evaluate an agent after the fact.
WINDTUNNEL asks a harder question:
Can an agent learn from its failures, change how it operates, and still be blocked from shipping if it remains unsafe?
What it does
WINDTUNNEL is an autonomous engineering and reliability system for AI agents.
It runs a tool-using agent through realistic simulated incidents, records the full execution trajectory, evaluates the result deterministically, reflects on failures, stores structured operational memory, and reuses that memory in later runs.
If memory alone is not enough, WINDTUNNEL treats the failure as structural and mutates the AgentSpec itself.
The loop is:
Observe → Learn → Repair → Compare → Freeze → Certify
The main demo uses an incident-response agent with tools such as:
query_metricsinspect_logsget_recent_deploymentsrestart_servicerollback_deploymentescalate
WINDTUNNEL evaluates not only the final answer, but also:
- which tools were used
- tool ordering
- required checks
- unsafe remediation
- escalation correctness
- completion
- latency
- regressions
Learning across runs
After a failed execution, WINDTUNNEL creates structured memory from the actual trace and evaluator result.
A learned rule includes:
- trigger/context
- recommended action
- action to avoid
- rationale
- confidence
- evidence count
- source runs
That memory is retrieved in later, distinct runs and can change the agent's behavior.
We also implemented memory revision so new evidence can narrow or correct an earlier learned rule instead of simply appending more text.
Structural agent engineering
WINDTUNNEL is not only a prompt optimizer.
When a serious failure repeats despite relevant memory being retrieved, WINDTUNNEL can classify the issue as structural and mutate the AgentSpec.
Supported mutation families include:
ADD_REQUIRED_CHECKADD_SAFETY_GATEREORDER_WORKFLOW_STEPCHANGE_ESCALATION_POLICY
For example, one repeated unsafe pattern caused WINDTUNNEL to make inspect_logs a permanent required diagnostic before remediation.
The mutation is driven by failure categories, RunTraces, evaluator evidence, and the current AgentSpec — not by hardcoded scenario IDs.
Candidate comparison and Regression Guard
WINDTUNNEL evaluates multiple candidate AgentSpecs on the same fixed regression suite.
In the committed benchmark:
- Baseline V1: 2/6 success, 12 unsafe actions
- V2: 2/6 success, 6 unsafe-action records
- V3: 3/6 success, 4 unsafe-action records
- Regressions: 0
- Critical safety regressions: 0
V3 was selected because it reduced unsafe behavior and improved success while preserving existing successes.
But selection is not the same as certification.
Any critical safety regression blocks promotion.
Freeze and sealed certification
Before the final holdout test, WINDTUNNEL freezes the selected AgentSpec and computes a stable SHA-256 hash.
The sealed scenarios are not used during learning, reflection, mutation, or candidate selection.
Final sealed result:
- 2/3 successful scenarios
- 1 unsafe action
- 1 policy violation
- PROMOTION BLOCKED
This failure is intentional and important.
WINDTUNNEL does not force a green result just because the agent improved. If an unseen test still exposes unsafe behavior, the agent does not ship.
Live browser run
Judges can also run a fresh AI-agent execution directly from the browser.
The live path uses TensorMux with glm-4-7-flash to choose tools, while WINDTUNNEL's simulator and evaluator remain deterministic.
The browser shows:
- run ID
- incident
- real model-driven tool calls
- actual simulator tool results
- deterministic evaluator checks
- unsafe actions
- reflection and learned lessons
- final promotion status
There is no heuristic fallback that fabricates a successful live result.
Generalization proof
To verify that the learning abstractions were not specific to incident response, we reused the same memory/retrieval/spec-application abstractions in a tiny refund workflow.
Results:
- Order A before learning: FAIL
- Order B with recalled memory: PASS
- Order B without memory: FAIL
The same learned-memory mechanism prevented an invalid refund in a different domain.
How we built it
WINDTUNNEL was built from scratch during the hackathon using AO (Agent Orchestrator) throughout the development process.
We used separate AO sessions for:
- deterministic simulator and baseline agent
- evaluation harness
- learning and persistent memory
- memory retrieval and revision
- structural mutation
- regression protection
- candidate comparison and certification
- final live browser execution and productization
Claude Code and OpenCode were used as coding agents inside AO.
This gave us an auditable development history with isolated tasks, iterative testing, and clear engineering milestones.
Challenges we ran into
The biggest challenge was avoiding a fake-looking optimization loop.
It would have been easy to hardcode a scenario-specific fix or report only improving aggregate scores.
Instead, we focused on:
- deterministic evaluation
- real execution traces
- persistent memory reuse
- structural mutations derived from evidence
- regression protection
- sealed holdout testing
- honest failure reporting
Another challenge was ensuring the live browser run remained real while keeping the committed benchmark reproducible.
We solved this by separating:
LIVE RUN — a fresh model-driven execution
from
EXECUTED PROOF — a committed, reproducible benchmark artifact
What we learned
The biggest lesson was that making an agent “better” is not the same as making it safe enough to deploy.
A system that improves average performance but hides regressions is not reliable engineering.
The most valuable part of WINDTUNNEL is therefore not the score increase itself — it is the decision boundary around what is allowed to ship.
What's next
Future work could extend WINDTUNNEL with:
- larger multi-domain evaluation suites
- richer candidate search
- broader tool/MCP integration
- learned mutation policies
- production trace ingestion
- team approval workflows
- more advanced cost/latency optimization
The long-term vision is: CI/CD for learning agent architectures — build, break, learn, repair, and certify before production.
Built With
- agent-orchestrator
- ai-agents
- ao
- claude-code
- glm-4.7-flash
- llm
- next.js
- node.js
- opencode
- react
- tensormux
- tool-calling
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.