Inspiration

AI agents are becoming capable of taking real actions through tools, but testing them is still much harder than testing traditional software. An agent can appear to work correctly in normal situations while failing badly when it encounters ambiguous instructions, prompt injection, unsafe tool usage, or adversarial inputs.

We wanted to build a practical "crash test" environment for AI agents — something similar to crash-testing a car before putting it on the road.

That idea became AgentCrashLab.

What it does

AgentCrashLab lets developers register an AI agent, run adversarial crash-test scenarios against it, inspect execution traces, identify failures, and compare different versions of an agent.

The built-in Customer Support Agent demonstrates realistic failure cases such as unsafe refunds, authorization bypasses, prompt injection, invalid emails, and repeated tool usage.

The platform provides:

  • Adversarial crash tests to deliberately challenge agents
  • Deterministic safety rules for critical failures
  • Optional Gemini evaluation for more nuanced behavior
  • Failure DNA containing the expected behavior, observed behavior, tool involved, trace, and remediation hints
  • Failure mutation to generate and test variations of a failing scenario
  • Version comparison to measure whether a patched agent actually becomes safer

A key part of the demo is comparing a deliberately vulnerable v1 agent with a stricter v2 and seeing critical failures decrease after the fixes.

How we built it

The frontend is built with React, Vite, TypeScript, and Tailwind CSS. The backend uses Node.js, Express, and Zod, with PostgreSQL and Prisma for storing runs, failures, and execution traces.

For background jobs, the project supports in-process execution on Render and Redis + BullMQ for local development.

Google Gemini is used server-side for scenario generation, failure mutations, and nuanced evaluation. However, safety-critical checks such as refunds without confirmation are handled deterministically rather than relying entirely on an LLM judge.

The architecture is:

React UI → Express API → PostgreSQL ↓ Job processing ↓ Safety Rules + Gemini

Challenges we faced

One of the biggest challenges was deciding which failures should be judged deterministically and which require an LLM. We did not want safety-critical behavior to depend entirely on an LLM's opinion, so we implemented deterministic rules for objective violations and used Gemini only where more nuanced evaluation was useful.

Another challenge was making failures actionable rather than simply reporting that an agent failed. This led to the idea of Failure DNA, which connects the failure to the tool involved, execution trace, expected behavior, observed behavior, and a remediation hint.

We also wanted to test whether a fix actually worked. This motivated the version comparison workflow, where developers can run the same crash-test suite against different agent versions and directly see the change in reliability.

What we learned

We learned that testing AI agents requires more than checking whether the final answer is correct. An agent can reach a seemingly valid outcome while using an unsafe tool, violating authorization, or taking an action that should never have been allowed.

We also learned that reproducible adversarial testing and detailed failure traces can make AI-agent debugging much more concrete.

What's next

We want to evolve AgentCrashLab into a more complete reliability platform for production AI agents.

Future improvements include stronger sandboxing, authentication and multi-tenancy, more agent/tool integrations, larger adversarial test suites, CI/CD integration, historical reliability tracking, and more sophisticated failure clustering.

The long-term goal is simple:

Before an AI agent interacts with real users or real systems, crash-test it first.

Built With

Share this project:

Updates

Submission history