Inspiration
AI agents are becoming capable of making decisions and calling tools on behalf of users. But testing an agent only by looking at its final response can miss dangerous behavior.
An agent might respond correctly while still making an unsafe tool call in the background.
We wanted to build a practical testing lab where developers could deliberately try to break an AI agent, understand exactly why it failed, and verify whether a fix actually worked.
That idea became AgentCrashLab.
What it does
AgentCrashLab is a crash-testing and regression platform for AI agents.
It runs adversarial scenarios against a sandbox Customer Support Agent, including prompt injection, authorization failures, unsafe refunds, contradictory instructions, ambiguous requests, tool failures, timeouts, and goal drift.
Each run produces an execution trace and can generate Failure DNA, showing the expected behavior, observed behavior, tool involved, evidence, and a remediation hint.
Developers can also mutate failing scenarios and compare different agent versions to verify whether a fix actually improves reliability.
How we built it
The frontend uses React, Vite, TypeScript, and Tailwind CSS. The backend is built with Node.js, Express, and Zod.
We use PostgreSQL with Prisma to store agents, runs, failures, and traces. Redis and BullMQ handle asynchronous crash-test jobs through a dedicated evaluator worker.
For AI-assisted evaluation, we use Google Gemini for scenario generation, mutations, and nuanced scoring when available. Safety-critical checks remain deterministic, such as detecting a refund without confirmation.
The architecture follows:
React UI → Express API → PostgreSQL Express API → Redis/BullMQ → Evaluator Worker → Safety Rules + Gemini
Challenges we ran into
One major challenge was making safety evaluation reliable.
Using an LLM alone to determine whether an action is safe can produce inconsistent results. We therefore separated objective safety checks from subjective evaluation.
Critical behaviors are checked with deterministic rules, while Gemini is used for cases where flexible reasoning is useful.
Another challenge was making failures useful to developers. Instead of simply showing “failed,” we built execution traces and Failure DNA to explain what happened and what should be changed.
Accomplishments that we're proud of
We are proud of building a complete attack → observe → diagnose → harden → retest workflow.
Our demo includes a Customer Support Agent with mock tools for order search, cancellation, refunds, and email. Version 1 is intentionally vulnerable, while version 2 applies stricter safety rules.
AgentCrashLab can run 16 adversarial scenarios, inspect critical failures, explore mutations, and compare agent versions to measure improvement.
Most importantly, the platform turns AI safety testing into something developers can repeatedly run and measure instead of relying only on manual testing.
What we learned
We learned that agent reliability is about behavior, not just responses.
A good-looking response does not guarantee that the agent followed authorization rules, used tools correctly, or avoided unsafe actions.
We also learned that combining deterministic checks with LLM-based evaluation is more practical than relying entirely on either approach.
What's next for AgentCrashLab
We want to expand AgentCrashLab with larger attack libraries, support for more agent frameworks, CI/CD integration, automated regression gates, stronger isolation, and deeper observability.
Our goal is to make testing AI agents as natural as testing traditional software:
Break it. Understand it. Fix it. Prove it's better.
Built With
- bullmq
- express.js
- google-gemini
- neon
- node.js
- postgresql
- prisma
- react
- redis
- render
- tailwind-css
- typescript
- vite
- zod
Log in or sign up for Devpost to join the conversation.