We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

AI systems are increasingly being used in areas where mistakes can have serious consequences. We wanted to explore a simple question: How do you know an AI system is actually following its safety and governance policies?

That led us to build RedTeam AI Labs, a platform that turns governance requirements into adversarial tests and uses those tests to expose weaknesses in an AI system.

What We Built

RedTeam AI Labs is an autonomous AI safety and governance testing platform. A user defines governance requirements, and the system generates a 30-test adversarial campaign across six areas:

Prompt Injection Jailbreak / Safety Evasion Sensitive Information Leakage Bias / Disparate Treatment Unsupported / Hallucinated Legal Claims Governance / Policy Compliance

The tests are executed against a synthetic AI Legal Intake Assistant and evaluated using a combination of deterministic checks and semantic AI evaluation. Findings are persisted in Supabase and presented with evidence, severity, remediation guidance, and reproducibility information.

A key design principle throughout the project is:

The LLM proposes. The backend disposes.

The AI can generate tests and provide semantic analysis, but the backend remains responsible for validation, state changes, persistence, and final severity decisions.

How We Built It

We built the application as a Next.js 14 modular monolith with TypeScript, Tailwind CSS, Supabase PostgreSQL, Zod, and the Vercel AI SDK with Google Gemini.

The system is divided into clear components for target adapters, test generation, deterministic evaluation, semantic evaluation, severity mapping, and reporting.

Campaign execution uses bounded chunks rather than one long-running request, allowing the system to operate reliably in a serverless environment. Every test execution is tracked with idempotency controls, while campaign metadata such as policy hashes, model versions, seeds, and target versions is stored for reproducibility.

The platform produces both machine-readable JSON audit data and human-readable governance reports in HTML and PDF.

What We Learned

One of the biggest lessons was that building an AI safety product requires more than simply asking an LLM whether another LLM is safe.

We learned to separate generation, evaluation, and authority. Deterministic checks handle exact conditions, semantic evaluation handles cases requiring reasoning, and the backend makes the final decision.

We also learned the importance of treating model outputs as untrusted data, designing for retries and partial failures, and making AI evaluations reproducible enough to support meaningful audits.

Challenges

The hardest parts were making the system reliable while still using AI dynamically. Model outputs can vary, APIs can fail, and serverless functions have execution limits. We addressed these challenges through structured Zod validation, chunked execution, idempotency, deterministic benchmark scenarios, and strict separation between untrusted model output and application logic.

Another challenge was making technical AI safety results understandable to non-technical reviewers. We solved this by building a visual findings inspector that shows the attack, target response, expected behavior, evidence, evaluation process, and final backend decision.

The result is a working prototype designed to make AI safety testing more practical, transparent, and auditable.

Built With

  • adversarial-testing
  • ai-evaluation
  • ai-governance
  • ai-red-teaming
  • ai-safety
  • api-testing
  • audit-reports
  • bias-detection
  • compliance
  • google-gemini
  • legal-tech
  • next.js
  • playwright
  • postgresql
  • privacy
  • prompt-injection
  • react
  • saas
  • security-testing
  • serverless
  • supabase
  • tailwind-css
  • typescript
  • vercel-ai-sdk
  • zod
Share this project:

Updates

Submission history