Demo Video
https://youtu.be/pqsBWAL0rRY Github Repo - https://github.com/HarmanPreet-Singh-XYT/RazeQA
(Note for judges: Please refer to the link above if an outdated preview appears in the player is of 3:05 minutes coz that is showing older video, actual is around 4:45 minutes.)
Inspiration
With modern AI coding agents like Claude Code, Cursor, and Copilot, developers ship features and refactors at unprecedented speed. However, QA remains the single greatest bottleneck in modern software delivery:
- Traditional CI suites test only predefined human assertions — they miss regressions in adjacent, nuanced user journeys that developers never anticipated.
- Autonomous code generation introduces silent regressions — an agent fixing a subtle backend edge case can inadvertently break an authentication modal or trigger visual layout shifts across responsive viewports.
- Reproduction and triage remain manual and slow — when unexpected errors occur, developers spend hours attempting reproduction without video replays, network traces, or visibility into the original design intent behind the changes.
We built RazeQA to bridge this gap: an autonomous closed-loop platform where intelligent agents interpret developer intent, dynamically navigate and test preview builds in live browsers, record undeniable forensic proof of regressions, and deliver instant remediation prompts directly to the pull request.
What it does
RazeQA is an autonomous closed-loop QA engine that bridges prompt intent to full-fidelity browser verification:
- Agent Intent Bridge: A background daemon and CLI plugin (
agent-bridge) hooks into local coding environments (Claude Code, Cursor, Copilot). It captures both the raw code diffs and the underlying developer intent, assembling user prompts and agent reasoning loops into a structured Intent Stream. - Ephemeral Sandbox Orchestration: On every Pull Request or on-demand check, RazeQA provisions an isolated, hardened container, clones the target branch, resolves dependencies, and boots the live preview build.
- Autonomous Vision & Browser Agents: Using Playwright coupled with multi-modal vision models, RazeQA dynamically explores affected pages, executes critical user flows, examines dynamic state transitions, and detects anomalies including layout shifts, responsive breakage, console errors, and failing network calls.
- Baseline Differential Analysis: Journeys execute concurrently against the PR branch and the
mainbaseline to reliably isolate new regressions from pre-existing issues. - Forensic Artifact Packaging: For every failure, RazeQA captures synchronized video recordings (
.mp4/.webm), Playwright trace archives (.zip), network waterfalls, and DOM accessibility snapshots. - Self-Healing Remediation Prompts: Posts native GitHub Check Runs and PR comments containing a structured root-cause analysis and an actionable AI Fix Prompt that coding agents can immediately execute to resolve the issue.
How we built it
RazeQA is designed around a decoupled, multi-engine architecture:
- Autonomous Orchestration Engine (Python / FastAPI / Docker / Playwright):
- Agent Orchestrator: Multi-agent planning routines synthesize comprehensive user journeys directly from git diffs and prompt intent logs.
- Vision & Exploration Layer: Employs Playwright Chromium with element bounding box analysis and perceptual visual diffing to explore arbitrary web applications without relying on brittle DOM selectors.
- Container Sandboxing: Hardened container lifecycle management handles dynamic port mapping, memory isolation, and zero-host-disk source retention.
GitHub App Integration: Native GitHub Check Runs API and PR commenting pipelines authenticate via short-lived RS256 JWTs and ephemeral installation tokens.
Control Center & Telemetry Dashboard (Next.js 16 / TypeScript / Tailwind CSS / Supabase):
Command Interface: Real-time run timeline, interactive video player with synchronized step-by-step telemetry, network inspector, and one-click fix handoffs.
Data Layer: Multi-tenant PostgreSQL schema with Row-Level Security (RLS), HMAC-verified webhook processing, and signed cross-origin artifact storage.
Developer Tools Suite: Built-in utilities featuring OpenGraph inspection, schema validation, header analysis, and payload verification.
Challenges I ran into
- Capturing True Developer Intent: Linking raw git diffs to human prompts required building a lightweight local daemon that connects unobtrusively to tool execution loops over WebSockets without interrupting local development workflows.
- Dynamic Browser Navigation & Deep DOM Scrolling: Standard automation tools fail when layouts contain complex nested scroll containers (
overflow-y: auto) or dynamic modals. We implemented accessibility tree traversal and element-targeted scroll inspection, giving vision agents complete spatial awareness. - Container Memory & Resource Isolation: Full-stack compilation steps (such as Next.js Turbopack and Tailwind builds) routinely exhausted memory in constrained environments. We resolved this through targeted Node heap tuning (
--max-old-space-size=2048) and dynamic memory allocation (3g+). - Secure GitHub App Authentication: We established a strict zero-persistent-secret model by implementing RS256 JWT generation, on-the-fly 1-hour installation token minting, and HMAC-SHA256 signature verification across all webhook events.
Accomplishments that I am proud of
- 🎬 True Zero-Config Video Forensics: Generates high-resolution video recordings, Playwright traces, and visual snapshots of real regressions on live PRs without requiring developers to write a single line of test code.
- 🔁 The Closed-Loop Feedback Cycle: Delivers an unbroken workflow from initial human prompt through coding agent implementation, autonomous browser verification, and automated fix delivery back to GitHub in minutes.
- 🛡️ Hermetic Test Suite: Validated with 666+ automated tests verifying security boundaries, memory limits, webhook handshakes, and sandbox isolation.
- 🚀 Full End-to-End Delivery: Glues together a local CLI, cloud backend, PostgreSQL/Supabase database, Next.js frontend, and GitHub App into a unified developer tool.
What we learned
- Intent transforms QA: Knowing why code changed makes dynamic test synthesis substantially more effective than analyzing static ASTs alone.
- Vision-language agents excel at exploratory testing: Merging DOM accessibility trees with visual screenshots enables AI agents to evaluate modern, interactive user interfaces with human-level discernment.
- Actionable proof builds engineering trust: Teams trust automated failures when accompanied by definitive visual evidence, network waterfalls, and reproducible trace packages.
What's next for RazeQA
- 📱 Cross-Browser & Multi-Device Testing: Expanding verification to mobile viewports (iOS Safari and Android Chrome emulation) along with full multi-browser matrices (Firefox, WebKit).
- ⚡ Autonomous Patch Commit Workflows: Adding an automated remediation flow where RazeQA tests its generated patch in a fresh sandbox and commits the verified fix directly to the PR branch.
- 🤝 Direct IDE Extensions: Providing dedicated plugins for Cursor, VS Code, and JetBrains to let developers run preview verification before pushing code upstream.
- 🧠 Repository Memory & Flake Analysis: Tracking repository flakiness profiles and critical user conversion funnels over time to optimize verification efficiency.
Key Links
- Demo Video: Watch RazeQA in Action on YouTube
- Presentation Deck: View 9-Slide Deck (Google Slides)
Built With
- claude
- docker
- fastapi
- github-api
- jwt
- next.js
- playwright
- postgresql
- python
- react
- supabase
- tailwind-css
- turbopack
- typescript
- websockets
Log in or sign up for Devpost to join the conversation.