Inspiration

QA engineers and developers still spend hours manually clicking through web applications for exploratory testing, retesting bug fixes, and regression checks after releases. We wanted to build an autonomous agent that could handle these repetitive yet critical testing tasks using AWS Strands Agents SDK and Amazon Bedrock.

What it does

GemmaQA is an autonomous QA agent that opens a real Chromium browser and tests authorized web applications. It autonomously explores pages, fills and submits forms, executes CRUD workflows, detects validation errors, recovers from failures, and generates structured QA reports with screenshots and activity logs.

The AWS Strands Agent coordinates the testing lifecycle by calling custom tools to create, start, pause, resume, and end QA runs. The AgentController uses Playwright to drive the browser, executing one safe action at a time while respecting authorization boundaries and safety checks.

Key capabilities include live operator controls (pause/resume/end run), form validation error recovery, test-data creation and cleanup for CRUD testing, and comprehensive evidence capture (screenshots, traces, activity logs, structured reports).

How we built it

  • AWS Strands Agents SDK + Amazon Bedrock for agent coordination and orchestration
  • Strands Agent tools wrapping the run_manager lifecycle (create/start/pause/resume/end)
  • FastAPI backend with SQLAlchemy and SQLite for run state management
  • Playwright for direct Chromium browser automation
  • React + TypeScript + Vite dashboard with live WebSocket activity streaming
  • Pydantic for structured data validation and safety policy enforcement
  • Docker for containerization with Playwright Chromium support
  • Optional Gemini API integration for enhanced page-by-page reasoning

The architecture separates the Strands coordinator (orchestration) from AgentController (execution) from Playwright (browser automation), with clear tool interfaces between each layer.

Challenges we ran into

  1. Form validation recovery - Handling cases where applications reject inputs (like invalid phone numbers) without getting stuck in retry loops. We solved this by treating validation errors as application feedback and allowing the agent to learn and adapt.

  2. Safety boundaries - Ensuring the agent only operates on authorized domains and respects destructive action permissions. We implemented a multi-layer safety policy with domain validation and explicit operator permissions.

  3. Strands tool integration - Wrapping existing run_manager functionality into tools that the Strands Agent could call while maintaining backward compatibility with direct API calls.

  4. Live activity streaming - Building WebSocket infrastructure to stream real-time testing progress to the UI while maintaining durable logs for post-run analysis.

Accomplishments that we're proud of

  • Successfully integrated AWS Strands Agents SDK with a real-world browser automation system
  • Demonstrated autonomous signup, CRUD operations (add/edit/delete contacts), and form validation recovery on the Thinking Tester Contact List application
  • Built a production-ready FastAPI + React system with live controls, WebSocket streaming, and comprehensive evidence capture
  • Implemented safety-first design with authorization checks, domain validation, and operator override controls
  • Created a flexible architecture supporting multiple LLM providers (Bedrock, Gemini, Ollama, Mock) through a clean abstraction layer

What we learned

  • How to design tool interfaces for AWS Strands Agents that wrap complex stateful operations
  • The importance of treating validation errors as application feedback rather than failures
  • Safety policy design for autonomous agents - authorization boundaries are critical
  • Real-time streaming architecture for live agent monitoring and control
  • Balancing autonomous exploration with operator oversight and safety stops

What's next for GemmaQA

  • AgentCore deployment - Deploy the Strands coordinator to Amazon Bedrock AgentCore Runtime for serverless operation
  • Multi-agent workflows - Parallel testing across different application sections
  • Enhanced reporting - Visual regression detection, accessibility checks, and performance metrics
  • Learning from history - Use previous run evidence to improve future exploration strategies
  • Cloud-native deployment - Horizontal scaling with container orchestration for concurrent test runs

Built With

Share this project:

Updates

Submission history