-
-
Light theme of TestGen webapp
-
Python file uploaded to test the code
-
The python is tested and it shows the test status here
-
All the test are shown
-
Coverage across runs
-
Why these tests were generated
-
Each test below shows its category, rationale, target lines, and whether mocking was detected
-
Shows the quality of code
-
Download of all test files is given
-
Run history for your 20 most recent runs that are saved on this device and restored after refresh or restart
-
It uses testing frameworks like pytest and hypothesis to create test cases & automatically writes out scenarios to test the code's behavior
-
Works in a coverage loop, generates tests, runs them, sees what parts of the code were missed & then generate more tests to cover those gaps
-
Setup and guide page
-
Tells execution environment, visual theme, and guide on how to read a test report
Inspiration
Writing unit tests is one of the most repetitive and time-consuming parts of software development. As an engineering student, I constantly found myself spending hours writing Happy Path, edge case, and negative tests just to get decent coverage. Existing tools either generate very basic tests or require heavy setup.I wanted an agent that could take this busywork off my plate — something that understands the code, generates meaningful tests, runs them, checks coverage, and improves itself when needed. The Agents for Humans hackathon theme of building agents that handle repetitive work in the background perfectly matched this idea. That’s how TestGen Agent was born.
What it does
TestGen Agent is an intelligent AI agent that automatically generates high-quality unit tests for Python code. It analyzes your source code, creates multiple types of tests (Happy Path, Negative/Error, Boundary/Edge, and Parametrized), optionally adds property-based tests using Hypothesis, runs the tests, measures coverage, and iteratively improves the test suite when coverage is incomplete. It runs quietly and only surfaces results or asks for input when needed — perfectly aligned with the “agent that handles repetitive work” theme.
How we built it
TestGen Agent is built using the Strands Agents SDK as the core orchestration layer.
- The main agent uses Strands tools to analyze code, generate tests, run them, and improve coverage.
- Test generation focuses on four key unit test types: Happy Path, Negative/Error, Boundary/Edge, and Parametrized tests.
- We integrated Hypothesis for property-based testing to make the generated tests stronger and more unique.
- A coverage-driven improvement loop allows the agent to detect missing lines and generate additional tests automatically.
- The system supports both a CLI and a Streamlit web interface, where users can upload files/folders or paste a GitHub URL.
- Tests are executed safely using pytest + coverage, with optional Docker sandboxing.
- We also added explainable tests (comments explaining why each test was generated) and a clean results dashboard. The architecture keeps the heavy logic in tools while letting the Strands Agent decide the workflow.
Challenges we ran into
- Making the agent reliably generate correct assertions (especially for negative and boundary cases) was harder than expected.
- Balancing determinism (rule-based generation) with LLM-powered flexibility took several iterations.
- Integrating the real Strands Agents SDK properly and handling cases with missing LLM credentials required careful fallback design.
- Getting good coverage improvement without creating flaky or failing tests was challenging.
- Designing a clean and impressive Streamlit UI while keeping the backend robust within limited time was difficult.
Accomplishments that we're proud of
- Successfully built a working agent on top of the Strands Agents SDK.
- Combined traditional unit test generation with Hypothesis property-based testing.
- Implemented a real coverage improvement loop that makes the agent feel intelligent.
- Created both CLI and web interfaces with support for file upload and GitHub repos.
- Made the system usable even without LLM credentials (deterministic fallback).
- Delivered a complete, demo-ready project focused on a real developer pain point.
What we learned
- How to properly structure an agent using the Strands Agents SDK (tools + system prompt + agent loop).
- The importance of having a strong deterministic baseline before adding LLM intelligence.
- Generating high-quality tests is much harder than just generating “some” tests.
- Good UX (clear progress, coverage visualization, explanations) significantly improves the perceived quality of an AI agent.
- Building for a hackathon requires ruthless prioritization — focusing on depth rather than too many features.
What's next for TestGen Agent
- Improve the quality of generated assertions using better analysis and feedback loops.
- Add support for more complex code (classes with dependencies, async functions, etc.).
- Introduce automatic test repairing when generated tests fail.
- Explore multi-file / project-level test generation.
- Make the agent available as a GitHub Action for continuous test generation.
Built With
- ai
- amazon-web-services
- automation
- code
- coverage
- developer
- docker
- generation
- github
- hypothesis
- llm
- openai
- pytest
- pytest-cov
- python
- sdk
- streamlit
- test
- testing
- tools
- unit
- web
Log in or sign up for Devpost to join the conversation.