Inspiration

Writing unit tests is one of the most repetitive and time-consuming parts of software development. As an engineering student, I constantly found myself spending hours writing Happy Path, edge case, and negative tests just to get decent coverage. Existing tools either generate very basic tests or require heavy setup.I wanted an agent that could take this busywork off my plate — something that understands the code, generates meaningful tests, runs them, checks coverage, and improves itself when needed. The Agents for Humans hackathon theme of building agents that handle repetitive work in the background perfectly matched this idea. That’s how TestGen Agent was born.

What it does

TestGen Agent is an intelligent AI agent that automatically generates high-quality unit tests for Python code. It analyzes your source code, creates multiple types of tests (Happy Path, Negative/Error, Boundary/Edge, and Parametrized), optionally adds property-based tests using Hypothesis, runs the tests, measures coverage, and iteratively improves the test suite when coverage is incomplete. It runs quietly and only surfaces results or asks for input when needed — perfectly aligned with the “agent that handles repetitive work” theme.

How we built it

TestGen Agent is built using the Strands Agents SDK as the core orchestration layer.

  1. The main agent uses Strands tools to analyze code, generate tests, run them, and improve coverage.
  2. Test generation focuses on four key unit test types: Happy Path, Negative/Error, Boundary/Edge, and Parametrized tests.
  3. We integrated Hypothesis for property-based testing to make the generated tests stronger and more unique.
  4. A coverage-driven improvement loop allows the agent to detect missing lines and generate additional tests automatically.
  5. The system supports both a CLI and a Streamlit web interface, where users can upload files/folders or paste a GitHub URL.
  6. Tests are executed safely using pytest + coverage, with optional Docker sandboxing.
  7. We also added explainable tests (comments explaining why each test was generated) and a clean results dashboard. The architecture keeps the heavy logic in tools while letting the Strands Agent decide the workflow.

Challenges we ran into

  1. Making the agent reliably generate correct assertions (especially for negative and boundary cases) was harder than expected.
  2. Balancing determinism (rule-based generation) with LLM-powered flexibility took several iterations.
  3. Integrating the real Strands Agents SDK properly and handling cases with missing LLM credentials required careful fallback design.
  4. Getting good coverage improvement without creating flaky or failing tests was challenging.
  5. Designing a clean and impressive Streamlit UI while keeping the backend robust within limited time was difficult.

Accomplishments that we're proud of

  1. Successfully built a working agent on top of the Strands Agents SDK.
  2. Combined traditional unit test generation with Hypothesis property-based testing.
  3. Implemented a real coverage improvement loop that makes the agent feel intelligent.
  4. Created both CLI and web interfaces with support for file upload and GitHub repos.
  5. Made the system usable even without LLM credentials (deterministic fallback).
  6. Delivered a complete, demo-ready project focused on a real developer pain point.

What we learned

  1. How to properly structure an agent using the Strands Agents SDK (tools + system prompt + agent loop).
  2. The importance of having a strong deterministic baseline before adding LLM intelligence.
  3. Generating high-quality tests is much harder than just generating “some” tests.
  4. Good UX (clear progress, coverage visualization, explanations) significantly improves the perceived quality of an AI agent.
  5. Building for a hackathon requires ruthless prioritization — focusing on depth rather than too many features.

What's next for TestGen Agent

  1. Improve the quality of generated assertions using better analysis and feedback loops.
  2. Add support for more complex code (classes with dependencies, async functions, etc.).
  3. Introduce automatic test repairing when generated tests fail.
  4. Explore multi-file / project-level test generation.
  5. Make the agent available as a GitHub Action for continuous test generation.

Built With

Share this project:

Updates

Submission history