Inspiration

Most bug reports still arrive as screen recordings, screenshots, or vague notes. The problem is that an engineer still has to reproduce the issue, inspect what happened, write a test, and remember to check it again later.

TaskTape Replay started from a simple question: what if a bug recording could become a reusable regression check?

What it does

TaskTape Replay is a macOS desktop app that lets Claude Code or Codex reproduce a local browser bug through a built-in MCP server. While the agent investigates, TaskTape captures the actions, screenshots, DOM snapshots, console logs, network failures, and trace evidence.

From that session, TaskTape creates a reviewable check. GPT-5.6 can replay the workflow on the real interface, compare the final screen with the expected outcome, and save a clear pass or fail result. The check can also be scheduled, exported as Playwright, or turned into a ticket-ready report.

How we built it

We built TaskTape with Electron, React, TypeScript, Playwright, MCP, and the OpenAI API. GPT-5.6 is used for structured workflow understanding, computer-use replay, and visual outcome evaluation.

Codex was used throughout the build process to research the market, design the product direction, implement features, test the desktop app, and prepare the final release.

Challenges

The hardest part was making the product feel real instead of like a demo. We had to make the app record useful evidence, replay workflows safely, handle schedules, keep local history, export readable tests, and package everything into a working macOS build.

Another challenge was scope. General desktop automation is huge, so we focused the demo on one strong use case: turning an agent-reproduced browser bug into a reusable regression check.

What we learned

A useful AI product should not just say “the agent completed the task.” It should leave behind evidence that a human can inspect and reuse.

For TaskTape, that means every check has a visible outcome, local history, screenshots, traces, and exportable Playwright code. The goal is to turn temporary debugging context into something that keeps protecting the product.

Built With

Share this project:

Updates

Submission history