Inspiration

I built Retrace after repeatedly seeing how painful it is to debug AI agents. When an agent makes a wrong decision several steps into a run, logs usually tell you what happened, but not in a way you can easily reproduce and experiment with. I wanted to make debugging agents feel more like debugging normal software.

What it does

Retrace records an AI agent's execution as a replayable trace, capturing LLM calls, tool calls, inputs, outputs, errors, latency, and costs. You can replay a run, fork it from a specific point, and experiment with changes without rerunning everything from scratch.

How we built it

I built Retrace around an execution and tracing layer that captures every step of an agent run and stores it as structured execution data. On top of that, I built replay and forking capabilities so developers can jump back into an execution and continue from any point.

Challenges we ran into

The biggest challenge was dealing with the non-deterministic nature of LLMs and external tools. Making executions replayable while preserving enough context to reproduce behavior required careful handling of state, inputs, outputs, and side effects.

Accomplishments that we're proud of

I'm proud that Retrace turns a complex, multi-step agent execution into something developers can actually inspect, replay, and experiment with. The idea of treating an AI run almost like a Git history—where you can replay, fork, and debug—is what I'm most excited about.

What we learned

We learned that observability alone isn't enough for AI agents. Developers need the ability to interact with past executions and understand why an agent behaved a certain way. Building Retrace also taught us a lot about execution state, reproducibility, and the practical challenges of debugging LLM-based systems.

What's next for Retrace

Next, I want to make Retrace deeper and more seamless across agent frameworks, with better debugging workflows, evaluation tools, and collaboration features. The long-term goal is to make replayable execution a standard primitive for building reliable AI agents.

Built With

Share this project:

Updates

Submission history