RunProof: Every Agent Run Leaves a Receipt

Inspiration

AI Agents can perform increasingly complex work, but their final answers reveal very little about what happened during execution. When an Agent becomes slow, fails, or is cancelled, developers are often left reconstructing its behavior from scattered console output.

Simply recording everything is not a safe solution. Prompts, responses, thread identifiers, credentials, and error messages may contain sensitive information.

We created RunProof to solve this gap: lightweight observability middleware that produces persistent, correlated, and privacy-safe evidence for every AI Agent run.

What it does

RunProof turns each Agent execution into a durable receipt.

For every run, it creates:

  • An agent.run orchestration span representing the complete Agent lifecycle
  • A codex.run child span representing the underlying model execution
  • A shared trace ID connecting the spans
  • Status and timing information
  • Completed, cancelled, and failed outcomes
  • Token usage and safe event counts
  • Sanitized error evidence when something goes wrong

The Trace panel allows operators to select a run and inspect its status, duration, span hierarchy, token usage, and identifiers without leaving the Agent Playground.

The evidence also survives beyond the interface. Sanitized traces are stored alongside Agent and Run data in the persistent launchpad.json store. Relationships are maintained through the run ID, trace ID, and parent span ID, allowing a run to be investigated after refreshing the page or restarting the application.

How we built it

RunProof is implemented as middleware inside our existing Agent platform rather than as a separate demonstration or replacement runtime.

The execution flow is:

React Playground
      ↓
Fastify API
      ↓
Agent Service
      ↓
Codex Runtime + BytePlus ModelArk

A dedicated Trace Service listens to lifecycle and runner events produced along this path. It builds correlated orchestration and model spans, sanitizes their metadata, and saves the resulting evidence through the platform’s existing persistent data store.

The frontend is built with React and TypeScript. It provides a Trace panel where users can compare recent completed, cancelled, and failed runs.

The backend uses Fastify and TypeScript. Agent executions run through an isolated container runtime, while BytePlus ModelArk provides the model endpoint. We tested the full workflow locally using Podman.

We designed tracing as a best-effort operation. If writing a trace ever fails, the Agent result remains unaffected. Observability must explain the system—it must never become a new reason for the system to fail.

Privacy and safety

RunProof records execution structure rather than sensitive content.

Before any trace is saved, its metadata and errors are sanitized. The data is sanitized again before being returned by the API, creating protection at both the persistence and presentation boundaries.

By default, RunProof does not store:

  • User prompts
  • Model responses
  • Conversation or thread identifiers
  • API keys or credentials
  • Raw environment variables

This gives teams useful operational evidence without turning the tracing system into a repository of private data.

Challenges we faced

One major challenge was representing every terminal outcome consistently. A completed run follows a predictable path, but cancellation and failure can occur between lifecycle events. We had to ensure that both the orchestration and model spans always reached the correct final state.

Another challenge was persistence. A trace is most valuable after something has gone wrong, so it cannot disappear when the live process ends. We integrated traces into the existing atomic JSON store and verified that they remain queryable after an application restart.

Privacy required careful design as well. Error objects and metadata can contain unexpected information, so redaction could not be limited to the user interface. We added sanitization before storage and again at the API boundary.

We also learned that token usage is more nuanced than a single number. Input usage can include conversation context and cached tokens, while output usage may include hidden reasoning. Presenting these values accurately requires distinguishing reported runtime usage from the visible length of a prompt or response.

Finally, integrating the containerized Codex runtime with ModelArk required us to manage environment configuration, persistent workspace mounts, cancellation, and recovery without exposing credentials.

What we learned

We learned that Agent observability is not just about collecting more logs. The difficult part is deciding what evidence is useful, how events relate to one another, and what information must never be stored.

Correlation is essential. A trace becomes meaningful when the Agent lifecycle, model execution, status, timing, and usage can all be connected to the same run.

We also learned that cancellation is a first-class outcome rather than an unusual error. A trustworthy Agent platform should explain not only what completed, but also what stopped and why.

Most importantly, we learned that observability should be designed as a trust boundary. It must provide enough evidence to understand the system while preserving user privacy and leaving the original execution behavior unchanged.

Accomplishments

We are proud that RunProof is integrated into a working Agent platform rather than presented as a mock-up.

The completed project includes:

  • Real BytePlus ModelArk execution
  • Correlated orchestration and model spans
  • Completed, cancelled, and failed run evidence
  • Persistent traces that survive restarts
  • Privacy-aware metadata and error sanitization
  • A working trace visualization interface
  • Containerized Agent execution
  • Type checking and production builds
  • 95 automated tests

What’s next

Our next step is to make RunProof useful across larger Agent deployments.

Planned improvements include:

  • Separating cached, uncached, visible-output, and reasoning-token usage
  • Cost estimation across different model providers
  • OpenTelemetry-compatible trace export
  • Search, filtering, and comparison across runs
  • Alerts for failures, unusual latency, and token spikes
  • Configurable retention and redaction policies
  • Additional persistent storage backends
  • Team dashboards and audit exports

Our goal is simple: whenever an AI Agent acts, RunProof should leave behind a safe and durable explanation of what happened.

Built With

Share this project:

Updates

Submission history