Inspiration

AI agents can retry conversations, but retrying external actions can create duplicate bookings, payments, or orders. We wanted to make these actions safe even when a worker crashes at the worst possible moment.

What it does

Crash-Safe Booking uses an Action Ledger to record intent before execution. Stable idempotency keys and provider reconciliation ensure retries recover the original booking instead of creating duplicates.

How we built it

We extended the provided Agent Launchpad while preserving its Agent lifecycle, Playground, persistent sessions, ModelArk execution, and disposable container Runtime.

Our middleware adds:

  • A durable Action Ledger in the Fastify control plane
  • Short-lived, Run-scoped tool capabilities
  • A transactional booking tool inside each Agent workspace
  • A mock provider with idempotency and result lookup
  • Controlled crash injection after provider acknowledgement
  • Automatic, startup, and manual reconciliation
  • A React evidence panel showing correlated action events
  • Docker-based disposable Agent execution
  • BytePlus ModelArk through the Responses API

Challenges we ran into

The hardest scenario was an ambiguous outcome: the provider accepted a booking, but the worker failed before the control plane recorded success. We had to recover that result without repeating the provider action.

We also needed to distinguish a legitimate retry from a conflicting request. The middleware binds each operation ID to the original input hash and returns 409 Conflict if the same ID is reused with changed booking details.

Accomplishments that we're proud of

Our demo proves that repeated requests and crash recovery still produce only one provider booking. The ledger clearly shows every execution, failure, reconciliation, and durable result.

Our working demo proves that:

  • Repeated identical requests return the same durable result
  • A controlled crash is recovered without creating a duplicate booking
  • Changed inputs under an existing operation ID are rejected
  • The Agent remains usable and controllable after recovery
  • The complete middleware decision trail is visible in the browser
  • All 18 automated tests and the production build pass

For the implemented booking adapter, the system guarantees at-most-one provider booking because the provider supports stable idempotency keys and result lookup.

What we learned

Reliable agents need more than conversational memory. Real-world actions require durable intent, controlled authority, idempotency, reconciliation, and clear operational evidence.

What's next for Crash-Safe Booking

Next, we want to support production booking APIs, payments, orders, multiple transactional tools, stronger isolation, observability, and configurable recovery policies.

Built With

Share this project:

Updates

Submission history