Inspiration

AI coding Agents can complete impressive work, but long jobs become hard to trust when several workers share one workspace. A browser can disconnect, a Runtime can die, a tool call can be unsafe, and an Agent can confidently claim success without proving the result. Track 1 asked for one meaningful middleware capability behind the supplied Agent platform. We built the dependable operating layer we wanted between a person's request and the Codex workers carrying it out.

What it does

A user describes one ordinary job in the existing Agent Launchpad conversation. The middleware then:

  • chooses the smallest justified team of 1–8 fresh, task-specific Agents;
  • assigns evidence-bearing checkpoints and serializes accepted shared-file writes;
  • delivers work through a durable NATS JetStream mailbox and ledger;
  • gives every coding Run an isolated transactional workspace;
  • promotes only accepted checkpoints and retries only unfinished work;
  • records real Runtime, policy, handoff, and recovery events in Glassbox;
  • stops an exact Codex process through Kill Switch while rejecting unfinished output;
  • blocks documented direct deletion actions through Bouncer before execution;
  • links observable events in an append-only SHA-256 Agent Flight Recorder; and
  • independently opens accepted web results at mobile and desktop widths through Proof Gate, recording browser errors, overflow, screenshots, and hashes.

The result is not another AI worker. It is middleware that keeps AI workers coordinated, inspectable, protected, and recoverable.

Why this is middleware

The important behavior runs behind the UI, between the Fastify request boundary, AgentService, JetStream coordinator, and Codex Runtime. The browser displays durable evidence; it does not invent orchestration or failure receipts. The organizer's Create Agent, lifecycle, workspace, Playground, and Codex paths remain intact.

Demonstration

We asked Launchpad to extend an existing TikTok creator tool with an A/B Hook Arena. Launchpad created three fresh roles: a clarifier, an implementer, and an independent reviewer. During implementation, we stopped the real active Codex process. Completed analysis remained accepted, unfinished changes were rejected, and only the remaining checkpoint moved to a fresh Runtime.

The first implementation passed ten checks, but the independent reviewer found that winner-storage failure could break the unrelated favourite path. Launchpad recorded an honest FAIL and created a smaller two-Agent repair team. During repair, Bouncer denied an attempted delete-and-recreate operation on a protected file; the Agent continued with a safe edit. The repair passed thirteen checks and independent review. Proof Gate then loaded the real creator tool, where Hook B could be selected, saved, and retained after reload.

How we built it

We preserved the MIT-licensed organizer starter at commit 8d0bd4f14ad1e453d984149aebcdd0bcb4f74178. The control plane is Node.js and TypeScript with Fastify. The browser uses React. NATS Server/JetStream 2.14.5 supplies file-backed streams and KV state. Application-level accepted-turn records, compare-and-set transitions, bounded retry, and content digests convert at-least-once delivery into one accepted business result per turn. Codex CLI performs the real Agent work. Playwright-based trusted-host verification supplies independent browser receipts.

Challenges we ran into

The hardest part was separating durable claims from durable evidence. JetStream is at-least-once, not exactly-once, so we added idempotency and accepted-turn checks instead of overclaiming transport guarantees. Recovery also had to reject partial workspace mutations, not merely retry a prompt. Finally, we kept worker self-reports separate from control-plane proof so a worker cannot approve its own visual result.

Accomplishments that we're proud of

  • Real multi-Agent planning and execution from one plain-language request.
  • Checkpoint-level recovery after active Runtime cancellation and coordinator restart.
  • Transactional promotion that prevents unfinished Agent work from reaching the shared workspace.
  • One causal Glassbox joining coordination, policy, Runtime, recovery, and proof evidence.
  • Live Bouncer denial followed by successful safe work.
  • Independent Proof Gate receipts for 375×812 and 1440×900 browser loads.
  • A reproducible verification path covering builds, server tests, live JetStream restart, cancellation, disconnect/reconnect, protocol reuse, and SSE recovery.

What we learned

Durability is only valuable when a person can understand and control it. The strongest middleware is not the one with the most orchestration features; it is the one that makes every handoff, accepted checkpoint, failure, retry, and proof legible without hiding the real Agent work.

Honest boundaries

The reproduced topology uses one file-backed NATS node. It survives process restart, but does not claim cluster, disk-loss, or machine-loss recovery. Recovery is checkpoint-level rather than token-level continuation. Bouncer enforces documented direct deletion patterns and is not a complete sandbox. The successful local real-model path uses this machine's authorized Codex login; portable deployment still requires an authorized provider configuration.

What's next

We would extend the same contracts to replicated JetStream, portable scoped model credentials, multi-instance SSE fan-out, and additional narrowly defined Runtime policies while keeping Glassbox and Proof Gate as the evidence boundary.

Built With

Share this project:

Updates

Submission history