Inspiration

We run a fleet of nearly fifty small business sites, and the operations work never ends: something is always broken, slow, or quietly losing orders. We could not hire an ops team, so we built one out of AI agents. Then we hit the question every agent builder hits: how much do you let it do on its own? Full autonomy felt reckless. Approval on every keystroke felt useless. The answer we landed on runs our private ops room today, and Fleet Command is that answer made public: agents do all of the work, and a hard human gate holds every consequential action.

What it does

Fleet Command is a one screen command center. Four specialist agents work a target site in sequence: SCOUT does reconnaissance on both the site and the market around it, AUDIT ranks the defects, MEDIC drafts the actual fix as a diff, and SHIP checks the deploy target's own domain and stages the deployment with a risk summary. Then SHIP stops. It holds at the approval gate until a human clicks Approve or Send back. Send back returns the fix to MEDIC with the objection attached. Approve releases the deploy. The hold is structural: SHIP has no code path that releases without the click.

The demo target is a fictional bakery site we broke on purpose (dead order form, no security headers, a six second load) so judges can watch the full loop without any real site being touched.

How we built it

Cloudflare Pages hosts the cockpit; a Pages Function (POST /run) runs each agent step against the Anthropic API, each agent with its own system prompt and the prior crew output as context. Two agents reach outside: SCOUT queries SerpApi for live Google results about the market the target competes in, and SHIP queries the name.com registrar about the target's own domain before it stages. Both are cached in Workers KV. Zero dependencies, no build step, no framework: hand rolled HTML, CSS, and JS on the front, one Function on the back, security headers on every response via middleware.

Three honest modes on three independent switches. With no model key set, /run serves a recorded run and the UI labels it "replay". With no SerpApi key, SCOUT gets clearly marked sample data and a chip reads "serpapi sample". With no name.com credential, SHIP's registrar check is labelled sample in the console the same way. Arm any key and that half goes live on the same code path. The app never pretends a canned answer is live; we think honesty is a feature judges can verify, so the mode is always on screen.

Challenges we ran into

Designing a gate that is real rather than theater: the hold lives in the mission flow itself, not in a dismissable dialog. Keeping the demo honest while still working with no key armed took a deliberate replay design. And a fun one: our first typewriter effect froze in throttled background tabs, so agent output now renders on wall clock time and a mission finishes even in a hidden tab.

Accomplishments that we're proud of

The whole thing is one screen with zero scroll, live on the edge, with a crew you can watch think. The pattern is production proven: the same engine runs our private ops room across dozens of live sites.

What we learned

Multi agent systems get useful exactly when you make their autonomy legible: who is working, what they produced, and where the human sits in the loop. The gate turned out to be the feature people trust, not the feature that slows them down.

What's next for Fleet Command

Real targets behind the same gate: connect a user's own site, let the crew run on a schedule, and grow the approval gate into a queue with an audit trail of every decision. The engine already does this privately; Fleet Command is the version everyone gets.

Built With

  • anthropic
  • claude
  • cloudflare-pages
  • cloudflare-workers
  • cloudflare-workers-kv
  • javascript
  • name.com
  • serpapi
Share this project:

Updates

Submission history