Inspiration
Every Brainbase agent is hand-written today: instructions, tools, skills, playbooks, evals, handoffs. But the best spec for an agent already exists. It's the coding session where you actually solved the problem, in Claude Code or Codex. We were throwing that away. So we asked: what if you could record the work once and ship it as a team of agents?
What it does
symbiote → Brainbase turns a recorded coding session into a live, graded Brainbase agent team.
- Record: work normally in Claude Code or Codex through symbiote (our session recorder). Any harness, any model, switch mid-task with
symbiote continue --with codex. - Compile in one click:
- Each harness segment becomes a Brainbase agent, and each switch becomes a handoff.
- The repo's
CLAUDE.md/AGENTS.mdbecome project rules, and the skills you used become skills or playbooks. - What the session actually did (commands, files, checks and their results) becomes a "How this was done" playbook.
- Jev's grading rubric becomes native Brainbase evals.
- Real tools: every agent gets the MCP servers we set up once in symbiote (AgentMail, Bland, Linear), with permission rules and a live log of every tool call.
- Run it, scale it, schedule it:
- Chain runs with Jev checking every handoff.
- Swarms: N teams in parallel, Jev ranks them and features the winner.
- Cron schedules, and public agent pages anyone can use.
- Artifacts, like Claude Artifacts, which deploy to Cloudflare in one command.
In our demo, a startup's website is built by an 8-team swarm, monitored every morning by an agent that opens it in a browser, emails us, calls us, and files Linear tickets, and fixed by a "bug fixer" team compiled from a session we recorded live.
How we built it
- App: Next.js 16 + React 19 + TypeScript on Bun, deployed on Vercel. Sessions come from symbiote's Supabase, and the app keeps its own state in Supabase Storage.
- The compiler (
src/lib/pipeline.ts) builds a deploy plan from the transcript: segments, handoff edges, redacted instructions, playbooks, evals and MCP entries. The CLI readsCLAUDE.mdand skills on the laptop, because the server can't see your repo. - Brainbase API, used end to end: agents, evals, playbooks, skills, orchestrations and schedule triggers, threads and events (SSE), the files API, and
mcp_servers. - Jev (TypeSafe) at every decision point: the handoff gate, the risk check, the model router, stuck-session watch, a judge that reconciles Jev with Brainbase evals, and swarm ranking. A receipts log records every call with its probabilities and cost.
- CLI (
hack ship | swarm | schedule | site ...) and Claude Code commands (/ship,/spawn), so any harness can ship an agent.
Challenges we ran into
- Brainbase doesn't hand off automatically. An orchestration handoff is a tool the agent must choose to call. So we built a chain runner: the app starts each next step itself, carries files between fresh sandboxes, and asks Jev whether the step is really done.
- Playbook writes fail once an agent has evals. We found this live and reordered the deploy so playbooks go first.
- Codex on Brainbase only accepts OpenAI models, so we map models per harness.
- CLAUDE.md and skills live on the laptop, not in the recording, so the CLI gathers a capped, redacted context pack and uploads it.
- Agent tools with real side effects (email, phone, tickets) need guardrails. We added per-server permissions, a scoped AgentMail inbox, a phone number only the owner can set, and masked logs.
Accomplishments that we're proud of
- An agent we never hand-wrote, running on a schedule: it opened our live site in Brainbase's browser, emailed the report, called us with a spoken summary, and filed 3 Linear tickets.
- An 8-team swarm that finished 8 of 8, ranked by Jev in one call (winner p=0.77 on Cloudflare** in one command.
- One-click compile of a real Claude Code → Codex session into a team with project rules, playbooks, evals and tools.
What we learned
- Recorded evidence beats self-report. Compiling from what actually happened (commans far better agents than asking a model to describe what it did.
- Calibrated judgment is what makes autonomy safe. Jev's probabilities decide when to continue and when to ask a human.
What's next
- Read the MCP list live from the symbiote account: connect any MCP once with OAuth, and for each agent.
- Re-shipping a newer session updates the same agent, so it keeps improving.
- Paid public agents, with Jev keeping each run's cost under its price.
- Every run's artifact auto-published to Cloudflare.
Honest notes
- Timeline: we started the evening before, with the organizers' OK. symbiote, our ree event and is a dependency. The ship, swarms, playbooks, MCP, schedules, artifacts,Cloudflare and the dashboard were built at the event.
- Payments: Brainbase Payments is provisioned but no purchase was made.
- Permissions: they're enforced through each agent's instructions.
Built With
- agentmail
- anthropic
- bland-ai
- brainbase
- bun
- claude
- claude-code
- cloudflare-workers
- codex
- jev
- linear
- mcp
- nextjs
- openai
- react
- stripe
- supabase
- typesafe
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.