Inspiration
We gave an agent harness real credentials and watched it work. It was impressive and unnerving in equal measure: every step went to the most expensive model, it could see every tool we owned, our data went wherever it decided, and the first time we learned it was about to send an email was after it had sent it. Agents today are a black box holding your keys. We wanted the opposite: an agent you can supervise — and that gets cheaper and more private because of it, not in spite of it.
What it does
Zephyr sits between any agent harness and the world. You type a goal; Zephyr runs it as a supervised task:
- Routes every step. A tiny decision model (Jev) classifies each step on two axes — private vs cloud, low vs high intelligence — and sends it to a local model, a cheap cloud model, or a frontier model. Most steps don't need the frontier.
- Narrows the tools. From a registry of browser tools, 1000+ Composio integrations and any MCP server you add, the agent is handed only the few tools this step needs. It cannot call what it was never shown.
- Guards what leaves. Sensitive spans are detected and replaced with
[[PII_n]]placeholders before any cloud call. Every outbound call lands in an egress ledger: destination, class of data, and the rule that allowed it. - Stops before the irreversible. Actions are classified by reversibility. Sending an email or submitting a form pauses the run and shows you the exact action. Approve, reject, or edit it — an edit is re-authorized and can only narrow the action. Text written in your name is scored by GPTZero first; a high AI-probability can add a human check, never waive one.
- Knows when it's done — or stuck. Jev judges completion from canonical session state.
blockedpauses for a human instead of burning turns. - Shows its work. A live, replayable trace of every decision, model call, tool call and approval, plus a side-by-side comparison against a single-LLM-call baseline (tokens, cost, latency). You can pause, resume and cancel any run.
It drives real browsers (Browserbase in the cloud, local Chrome for private data), and runs parallel browser workers as a swarm.
How we built it
A TypeScript pnpm monorepo: an Express 5 API, a React + Vite dashboard, and a shared package of zod schemas that both sides import, so the event stream is one contract.
- Harness-agnostic core. The harness (Hermes, from Nous Research) runs as a subprocess
over the Agent Client Protocol. It is pointed at our OpenAI-compatible model gateway
(
/v1) and our MCP server (/mcp), so every model call and every tool call passes through Zephyr's policy whether the harness likes it or not. - Jev as System One. Jev (TypeSafe AI, via the Vercel AI Gateway) cannot generate text. We hand it state plus typed questions and get back a choice with calibrated confidence. It picks routes, tool families, tools, browser targets (by index into an element table built from the accessibility tree) and completion. Below a confidence threshold, or when its time budget runs out (2 s for an interactive browser step), a deterministic rule decides instead — and the trace says which one did.
- An exact-action broker. Tools enter a reviewed registry with a reversibility class. The broker authorizes the exact proposed action — arguments, destination, data labels — not the tool in general. Unknown tools from a newly added MCP server are discovered but stay unavailable until classified: fail-closed.
- Providers behind capabilities. Playbooks ask for
browserortext.model, never a vendor. Every provider has a mock twin, and a missing key downgrades live → mock instead of crashing, so the full demo runs from a fresh clone with no keys at all. - Streaming and persistence. Runs stream over SSE with monotonic ids and
Last-Event-IDreplay; events persist to SQLite so a finished run rebuilds from history. - Invariants as tests. Eleven
check:*suites pin the safety properties (secret egress blocked, revised approvals re-authorized, fail-closed tools, outbound-text check is escalate-only), run in CI with the typecheck and web tests.
Challenges we ran into
- Designing for a model that can't talk. Our first instinct was to prompt Jev. It takes criteria, not prompts, and returns an index, not prose. Rebuilding browser control around "here is a numbered table of elements — which one?" was the unlock, and it made the system safer: a model that can't write text can't be talked into inventing an action.
- Supervising a harness we don't control. Hermes has its own loop and its own ideas. Putting it behind our model and MCP gateways — and correlating its calls back to the right run — is what made policy enforceable rather than advisory.
- Four people, one event contract. Backend and frontend drifted mid-weekend: the UI was still calling graph-synthesis endpoints the backend had replaced with the agent playbook. Shared zod types caught the shape errors; an end-to-end audit caught the missing routes.
- "Safe" is a separate question from "authorized". Jev being unable to generate an action says nothing about whether the action it picked is allowed. Every path still goes through deterministic authorization, and we wrote tests so nobody can shortcut that.
- Honest mocks. Making the keyless demo complete without ever labelling a mock as live took real discipline in the UI and the provider layer.
Accomplishments that we're proud of
- A run that pauses on an irreversible action, shows the exact payload, and resumes on approval — against real services, from a plain sentence.
- An egress ledger that turns "your private data never reached the cloud browser" from a claim into a row you can point at.
- A working comparison that shows what supervision costs and saves versus one big LLM call.
- The whole thing boots and demos from
pnpm install && pnpm devwith zero keys.
What we learned
- Small, calibrated, non-generative models are a better fit for control decisions than LLMs: faster, cheaper, thresholdable, and not injectable.
- Classify by reversibility, not by vague "risk". It is the question a human actually needs answered before an agent acts.
- Make the safe path the only path: gateways the harness can't route around beat guidelines the harness is asked to follow.
- Build the mock twin first. It is demo insurance and it keeps four people unblocked.
What's next for Zephyr
- More harness adapters (the adapter boundary is already there) so the same policy layer supervises any agent.
- Editing a node of a running workflow, with the same re-authorization rules as revised approvals.
- Per-user auth and policy packs — budgets, approval thresholds, data residency — for teams.
- Caching resolved browser targets so repeat runs need no model call at all.
- Using the ledger and approval history to learn which steps never needed the frontier.
Built With
- agent-client-protocol
- browserbase
- claude
- composio
- express.js
- gemini
- github-actions
- gptzero
- hermes
- jev
- model-context-protocol
- playwright
- pnpm
- react
- react-flow
- server-sent-events
- sqlite
- stagehand
- tailwindcss
- typesafe-ai
- typescript
- vercel-ai-gateway
- vercel-ai-sdk
- vite
- zod


Log in or sign up for Devpost to join the conversation.