-
-
The Conduit landing page: a modern, dot-grid canvas demonstrating our vision for a unified multi-agent control center.
-
The primary control center. Run multiple coding agents like Claude Code and Aider in real, persistent PTYs.
-
Meet Luna (The Keeper): Conduit's orchestrator brain. Chat with Luna to securely delegate tasks to your entire autonomous agent fleet.
-
The Safety Gate: The Strands Agent on Bedrock intercepts a destructive command, freezing the terminal until you approve.
-
Multi-agent group chat. Agents can collaborate and communicate securely within their isolated workspaces.
-
The Settings modal: Securely manage your API keys, AWS credentials, and configure local agent environments.
-
The Activity Panel: Monitor your agent fleet's real-time events, token usage, and active tasks at a glance.
-
Built-in project Wiki. Maintain your project's context, architecture, and shared knowledge for all agents to access.
-
Download center: get the native Windows desktop app for the full hardware-accelerated experience.
Inspiration
The autonomous coding revolution arrived before the tooling did.
It starts reasonably. You install Claude Code because everyone says it's good, and they're right. A week later you add Codex, because the two are good at different things. Then Gemini CLI for a second opinion on architecture. Then aider driving a cheaper hosted model, because some work is repetitive and doesn't need your best agent. None of these decisions is wrong. Each one makes you faster.
And then one evening you look up and there are six terminals on your screen, all of them producing text, and you realise you have stopped reading any of them.
You can run ten coding agents. You cannot watch ten.
The failure is quiet
We expected the problem to be dramatic — an agent doing something catastrophic, like deleting your database at 3am or force-pushing over a colleague's branch. Those happen, and they matter. But they aren't the common failure, and they aren't the expensive one.
The common failure is quiet. An agent takes a wrong turn at 11:40pm — misreads the schema, picks the wrong abstraction, or decides a file is unused — and then builds confidently on that wrong turn for fifteen minutes while you're reading a different window. It doesn't crash. It doesn't error. It produces plausible, well-formatted work on a false premise, and every minute it keeps going makes the cleanup more expensive than the work was worth.
By the time you notice, you're not reviewing a small code change. You're doing an archaeology dig.
The information was on screen the whole time. It just scrolled past in a stream no human can actually read, in one of six windows, sandwiched between four hundred lines of routine output that meant nothing. The bottleneck was never computing power. It was noticing.
The industry's answer makes it worse
The prevailing response is to make agents more autonomous. Longer leashes, fewer check-ins, bigger task scopes. Let it run for an hour and review the result later.
That trades a supervision problem for a blast-radius problem. An agent with unchecked filesystem access and a plausible-sounding plan can do a remarkable amount of damage very politely, and "review the result" only works if the result is small enough to review. An hour of confident work on a wrong premise is not.
It also misunderstands where the human is useful. Reading 4,000 lines of agent output to find the twelve that mattered is not a job a person should have. Deciding whether an agent should drop a database table is exactly a job a person should have. The tooling had these backwards: it gave humans all of the reading and none of the decisions.
The bet we made
Let the agents work. Make noticing cheap.
Keep the human deciding, and build the machinery so the only things that reach them are things that actually need a decision. Everything else — the visual redraws, the progress chatter, the files it opened and closed — gets read by something else and thrown away.
That's Conduit. It's not just an autonomy play and it's not a wrapper. It's a supervision layer with strong opinions about what a human's attention is for.
What it does
Conduit is a control centre for running several coding agents at once, with a supervision layer that reads everything they print and interrupts you only when it matters.

Six agents, in real terminals
Claude Code, Codex, Gemini CLI, OpenCode, GPT-OSS on Groq and Nemotron on OpenRouter run side by side, organised by project.
These are real terminal sessions, not just log views. You can click into any of them and type, and Conduit is simply another writer on the same terminal — the agent doesn't know the difference between a keystroke from you and one from the supervision layer. That matters more than it sounds: the moment an agent gets stuck in a way the tooling didn't anticipate, you can take over by hand instead of restarting it.
Five run as real terminals. Codex runs differently, as a background service emitting structured data, because it isn't a terminal program and pretending otherwise would mean scraping a transcript it already gives you cleanly.
Five layouts: one big pane, 2-up, 3-up, grid, or a free canvas. You'll use different layouts for different work, which is why there are five and not one.
Each agent has a status you can trust. Claude and Codex report exact states —
running, awaiting_input, idle — because they emit lifecycle events we hook into. The
other four only report whether they are alive or dead. The interface does not pretend those are the
same thing, because a status indicator that claims to know more than it does is worse than one
that admits its limits.
Luna, the Keeper

At the centre is Luna — a top-level agent with twelve MCP (Model Context Protocol) tools spanning every project, not just the one you happen to be looking at.
Ask her what's running and what's blocked and she answers from live state, not from a summary she wrote earlier. Tell her to put an agent to work and she does — through the same approval machinery everything else goes through.
Luna is the name we use; the wake word is yours to set. She answers by voice through Amazon Nova Sonic over a fast, two-way stream. We measured it at 603 milliseconds from the end of your sentence to a spoken reply with a real tool call behind it. That number is the difference between "voice control" as a slow demo and voice control as something you'd actually use while your hands are on the keyboard.
Every line gets read twice
This is the core of it.

First by fast pattern matching. Twenty-three rules run instantly on the raw text, with no slow AI model call at all. Eleven rules catch a prompt waiting for a human — like [y/N] or
(yes/no). Twelve catch a command that is expensive to undo:
rm -rf, DROP TABLE, git push --force, git reset --hard, kubectl delete.
This path exists precisely because it is instant. A destructive command must not wait on a network request. If the only thing standing between an agent and your database is a cloud model that takes two seconds to reply, that's not a safety mechanism, it's a hope.
Then by the Supervisor — an agent built on the AWS Strands Agents SDK. It reads the output in chunks to avoid spamming the system, and sorts what it sees into five buckets: progress, a question, a blocker, something risky, or noise.
Noise is thrown away and never reaches you. That deletion is the feature. A supervisor that forwards everything has just moved the reading problem somewhere else.
A gate is a decision, not a notification
When something risky appears, the agent is frozen where it stands.
Not warned about afterwards. Not logged for later review. Not surfaced as a popup you can dismiss. The process is paused at that exact point, its terminal is brought in front of you, and it waits for a human.
You have three options. Approve and it continues. Reject and Conduit safely interrupts it, then tells it to stop. Or answer the prompt yourself, in your own words, which handles every case the first two don't.
Approving and rejecting are equally one click. That's deliberate: a safety valve that is tedious to refuse gets approved by reflex, and a reflexive approval is the same as no safety gate at all.
Measured end to end on a real destructive command: an agent was asked what command deletes
the build directory, answered rm -rf build, and the gate fired about four seconds
later — on the agent's own output, not on anything we typed. Rejecting it cleanly
interrupted the terminal.
They talk to each other

Group Chat is one conversation per project. What you type reaches every running agent;
@name targets one. The Supervisor's summaries land in the same stream, labelled by what
kind of thing happened, so the machine's observations and your instructions live in one
place rather than two.

Claude agents also message each other directly, over a local MCP server.
message_agent sends a teammate something; list_teammates shows who's on the project. A
research agent can hand findings to an implementation agent without routing through you —
and it shows up in the activity feed, so "without routing through you" doesn't mean
"invisibly".

And they remember. The wiki is long-term project memory that agents write to and read back, so a new agent starts knowing what the last one decided instead of starting from nothing. This is the difference between six agents and one agent used six times.

A shared folder is how they hand actual files over — a spec one agent wrote, a fixture another needs.
Catching up on what you missed

Everything that happened while you were heads-down is one filterable screen: file changes, agents starting and stopping, messages, decisions. "What did I miss" is one place, not five separate windows.

And a usage panel shows how much of your Claude and Codex limits you've spent this session and this week, with reset times — because the other way you find out is by hitting the wall mid-task.

⌘K jumps anywhere. ⌘J opens Luna. ⌘1–5 focuses an agent. In a gate, y approves,
n rejects, Esc closes.
All of it works seamlessly on mobile screens, because agents keep running when you leave your desk.
Why this is a non-obvious use of Strands
Most "an AI watches your agents" patterns stop at summarisation. A model reads the output and writes a nicer version of it. The summary is advisory. The agent carries on regardless.
That's a reasonable thing to build, and it's also the thing everyone builds, because summarising text is what a language model obviously does best.
Conduit's Strands agent is wired so the agent it's watching cannot carry on. Two design decisions do that, and both were more work than the obvious version.
1. The classification is load-bearing
When the Strands agent classifies something as risky, the system doesn't just show a warning. It holds the agent's process at that point and blocks until a human resolves it.
The model's judgement actually controls the flow of the application, not just a notification tray. That's a meaningfully different posture. It means a mistake has a cost — a false positive stops real work — which forced us to take accuracy seriously in a way an advisory system never would. It also means the fast pattern-matching path isn't redundant: the model's judgement is allowed to be slow because the instant path has already caught the most obvious dangers.
A summary you can scroll past would not need any of this.
2. Its only write tool is a request for permission

The Strands agent has exactly one tool that can reach another agent: proposing a plan.
Calling it does not send anything. It returns "Plan submitted for human approval. Await their decision" and creates a pending record. The human sees the reasoning, the target agent, and — critically — the exact message that would be sent, character for character. Not a paraphrase of the intent. The literal text.
The single code path that delivers anything to an agent is the human approval button. Both outcomes are logged, and a rejection carries a reason the Supervisor sees on its next turn, so it doesn't re-propose what you just refused.
That inverts the usual arrangement. Normally a tool call is the action, and safety is just a prompt instruction asking the model to be careful — an instruction a model can sometimes ignore, especially under a long context. Here the tool call is a request, and the permission is enforced in the software architecture itself, where the model can't bypass it.
The agent is architecturally incapable of acting alone, rather than asked nicely not to.
We tested this at the agent's terminal rather than through the API, because the question worth answering is not "does the endpoint behave" but "does the instruction leak":
BEFORE approval: reached the agent terminal -> no
shown to the human in chat -> yes (a proposal)
AFTER approval: reached the agent terminal -> YES
That distinction caught us out, and it's worth admitting: our first test read Group Chat and reported the instruction had leaked. It hadn't. The text was in Group Chat because the proposal is displayed there, which is correct behaviour. Only measuring at the terminal answered the real question. We nearly shipped a fix for a bug that did not exist.
3. The same asymmetry, applied to voice

Rejecting a safety gate out loud has no checks at all. Stopping something is always safe, so there is nothing to guard.
Approving one requires four conditions, checked securely on the server:
- The command was read aloud to you by the system, for this exact gate
- That was under sixty seconds ago
- Your own recorded speech since then contains "approve", un-negated — a bare "yes" is not enough
- The gate is still open, with its text unchanged
Say "yes" and you hear: "Say 'approve it' out loud and I will. A yes on its own is not enough for this."
And on the simpler voice path — browser or Whisper transcription — approving isn't merely disabled. There is no approve action in that router's logic. No transcript can produce one, because the code that would represent it doesn't exist, and the unit tests assert its absence. That path sees a single sentence with no memory of what was read out; there is nothing it could meaningfully verify, so it isn't allowed to try.
Safety you cannot turn off by accident is safety that survives contact with a deadline.
How we built it

Two processes, and the split is the whole point
The daemon (background service) owns every agent process, the status engine, the watcher and Luna. The web server serves the user interface and relays messages — it owns nothing.
Restart the interface and your agents keep working. It's a property you can test in about ten seconds, and we do test it: kill the web server, restart it, and the agents are still there with their chat history intact. It falls directly out of refusing to let the UI hold a process directly, which is a rule that's easy to state and easy to break with one convenient shortcut in the code.
The two talk over a local socket that reconnects on its own. The daemon also exposes a local API, which is what Luna and the voice tools drive.
The Supervisor: one SDK, two providers
The Supervisor is a Strands Agent with specific tools to report updates and propose plans, plus four
read-only tools for reading project context. It runs on Amazon Bedrock, falling back to
Anthropic when Bedrock can't serve the request — both through the AWS Strands SDK.
That detail earned itself the hard way. Originally the fallback was a manual API call. It worked, and it was wrong: it meant the SDK left the running system precisely when the primary provider was down. And on a new AWS account — where Bedrock limits are very strict — that wasn't an edge case. It was the common case. The framework the project is built on would have been absent most of the time, and the only way to find out would have been to read the code.
So we rewired both providers through the Strands SDK, and added a health check to prove that the last successful classification actually went through Strands. "The Supervisor works" and "the Supervisor works as a Strands agent" are different claims, and a submission should be able to prove the second one in the software rather than assert it in a README.
The Anthropic path also walks a model ladder — trying the largest model first, then stepping down to smaller ones — because a subscription token is routinely refused on the larger models when busy, but allowed on the smaller ones. Hitting a rate limit on the top tier gracefully steps down rather than giving up.
The agent runtimes
A central registry is the single source of truth: for each agent, it tracks the binary that must be installed, the command to install it, and any environment variables it needs. Adding a new CLI is a single edit.
Every agent is preflight-checked before it spawns. A missing installation or unset API key fails immediately, providing the command that installs it, rather than opening a terminal that lands in a shell and looks alive while having silently printed an error. That distinction — refusing clearly versus failing invisibly — is most of what makes a multi-agent tool trustworthy.
Output is buffered so a browser attaching after an agent started gets a replay rather than a blank pane, and a browser that attaches before it runs is remembered and bound when it starts.
aider, specifically
The two hosted models run through the popular aider tool, with two flags that are deliberate and that we'd
ask a reviewer not to "fix":
--no-auto-commits, because aider commits to Git after every edit by default, which would slip changes past the approval gates and into your history.--yes-alwaysis NOT set. Setting this would auto-approve aider's own confirmations and defeat the Conduit gates entirely. Those confirmations are exactly what our safety matching is watching for.
Storage
Data is stored as plain JSON files in a local folder. No complex database, no migrations. One file per project, one for group chat, one for audit logs, plus folders for the wiki and shared files.
When something is wrong you can read the entire state of a project with a simple text editor. We think that's worth more than complex query performance at this scale, and it made every debugging session in this build dramatically shorter.
Two processes write to the project file, so we use a strict file lock to prevent corruption. Measured with two processes doing 150 writes each: 0 lost out of 300, valid JSON at every read, never observed empty.
The interface

The interface is built with React, and bundled into a native Windows desktop app using Electron. We provide a real installer and portable archive. The exact same UI serves from the same codebase in a browser, so there's one codebase and one set of behaviours.


The download page reads the files on disk and reports their real sizes. A platform with no build says "Coming Soon" rather than offering a broken link — which it did, briefly, until we noticed it was advertising filenames typed by hand that matched no build that had ever existed.

Challenges we ran into
Bedrock has four walls and they all look identical
Getting the Supervisor onto Amazon Bedrock meant discovering a complex maze of requirements:
- Old model IDs are retired. You must use specific "inference profiles", and the old IDs will just fail with obscure messages.
- The profile needs its own line in the security policy — and the underlying foundation models, in all three regions it routes to. A policy naming only your region still fails.
- Anthropic models are AWS Marketplace products. Enabling them requires a successful API call from an authorized account, which enables the model account-wide forever.
- A new account's Bedrock quota is provisioned at zero. Not low. Zero.
Every one of those surfaces as a generic "Access Denied" or rate limit error. We kept confidently attempting
the wrong fix, so we wrote a diagnostic script (npm run check:bedrock) that distinguishes them
and prints the specific next step. It was the single highest-leverage hour of the build.
Terminals are not simple text boxes
Every command-line tool behaves differently, and the differences are all in the places you don't think to look.
Claude Code's interface reads text pasted all at once as a literal newline, not an "Enter" keypress. So a message injected as one write appears in the input box and never submits. Message delivery in Conduit sends the text, waits 150ms, then sends the Enter key as a separate keystroke.
Its first-run trust prompt is an arrow-key menu whose default row is "No, exit." A freshly started agent would sit there forever waiting for a human who doesn't know to look. Conduit recognises it and automatically answers — Down, then Enter, with measured delays, because the terminal needs time to redraw between keystrokes and sending them together highlights the wrong row.
And a [y/N] scrolling past in output is not always a prompt waiting for you. Matching it
too eagerly typed characters into a terminal that wasn't asking anything. There's now a 400ms
settle window: the stream must go quiet before a prompt counts as a prompt, and then we look
only at the end of it.
A safety fix broke the feature it was protecting
We added a guard to stop messages being typed into a blocking prompt. It refused delivery to
any agent that was awaiting_input or idle.
Those are the normal states of an agent that has just finished something — which is
exactly when you send the next instruction! We measured it: an agent that had answered a question
sat at awaiting_input, and delivery failed. That's the path Luna, broadcasting, and approved plans all use.
Finishing a task made an agent unreachable. The guard now safely keys off a pending gate, which is the narrow case it was actually written for.
The packaged app can fail silently and completely
Any terminal opened from an Electron application inherits a specific environment variable (ELECTRON_RUN_AS_NODE). That
variable makes the Electron app run as a plain background script and exit — no window,
no error, no log entry. Launching Conduit from such a shell appears to do nothing at all.
This is not obscure: the editors and coding tools a person likely to run Conduit from are themselves Electron apps (like VS Code). Our own desktop test had been deleting that variable before launch — which is exactly how it passed 31/31 tests while a real manual launch did nothing!
Codex changed its protocol underneath us
Codex agents stopped starting entirely, reporting only "Failed to start". The background log had the real answer: the newest version of Codex removed a specific setting we relied on, and rejected the startup entirely. We moved to the new standard setting, which is the right value anyway — Codex asks before acting, and that question becomes a Conduit safety gate.
Accomplishments that we're proud of
- A safety valve that is built into the architecture, not just a suggestion. The Strands agent's only write tool creates a proposal. There is no code path from the model to an agent that doesn't pass through a human, and we verified it at the terminal rather than trusting a code read.
- Luna. An orchestrator that holds live context across every project, answers in under a second by voice, and dispatches work through the same approval machinery as everything else — no hidden shortcuts.
- Gates that actually stop things. In-process pattern matching, no model call in the path, firing about four seconds end to end on a real destructive command produced by a real agent.
- Refusing to let the SDK fall out of the system. The fallback runs through AWS Strands too, and our health checks prove it.
- Verification that argues with us. 292 unit and 61 end-to-end checks, plus layout, abuse, lifecycle, multi-agent, desktop and live-voice suites — including one that drives a real browser at four widths and fails on a tap target two pixels too short. It caught three regressions in this build that a human review had already passed.
- A desktop app that cleans up after itself. 31 checks against the packaged build, including one asserting that quitting leaves no orphaned background agents — because an orphan keeps spending your API budget after you believe you've quit.
What we learned
Verification has to be strict about its own results. Our multi-agent test suite reported four failures that were all mistakes of the test itself, not the software. It went from 18/4 to 26/0 without a single product change. A test that fails for its own reasons is worse than no test — you waste time and learn nothing.
Where you measure decides what you conclude. The plan-approve false positive is the clearest example: the same system, tested one layer up, told us the opposite of the truth.
Deriving beats declaring. Our icons were five copies of one massive image — so a browser asking for a tiny icon downloaded most of a megabyte. The download page advertised sizes typed by hand. Both are the same mistake: asserting a fact next to the thing that already knows it. Both now derive automatically from the source.
Multi-agent development needs two things to be viable: something that holds the context no single agent has, and human-in-the-loop safety as a load-bearing wall rather than a setting. We built both, and the second is the harder and more interesting problem — because the moment safety is a setting, it becomes a setting people turn off.
What's next for Conduit
- Bedrock quota. The Supervisor runs on Anthropic today because the AWS account's Bedrock quota is provisioned at zero. Same SDK, same tools — it moves back with no code change once quota lands.
- Reasoning about silence. Today the Supervisor only reasons about output. An agent that has gone quiet is caught by a status timer, not a judgement. Making absence something the agent reasons about — a stalled agent, a loop producing nothing, a task that should have finished — is the next genuinely useful step, and we're not claiming it yet.
- Cloud workspaces. Moving the background daemon onto temporary cloud instances, so you manage agents on disposable cloud machines from a local UI.
- Granular safety policy. Letting you define your own patterns and safety rules for the Supervisor rather than shipping only ours.
- A real multi-user story. Today it assumes one human: one shared login, and an unauthenticated local API. Both are documented, and both are the wrong trade-off the moment a second person is involved.
- macOS and Linux builds. The Windows installer and portable archive are real; the others are configured but unbuilt, and the download page says so rather than offering a dead link.
Built With
- amazon-bedrock
- amazon-web-services
- docker
- electron
- express.js
- node.js
- react
- strands-agents-sdk
- typescript
- vite
- websockets
Log in or sign up for Devpost to join the conversation.