A coding agent today can write, run, test and fix software on its own. Point the same agent at an ESP32 and the loop breaks at the first step: it can write the firmware, but it can't see it run. It doesn't know if the camera came up green, whether the frame rate held, or that the board rebooted twice while you were reading the log. So you become its eyes: you flash, you watch, you report back, and one wrong setting sends you around again. Every gain agents brought to software stops at the edge of the cable.

OpenHardware puts the board in front of the agent. A browser page, an ESP32-S3 on a USB cable, and a set of WebMCP tools that let ChatGPT's built-in browser read live telemetry, run measured experiments, build firmware on top of a baseline that never stops reporting, and flash it, with every write to hardware waiting on a human.

Why this is a strong fit for WebMCP

WebMCP tools run inside the page, on the user's machine. That one property is the reason this project can exist. A cloud MCP server can't see a USB port; a browser tab can. So the same tab that draws the telemetry for the human is the tab whose JavaScript holds the serial link, and the tools the agent calls are functions in that tab. The agent's tool calls have physical consequences, and the human is watching them land on the same screen.

It's also the shape WebMCP was designed for: one instrument, two operators. The human sees the charts, the frames, the attempt history and the gate. The agent sees structured summaries of the same data through tools. Neither is working from a description of what the other did. They're looking at the same board.

How it creates a better user experience

The firmware developer's loop today is edit, plug in, flash, squint at a serial monitor, guess, repeat. Minutes per cycle, and the agent is locked out of all of it. With OpenHardware the loop is a sentence. Ask for "the maximum resolution that holds 12 fps under 50 °C" and the agent reads the board, applies a resolution, soaks, measures frame rate and die temperature, records the attempt, and climbs until a limit breaks. Every attempt is a card in the page, every number was measured seconds ago on this board, and what the agent learns is written into a limits panel in plain language that persists for the next session.

When the agent writes firmware, it builds on top of the harness rather than replacing it, so the telemetry stream survives its own flash. If something comes up wrong, the agent sees the frame and the numbers in the same moment and debugs with live data instead of a guess. Every build is recorded with its source, so the human can code-review before anything touches the board, and a flash waits at the gate for approval.

For anyone without hardware, the page ships a simulated board with scenes (a board that runs hot, a lost link, a missing camera). It's badged amber so it can never be mistaken for the real thing, and every tool result says which one it came from.

What people and agents can do together that was difficult or impossible before

Measure instead of recall. Two boards with the same part number have different amounts of contiguous PSRAM free. The agent now reasons from what this board reported, not from training data. Iterate without becoming the agent's eyes. The human declares what's wired to which pins, states the goal, and approves. The agent runs the experiments. Debug after a flash. Because the harness is never overwritten, the agent still has telemetry and camera frames after its own firmware boots. The failure mode where the agent flashes and goes blind is gone. Review before it's real. Reads are always available to the agent. Writes to hardware are gated, and the source of every build is one click away, so trust is earned attempt by attempt rather than assumed.

How we implemented WebMCP

The page registers tools with document.modelContext.registerTool. Read tools are annotated readOnlyHint and stay registered whenever the page is open: board identity and capabilities, telemetry over a window (count, min, max, mean and slope per field, never raw samples), learned limits, a frame capture that renders in the camera panel, the available builds, the wiring list, and the current work order with its attempts. Write tools register only while a board is linked and unregister on link loss: camera on/off and config, run_experiment (apply, soak, measure, record in one call), watch_for (block on a telemetry condition, honoring the abort signal), record_limit, and flash_image, which routes through the same approve/hold gate a human uses. Approve and hold are deliberately not tools.

Every result carries its source, sim or usb, and the link state. The page infers agent presence from tool calls and narrates it next to the link badge, without ever changing the badge's color. If document.modelContext is absent, the page works exactly as before with a single hint that tools are available in a WebMCP-enabled browser.

Underneath: a static frontend with no build step, a framed line protocol (OHW1) over Web Serial with CRC and four-per-second heartbeats, an ESP-IDF harness built in Docker, and a simulated board that speaks the same protocol so the tools are tested against it in CI.

Built With

Share this project:

Updates

Submission history