Inspiration
Every "AI for observability" demo we'd seen put the agent in a sidebar. You type a question, it answers in a chat bubble, and the dashboard sits there untouched — two separate surfaces that happen to be about the same outage.
That is not how two engineers debug an incident. They look at the same screen. One drags the time window, the other says "there, that span," and neither narrates what they just clicked.
WebMCP makes that possible for the first time: the page can hand an agent real tools instead of asking it to read pixels. So we built the thing the sidebar demo is a substitute for.
What it does
Sightline is a single-screen incident console. A production incident is in progress: checkout latency is spiking. A human on-call engineer and an external AI agent — ChatGPT, or Chrome with WebMCP enabled — investigate it at the same time, on the same screen.
There is no chat panel in the app and we call no model API. The page registers nine WebMCP tools and renders state; the agent is whatever host you open it in.
The agent queries metrics, pulls traces, searches logs, correlates a deploy, pins findings with confidence and source references, proposes a rollback the human must approve, and drafts a postmortem from what was pinned. Every call visibly changes the screen, every pane says who last touched it, and every call can be opened to read the exact arguments and the exact JSON that came back.
The handoff. Drag the chart handles to 14:15–14:30, click a trace, then ask the agent what you're looking at. get_current_view reads your window, your open trace and your filters, and it continues your investigation instead of restarting it. The header stamps the second it happened.
The incident. checkout-service p99 jumps from 181ms to 3.5s at 14:20 while p50 moves 1.25× — the tell that requests are queueing, not computing. 96.2% of trace time sits in db.connection.acquire. The pool logs 18 acquisition timeouts. dep-1104 shipped eight minutes before onset and its diff reads hikari.maximumPoolSize 50 → 10. Two red herrings sit in the data — an unrelated cache deploy, and a deploy that lands 35 minutes after onset — plus a downstream service that looks sick but is only a victim.
How we built it
Vite, React, TypeScript, Tailwind, Recharts, Zustand. No backend, no database, no API keys.
One store, no parallel code path. Tools are not a second implementation. filter_traces calls setTraceFilter; so does the input in the trace pane header. Every action carries an actor, which is how the interface can attribute what happened. Anything the agent can do, a human can do, and each sees the other's work immediately.
Nothing is mocked. Every value a tool returns is computed from the fixture series by real code — median baselines, a change-point detector, span aggregation, log pattern grouping, deploy proximity scoring. If query_metrics reports a 3498.6ms peak at 14:49, that is Math.max over the series. The fixture is generated deterministically — no Math.random, no Date.now, hash-based noise — so the same minute always produces the same value and the demo is reproducible.
Errors teach the caller how to retry. Not "No results" but:
No traces for
checkout-serviceat or above 3000ms in 13:30-14:00. The slowest trace in that window is 197.9ms — lowermin_latency_msto 158 or below to see it.
The approval gate is structural. propose_rollback writes a proposal and returns awaiting_human_approval. The only function that applies a rollback is store.approveRollback, and its only caller is the Approve button's onClick. A test runs every registered tool with arguments plausible enough to do damage while a proposal is pending and asserts the rollback stays unapplied — and checks its own tool list against the registry so a new tool can't skip it.
The interface has one signature idea and one motion. Every pane header carries a provenance stamp naming who last changed it and how. When a tool mutates a pane, that pane gets one sharp flash in the colour of whoever caused it. Nothing else animates.
Challenges we ran into
A rejected tool promise throws away your error message. The WebMCP spec discards the rejection reason and reports a bare failure to the agent. We had written careful retry hints into every error path — all of which would have vanished. Tools never throw; they resolve with isError and the hint survives.
Origin-Agent-Cluster: ?1 or nothing works. registerTool() rejects with a SecurityError unless the document is origin-keyed, and it fails quietly enough that you'd blame your own code. We send the header in dev, preview and production, and the app surfaces window.originAgentCluster in its diagnostic so the failure names itself.
The IDL and the implementation disagree. The spec declares executeTool(tool, object inputObject), but Chrome wants the arguments as a JSON string and answers "Failed to parse input arguments" for an object — and it resolves to a string, not the ToolResult. We only found both by calling the API against a live browser.
The hardest problem was not code. From the single prompt "Checkout p99 is spiking. Find out why.", the agent reached the root cause unaided on the first attempt — the deploy, the pool size, the p50-versus-p99 reasoning, both red herrings dismissed. And then it wrote the answer in the chat window and pinned nothing.
It had treated the task as a question to answer. Nothing in our tool surface told it that the on-call engineer is looking at the console, not at the conversation — so a conclusion stated in a reply reaches no one. The fix was entirely in tool descriptions and in what each read tool returns: every one now ends with a next step computed from what that call actually found, and the tool the agent is holding when the answer lands spells out that describing a fix in conversation puts nothing on anyone's screen.
That iteration, not the app, was the real work.
Accomplishments that we're proud of
- An agent solves a deliberately noisy incident from one sentence, with no hints, and rejects both red herrings on the way.
- The safety gate is structural, not decorative — and there is a test that proves no tool can reach it.
get_current_viewmakes the handoff real in both directions: the agent continues from a window you dragged and a trace you opened.- The protocol is inspectable. Any call opens to show the exact arguments and the exact JSON reply.
- The approval gate shows pre-flight checks computed from the deploy — restore target, later deploys, migrations in the diff — rather than reassuring constants. An approval gate that displays a check it never ran is worse than one that shows nothing.
What we learned
Tool descriptions are the product. The schemas took an afternoon; getting an agent to behave like a colleague rather than a search engine took the rest of the day, and every fix was a sentence, not a function.
Agents infer the shape of the job from the tools you hand them. Ours had all the information it needed and still stopped at an answer, because nothing said the work included recording it. Once the tools said so, it did.
And a detector and an alert threshold answer different questions. A service can be alerting with no change point in its series — when two of your tools disagree, the fix is to say why, not to move the threshold until they agree.
What's next for Sightline
Real telemetry backends behind the same tool contracts — the analysis layer already takes series and returns summaries, so the fixture is the only thing to replace.
Multi-incident state, so get_current_view can hand over an entire war room rather than one board.
And a second operator: two humans and an agent on the same console, with provenance stamps that already know how to name three parties instead of two.
Built With
- netlify
- react
- recharts
- tailwindcss
- typescript
- vite
- vitest
- webmcp
- zustand
Log in or sign up for Devpost to join the conversation.