Inspiration

Every always-on microphone has a second set of people in front of it: the ones who never agreed to anything. The colleague at the next desk. The diner at the next table. The family member in the background. Bee is a good product for the person wearing it, and the developer surface is built entirely around that person. Nobody else in the room has a state in the system at all.

I went looking for the API that would let me fix that, and found that it does not exist. Bee's transcript endpoint returns acoustic cluster labels — SPEAKER_0, SPEAKER_1, sometimes an empty string — and nothing else. No participant identity. No voiceprint enrolment. No flag marking which cluster is the person wearing the device. That is checkable in the official @beeai/cli v0.7.3 source, and I quoted it with file and line in the research log rather than paraphrasing it.

So consent cannot be looked up. It has to be maintained. That is what Bystander is.

What it does

Bystander sits between Bee's capture pipeline and anything downstream that wants to read it — a summariser, an agent, a vector store.

  1. A consent ledger per speaker cluster. Each acoustic cluster maps to an enrolled participant with a role and a consent status.
  2. UNKNOWN is never treated as consent. It behaves as refusal. This is the one invariant the whole project rests on, and a test fails if isConsented() ever returns true for it.
  3. Whole-speaker suppression, verified by absence. An unconsented speaker's turn is not trimmed or paraphrased — it is replaced, and the tests prove the original text appears nowhere in the returned payload.
  4. Boundary-aware PII scrubbing for consented speech, because a consenting speaker can still name a third party who did not consent. Emails, phone numbers and named entities go, with word boundaries so Ann does not match inside annual and Bob does not match inside bobcat.
  5. A loud refusal, not a quiet guess. If a conversation contains unconsented speech the summariser declines and says why, with a machine-readable code: REFUSAL_UNCONSENTED_PARTICIPANTS, REFUSAL_CONSENT_REVOKED, REFUSAL_AMBIGUOUS_ATTRIBUTION.
  6. An audit trail a wearer can act on — category, cluster, span offsets, sizes and the replacement token for every removal.
  7. All of it over MCP Streamable HTTP at spec version 2025-11-25, so an assistant consumes governed context instead of raw ambient audio.

It is an application, not a diagram

The consent layer above is the argument; this is the product it ships as. Four views, hash-routed:

  • Capture — the pipeline: raw turns in, redactions in place, the gate's verdict, and the audit trail for the selected conversation.
  • Consent Ledger — enrol a speaker cluster with a name, role and consent status; toggle or un-enrol one and watch the capture view's gate flip to REFUSED. Every change writes an immutable event row.
  • Audit & Compliance — the stored records read back out of SQLite, with JSON and RFC 4180 CSV export, and a purge behind a confirmation dialog.
  • Settings & Transparency — retention window, the live Bee connection status and endpoint, the MCP endpoint and protocol floor, and the storage invariants read back with PRAGMA rather than asserted in the page.

The ledger is SQLite, so consent survives a restart — the event log on the ledger view shows the seed and every sweep since. Purge matters more here than it looks: Bee's conversations endpoint is read-only, so a revocation can never delete anything upstream, which makes a local purge the only place the word "forget" can mean anything at all.

63 tests across 13 suites. node ops/mcp-client-demo.mjs opens a real Streamable HTTP session against the running server, negotiates 2025-11-25, lists the four tools and prints a live refusal with its reason code — that script's stdout is the artefact, not a screenshot of it.

The part I would rather not have had to write

The first working version of the redactor put every removed span into the audit record as originalText. That audit is returned to callers. So suppressed bystander speech travelled straight back out inside the record that claimed to have removed it — and the entity reason string leaked a second copy.

The absence test passed the entire time. It serialised only the two parts of the result that were never at risk.

That is the interesting failure, and it is worth more to a reader than a clean story: an absence assertion is worth exactly as much as the surface it serialises. A removal record now carries a character count and a word count and nothing else. Both absence tests serialise the whole returnable payload. A new test asserts that no record has an originalText field, and that no record contains any four-word run of the source speech. Both new guards were confirmed by reinstating the leak and watching them fail, then reverting. The demo video was re-recorded against the fixed service, because the old cut had the leak visible on screen in the audit table.

How I built it

TypeScript throughout. Express for the REST surface, the official MCP TypeScript SDK 1.30.0 for Streamable HTTP, Vite for the dashboard, node:test via tsx for the suite, better-sqlite3 for the ledger — 63 tests across 13 suites.

There is no Bee device on this machine, and the submission says so everywhere rather than in a footnote. Every live call to the developer API is real HTTPS and returns a real HTTP 401. The demo then runs on an offline fixture set built to match the official CLI schemas, and the dashboard badges that state on screen instead of hiding it.

Everything claimed is reproducible from a checked-in script:

  • node ops/probe-bee-live.mjs — the live host, twice in one process, with and without Amazon's private root CA.
  • node ops/probe-protocol-version.mjs — our own protocol floor, every sub-floor version raised to 2025-11-25.
  • node ops/probe-bee-mcp-tools.mjs — Bee's own MCP server, measured.
  • ops/test.cmd — the suite plus a production build.

Challenges I ran into

The production API does not work with a standard HTTP client. It terminates TLS with an Amazon internal private CA, and nothing on docs.bee.computer mentions it. A plain fetch dies before any HTTP status with SELF_SIGNED_CERT_IN_CHAIN. The certificate is embedded in the official CLI; I extracted it and configured an https.Agent with it. Friction log Entry 1.

The MCP protocol floor was the surprise. The rules set 2025-11-25 as a minimum, so rather than assert it I measured it — and our own server was silently serving whatever older dialect a client asked for, because the SDK defaults to 2025-03-26 and offers no floor option. Fixed by rewriting sub-floor initialize requests, with a probe that proves it.

Then I pointed the same probe at Bee's own MCP server, and it does the same thing: ask it to initialise at 2025-11-25 and it answers 2024-11-05 — four published dialects below the request — as a successful handshake, with no error and no warning. It also exposes 34 tools, which I counted from a live tools/list rather than from the docs, and which is how I found my own product feedback had said 35. Both findings are in the friction log with reproducing output.

And the audit leak above, which is the one that mattered.

Accomplishments I am proud of

The invariant holds under adversarial testing, and I tested it adversarially rather than hopefully. Every guard in this project has been deliberately broken, watched to fail, and reverted — that evidence is checked in, failing output and exit codes included. When a defect got past the suite anyway, the fix was to change what the test looks at, not to widen the claim.

What I learned

Consent is not a boolean you attach to a recording. It is per-person state that has to survive the total absence of identity, and a platform that returns anonymous clusters has quietly made that the developer's problem.

Also: partial serialisation in an absence check is a trap, and I will not write one again.

What's next for Bystander

The four things this section used to list as future work are now built, so the honest next steps are smaller and harder. A token-holder path finished end-to-end for someone who actually owns a paired device, which I cannot verify from here. A test proving the token never reaches a response body, not just that it never reaches the status object. Foreign keys actually declared rather than only enabled by pragma. And the thing I would most like to hand to the Bee team instead of building around: a speaker enrolment endpoint, so that consent is something the platform can answer rather than something every developer has to re-invent in a local database.

Built With

Share this project:

Updates

Submission history