-
-
The answer, withdrawn mid-sentence. Not a mockup: the Echo Show card driven by a real notifications/resources/updated frame.
-
The physio signs one change, the MCP server pushes it to Ray's screen, and the private record never reaches Ray's token.
-
The physio's screen: one record, one choice, one button — and a receipt that measures the correction instead of asserting it.
-
The demo as code on the real server: Ray's device is told a private record doesn't exist (-32002); the retraction chains to v2.
-
34 safety properties asserted as failures. If one of these passes when it should fail, the clinical-safety claim is false.
-
Public, no account: the live /verify replays every version chain and admits nothing about the records Ray may not hear.
Ray is 68, five days home after a hip replacement, and Alexa is about to tell him the wrong thing about his leg. Unsay makes it stop mid-sentence. An MCP server (spec 2025-11-25, Streamable HTTP) for the Alexa+ track whose care plan is live resources, not a PDF. Live: https://api.unsay.edycu.dev | Repo: https://github.com/edycutjong/unsay | Public proof: https://api.unsay.edycu.dev/verify | 344 tests, 34 safety assertions, 200/200 retractions inside the speech window.
Inspiration
A discharge summary is a photograph of a plan. The plan keeps moving after the photograph is taken.
Ray Dunn is 68 and five days home after a hip replacement. He has four sources of instruction: the surgical team's weight-bearing status, the physio's exercise progression, the GP's anticoagulant stop date, and his daughter Nell who does the shopping. They change, they contradict each other, and nobody tells the others. He asks the Echo Show by the bed, because it is the only thing at eye level when you cannot bend.
Every assistant available to him today answers from a snapshot — a PDF, read aloud with total confidence, with no idea it went stale two days ago. A 68-year-old acting on a two-day-old weight-bearing instruction is how people end up back in theatre.
The realisation that started this build is narrow and it is not about healthcare: we have spent two
years making assistants more confident and almost no time making them correctable. There is a
protocol field for it. MCP has resources/subscribe and notifications/resources/updated, which
means a resource an assistant is already speaking from can change underneath it. We found almost
nothing in the wild using it that way, and it is the one property that turns a confident assistant
into a safe one.
The second thing that started it is annotations.audience — a two-line field in the MCP spec with
exactly two legal values, "user" and "assistant". Of 129 project ideas pooled for this hackathon,
none touched it. Here it is the entire clinical-safety architecture: some things that are true about
Ray must shape the answer and must never be said to Ray.
What it does
Unsay publishes a patient's recovery plan as live MCP resources instead of a document, so an assistant reading from it can be corrected — including in the middle of a sentence.
Ask "can I put weight on it yet?" while the physio is updating the plan, and this is what
npm run e2e actually prints, from the run that wrote docs/proof/live_run.jsonl:
ALEXA "You can put about half your weight on it—"
"Wait — don't do that. What I just told you is out of date. I said "Partial weight-bearing, about half your body weight through the operated leg." Sarah Okafor, physio changed it just now: "Full weight-bearing as tolerated." In plain terms, full weight-bearing as tolerated means you can put as much weight through that leg as is comfortable."
"Take it slowly the first time, and have someone nearby." ← shaped by
care-internal://ray/risk; the reason is never spoken
The retraction itself — everything from "Wait" to "comfortable" — is not written in a slide and
not typed into the demo script. It is produced by renderRetraction() in src/retraction.ts from
the two record versions and handed to the host on _meta['unsay/retraction']; npm run e2e prints
the server's output, and test/docs.test.ts recomputes the copy on the landing page and fails if
the two ever differ. The last line is printed by scripts/e2e.ts to show the shaped answer — the
sentence Ray hears because of a fact he is never told.
The order inside the retraction is the safety property — withdraw first, then name what is being withdrawn, then the new value with its author and its age, then a gloss of the clinical language — because a frightened person acts on the first clause they hear.
Three things are happening there, and only the first is obvious.
1 · The correction reaches speech. The physio's write goes through a signed POST /write, the
server fires notifications/resources/updated on care://ray/weight_bearing, the subscribed host
re-reads, and the retraction begins. Measured 200 times over loopback Streamable HTTP: p50 1.7 ms,
p95 3.3 ms, max 8.5 ms end to end, signed write → the client holding the new value. Measured 200
times again against the deployed server over the public internet (client in Indonesia, server in
us-west2, so both legs cross the Pacific): p50 533 ms, p95 668 ms, max 1,148 ms. Both land
inside the assumed 3,400 ms speech window in 200/200 runs.
Those milliseconds move on every machine and between two runs on the same one — quote whatever
docs/proof/bench.txt says on the day, not this sentence. What does not move is the pair of
lines underneath the table: 200/200 inside the window, and a host was subscribed for every run.
npm run bench exits non-zero if either ever stops being true. The loopback figures are never quoted
as a production number, and the deployed figures are docs/proof/bench.remote.txt; the 3,400 ms is an assumed speech rate (12 words at ~150 wpm),
not a measurement of Alexa+ TTS. The benchmark prints both caveats itself, along with the fact that
the change bus is in-process.
2 · Not everything true is speakable. Ray's record also holds "fall risk: HIGH. Lives alone
Monday to Thursday. Family disputes the discharge plan." That is why Alexa said "have someone
nearby" — and Ray never hears the reason. The graph is split by annotations.audience, but the
annotation is not the control, because the MCP spec places no obligation on a client to honour
it. The partition is enforced server-side: two URI schemes (care:// and care-internal://) behind
two OAuth scopes, checked in LiveResourceStore.read() before any content is returned. A
user-scoped principal asking for care-internal://ray/risk gets -32002 Resource not found — the
same answer as for a URI that does not exist, so the partition does not leak by existence either.
resources/list under a user scope returns no care-internal:// URI on any cursor page, and
GET /verify without a token will not admit those chains exist.
3 · Stale facts announce their age. The GP's anticoagulant stop date passed nine days ago and
nobody updated it. Instead of reading it out flatly, the resource says so:
[STALE — last changed 9 days ago by Dr Mensah, GP; say this age aloud]. The record's own sentence
names a stop date that has already gone by, and npm run verify §9 asserts that it has — on the
clock a stranger's npm start actually runs, not a pinned one.
Underneath, every revision is a link in a SHA-256 chain — SHA-256(prev ‖ value ‖ writtenAt ‖ authorId)
— so a retraction is auditable and a tampered version is located, not merely detected. Anyone can
check it with no token and no account at GET /verify, which on a fresh npm start replays
5 chains · 6 versions · all intact.
Who it is for. Ray, post-operative and voice-first by necessity, who will never know what an MCP resource is. Sarah the physio, whose cost of publishing a change has to be near zero or she won't. Nell the daughter, who needs to know the plan changed without being told the clinical detail.
What it is not. Not medical advice — Unsay relays what a named clinician wrote, with its age and its author, and generates no guidance of its own. Not a medical device, and not clinically validated.
Potential impact
The moment this is about. Weight-bearing status after hip or knee arthroplasty is not a fact, it is a status, and it moves: non-weight-bearing, then partial, then full as tolerated, on a schedule the surgeon or the physio changes when they see how the patient is doing. The patient is at home. The instruction that reaches them travels by a phone call to a landline, a sheet printed on the day of discharge, or a letter. When the status changes on a Tuesday, the sheet on the fridge still says Monday's — and the person acting on it is post-operative, frequently alone, and has no way to tell that the sentence they are reading has been superseded. That is the failure: not a missing instruction, a stale one, followed with confidence.
Who is helped, specifically. Three people, and they are helped differently.
- The patient at home, in the hours between one clinical contact and the next — which is where almost all of recovery happens and where none of the existing delivery channels reach. What he gets is not more information; it is the ability of the machine beside his bed to be wrong out loud and then correct itself, with a name and an age attached.
- The clinician who changed the plan. Today, publishing a change means a phone call that has to be answered. Sarah's screen is one field and one button, and it hands back a receipt that reports what actually happened — the version, the hash, and how many hosts were subscribed at that moment, counted rather than asserted. A clinician receipt that renders "a host was corrected" from a constant is exactly the lie this build's own gates exist to catch.
- The family member who needs to know the plan moved without being handed the clinical reasoning
behind it. That is not a permissions checkbox bolted on later; it is the same audience partition
the assistant runs on,
prompts/get brief_carer, asserted speakable-only.
How many. We are not going to put a number here that we did not measure. An arthroplasty volume, or a readmission rate attributable to instruction error, would have to come from a named source with its year and its denominator — the National Joint Registry's annual report, NHS England discharge statistics, HCUP in the United States. That work has not been done, and inventing it would falsify the one claim this project actually makes, which is that its numbers are checkable. What can be said without a citation is the shape of the population: it is every patient discharged on a plan that is expected to change, which is what "recovery" means. The number is not small, and we are not going to pretend to know it.
What Unsay replaces, precisely. Not the clinician, and not the record system. The delivery of a
changed instruction to the place the patient actually asks the question. Unsay makes the fact itself
correctable: the assistant that already answered "about half your weight" is told, in the same breath,
that it was wrong, by whom, and how long ago. npm run bench measures that whole path against the
deployed server, across the Pacific and back, at p95 668 ms over 200 runs — five times inside the window a sentence gives you — and
npm run probe:resume proves it survives the Wi-Fi dropping: three revisions written with nothing
listening, all three replayed on reconnect after an 806 ms outage. A domestic Echo Show loses its
stream constantly, and a correction that is only ever pushed is a correction that can be silently
lost.
Who buys it. A discharge or therapy team inside a hospital, not a consumer. They already own the
plan and already carry the risk of a superseded instruction being followed. What they get that they
do not have today is a channel with a receipt: every change signed over the raw bytes, attributed to
the verified principal in an append-only audit log, hash-chained so a retraction is auditable, and
replayable by anyone at GET /verify holding no token and no key. The adoption cost is unusually low
for health software — one runtime dependency, no build step, no database to stand up, and
./scripts/fresh_clone_check.sh proves a stranger's clone runs every documented command verbatim.
Beyond the hackathon — shown, not asserted. The mechanism is domain-free and it already runs
somewhere else. packages/live-resources is the generic half — versioned resources, the
notifications/resources/updated wiring, the hash chain, the audience/scope partition — extracted on
day 9 with its own MIT licence. Its 29 tests import only its public entry point and run against an
incident-response scenario it was not extracted from: a status page an on-call engineer reads
while the root cause is still being written, with an incident-internal:// lane for the half that is
under legal review. That is a second vertical whose tests pass today, on this commit, not a paragraph
claiming the idea generalises. The same shape fits an on-call runbook that changes mid-incident, and a
price or an inventory count an agent is quoting while it moves. Each needs its own two schemes and two
scopes; nothing else changes.
Any MCP server publishing mutable state needs those four things — versioning, revision notification, an integrity chain, and an audience partition it can enforce. We looked for one that does them and did not find one; we are reporting a search, not a market survey.
Dated commitments after 2026-10-23.
- Publish
@unsay/live-resourcesto npm. It is extracted today and not published; do not read the extraction as a release. - File F-002 and F-003 to the MCP specification repository, the six SDK entries to
modelcontextprotocol/typescript-sdk, and F-013 to the MCP Apps extension, by 2026-09-20. The send-by date is written intoFRICTION.mditself, because a draft with no send date is a loss in progress. - Put the fresh-clone demo in front of five clinicians who did not build it, and publish what they say — including if what they say is that it is useless. Not done yet: as of 2026-09-29 no one outside the team has used it.
How we built it
One MCP server, one runtime dependency, no build step. Node ≥ 22 with native TypeScript
stripping, @modelcontextprotocol/sdk@1.30.0, and nothing else in dependencies. The source a judge
reads is the source the process runs.
MCP is the engine, not a wrapper. ARCHITECTURE.md is generated from the code by
npm run docs:arch and counts 10 request handlers and 3 notification senders, listed only if a
handler or a sender exists in the source. Those handlers carry sixteen distinct product behaviours:
| # | Surface | What Unsay does with it | Without it |
|---|---|---|---|
| 1 | resources/subscribe / unsubscribe |
The correction has somewhere to arrive | Poll on a timer, or invent a webhook the host does not speak |
| 2 | notifications/resources/updated |
The frame that fires the retraction | A second channel the assistant cannot receive |
| 3 | notifications/resources/list_changed |
A new care domain appears without re-handshaking | Client restart, or a stale list for the session |
| 4 | annotations.audience |
The speakable / reasoning-only split (hint; enforcement is server-side) | Two servers, two connections, a private convention |
| 5 | annotations.lastModified |
Self-announcing staleness | A bespoke _meta field every client must learn |
| 6 | annotations.priority |
Weight-bearing outranks contact details | Ordering smuggled into the body as prose |
| 7 | resources/templates/list, RFC 6570 |
One template addresses every version of every domain, per scheme | Enumerate every URI, unbounded |
| 8 | completion/complete with context.arguments |
{version} completes against the already-resolved {domain} — the documented purpose of that field, and we found no other server using it |
Client-side version guessing |
| 9 | resources/read, text and blob |
A 24 kB physio clip beside the text plan, under the same partition | A separate media API the assistant cannot verify |
| 10 | resources/list with cursors |
Paginated by last URI emitted, so a concurrent publish cannot drop a record into the gap | An offset that silently skips a resource mid-correction |
| 11 | prompts/list / prompts/get |
brief_carer — a briefing for Nell that provably cannot emit reasoning-only content |
A prompt assembled client-side out of whatever it could read |
| 12 | logging + notifications/message + logging/setLevel |
A second, cheap revision channel a host can turn down — and it is honoured, not merely served | One channel, no way to quiet it |
| 13 | Streamable HTTP + Last-Event-ID |
A correction written while the Wi-Fi is down is replayed on reconnect | A silently lost correction — the exact failure the product exists to prevent |
| 14 | OAuth 2.0 resource server + /.well-known/oauth-protected-resource (RFC 9728) |
Scope-per-URI-scheme makes the audience split enforceable | Trusting a client to honour an annotation, which is not a safety property |
| 15 | Server instructions + the whats_changed tool |
Graceful degradation on a host without native subscriptions — and the fallback carries the previous value, so it can retract rather than merely restate | No fallback when the host lags the spec |
| 16 | MCP Apps (ui://unsay/echo) + an Agent Skill (skills/unsay-care-plan/SKILL.md) |
The Echo Show card delivered through the protocol as text/html;profile=mcp-app, and the retraction protocol as a Skill whose rules test/server.test.ts asserts have not drifted from the server's own instructions |
A card served off a static route and a contract that lives only in prose |
Row 16 is the track's own "creative, not obvious" example list answered directly: the Alexa+ rubric names MCP Apps and Agent Skills by name. The track accepts an Agent Skill or a self-hosted MCP server; Unsay ships both.
The removability test. Take MCP out and there is no application left — not a degraded one, none. You would need a pub/sub bus the assistant does not speak, a bespoke audience convention every client must be taught, a custom versioned-URI scheme, a second media API, a hand-rolled resumable transport, and an authorization layer with no standard to point at: six systems and at least four extra libraries, to reimplement badly what one spec revision already defines. A "Unsay without MCP" is a JSON file on a server — which is the discharge PDF we exist to replace.
Security, because this is a health record.
- Two scopes, checked at the store.
care.read.usercannot reach acare-internal://record on any cursor page, over HTTP, ever — and not on a resumed stream either, which is where it once could (see Challenges, F-014). - AES-256-GCM at rest with the AAD bound to
patient|domain|version|audience, so the partition survives an attacker who owns the database and simply moves the bytes.npm run verify§7 pastes the ciphertext ofcare-internal://ray/riskintocare://ray/weight_bearingand asserts it refuses to open — and that the chain reports it broken at that exact version, and that the same bytes still open under their own identity. - The write path is HMAC-SHA256 over
timestamp . rawBody, beforeJSON.parse. Verifying a re-serialised parse is how signed webhooks get forged, and putting the timestamp inside the MAC is what stops a captured request being replayed under a fresh one. Unsigned, mutated, wrong-key and stale requests all return an identical{"error":"unauthorized"}naming no cause, and every refusal leaves an audit row that carries no credential. - The audit row names the verified principal, never the body's claimed
authorId— acare.writeholder can still claim any author on the record, and the trail says who actually presented a credential.
Proof, not assertion. Five commands, no flags, from an empty clone — there is no MOCK=,
OFFLINE=1 or --dry-run anywhere in the repo, and scripts/check_submission_readiness.py fails the
build if one appears:
| Command | What it proves | Receipt |
|---|---|---|
npm run probe |
subscribe → notification → new value, over a real transport | docs/proof/probe_subscribe.json |
npm run verify |
34 assertions — 15 in process, 19 against a real HTTP server — that the reads and writes which must fail, do | docs/proof/verify.json |
npm run e2e |
the whole demo as code against the same server npm start runs |
docs/proof/live_run.jsonl — 20 protocol frames + a summary line |
npm run probe:resume |
three corrections written while the stream is down are all replayed on Last-Event-ID |
docs/proof/resume.json |
npm run bench -- --n 200 |
p50/p95 end to end, and 200/200 inside the speech window | docs/proof/bench.txt, bench.json |
Plus 344 tests (npm test), tsc --strict clean, and ./scripts/fresh_clone_check.sh, which
clones to a temp directory with empty state and runs every documented command verbatim — including the
curl walkthrough against a real npm start — because two earlier projects in this builder's history
shipped 458 and 404 passing tests over demos that were broken from a fresh clone.
Three single-file web surfaces (index.html, echo.html, clinician.html) with no framework and
no imported script: each carries its own ~40-line MCP-over-fetch client and SSE framer, and all
three are served from the same origin as /mcp, because a browser blocks a cross-origin request
before it leaves and a judge should need one URL, not two. 58 of the 344 tests slice echo.html's own
client out of the page and execute it against a live server — the transport a demo video films was
the one transport nothing ran. With no server answering, the pages fall back to committed seed values
and say SEED · NOT LIVE on the device screen itself, where presentation mode cannot hide it.
AWS deployment is not built. The AWS KMS provider is SigV4-signed and shaped and has never executed against a live key, because the Amazon account is under a payment-verification hold (F-004). See Track & mini-challenge declaration and the product-feedback field. Do not add an AWS sentence here unless a deployed resource exists.
Challenges we ran into
1 · The safety field is advisory (F-002). annotations.audience is defined in the spec with two
legal values and no obligation on a client to honour it. A compliant client may read an
audience: ["assistant"] resource and speak it verbatim. A safety-relevant annotation a client may
ignore is a documentation comment, not a control. The whole architecture turns on the response: split
the graph into two URI schemes behind two OAuth scopes and enforce it at the resource server, so a
non-compliant client cannot leak internal context because it is never sent any.
2 · The annotation does not even arrive on resources/read (F-005). It arrives on
resources/list and is silently discarded on read — TextResourceContents and
BlobResourceContents in the shipped typings carry uri, mimeType, _meta and the payload and
nothing else, and both are $strip, so the client's schema parse deletes the field with no error and
no warning. list teaches you the field works; read quietly proves it does not. The staleness
signal had to move into the text body ([STALE — …]) and into _meta, because those are the only
channels that provably reach the model.
3 · A resumed stream walked around every guard we had built (F-014, 4 h, and it was a live leak).
We had guarded both live channels at send time, re-reading the session's current principal on every
notification, because a session's scopes can narrow within its lifetime. Then we dropped the stream
and reconnected with Last-Event-ID under a narrowed token of the same subject — and the resumed
stream returned a care-internal:// URI verbatim: its version, its author, its timestamp, both
hash-chain links. The reason is in the interface, not our wiring: replayEventsAfter(lastEventId,
{ send }) is handed an event id and a sink and nothing about who is asking now. A server that has
correctly re-authorized every send has still not re-authorized anything a client can ask for again.
Fixed by filtering inside our own MemoryEventStore with a canReplay predicate that fails closed on
any URI the partition did not issue — which works only because we own both the HTTP handler and the
event store. A server using a third-party EventStore has no seam at all.
4 · Resumability protects only a stream that has already delivered something (F-010, 3 h). The
SDK calls writePrimingEvent on the POST response stream and never on the standalone GET stream, so
a client that has received nothing holds no event id and reconnects with no Last-Event-ID — the
missed notification is silently lost. The failure inverts the guarantee: you configure an event store,
you watch it fill, and the one window where a domestic Wi-Fi blip actually loses a correction is the
window before the first event, which is most of the time on a freshly opened screen.
5 · Route-level OAuth cannot express what an MCP server needs (F-006). MCP multiplexes every
resource behind one POST /mcp. requireBearerAuth({ requiredScopes }) is a flat list checked once
per route — but a read of a speakable fact and a read of a never-speakable one are the same HTTP
request to the same path, differing only in a JSON-RPC param. Route-level scopes can only say "this
token may talk to this server at all", which for a server whose entire point is a safety partition is
the wrong granularity by one whole level. Authorization moved into the protocol handler, off
extra.authInfo.scopes.
6 · A cursor nobody validates (F-007). On the day the server still ignored params.cursor, a
garbage cursor produced a complete, successful, schema-valid response: BOGUS CURSOR : accepted,
returned 6 resources. Nothing in the spec or the SDK catches it. We had to invent the missing
semantics — the cursor carries the last URI emitted rather than an offset, so a concurrent publish
cannot drop a resource into the gap; an unrecognised cursor fails with -32602 rather than resetting
to page one and looping a client forever.
7 · A guard that could not be tested (F-012). Suppressing list_changed for the server's initial
state was a queueMicrotask flag — and a notification sent before a transport is attached fails
silently, so a correctly suppressed notification and a wrongly sent one look identical. The test
written to prove the guard worked would have passed with the guard deleted. It was rewritten to arm on
the first resources/list actually served, which is both the notification's own meaning and
observable.
8 · The bug every gate was structurally unable to see. The live entrypoint seeded its store on a
pinned clock and read the wall clock, so npm start — the one command every document sends a judge to
— served [changed -32d ago by …] while the whole suite, 184 tests at the time, stayed green. The
fix was a ninth verify section that takes no injected store and no injected clock, which is the only
way to test the thing a stranger runs.
9 · The Amazon account itself (F-004, multi-day, still open). An Amazon payment-verification hold on an expired debit card blocked the Ring Developer Portal and, with it, the Ring and Fire TV tracks entirely. Alexa+ was buildable on day one only because a self-hosted MCP server needs no Amazon credential. It is also why there is no live AWS integration in this submission — the hold was still unresolved on 2026-09-29.
10 · The residual risk we could not engineer away (SPEC T-7). A reasoning host legitimately holds
both scopes — that is the point, internal facts exist to shape the answer. The partition protects Ray
against a user-scoped client; it does not protect him against the reasoning host choosing to
speak. scripts/e2e.ts grants its host principal both scopes, and a judge who opens that file finds
this paragraph already waiting in docs/SPEC.md. Closing it needs the speech surface to be a
different principal from the reasoning surface, which is a host-side property MCP gives a server no
way to require.
Accomplishments that we're proud of
The retraction is real and it is reproducible by a stranger. npm run e2e runs the exact sequence
the video shows against the same HTTP server npm start runs — real Bearer token, cursor-paginated
listing over 3 pages, a real 24,044-byte audio blob, the MCP Apps card read over the protocol, the
partition attacked from Ray's own host, a tampered write refused, the physio's correction arriving
through the signed write path — and drops 20 protocol frames plus a summary into
docs/proof/live_run.jsonl. No flag anywhere disables the thing being judged.
The safety properties are asserted as failures, not as happy paths. npm run verify is 34
assertions — 15 in process, 19 against a real HTTP server — that the reads and writes which must
fail, do: the partition, the non-leaking of existence, the version chain and its located tampering,
staleness on the entrypoint's own clock, an audience a later write cannot flip, OAuth over real HTTP,
the AAD ciphertext-paste attack, and four ways of forging a clinician write. If one fails, the
clinical-safety claim is false and the build should not be submitted.
A correction survives the network dropping under it. npm run probe:resume subscribes, drops the
SSE stream the way a proxy timeout does, writes three revisions while nothing is listening,
reconnects with Last-Event-ID, and asserts all three are replayed — checked against the
server-side timestamp of the resume request, so a live send cannot be mistaken for a replay. Three
and not one on purpose: a replay that delivered only the last missed notification would pass a
single-revision probe and lose the middle of a correction sequence in the field.
We found a leak in our own build and wrote it down rather than quietly patching it. F-014 is the entry we would most have liked not to write: a resumed stream returning an internal URI past every guard we had. It is in the friction log with the terminal output, the interface that made it possible, and the upstream ask that would close it for everyone else.
The documentation cannot drift into fiction. ARCHITECTURE.md is generated from actual handler
registrations. docs/SPEC.md carries 17 invariants, a coverage-gap table, a "what we do not defend
against" section and a "not built" section, all longer than a marketing document would allow. The
landing page's p50/p95 are written into it by the benchmark itself, rounded up, and a test fails
if anyone edits one by hand. FRICTION.md closes with What this build does NOT do — an explicit
inventory of every absence, because a judge who finds one of these on their own stops believing the
rest.
Fourteen friction entries, thirteen of them upstream-filable. F-005 through F-014 were found by
building, not by reading; seven are specification or extension asks and six are SDK asks, each naming
the file, the terminal output or the measurement it came from, and several are one-line fixes
(writePrimingEvent on the GET stream; a subscribe alias; isTextResourceContents() guards).
What we learned
A protocol field that looks like a control is more dangerous than no field at all.
annotations.audience reads exactly like a security boundary and is specified as a hint. The correct
response is not to avoid it — it is to keep emitting it for compliant hosts and make it unnecessary,
by ensuring a client that ignores it never receives the bytes.
A guard on the send path is not a guard on the read path. F-014 cost four hours and it was a live leak. Every notification was correctly re-authorized; resumability then handed the same bytes back to a narrowed principal through an interface that receives no identity at all. Any channel a client can ask again for is a delivery, and delivery carries the original send's authorization obligations.
Delivery is not freshness. MCP guarantees a notification is delivered and says nothing about what a client does with it. "Live resources" without a client-side re-read obligation is a delivery guarantee wearing a freshness guarantee's clothes. That gap cannot be closed from the server, which is why the fallback tool exists and is exercised in every e2e run — an unexercised fallback is indistinguishable from a missing one.
Write the limitation down on the day, in the repo. Every friction entry here has a date because it
was written the day it was hit. The entries assembled from memory would have been three sentences; the
ones written on the day carry the terminal output, the file in node_modules, and the measurement.
Unit tests do not test the sequence a human types. 344 tests pass in under nine seconds and none of
them would have caught a demo that needs state you forgot you had, or an entrypoint reading a
different clock from the one it seeded on. fresh_clone_check.sh — clone to a temp directory with
empty state, install, run every documented command verbatim — catches a different class of failure
entirely, and it is the one that decides whether a judge's clone works.
Measuring the thing you claim is cheap, and quoting it honestly is the hard part. The benchmark prints its own caveat: loopback, and a speech window that is an assumption rather than a measurement of Alexa+ TTS. It exits non-zero if a run stops landing inside that window, or if a host was not actually subscribed for every run.
What's next for Unsay
- Close the SPEC T-7 gap with a two-principal host pattern — the speech surface holding only
care.read.user, the reasoning surface forbidden from emitting text directly. It is the one design change that would make the partition complete rather than partial. - File the friction log upstream. F-002 and F-003 to the MCP specification repository, F-013 to
the MCP Apps extension, and the six SDK entries (F-001, F-006, F-009, F-010, F-011, F-014) to
@modelcontextprotocol/sdk. The first of these, F-010, is drafted as a patch with a reproduction and has not been filed yet. - Publish
@unsay/live-resourcesto npm. It is already extracted:packages/live-resourcesis the generic half (versioned resources, the hash chain, the audience partition, the notifier) with its own MIT licence and 29 standalone tests importing only its public entry point, against an incident-response scenario it was not extracted from. It is consumed here from source by relative import and is not published — do not read the extraction as a release. - Re-measure from where the users are. The server is deployed (
https://api.unsay.edycu.dev, Railway us-west2) andnpm run bench -- --urlalready measures it over the public internet; next is running that client from a clinician's network and an Echo's, not a laptop. - Persistence, an authorization server, and per-clinician write authenticity — today the store, the event store and the audit log are in memory, tokens are HS256 self-issued, and one shared write secret proves a holder wrote this, not which clinician.
- A second patient. The store is keyed
patient/domainand every principal is scoped to the whole store; multi-tenancy needs per-patient authorization that does not exist yet. - Real users. Zero external users today; the plan is 5+ physios and carers running the fresh-clone demo and answering three questions, with the answers published as given.
We have spent two years making assistants more confident and almost no time making them correctable.
Ray deserves an assistant that can stop mid-sentence and say "actually, that changed." Everything
above is reproducible from an empty clone in five commands, and everything this build does not do
is written down at the end of FRICTION.md rather than left to be discovered.
— Edy Tjong (edycutjong)
Built With
- aes-256-gcm
- agent-skills
- alexa
- alexa-plus
- hmac-sha256
- html
- javascript
- json-rpc
- mcp
- mcp-apps
- model-context-protocol
- node.js
- oauth2
- rfc-6570
- rfc-9728
- server-sent-events
- sha-256
- streamable-http
- typescript
- vitest
- wav
Log in or sign up for Devpost to join the conversation.