TAIL RISK

TAIL RISK is one web page that tells an organ-transport coordinator whether a commercial flight is still operationally safe to rely on, and that takes the recommendation away from the assistant acting on it the moment the evidence behind it stops being true.

A donated kidney is available in Chicago at 19:45 and has to be in a Boston theatre by 04:15. Four commercial flights connect ORD to BOS inside that window. An assistant asks the page which one to use.

The earliest, UA 417, is the obvious pick on a timetable and the wrong one. Its aircraft is over Colorado, 87 minutes late into O'Hare, and 76 of those minutes cannot be absorbed by the scheduled turn:

UA 417   BROKEN   −76 min   binding: AIRCRAFT_ROTATION
         cargo acceptance +20 · aircraft rotation −76 · destination delivery −18

UA 623 works, at +33 minutes. The assistant commits, and what comes back is not a message. It is a certificate: the margin, the binding constraint, the tail number, the inbound sector, a validity horizon, and a fingerprint over all of it.

Five minutes later the carrier swaps the aircraft. N24211 becomes N77518, which is still inbound from Houston. Nothing about the timetable moves. The mission margin goes from +33 to −11.

And this is the part the project exists for:

PLAN_SELECTED                          PLAN_INVALIDATED
  select_transport_plan   withdrawn      select_transport_plan   withdrawn
  get_selected_plan       registered     get_selected_plan       withdrawn
  revalidate_plan         registered     revalidate_plan         withdrawn
  explain_invalidation    withdrawn      explain_invalidation    registered
  replan_mission          withdrawn      replan_mission          registered

get_selected_plan and revalidate_plan are not disabled and they do not return an error. They are unregistered. toolchange fires. The assistant holding that plan discovers the world moved because its own tool list changed under it, and the only route forward is replan_mission, which re-evaluates against current evidence and lands on UA 1051 at +15 minutes, TIGHT.

An agent's previous decision had a validity condition, reality broke it, and the web application withdrew the capability before the agent could act on stale information.

The problem is real, and it is not a data problem

The FAA notes that organs and biological material travel on commercial airlines virtually every day. UNOS documents cargo acceptance cutoffs of 60 to 120 minutes and cargo offices that are not open around the clock.

A coordinator choosing between two flights is reading a timetable. A timetable does not say which aircraft is assigned to a flight, and it does not say that the aircraft is currently in the air somewhere else. That single fact is what turns a comfortable +58 minutes into −76, and no amount of delay-probability modelling recovers it, because it is not uncertainty. It is arithmetic that nobody did.

Why this is a strong fit for WebMCP

Most WebMCP surfaces are a menu: the same tools, always present, and an agent picks one. This one is a statement about the present. Which tools exist encodes what is currently true.

That matters here more than it would in a shop. Time-critical transport is a domain where the expensive failure is not a bad decision, it is a decision that was good and quietly stopped being. There is no channel for pushing that fact to a host agent mid-conversation, and a page that only answered questions could not raise it. Removing the capability is the push. registerTool's AbortSignal plus toolchange is the only mechanism on the web today that lets a site interrupt an agent's plan without being asked, and it is the whole reason this product is a page with tools rather than a report.

What people and agents can do together that was difficult before

Hold a plan whose validity the page keeps checking, and be interrupted by the page when it stops holding.

A coordinator with a phone can refresh a flight tracker. An agent in a chat window cannot be interrupted: between the moment a plan is chosen and the moment the shipment reaches the counter, nothing reaches it. TAIL RISK re-evaluates every option on every evidence change, rechecks the certificate's basis against the new world, and when the fingerprint no longer matches it withdraws the capabilities whose existence implied the plan was still good. The agent then has a structured explain_invalidation giving the field-by-field difference, and exactly one way forward.

The division of labour is the point. The person is the one who knows the deadline and who will call the airline. The agent is the one who can hold four candidate flights, three constraints and a validity horizon in mind at once and act at the moment something changes. Neither is doing the other's job.

How I implemented WebMCP

  • Twelve tools, and the set of them is a pure function of mission state. surfaceFor(state, hasInvalidation) in packages/domain/src/machine.ts is the whole rule. ToolLifecycle owns one AbortController per registered tool, so unregistering is aborting a signal and the tool surface is derived state rather than something maintained by hand.
  • Reconciliation is driven by state transitions, never from inside a handler. reconcile() throws if called during a tool execution. A transition that lands mid-call is deferred and applied the moment the handler returns, so the surface never silently stops tracking the machine exactly when an agent is the one driving it.
  • One measured platform behaviour that cost a working build. When a handler's own state change unregisters that handler's tool, aborting the registration inside a finally lands in a microtask that is still within Chrome's await on the handler's promise. The agent then receives UnknownError: The operation failed for an unknown transient reason for a call that in fact succeeded. guard() therefore schedules the reconcile with setTimeout(0), the next macrotask, so the result reaches the caller before its tool stops existing. Measured in Chrome 152 by driving the deployed page with Playwright and calling every tool through document.modelContext.executeTool, which is how the failure was found at all.
  • Every direct document.modelContext call lives in one package. @tr/webmcp is the only code that touches the platform. It also enforces Chrome's documented budgets at registration time (30 characters of name, 500 of description, 150 per parameter description, roughly 1.5 KB of output) and records overruns rather than leaving them to surface as unexplained tool-selection drift.
  • No free-text field anywhere in the surface. Every string an agent can send is an enum, a three-letter airport code, a flight number under ^[A-Z0-9]{2,3} ?[0-9]{1,4}$, an option id under ^O[0-9]{1,2}$, or a bounded ISO instant. Every string an agent receives is an identifier, an enum, an ISO instant or a number. untrustedContentHint: false is a claim this build can make honestly, because no external prose enters the system in the first place.

The decision is arithmetic, and it names its own bottleneck

No language model decides whether a mission works. An agent chooses which flights to ask about and explains the answer to a person; the answer is computed:

acceptance_margin       = latest_acceptance − package_at_cargo
rotation_margin         = sched_departure − (projected_inbound_arrival + turnaround)
propagated_delay_floor  = max(0, inbound_delay − max(0, scheduled_turn − turnaround))
delivery_margin         = deliver_by − (projected_arrival + release + ground)

mission_margin          = min(acceptance, rotation, delivery)

The engine always returns which constraint bound. Collapsing that to a score is the failure this product exists to avoid: a coordinator told the binding constraint is aircraft rotation calls the airline, and one told the risk is 74% calls nobody.

Empirical P50/P80/P95 margins sit beside the deterministic figure and never override it. UA 623 is VIABLE +33 and simultaneously TIGHT AT P95, because 41 minutes of historical additional delay over 126 observations takes it to −8. Where history is thin the answer is INSUFFICIENT_HISTORY; where no aircraft is published the verdict is UNKNOWN rather than a number. Missing evidence never becomes operational confidence.

TAIL RISK. A schedule tells you when a flight should leave. This tells an agent whether the mission still works.

Built With

Share this project:

Updates

Submission history