Inspiration

Agents plan trips, deliveries and events, and they read a weather forecast as a fact. Six models disagree about Saturday, and every weather tool hides that. Ilma shows the agent how much to trust the forecast right now, with proof, and lets the human and the agent work on the same page.

What it does

ilma.io is a live multi-model forecast site that is also a WebMCP server. It registers 19 tools on document.modelContext.

The human's gesture is agent-visible state. Drag across the chart, click two hours, or tap a day: the picked window goes into every tool description (the tools are re-registered on each change) and into every result. assess_plan with no arguments assesses what the human selected. get_current_view resolves "here" and "that weekend".

The agent works on the page, not only in the chat. mark_hours paints its reasoning onto the human's chart. select_window makes the same gesture the human makes. set_plan writes a plan card both parties edit; human-edited lines survive a rewrite and come back flagged. share_plan returns one link that reopens place, plan, highlights and selection for another person or another agent. snapshot_chart returns a PNG of what the human sees.

The human stays in control. Ilma, a small avatar, narrates each tool call. An activity log lists every action in plain words with Undo on the newest change of each kind. Read-only tools carry readOnlyHint; anything the human typed carries untrustedContentHint, so a plan card cannot prompt-inject the agent.

The numbers are earned. get_probability, get_rain_probability, get_stability and when_to_decide come from a pre-registered benchmark that has snapshotted every major forecast source every five hours since July and scores them nightly against real stations. why_trust_this returns the verification record itself: matched pairs, confidence intervals, p-values against Foreca, the Finnish Meteorological Institute and Google. Event presets (sauna, sailing, hike, wedding, drone, paint, harvest) give agents semantics instead of thresholds.

How we built it

One tool table in the page. It registers on the open standard, re-registers when page state changes, and also drives the site's own chat agent (GPT-5.6 Luna), so three doorways share one toolset. The forecast blend and the verification pipeline run on a VPS (Python, FastAPI, SQLite, 25M+ forecast rows, 32 cities in 5 countries). The site is static on Vercel.

Challenges

WebMCP has no "push state" call, so page state travels in descriptions and results, with re-registration through an AbortSignal. Calibration had to come before tools: probabilities are verified frequencies, not model output. Judges may test in browsers without native WebMCP, so every surface feature-detects and fails closed.

Accomplishments

The blend beats Foreca, FMI and the best single model over days 1 to 7, each Holm-significant, and the site's headline claim is gated on that significance and recomputed nightly. If we stop being better, the site stops saying so.

Built With

Share this project:

Updates

Submission history