Inspiration

Every winter, duty teams across the UK uplands watch the same fragmented
signals — a Met Office amber warning here, a National Highways closure there,
a rescue log from years ago — and still get surprised when the A66 closes at
23:40. The first signal was visible at 04:00. That 19-hour 40-minute gap
between an available signal and a usable decision is what Bothy was built to
close. Not by replacing the human on duty, but by giving them an
evidence-backed draft they can sign — or refuse.

## What it does

Bothy is an evidence-backed, human-approved winter access decision agent for UK upland roads. It monitors routes on a deterministic beat grid, cites every
claim to a timestamped source (Met Office, National Highways, EA river gauges, DfT sensor feeds), and produces a draft assessment with a causal chain. When
the risk score crosses a corridor's threshold, the agent calls
create_human_review — but only if a provenanceGuard Strands hook
confirms every claim has a source and timestamp. No cited sources means no
review. A named duty officer then signs the case in one action. Subscribers
receive a forwardable, signed /case/:id link tagged with the lead time vs.
outcome. Every approval is auditable.

## How we built it

  • Strands as the agent loop, with provenanceGuard blocking
    create_human_review calls without timestamped cited sources.
  • A Node/Express + SQLite agent service exposing
    /api/scenario/{backtest,live,flood}/... plus /api/subscriptions,
    /api/digest, /api/assessments/:id/decision.
  • A Next.js 15 web app at bothyapp.netlify.app with a Replay timeline,
    Live desk, /case/:id shareable pages, and a Strands trace viewer.
  • A HyperFrames 90-second demo composition (1:30 at 1920×1080) rendered
    to MP4 using a single paused GSAP timeline driven by live API data.
  • Traefik reverse proxy on the VPS, Netlify for the web, all wired
    through RESEND_API_KEY, DIGEST_TOKEN, and PUBLIC_APP_URL.

## Challenges we ran into

  • No LLM dependency at submission time — the agent must run without paid inference. Solved with a scripted fallback that emits the same causal-chain
    shape the Strands loop would produce, plus a "no create_human_review" guard so the gate is never bypassed.
  • The "19h 40m" number had to be real — not a marketing claim. Solved by replaying the A66 Brough-Bowes incident from 12 Feb 2026, scraping cited
    signals (W1, W2, sensor, EA, DfT) and computing the actual time between
    first cited signal (04:00 YELLOW warning) and reported outcome (23:40
    closure).
  • VPS rebuild initially failed — the in-place snapshot was 4 weeks old
    (pre-Strands), and npm ci rejected the Strands express@^5 peer. Fixed with a one-line Dockerfile change (--legacy-peer-deps) and a fresh source rsync.
  • Demo video pacing — v1 was 4:30, felt like a slideshow. Compressed to
    90s while keeping all five beats: pain-point stat, problem, rest→wake with
    animated A66 corridor map, approve gate, case share, flood wedge, closer.

## Accomplishments that we're proud of

  • The provenance guard is real, not a comment. The Strands hook
    literally cancels any create_human_review call missing timestamped cited
    sources. This is the kind of agent safety that hackathons usually wave at.
  • Every number on screen is live from the API. The "19h 40m" lead time
    is the actual replay delta. The "2 neighbours on watch" chip is the live
    subscription count. The Strands trace panel is real output from
    strands:qwen-hf.
  • A one-action human gate that the agent cannot bypass. The dashboard
    audit line includes Demo Officer · Hackathon Take with a timestamp, and the
    case page is forwardable.
  • A working demo video built with the toolkit's own primitives —
    beat-timeline, beat-accent, camcorder-hud, news-ticker, plus a custom
    animated A66 corridor map with pins dropping on each beat.
  • Shipped to production — bothy-agent on the VPS behind Traefik at
    api.bothy.trustfall.xyz, web at bothyapp.netlify.app, both healthy and
    verified live.

## What we learned

  • Agent safety primitives need to be at the hook level, not in prompts.
    A provenanceGuard that cancels an unsafe tool call is enforceable; a system
    prompt that says "please cite sources" is not. Bothy ships the former.
  • A 90-second demo with five live beats beats a 4:30 demo with twelve.
    Compression forced us to keep only what was structurally important — the
    rest→wake arc, the gate, the case page, the flood wedge. Everything else was
    decoration.
  • Determinism matters more than fidelity. The dashboard stills, the
    audit receipt, the case page — all are captured from the running system at
    recording time, not mocked. That's why the demo is honest.
  • Hackathons reward honest constraints. We could have run with a paid
    LLM, but the demo is stronger for showing the gate works without one.

## What's next for Bothy

  • Wire up RESEND_API_KEY + DIGEST_TOKEN + PUBLIC_APP_URL so the
    digest email pass actually sends (currently queues + logs only).
  • Bring Resend live for forwardable case digests — parish clerks, school offices, and food banks as the first named audience.
  • Add river-flood stratification for the Eden Valley and other inland
    corridors — same ledger, new wedge, already prototyped as /watch?case=flood.
  • Promote the demo render to 1440p via HyperFrames Cloud for
    hero-quality submission.
  • Open the audit feed — duty teams should be able to query their own
    corridor's decisions across seasons, not just see today's.

Built With

  • ai-agents
  • amazon-web-services
  • civic-tech
  • emergency-response
  • human-in-the-loop
  • open
  • source
  • strands
  • winter-operations
Share this project:

Updates

Submission history