Inspiration

Cyanobacteria blooms are invisible until they're dangerous. A lake can look perfectly calm on Monday and be shut for drinking water and swimming by Friday — toxins you can't see, from organisms measured in cells per milliliter. Water managers still fight this mostly the old way: drive out, fill a bottle, wait days for a lab. Accurate, but slow, sparse, and impossible to run continuously across every urban waterbody in Europe. The OneAquaHealth mission — connecting ecosystem health to human well-being in cities — gave that frustration a shape: what if the warning came before the water changed color, for any lake, with the honesty to say how sure it is? That's the project we built.

What it does

BloomCast is an early-warning system for toxic algal blooms. Pick any point on Earth and get a live bloom-risk probability with a 3–7 day outlook, confidence intervals, and the reasons behind it — calm winds, warm anomaly, rising chlorophyll — shown as plain driver bars, not a black box. For 24 pilot waterbodies (five in EU research cities: Coimbra, Ghent, Benevento, Oslo, Toulouse) it goes deeper with a full 32-feature assessment. Around that core sit the resilience tools: a counterfactual Sandbox ("what if it's 2°C warmer, or nutrient load drops 30%?" — the actual scenario grid), a Replay Theatre of past confirmed blooms, an urban-stream wash-off nowcast, threshold alerts that reach your phone, AI situation briefs built to quote the live assessment (with a validation step that rejects anything ungrounded), and FHIR health-standard alert bundles a hospital system can ingest. Every number carries its pedigree — which model scored it, what each input was, how fresh the satellite reading is.

How we built it

One Docker container on Render's free tier runs everything: a FastAPI backend serving a statically-exported Next.js frontend. The model is a LightGBM + tiny numpy-CNN ensemble trained on 23,570 real in-situ samples (CAML program), reaching 0.85 out-of-fold AUC on the weather-only variant — and a CI honesty gate refuses any synthetic export from ever shipping. Live weather comes from Open-Meteo; spectral context comes from real Sentinel-2 scenes we fetch ourselves (22 pilots with measured scenes, cloud-gated and water-masked) served as honestly-dated priors — never a fake "live satellite" claim. Citizens report blooms through an anonymous form; stewards validate; validated sightings become capped training weights with a published share, and the influence ledger proves each instance. Postgres on the free tier persists it all where configured, with local SQLite otherwise — and the whole stack costs $0 with no credit card anywhere.

Challenges we ran into

The free tier fought us at every step — no cron, no workers, no background push — so the architecture bent toward honesty instead: page-open alert watches that admit their limits, stale-while-revalidate caches, and a browser-fetch escape hatch when the server gets throttled. The satellite pipeline had real bugs: unsigned requests rejected with HTTP 409s, and a coordinate-system bug that read empty ocean windows and misreported them as "cloud-covered." Our own FHIR validator passed bundles that the public HAPI server correctly rejected — an alert about a lake can't hang off the patient field, a distinction no self-check would ever teach us. And the ~3,500-row weather join for training had to be fetched in polite chunks because Open-Meteo throttles sustained scraping.

Accomplishments that we're proud of

  • A real-label model serving real users, guarded by a test that makes synthetic shipping impossible.
  • Measured satellite priors for 22 pilots, each labeled with its scene date — plus the discipline to leave cloudy pilots on a labeled fallback instead of faking it.
  • A FHIR bundle accepted and stored by a public FHIR server (retrievable by ID), after fixing a genuine interoperability bug the round-trip exposed.
  • 107 green tests, a complete API spec generated from the app itself, and five EU pilots where the judges' project lives.
  • A citizen loop that actually moves the model — capped, published, audited — instead of collecting reports into a void.

What we learned

Our biggest lesson: honesty scales better than accuracy theater. Every shortcut we refused — silent zeros, fake live satellite, unvalidated FHIR, exaggerated claims — came back as a feature: provenance labels judges trust, fallbacks that degrade gracefully, and a scorecard that says 0.00% instead of hiding. We also learned that free-tier constraints are a design gift: they forced caching, idempotency, and offline-first thinking we'd have skipped on paid infra. And that no validator replaces the real server — HAPI taught us more in one 422 than our test suite did in a hundred passes.

What's next for BloomCast

  • Clearer skies: retry cloud-blocked pilots (Toulouse, Erie) and re-run the spectral refresh on a rhythm, so priors never go stale.
  • Close the spectral loop: retrain with the measured satellite archive joined into the training frame, so the full model learns from real pixels, not just real weather.
  • More EU waters: the pilot pattern is proven — new cities are a data entry, not a project.
  • A real dispatcher: the day there's an always-on host, the existing alert-check endpoint becomes a server-side watcher and the page-open limitation disappears without redesign.
  • Hospital pilots: the FHIR path is proven against a test server — the next step is a real public-health inbox receiving real alerts.

Built With

Share this project:

Updates

Submission history