Inspiration

Every city has a stream that people walk beside every morning. Then one dry day the water turns grey and starts to smell, because somewhere upstream a washing machine or a toilet has been plumbed into the surface water drain. In the UK alone, an estimated 150,000 to 500,000 homes have a misconnected drain (CIWEM).

Knowing that a stream is dirty is the easy part. Finding which pipe is responsible is hard: a stream has dozens of outfalls, the discharge comes and goes, and field crews walk the bank checking one pipe at a time. A pipe that looks clean today can be dirty tomorrow. Meanwhile the people who notice first, walkers, parents and schools, report into a void and never hear back, and what they saw never reaches the search.

This is about to matter much more. The recast EU Urban Wastewater Treatment Directive (2024/3019) requires integrated urban wastewater management plans, covering storm overflows and urban runoff, for every agglomeration of 100,000 population equivalent and above by 2033. We wanted to give those cities a way to turn scattered citizen signals into a located source, and to close the loop with the people who reported.

The name comes from the white-throated dipper, a small songbird of fast, clean streams that walks underwater to feed. Ecologists use it as a living indicator of stream health. Our logo is a dipper shaped like a map pin, dipping into the water: a clean stream sentinel that points to where the pollution enters.

What it does

Dipper turns citizen reports into a search that finds the polluting pipe, then hands the case to the people who act on it.

Who What they get
Citizen (phone, no account, works offline, in the city's language and English) Reports what they see or smell in under a minute. Instead of silence, gets one short, nearby check that sharpens the search, and sees what their check changed.
Investigator (utility or municipality) A case with a probability for each explanation and each candidate outfall, shaded along the real stream. The next best checks, each with both outcomes spelled out before anyone leaves the office.
Public health officer A drafted advisory that keeps what was observed apart from what is inferred, the possible risk, and what still needs a lab to confirm. Nothing is published until a person approves it.
Health information systems An HL7 FHIR R4 bundle built on the OneAquaHealth implementation guide, with 0 errors in the official HL7 validator, pushed to any FHIR server.

End to end: report → case → next best check → source localized → handover in FHIR → advisory → fix verified. The utility is always asked to confirm the pipe (dye test, smoke test or CCTV) before anyone digs, because a localization is a probability, not proof.

How we built it

Reports become evidence. Each report is snapped to a stream network built from OpenStreetMap with NetworkX: flow direction, culverts, and gaps bridged where the mapped line breaks. Six real streams are mapped today, in Coimbra, Oslo and Ghent. Weather comes from Open Meteo.

One belief over what and where. Dipper keeps an exact joint probability over six explanations \(h\) (foul sewage, wet weather overflow, chemical discharge, sediment, algal bloom, benign) and every candidate entry point \(j\), plus "upstream or unmapped". Weather sets the priors: dry days favour a misconnection, rain favours an overflow. Each observation \(e_i\) updates the belief:

$$ P(h, j \mid e_{1:n}) \;\propto\; P(h, j)\,\prod_{i=1}^{n} \Big[(1-\varepsilon)\,P(e_i \mid h, j) + \varepsilon\,P(e_i)\Big]^{w_i} $$

The outlier weight \(\varepsilon = 0.08\) bounds how much any single observation can move the belief, and \(w_i = 1/(1 + 0.5k)\) tempers the \(k\)th report near earlier ones, so twenty reports from one bridge count for far less than twenty independent ones. A look at the stream is positive with probability \(\text{pres}\cdot\text{det} + (1 - \text{pres}\cdot\text{det})\cdot\text{fa}\), where presence depends on whether the source is upstream, whether it is discharging at that hour, and whether a burst was seen in the last two hours. That is what makes a clean result count: it lowers every source upstream of it without ever ruling an intermittent source out.

The next best check. Every possible check (a citizen look, an outfall look, an ammonium strip, a lab sample) at every point is scored:

$$ \text{score} = \text{EVSI}(\text{advisory}) + 1.2 \cdot \Delta H - \text{cost} - 0.1 \cdot \text{delay} $$

EVSI is the expected value of the check for the advisory decision (a missed advisory costs four times the exposure stakes of the places downstream), and \(\Delta H\) is the expected number of bits the check removes from where the source is. Citizen missions are also charged 0.08 per km for the walk from where the person reported, so a nearly as useful check nearby wins over a slightly better one across town.

Photos help, people decide. A photo is redacted on the server first: EXIF (including GPS) is stripped and faces are blurred with OpenCV. Only then does a vision model (Claude or Gemini, with structured output) read the visual indicators. The photo counts as a separate, weaker observer; if it disagrees with the citizen, the citizen is asked, and their answer is never overwritten.

The stack. The engine is pure Python with NumPy. The API is FastAPI with role based staff accounts, hashed tokens, rate limits, a strict Content Security Policy and an append only SQLite event store that rebuilds every case by replaying its evidence. The web app is a React and MapLibre progressive web app in four languages that queues reports offline. The FHIR profiles are written in FHIR Shorthand on the OneAquaHealth IG (nine Dipper profiles on LocationOah and GroupOah), compiled with SUSHI and validated with the HL7 validator in CI. Everything ships as one Docker container, deployed on Render. CI runs 84 tests, an axe accessibility audit (WCAG 2.2 AA) at desktop, phone and 320 px widths, FHIR validation, an image build and a smoke test.

SourceBench. To check that the search actually helps, we built a simulator: 7 stream networks (5 real OneAquaHealth streams and 2 synthetic), 40 trials each with identical truths for every strategy, a budget of 25 checks, and a simulated world that is deliberately noisier than the model, with bursty discharges.

Strategy Localized correctly Median checks when it did Mean cost
Walk the bank, outfall by outfall 28% 20 1.49
Bisect the stream 43% 13 0.95
Dipper 56% 9.5 1.30
Greedy information gain, ignoring cost 74% 8 4.92

These are simulated results under stated assumptions, not field performance.

Challenges we ran into

  • Clean does not mean innocent. Misconnections are intermittent, so a naive model treats a clean look as proof and rules out the real source. Modelling activity by time of day and a two hour persistence window after a sighting made clean results informative without making them final.
  • The cheapest check is not always the best, and neither is the most informative. Pure information gain localized more often in SourceBench but spent about 3.8 times more, mostly on lab samples. Tying the score to the actual advisory decision, and tuning the search weight on a sweep (0.6, 1.2, 2.0, 3.0), gave the best trade off.
  • Real maps are messy. OpenStreetMap streams come in fragments, disappear into culverts and have gaps. The graph builder bridges gaps, drops stray fragments and keeps culverts as unobservable stretches, so checks are only proposed where someone can actually look.
  • Anyone can report, so public evidence must be bounded. Correlated reports are tempered, a citizen can answer only the mission they were given and only once, and a case needs at least one trained or staff check before it can be handed over.
  • Standards do not always fit. The OneAquaHealth cohort value set has only Age and Sex, while Dipper's exposed group is defined by place, so we used our own code. It is the one remaining validator warning, and we documented it.
  • Rate limits behind a proxy. On the live deployment, our rate limits first keyed on a load balancer address that changed on every request. We switched to the client address the hosting edge always overwrites, read only from trusted internal proxies, and tested it through the production mount.

Accomplishments that we're proud of

  • A complete loop that runs live: a citizen report on a phone, a search that localizes the source, a validated FHIR handover and a published public advisory, all in one deployed container.
  • In simulation, twice the success rate of walking the bank, with about half the checks and at lower cost.
  • 0 errors in the official HL7 validator on the OneAquaHealth IG, and 0 axe violations at three screen widths.
  • Advisories that never overclaim: every one separates observed, inferred, possible risk and needs confirmation, and a person approves it.
  • A citizen experience that answers back: people get a mission and see how much their check narrowed the search.
  • Privacy by design: anonymous citizens identified only by a keyed pseudonym, private report tokens never sent in URLs, faces blurred before any model sees a photo, and no IP addresses in logs.

What we learned

  • Negative evidence is evidence. In the live replay, four of the six checks that found the source came back clean. A search moves forward on clean results, as long as the model respects that discharges come and go.
  • Value of information needs a decision. Entropy alone chases certainty; tying it to the advisory decision keeps the recommendations focused on what changes an action.
  • Explaining beats ranking. Spelling out both outcomes of a check ("if clean, it is most likely here; if polluted, it is most likely there") lets an investigator judge a recommendation instead of trusting a score.
  • Humility has to be built in. In SourceBench about 1 in 7 localizations points at the wrong outfall, so the FHIR handover always includes a request to confirm the pipe before repair.
  • Designing around a standard early pays off. Building on the OneAquaHealth IG from the start shaped a cleaner data model than bolting on an export at the end.

What's next for Dipper: Find Urban Stream Pollution at Its Source

  • A pilot with a utility or municipality, replacing the synthetic candidate outfalls with a real outfall inventory.
  • Expert review of the likelihoods and calibration against lab results.
  • More evidence sources: overflow telemetry and low cost sensors as observers, and the OneAquaHealth distance to sewage station and faecal risk scores as priors for each outfall.
  • Better photos: number plate redaction and an expert labelled photo set.
  • Closing the loop further: native review of the translations and notifications to citizens when their case changes.

Built With

Share this project:

Updates

Submission history