Inspiration

City streams are some of the most overlooked water in a city. People walk past them every day, but almost nobody writes down what they see: trash on the bank, a pipe that should not be there, water that turns brown after rain. Scientists cannot walk every meter of every stream every week, so most of that story is lost.

Citizen science apps try to fix this, but they usually run into the same four problems. People try them once and never come back. Volunteers check the same easy spots again and again, so long parts of the stream are never seen. Old checks pile up and say little about the stream today. And a single person's answer is hard to trust without a second opinion.

We kept asking one question: games like Pokémon GO already get people to walk to specific places, over and over, for weeks. What if the game rules were the sampling plan?

What it does

StreamRealm is a territory game played on real urban streams in Coimbra, a OneAquaHealth research city. Each stream is cut into 100 m tiles. Players join one of three teams (Otters, Frogs or Kingfishers), walk to a tile, and claim it by doing a quick stream check: a safety confirmation, one photo upstream, one photo downstream, five simple questions (colour, smell, foam, trash, flow) and an overall feeling.

Every game rule quietly pushes players towards the data scientists need:

  • Fog tiles (never checked) give an explorer bonus, so players go where nobody has looked.
  • Land fades. A tile stays fresh for a week, then fades, then goes back to neutral:

$$ \text{state}(a) = \begin{cases} \text{fresh} & a \le 7 \text{ days} \ \text{fading} & 7 < a \le 14 \text{ days} \ \text{neutral} & a > 14 \text{ days} \end{cases} $$

where $a$ is the age of the last check. Defending your land means checking it again, so the data stays fresh.

  • Attacking an enemy tile is a full new check. If it agrees with the old one, the first check becomes confirmed. If it disagrees, a dispute opens, and a third check or a scientist settles it.
  • Storm quests double the points when heavy rain is forecast, because that is when sewers overflow and streams change the most:

$$ \text{points} = m \cdot \left(b + \sum_i \text{bonus}_i\right), \qquad m = \begin{cases} 2 & \text{storm quest active} \ 1 & \text{otherwise} \end{cases} $$

  • Treasures (pipes, trash, wildlife, strange plants, algae) become a work list for scientists. When a problem is marked as fixed, the team's kingdom gets healthier, so players see real fixes show up in the game.

Every check also gets a second opinion. Simple photo rules catch blurry, dark, re-used or old photos, and an OpenAI vision model suggests answers ("AI thinks: a lot of trash, 99%"). The AI never changes anything on its own. The player taps Use this or Keep mine, and both answers are stored, so scientists can see exactly where people and the computer disagree.

On the other side, scientists get a dashboard with coverage, data age and confirmed checks, a dispute queue, the treasure work list, and one-click exports as CSV, GeoJSON and a FHIR R4 bundle that uses the HL7 Europe OneAquaHealth profiles. Here, coverage means the share of tiles with a recent check:

$$ \text{coverage} = \frac{\left|{\text{tiles checked in the last 14 days}}\right|}{\left|{\text{all tiles}}\right|} $$

How we built it

  • App: Expo (React Native) with expo-router, running on the web first and on Android and iOS from the same code. The map uses MapLibre GL on the web, with real stream geometry from OpenStreetMap.
  • Server: FastAPI with SQLModel and SQLite. It handles tiles, checks, disputes, quests, the leaderboard, storm detection from the Open-Meteo rain forecast, and the exports.
  • Photo check: heuristics (blur, brightness, duplicate photos, photo age from EXIF) plus an optional OpenAI vision call with a strict JSON schema. If the AI is slow, missing or fails, the app falls back to the photo rules instead of breaking.
  • Look and feel: a "Kingdom UI" kit of wooden panels, stone tablets, ribbons and chunky 3D buttons drawn in code, plus a pack of game images that we processed with our own script (clean edges, trimming, team colours and three screen densities).
  • Demo tools: a Dev Panel lets judges play from a laptop. You can walk with W A S D, teleport, jump days ahead with Time Warp, force a storm, and let bot players make moves.

Challenges we ran into

  • Making the game honest. It is easy to make a score that looks scientific. We had to keep reminding ourselves that the tile health score is a game indicator, not "safe to swim" advice, and to label it that way everywhere.
  • Keeping AI in the passenger seat. Our first instinct was to let the AI correct answers. We decided against it: the AI only suggests, the player decides, and both answers are kept. That turned a risk into useful data about where people and models disagree.
  • Slow AI responses. A real vision call on two photos takes 10–20 seconds, longer than our first timeout. We had to design the waiting screen and the fallback carefully so a slow call never blocks a player.
  • Fitting citizen checks into FHIR. The OneAquaHealth guide has no profile for visual citizen checks, and its observation profile expects final results. So unconfirmed checks go out as preliminary base R4 observations, and some answer codes are local. The guide's build page was offline, so we built the profiles from the source to validate against.
  • Maps on a bad day. The base map tiles sometimes stopped loading. We added an automatic switch to plain OpenStreetMap tiles so the game always shows the stream, even when the pretty map does not.

Accomplishments that we're proud of

  • The game rules and the sampling plan are the same thing. Exploring, defending, attacking and storm hunting all produce data scientists actually want: wider coverage, fresher checks, independent second checks and data at the most important moments.
  • A full loop from player to scientist and back: a player reports a pipe, a scientist marks it fixed, and the team's kingdom grows.
  • The AI check works end to end with real photos. It spotted trash piled on a bank and a grate at an outlet, and it rejected a photo of a desk as "not a stream photo".
  • The FHIR export passes the official HL7 validator with 0 errors.
  • It looks and feels like a real game, while the scientist dashboard stays plain and professional on purpose.

What we learned

  • Good citizen science design is mostly about where people go and when they come back. Game mechanics are a surprisingly precise tool for shaping both.
  • A second opinion is more useful than an automatic correction. Keeping both the human answer and the AI suggestion makes the data more trustworthy, not less.
  • Standards like FHIR are worth the effort, but you learn quickly where real-world data does not fit the profiles yet.
  • Honesty is a feature. Saying clearly what is a demo, what is a game indicator and what still needs testing makes the project easier to trust.

What's next for StreamRealm

  • A real pilot in Coimbra, to measure what really matters: do people come back, how much of the stream gets checked, and how fresh the data stays.
  • Connecting to the official OneAquaHealth Citizen Science App and Resilience Map, with shared tiles and single sign-on.
  • Automatic face and licence-plate blurring on uploaded photos.
  • Portuguese and other languages, proper accounts, and school and club leagues.

Built With

Share this project:

Updates

Submission history