Inspiration
My day job is cloud security engineering. Most of it comes down to one question: can this signal be trusted enough to act on? A login from a new country, a key used at 3 a.m. You don't throw the alert away and you don't act on it blindly. You look again, add context, and put a person in the loop when the stakes go up.
Citizen stream data has the same shape. The Track 3 brief says citizen observations can be inconsistent and error-prone, and that AI should support assessment without replacing human judgment. I read that as a triage problem, not a prediction problem.
One paper set the boundary for everything I built. McMurray et al. (2026, PLOS ONE) scored four urban streams with the NRCS Stream Visual Assessment Protocol and found the visual scores tracked measured physical features well but had few significant correlations with water quality. So a visual assessment can tell you a lot about habitat and almost nothing about whether the water is safe. Any tool that implies otherwise is lying to the volunteer.
What it does
A volunteer answers eight plain-language questions ("Are the banks holding firm?"), each with the scientific term shown underneath ("Scientists call this bank stability"), plus smell, pipes, recent rain, minutes on site, and whether a photo is attached. Before submitting, they press Check my answers.
Riffle then runs four kinds of check:
- Consistency. Answers that contradict each other, like a dry stream bed with water rated clear to the bottom.
- Effort. Every question given the same score, or eight checks done in under four minutes.
- Site history. An answer far from what this site usually reports, using a modified z-score against at least five past visits.
- Context. Rain in the last 48 hours makes cloudy water expected, so a clarity drop after rain is softened and the reason is stated.
Each flag shows its explanation, a prompt written as a question, and exactly how much it added to the review priority. The volunteer can change the answer or press I re-checked, keep my answer. Keeping it halves that flag's weight. The flag stays on record.
Below 20% priority the report is accepted. Between 20% and 50% it is accepted with a note for researchers. At 50% or more it goes to the reviewer queue, where a person approves or rejects. Rejection needs a written reason.
Separately, smell, outfall and algae answers generate One Health notes: sewage signs mean avoid skin contact and keep dogs out; a thick bloom means some algae produce toxins harmful to pets and people. Every note ends with the same line: a visual check cannot confirm water quality either way. These reports also enter the queue, because a trustworthy sewage report still needs someone to act on it.
Every submission exports as an HL7 FHIR R4 Bundle: a Location, twelve Observations and a Provenance resource.
How I built it
Plain JavaScript modules, HTML and CSS. No framework, no build step, no API keys, so a judge can clone and run it in under a minute and a city could host it on any static server.
The engine is a set of pure functions. It takes the assessment, the site's history and the set of re-checked flags, and returns flags, a priority and a decision. It never mutates the input, and a test proves that. Priority is a noisy-OR, 1 − ∏(1 − wᵢ), which I chose because each weight stays readable on its own and the total can't pass 100%.
The FHIR export is where the human-in-the-loop design pays off. FHIR already has the vocabulary: Observation.status is preliminary until a reviewer approves (final) or rejects (entered-in-error, kept, never deleted). Provenance lists the volunteer as performer, Riffle as assembler and the reviewer as verifier. A downstream system can tell machine-triaged data from human-verified data without reading any Riffle-specific field.
Accessibility was part of the build, not a pass at the end: real fieldsets and legends for every question, keyboard-operable tabs, an aria-live feedback region, and colour never used as the only signal.
Challenges I ran into
The test suite caught two bugs I would have shipped.
The first one surprised me. On a 1 to 5 scale, a site that reports 4 for habitat on seven visits and 3 on one has a median absolute deviation of zero. The fallback formula then gave a one-point change a modified z of −6.4, far past the 3.5 cut-off. Every ordinary visit would have been flagged as an outlier. I added a second condition: the answer must also be at least two points from the site median.
The second was floating point. One flag with weight 0.2 produced a priority of 0.19999999999999996, which fell just under the 20% threshold. I now round before comparing.
The harder challenge was restraint. Every check I added makes the tool more accurate and the volunteer's experience a little worse.
Accomplishments that I'm proud of
The contradictory demo case raises six flags and an 82% priority. Re-checking and keeping the first answer drops it to 76%, still in the queue, with the volunteer's decision recorded in the FHIR Provenance. That chain is the whole idea working end to end: the machine asks, the person decides, the record shows both.
The rain context too. The same cloudy, fast-flowing reading at the walled demo site scores 36% on a dry week and 10% after rain, and the volunteer is told why.
Fourteen automated tests cover the engine and the FHIR mapping, and a headless-browser run checked the full flow from form to reviewer approval to export.
What I learned
The argument against my own project deserves a straight answer. Participation in citizen science is already fragile; Track 1 exists because people drop off. A tool that questions volunteers could make that worse. I tried to soften it with prompts phrased as questions, a one-click "I re-checked", and checks that only fire on combinations that really are unlikely. But I haven't tested it with a single real volunteer, and until someone does, I can't claim it helps more than it discourages.
And the weights are hand-set. 0.4 for a dry bed with clear water, 0.25 for firm banks with no plants: these are my judgments, written down and visible, not learned from data. That is honest, and it is also a limitation.
What's next for Riffle
- Map the demo questions to the OneAquaHealth Citizen Science App fields. My protocol is modeled on public visual assessment protocols, not the app's own schema.
- Calibrate the weights against paired volunteer and expert assessments of the same site on the same day.
- Send bundles to a FHIR server instead of downloading them, then feed approved data into city dashboards.
- Add reviewer authentication and pseudonymous volunteer IDs. The prototype stores everything in the browser and has no login, which is fine for a demo and not for real data.
Track: Track 3, AI-Supported Assessment
Reference: McMurray et al. (2026). Evaluation of stream visual assessment protocol with measured water quality parameters in urban streams. PLOS ONE. https://doi.org/10.1371/journal.pone.0351972
Built With
- css3
- fhir-r4
- github
- hl7-fhir
- html5
- javascript
- node.js
Log in or sign up for Devpost to join the conversation.