🌊 Second Look
See what your creek is hiding.
🎯 Track 3, AI-Supported Assessment. The track says citizen observations can be inconsistent and error-prone. Second Look measures that, per person and per feature, and saves the measure with every observation.
Second Look adds one thing OneAquaHealth's citizen science workflow does not have yet: evidence about how well each observer sees. A two-minute photo test teaches volunteers the creek damage most people miss, scores them on it, and attaches that score to every observation they make, in OneAquaHealth's own data standard.
TRAIN → CHECK → VERIFY → RECORD → ACT
👉 Try it, no camera needed: https://second-look-79t.pages.dev 🧑⚖️ Judges start here: https://second-look-79t.pages.dev/judges
🚨 The problem
Most of us judge a creek the way we judge a park.
Tidy and green looks healthy.
The water sparkles.
The path is clean.
And the creek runs through a concrete channel.
At the first hackathon workshop, OneAquaHealth's project lead named exactly what volunteers miss: concrete margins, deepened channels, and attractive invasive plants. They notice smell, foam and colour. They walk past the rest.
So the best-looking creek can get the best rating and deserve the worst.
Professional river surveyors solved this long ago. In the UK's River Habitat Survey, only surveys from accredited surveyors are entered on the database, and accreditation means passing a test.
Volunteers never had that test.
Their observations arrive with no mark of how far to trust them.
🔎 What Second Look does
- 🖼️ One question first. Two creek photos: Which creek is healthier? Most people pick the tidy one. It is a concrete channel.
- 🎓 A two-minute lesson on the four things people miss: built banks, a dug-out channel, plants that do not belong, and pipes. Every photo marks the giveaway.
- 📝 A 16-photo test with known answers: four per feature, two with the feature and two without, so guessing scores about half.
- 📊 A score per feature, like 4 of 4 on built banks, that travels with every creek check for 90 days.
- 🏞️ A guided creek check that asks the official OneAquaHealth app's own questions, in its own words, in six languages, one question per screen, and works offline in a gully with no signal.
- 🌦️ Smart follow-ups, chosen by code, never by a model. At most two, from your answers, your score and the weather: > It has not rained here for 9 days. Is anything coming out of that pipe?
- 🤖 An AI checker that has to earn the right to ask. More on that below.
- 📋 A record a city can trust. Every answer sits beside the score of the person who gave it.
- 🏙️ A city page that turns records into what each creek needs, in OneAquaHealth's own restoration measures, and lists the pipes worth testing.
🧠 Why the AI can only ask
People look. Code checks. AI may only ask.
That separation is the core of Second Look.
Four AI models took the same 16-photo test as people, three runs each. A model passes a feature only if it gets all four photos of that feature right in at least two of three runs. Each model passed some features and not others, and the pass table is committed in the repo.
The model is free to:
👀 look at a photo 🏷️ propose a flag 💬 attach a short note
But before a flag can reach a person:
✅ the model must have passed that exact feature on the same test people take ✅ the gate must accept the flag's shape, feature and note ✅ the person must already have answered
Then it may ask once: The checker noticed something. Look again? Keep or Change. The person decides.
The model never writes an answer, never sets a label, and never puts its own words in front of a person as fact.
It cannot talk its way past the gate.
🛡️ The trust envelope
🔐 No accounts, no names, no emails, no free text in the test, no addresses in any log
📵 Works offline at the creek and sends later
🧾 Every record gets a receipt in a hash-chained audit log, checkable at /verify
⏱️ The study plan was tagged before any participant, and timestamped in Bitcoin with OpenTimestamps
🔍 Every number in the README is checked by code against results/ in CI
🌍 Every photo and footage clip is openly licensed and credited
🩺 Every health or ecology sentence comes from an approved list with its source, and none states a risk for a specific site
📋 The record
Every check becomes FHIR R4 under OneAquaHealth's own implementation guide: nested Locations for creek, reach and spot, an Observation per answer, the volunteer as a pseudonymous Practitioner whose qualification is their dated per-feature score, and Provenance linking every answer to the score of the person who gave it.
A lab result is trusted because its quality checks travel with it.
Now a volunteer's answer carries the same thing.
✅ Validated with the HL7 validator against OneAquaHealth's package (hl7-eu/oah, pinned commit b907cf0), terminology checks on: 0 errors
🏙️ What a city gets
🌿 What this creek needs, in OneAquaHealth's own restoration measures from their 2026 Policy Brief: replant both margins with native vegetation, fix leaking sewers, reconnect the floodplain, remove barriers, take out the concrete.
🚰 Pipes worth testing: a pipe reaches this list only when two separate contributors who both passed the pipe feature saw it running after dry weather. It becomes a FHIR referral, a reason to take a sample, never a finding about the water.
🧪 The way back: an example shows how a lab result, including the microbial metagenomics on OneAquaHealth's roadmap, would return to the same record.
⬇️ Downstream notes: a finding on one reach adds a plain line to the reaches below it.
🐾 A health card with one action each for the person, the pet and the city, every sentence from a cited source.
🎬 Demonstrated on real creek footage
There is no staged field visit. Every creek check in the demo is a desk check on openly licensed real creek footage, labelled as one:
📹 46 frames from 5 videos in 3 countries, screened for people and text 🤖 The models judged every frame 🚧 The gate stopped 29 of 64 candidate flags, because a model may flag only a feature it passed 🗺️ Three walks let a judge do the full creek check from a desk and get a validated record at the end
🧨 We also tried to break it
A demo that works once is not enough.
🙈 The answer key leak. Our own review found the walk pages shipping all 16 test answers in the page's code. Fixed, and a build check now fails if an answer ever reaches the browser. Judge mode stays shut until the study's lock for the same reason.
🧮 Our own idea failed. We expected scores to make group votes more accurate. Our simulation said four photos per feature is too coarse to weight votes with. We dropped the claim and published the simulation.
🔤 YAML ate our answers. The first paid model run told the models the allowed answers were true and false, because unquoted yes and no are booleans in YAML. Caught, quoted, pinned by a test, and rerun.
🔁 Lost answers. A refresh resumes the test, and answers lost on a bad connection are resent before the score appears.
🧬 Mutation testing on the gate, the follow-up selector, the scoring and the FHIR emitter, so the tests catch real bugs and not just pass.
🖼️ Adversarial frames. Blank frames, indoor scenes and screenshots must produce no flag.
🧑⚖️ Simulated judges. Agents used the live site the way a judge would, on a phone and a laptop. Everything they confirmed was fixed or listed in the README as a known weakness.
🗺️ How OneAquaHealth powers Second Look
OneAquaHealth is not a decorative integration. It is the standard we write to and the workflow we extend.
📱 Citizen Science App: its questions, in its order and its own six translations, including the closing question on how you feel
🧬 Implementation guide: Location and Observation profiles, value sets, UCUM units, a Questionnaire through their form extension, nested Locations with partOf
✅ HL7 validator with their package, in CI, for every record we emit
🗄️ Sandbox: tagged conditional writes and a Library entry describing our data set (their sandbox name stopped resolving on Sep 23; the ledger, validator output and screenshots from before that are in the repo)
🌿 Policy Brief: the restoration measures on the city page
🏙️ Follower city recipe: make new-city sets up a new city in one command
🌐 One Digital Health and FAIR: scored with their own template
🧰 Built with
App ▲ Next.js 💙 TypeScript ⚛️ React 📱 installable PWA with an offline queue
Platform ☁️ Cloudflare Pages 🛠️ Cloudflare Workers 🗃️ D1 💸 free to host, no card needed
Standards and data 🧬 FHIR R4 🏥 HL7 validator 🌊 OneAquaHealth implementation guide 🌦️ Open-Meteo 🌱 Cal-IPC 🐸 iNaturalist 🎞️ Wikimedia Commons
AI and evaluation 🤖 Anthropic Claude vision models 🐍 Python 🧪 pytest 🎭 Playwright 🧬 mutmut ⏱️ OpenTimestamps 🔌 a read-only MCP server
How it was built: with AI coding tools (Claude Code and Codex). The video's illustrations were made with ChatGPT image generation and its narration with ElevenLabs; everything that shows the app or a creek is real.
🔬 Reproducible without trusting us
🧑⚖️ make judge-check runs with no API key and no network: the tests, FHIR validation of the committed records, the web build, the audit log check and a secrets scan.
♻️ make reproduce re-grades every number in results/ from the raw model replies and seeds, and fails if one has drifted.
✅ Over 2,000 automated tests, and every README number checked in CI.
⏱️ /verify shows any record's receipt, its place in the audit chain, and the timestamp proof.
🌱 Contributing back
📬 To OneAquaHealth's implementation guide (hl7-eu/oah): a citizen-contributed example with a proposal for carrying observer quality, plus issues for what our validator runs found.
🗣️ To the organizers: 14 places where the official app's translations say something different from the English, such as the Italian and French pipe questions asking about rainwater instead of polluted water.
💡 What we learned
A citizen observation is only as useful as the trust you can put in it.
Without a measure of the observer,
"no concrete on this bank"
is a guess.
With:
📊 a dated score on that exact feature,
🧬 a record that carries it in the city's own standard,
and 🧑 a person who answered before any AI spoke,
the same sentence becomes something a city can act on.
That led to the design we ended up trusting:
People look. Code checks. AI may only ask.
🚀 What's next
📸 A larger, rotating photo bank, so the score sharpens with every retake and cannot be memorized 🏙️ A field pilot with a OneAquaHealth city 🧬 A citizen observer profile in the guide, if its authors want one 🌿 Local invasive plant lists for every pilot region 🧪 Real lab results flowing back to the volunteers who flagged the pipe 📱 An optional two-minute training step inside the official app
🌊 Second Look
Not another app that collects more observations.
An app that tells a city how far to trust each one.
Built With
- claude
- claude-code
- cloudflare-d1
- cloudflare-workers
- codex
- fhir
- fhir-shorthand
- hl7
- inaturalist
- mcp
- next.js
- oneaquahealth
- open-meteo
- opentimestamps
- python
- react
- typescript
Log in or sign up for Devpost to join the conversation.