Inspiration
Freshwater invertebrates are a reading instrument that costs nothing. The mix of creatures under a stream rock shifts weeks before a lab panel registers a problem, and anyone willing to turn that rock can take the measurement. Projects like miniSASS have already proven people will do this, at scale, for years.
So the shortage is not enthusiasm, and it is not science. We kept landing on what happens after the reading. A citizen records something real, and then it stops. Nobody downstream knows whether that particular record was careful or careless. Nothing translates an ecological signal into what it might mean for the people and animals living alongside that water. And almost nothing reaches the systems that could act, because health services speak FHIR and a citizen-science app speaks CSV.
That gap is what Bahari was built for. We deliberately did not try to beat existing tools at photo identification. We took the reading they already produce and carried it the rest of the way.
What it does
A citizen picks a stream, records what they find, and gets an interoperable health record. Eight things happen along that path.
Guided field assessment. Real monitoring points from the GBIF open biodiversity database, on a map colored by latest health, each with a trend line so a returning observer sees change. Recording is tapping plain-language creature groups, with a tip for anyone unsure.
An AI that answers second. The citizen identifies first. Only then does Amazon Nova Pro look at the photo, and it never receives the citizen's answer. It returns its own identification, its visual cues, and a confidence value, and Bahari reveals whether the two agree. Disagreement shows the distinguishing cues and asks the citizen to look again. Below 0.5 confidence the model declines to guess. The model cannot alter a count, and no score is ever computed from an unconfirmed suggestion.
A published ecological score. Biotic-index scoring in the miniSASS and ASPT family, with family weights annotated against their sensitivity ordering, and the arithmetic shown on screen rather than hidden.
A separate reliability score. This is the piece researchers asked for. It rates the submission rather than the stream, across location evidence, groups examined, photo evidence, count plausibility, and time spent. Weak data gets flagged instead of discarded.
Honest One Health translation. Qualitative considerations for human and animal health, each stating why it was raised, what it rests on, and its limits. Bahari never claims water caused a disease. It is an indirect indicator and says so.
A FHIR record that follows the official IG. The assessment exports as a FHIR R4 Bundle shaped to the OneAquaHealth implementation guide: the monitoring site as a Location, readings as Observations against the OAH indicator profiles, the sample as a Specimen, and One Health as a qualitative Observation. An evidence inspector reads the generated bundle and shows the profile mapping, the references between resources, and the human and model identifications as separate provenance. It validates with zero structural errors on the public HAPI FHIR R4 server (structural validity and IG alignment, not full IG profile conformance, which a public server cannot check because it cannot load the draft OAH IG).
A researcher trust console. Assessments aggregate into a catchment view, and reliability works as a control rather than a label. Turning on trusted-only changes the counts and reorders the most-stressed-first list, so a weak-evidence alarm cannot outrank a well-evidenced one. Excluded records are marked needs recheck with their specific reasons.
Local-first field operation. Assessments save to the device as they are entered. The app opens and resumes with no network. Records and photos stay on that device, with no account and no cloud sync, until the person chooses to export. The deployed site is entirely static.
Interface is English and French throughout, honors reduced-motion preferences, and holds AA contrast in both light and dark themes.
Tracks
Bahari fits four of the seven tracks: Track 1 (Citizen Science UX), Track 2 (Data-to-Insight), Track 3 (AI-Supported Assessment), and Track 7 (Digital Health Standards).
- Track 1: a guided, plain-language flow with simplified ecological terms, inline tips, a per-stream trend for repeat engagement, and English/French throughout.
- Track 2: a researcher dashboard with a map, aggregate health distribution, trends, and a most-stressed-first table, plus a One Health insight layer.
- Track 3: an explainable, human-in-the-loop AI that answers second, shows its reasoning and confidence, and defers below a stated threshold.
- Track 7: a FHIR R4 export aligned with the official OneAquaHealth HL7 FHIR Implementation Guide.
How we built it
React 18, Vite 6, Tailwind 3, and Leaflet, with no backend in production. The deployed artifact is static files, which keeps hosting close to free for a community or municipality.
Scoring, reliability, One Health rules, and FHIR construction live in pure modules under src/core/ with single ownership. One module owns the ecological score. One owns reliability factors and weights. One owns bundle construction and validation. UI code reads those outputs and is not permitted to patch them, which is what let us change presentation repeatedly without putting the science at risk.
Amazon Nova Pro runs through Bedrock's Converse API. Because Bedrock requires SigV4 and has no browser-safe key, the call goes server-side through Vite middleware during local development, and the browser posts image bytes and format to it with no taxonomy hint attached. A deterministic stub provider sits behind the same interface, so the deployed static build and any credential-less environment still exercise the full flow. The AWS SDK stays in devDependencies and never enters the client bundle, which we verify by scanning the build output.
Local-first storage uses IndexedDB for records and photo blobs, with localStorage and in-memory adapters behind one interface for constrained browsers. A generated service worker precaches the shell with a deterministic build id derived from emitted asset hashes. Updates activate on reload only, after the current assessment is durably saved, and old caches are retained until the reloaded client confirms readiness.
147 tests across 16 files cover the science modules, the confidence policy, reliability, FHIR shape, accessibility via axe, and exact English and French dictionary parity.
Challenges we ran into
The FHIR modeling was wrong before it was right. Our instinct was to express the One Health layer as a RiskAssessment. Base R4 restricts RiskAssessment.subject to Patient or Group, and our subject is a stream. We briefly modeled the stream as a Group, which validated but was semantically false, since the OAH guide uses Group for human cohorts. The honest answer was a qualitative Observation with the Location as subject, matching the Observation-centric design the IG actually specifies. We also had a placeholder bahari.example code system in the first pass, which would not have survived a FHIR-literate reviewer.
Bedrock has no browser-safe credential. There is no equivalent of a restricted API key, so a purely client-side call was never possible. Rather than add a backend and lose static hosting, we put the call behind local dev middleware with a provider seam and a deterministic fallback. The consequence is honest and worth stating: the deployed site runs the stub, and the live model is demonstrated locally.
Blind ordering without a longer flow. A genuinely blind second opinion usually means a separate answer-first step, which would have added a screen and duplicated the counter. We used the citizen's existing positive count as the hidden answer and withheld it from the provider, which preserved five steps while making the comparison meaningful.
Review found a count-loss bug we would have shipped. The trailing 600ms save read the current record through a shared reference when the timer fired. Resuming a different record inside that window wrote the wrong record and lost the outgoing counts. Each pending save now captures its own immutable record and id, and every transition flushes or deliberately cancels it.
Honest competitive positioning. We searched the public field and found about fifteen projects with direct event evidence and more than forty probable ones. AquaPlot in particular covers much of the same route. That forced us to rewrite our own claims. We are not first to combine citizen science, explainable AI, One Health, reliability, and FHIR, and saying otherwise would be easy to disprove.
Accomplishments that we're proud of
The FHIR export follows the official OneAquaHealth guide rather than inventing a plausible-looking shape, and the evidence inspector lets a reviewer check that claim in the interface instead of taking our word for it.
Reliability changes a decision. Many projects display a trust number. Switching trusted-only in Bahari reorders which streams a researcher looks at first, and tells them exactly what evidence is missing from the records that dropped out.
The blind second opinion is a real guardrail rather than a label. The model answers without seeing the human answer, cannot write to a count, and declines below a stated threshold. Agreement therefore carries information.
It works where the work happens. No network, no backend, no account, no hosting bill, with records kept on the citizen's own device.
And the restraint. The One Health layer states its basis and its limits on every consideration, and the validation wording claims structural validity and IG alignment rather than formal profile conformance, because that is what we actually verified.
What we learned
Adversarial review paid for itself. Automated review rounds surfaced ten blocking design gaps and then ten implementation findings, including the count loss, that passing tests and a working demo never revealed. Everything green is not the same as everything correct.
Reading the specification beats pattern-matching on it. RiskAssessment looked right and was wrong, and the only way to find that was in the IG's own text.
Competitor research changes the product and not just the pitch. Finding AquaPlot pushed us toward the static, keyless, decision-oriented framing that is genuinely ours.
Honest AI design is a feature, not a tax. Every constraint we accepted, blind ordering, confidence gating, no writes to counts, made the output more credible rather than less useful.
And offline is mostly a correctness problem. Caching was straightforward. Not losing a citizen's work during a save, a resume, or a service-worker update was the hard part.
What's next for Bahari
Explicit location capture, so the reliability score can use real submission evidence rather than treating it as absent.
Formal IG profile validation against a server with the OneAquaHealth guide loaded, which would let us upgrade the wording from alignment to conformance.
An optional, consent-first sync for groups who want shared catchment data, built behind the storage interface so local-first stays the default rather than a fallback.
More locales, added only with native review. We would rather ship two trustworthy languages than seven machine-translated ones.
Field calibration with a monitoring group, comparing citizen submissions against paired lab results to tune the reliability weights against real outcomes instead of reasonable assumptions.
Built With
- amazon-bedrock
- amazon-nova
- fhir
- gbif
- indexeddb
- javascript
- leaflet.js
- openstreetmap
- react
- tailwindcss
- vite
- vitest
Log in or sign up for Devpost to join the conversation.