Inspiration
Citizen science can generate valuable environmental observations, but asking people to assess water quality is difficult when ecological terminology is confusing and observations can be inconsistent.
We wanted to build a system where AI does not replace the citizen—it provides a calibrated second opinion. StreamPrior combines real laboratory measurements, rainfall data, and citizen observations to answer a simple question: “Does what I am seeing make sense for this location and date?”
This led to our core principle: the AI provides context; the human makes the decision.
What it does
StreamPrior is an AI-supported stream assessment platform built for the IEEE OneAquaHealth Global Hackathon.
A citizen can:
- Select an OAH monitoring site or location.
- See the expected water clarity for that location and date.
- View calibrated probabilities for clear, cloudy, or muddy conditions.
- See recent rainfall and the number of real laboratory samples supporting the expectation.
- Report what they actually observe.
- Receive an explainable “second look” verdict: consistent, plausible, or unusual.
- Choose whether to keep or change their original observation.
- Export the finalized observation as a FHIR R4 Bundle, with the human decision recorded in
Provenance.
StreamPrior also includes a Storm Watchlist that ranks live OneAquaHealth sites by their probability of muddy water and highlights locations that should be rechecked after storms.
How we built it
The model is trained using 16,177 real laboratory suspended-solids measurements from 150 stations, combined with historical rainfall data from Open-Meteo.
The laboratory data comes from Hub'Eau's MES dataset, which provides suspended-solids measurements used as a proxy for stream clarity.
We use grouped 5-fold cross-validation so that stations do not appear in both training and test sets. The resulting model produces calibrated probabilities rather than a simple binary prediction.
The platform is built around:
- Python 3.12 for the backend and ML pipeline
- Hub'Eau for laboratory measurements
- Open-Meteo for rainfall data
- OneAquaHealth API for live monitoring sites
- FHIR R4 for interoperable observation exports
- Leaflet/OpenStreetMap for geographic visualization
- Automated verification and reproducible evaluation
The trained model, held-out predictions, evaluation metrics, and verification checks are included in the repository.
Challenges we ran into
The biggest challenge was building something that was useful without pretending the model was more reliable than the evidence allowed.
Our training data comes from south-west France, while OneAquaHealth operates across multiple countries. We therefore had to explicitly distinguish between what the model demonstrated and what remained an extrapolation.
Citizen labels were another challenge. Public citizen observations were not available for broad validation, so we did not claim that the model beats rainfall or laboratory data for citizen observations. Instead, we included a small pilot and exposed its limitations directly.
We also had to handle the interoperability challenge of converting observations into a meaningful FHIR R4 structure while preserving the human decision through Provenance.
Accomplishments that we're proud of
- Trained on 16,177 real laboratory samples across 150 stations.
- Achieved 0.832 AUROC for detecting muddy conditions.
- Improved AUROC by +0.196 over the generic 20 mm rainfall rule for muddy conditions.
- Achieved a calibration error of 0.009 for the muddy target.
- Loaded 106 live OneAquaHealth sites with coordinates.
- Implemented a live Storm Watchlist for resilience monitoring.
- Built an explainable human-in-the-loop assessment workflow.
- Implemented FHIR R4 export with 0 base-validator errors.
- Added automated claim verification with 14 PASS, 0 FAIL, and 1 expected SKIP.
- Most importantly, the system explicitly communicates its limitations instead of hiding them.
What we learned
We learned that environmental AI is not just about achieving a high prediction score. Calibration, provenance, uncertainty, geography, and human interaction matter just as much.
A model can perform well on held-out laboratory stations while still having limited evidence for different countries or for citizen-generated labels. Making those boundaries visible is essential when building systems intended to influence environmental decisions.
We also learned that interoperability should be considered from the beginning. Recording the citizen's final decision and exporting it through FHIR makes the observation more useful beyond the application itself.
What's next for StreamPrior
The next step is expanding validation beyond the current south-west France training region.
We want to:
- Collect a larger dataset of citizen observations paired with laboratory measurements.
- Validate the model across different countries and climates.
- Improve calibration for high-probability predictions.
- Investigate better proxies and direct measurements for urban-stream clarity.
- Expand the Storm Watchlist into a broader early-warning and resilience system.
- Strengthen FHIR integration with OneAquaHealth profiles when the required profile information is available.
- Continue improving the citizen experience so environmental monitoring becomes easier, clearer, and more engaging.
The long-term goal is to turn individual observations into a reliable environmental intelligence layer connecting ecosystem health, biodiversity, and human well-being.
Built With
- api
- fast-api
- gradio
Log in or sign up for Devpost to join the conversation.