Inspiration

Patients have longitudinal health records. Ecosystems don't. Water-quality evidence is scattered across papers and portals, rarely traceable to its source, and easily mixed with live or unverified data. We asked what it would look like if a lake had a health record, with every value tied to where it came from.

What it does

TraceIT is a provenance-first health record for aquatic ecosystems. It is seeded with a real published dataset, not synthetic data: Man Sagar Lake, Jaipur (4 sites x 4 months, Nov 2019 - Feb 2020; Sharma & Choudhary, 2021).

  • Provenance-linked record: per-site, per-month measurements (pH, DO, BOD, COD, TDS, nitrate and more), each citing its source.
  • Explainable anomaly screening: an ensemble of Transformer reconstruction error, forecast error, an Isolation Forest score, and a transparent branch using published screening thresholds. Every score can be broken down into its parts.
  • Multivariate analysis: PCA and clustering across sampling sites.
  • Strict data separation: published history, live context (USGS, GBIF, Open-Meteo) and citizen observations are stored and shown separately. Live data is never written into the historical record.
  • Interoperability: observations can be exported in a FHIR-style format.
  • Citizen-science workflow: accounts, saved records, logged interventions and citizen observations.

How we built it

  • Backend: FastAPI, SQLAlchemy and Alembic on PostgreSQL (Supabase), with JWT auth and scrypt password hashing.
  • ML: a PyTorch Transformer encoder (3 layers, 4 heads, 48-d latent, with reconstruction, forecast and domain-health heads) plus a scikit-learn Isolation Forest.
  • Frontend: a vanilla HTML, CSS and JavaScript dashboard.
  • Deployment and testing: Render, and a pytest suite covering auth, aggregation, clustering, triage and neural scoring.

Challenges we ran into

  • Small data: the site record has only 16 observations, too few to claim a validated deep model. We say this in the model card. The neural branch uses only the three features shared with the national CPCB benchmark (pH, DO, BOD) and runs in research/demo screening mode.
  • Keeping data types separate: published history, live context and citizen input needed clear separation in both the data model and the UI.
  • Deployment: IPv6-only database endpoints, Postgres driver mismatches and free-tier limits cost us real time.

Accomplishments that we're proud of

  • A real, cited, traceable record instead of synthetic demo data.
  • Anomaly scores that explain themselves and don't pose as a regulatory classifier.
  • A working end-to-end flow, from the published record to a citizen adding an observation, live on the web.

What we learned

Provenance matters more than model size. Being honest about what a small dataset can and can't support made the project stronger. A high score means a reading is statistically unusual. It does not establish a cause, legal responsibility, an ecological diagnosis or a health risk.

What's next for TraceIT

  • Train and validate the model on the CPCB 2012-2023 benchmark (the pipeline is already in the repo).
  • Extend to more lakes and rivers.
  • Add photo evidence for citizen observations.
  • Add expert review of observations.
  • Implement the full HL7 FHIR resource set. a

Built With

Share this project:

Updates

Submission history