Inspiration

Rapid urbanization and climate change pose immediate threats to urban freshwater ecosystems. While massive amounts of environmental data, scientific papers, and citizen observations exist, they are completely fragmented across PDFs, websites, video content, and social streams. We were inspired by a simple question: What if anyone—from a local student to a city policy maker—could talk directly to our water systems using natural language?

We built AquaAsk to democratize access to environmental intelligence, turning chaotic, multi-modal data streams into a unified, high-value conversational search engine for the "One Health" mission.

What it does

AquaAsk is an advanced, multi-modal Retrieval-Augmented Generation (RAG) search engine designed specifically for urban water sustainability. Users interact with a clean, immersive search bar to ask complex ecological questions.

Behind the scenes, the engine dynamically searches local document repositories, processes web URLs, parses YouTube video transcripts, scans emails, and hooks into live internet search feeds when real-time validation is needed. It synthesizes these inputs to instantly generate comprehensive answers complete with verifiable, inline markdown citations—bridging the gap between data and actionable community insight.

How we built it

We separated the development into an agile, high-performance architecture:

  • Frontend: A bespoke, immersive dark-themed user interface utilizing custom semantic search modules, engineered for immediate user engagement and minimalist data friction.
  • Backend Framework: Built as a production-ready, asynchronous Python backend powered by FastAPI and orchestrated with LangChain (langchain-xai and langchain-google-genai).
  • Vector Architecture & Embeddings: Documents are broken down using a RecursiveCharacterTextSplitter. We generate deterministic SHA-256 content hashes for idempotent database upserts, embed them using Google Gemini Embeddings (models/text-embedding-004), and index them in a local, persistent ChromaDB store using HNSW cosine distance tracking.
  • Hybrid Search Engine: Implemented an automated routing threshold. If the top local vector distance exceeds 0.45, or if local data is missing, the engine automatically triggers a live web search API fallback with strict 5-second timeout blocks.
  • Generation Layer: Powered by xAI Grok (ChatXAI) configured with custom system prompts enforcing zero hallucinations, strict grounding, and mandatory inline metadata citation tracking ([Source: origin_name]).

Challenges we faced

  • Universal Parsing Conflicts: Standardizing incoming text structures from fundamentally different mediums (like unstructured web layouts, timestamped video captions, and heavily formatted data tables) required building a strict DataIngestionManager with custom text-cleansing guardrails to prevent token bloat.
  • Dynamic Fallback Thresholds: Finding the perfect sweet spot for the cosine distance cutoff (0.45) required extensive iteration. Setting it too low ignored valid local data, while setting it too high failed to bring in live web context when stream metrics changed over time.
  • Frontend-to-Backend State Bridging: Transitioning our custom static frontend components to asynchronously POST queries to the FastAPI /api/search endpoint required robust asynchronous javascript handling to update our sleek answer card in real-time.

Accomplishments that we're proud of

We successfully constructed a fully functioning, end-to-end multi-modal RAG engine in an incredibly tight development window. We are immensely proud of combining cutting-edge models like xAI Grok and Google Gemini into a unified pipeline that turns complex scientific context into human-readable, deeply educational ecosystem narratives.

What we learned

We gained deep expertise in managing high-fidelity document metadata through complex vector parsing pipelines. More importantly, we learned that for artificial intelligence to safely support environmental monitoring, human-in-the-loop explainability and verifiable inline citations are non-negotiable foundations for building civic trust.

What's next for AquaAsk

We plan to scale AquaAsk by fully integrating Digital Health Standards (FHIR models) to interconnect localized river telemetry data straight into municipal public health registries. We also aim to implement local community gamification features, rewarding users with digital stewardship badges when they actively paste and contribute verified stream metrics into the platform.

Built With

Share this project:

Updates

Submission history