Inspiration

My inspiration came from watching a reel about a community living near a large data center where residents discussed concerns about local water availability and water quality. Regardless of whether that story represented every data center, it raised an important question: as AI continues to grow, who is thinking about the water required to cool the infrastructure behind it?

Modern AI data centers consume enormous amounts of electricity, and many rely on water-intensive cooling systems to keep thousands of GPUs operating safely. The industry has made real progress optimizing compute performance and energy efficiency, but water consumption remains a less visible, equally critical sustainability challenge.

I wanted to build a system that doesn't just monitor infrastructure, one that continuously learns from operational history and helps operators make smarter, more water-efficient cooling decisions. That idea became Aqua Rack.

What it does

Aqua Rack is an Agentic Digital Twin for AI data centers that reduces cooling-water consumption through intelligent, memory-driven decision-making.

The platform combines:

  • Live laptop telemetry (CPU, GPU, RAM, disk, temperatures)
  • An OpenDC-powered digital twin that derives a synthetic multi-rack fleet from that real telemetry signal
  • Live weather conditions (Open-Meteo)
  • A CoolProp-based thermodynamic water estimation model
  • Persistent operational memory in CockroachDB, including a closed-loop Episode → outcome → confidence system
  • A multi-agent LangGraph reasoning pipeline running on Ollama and Groq

Instead of only displaying dashboards, Aqua Rack continuously:

  • Observes infrastructure.
  • Retrieves similar past episodes and their real outcomes from CockroachDB's distributed vector index.
  • Predicts water demand and thermal risk.
  • Recommends — and, in closed-loop mode, actively executes — a cooling strategy.
  • Records the decision as a new episode.
  • Resolves the outcome ~15 minutes later and updates that strategy's confidence score.
  • Reuses what worked, and explicitly avoids strategies that previously failed.

Its mission is simple: help AI data centers use water more intelligently, and get measurably better at it over time.

How I built it

I designed Aqua Rack around a hybrid digital twin architecture. A Python telemetry agent continuously collects real-time hardware metrics from my laptop while running AI workloads. OpenDC then derives a full synthetic multi-rack data center from that single real signal — one rack mirrors the laptop exactly, and the rest apply per-rack hardware-profile multipliers plus a time-varying workload-drift curve, so the whole fleet evolves realistically tick to tick.

These live inputs combine with real-time weather data to estimate cooling-water demand through an engineering-based thermodynamic model (CoolProp).

All telemetry, incidents, recommendations, and — critically — outcomes are stored in CockroachDB, which is the system's persistent memory. A LangGraph state machine (Monitor → Predictor → Optimizer → Action → Reflect → Explainer) retrieves similar historical episodes using CockroachDB's native distributed vector index before every decision, reasons over them with Ollama (with Groq as a fallback), and — in closed-loop mode — actually executes the recommendation against the twin rather than just suggesting it.

The part I'm most proud of architecturally

Every decision creates an Episode row with the outcome fields left empty. An async job resolves that outcome roughly 15 minutes later by comparing before/after telemetry, computes a reward, and updates a Beta-distribution confidence score for that strategy. The next time a similar situation comes up, the agent's stated confidence is grounded in real track record, not just an LLM's self-assessment.

Current Implementation

  • Real Device: Laptop (RACK-001) with live telemetry collection
  • Digital Twin Fleet: 99 synthetic racks (RACK-002 to RACK-100) with hardware-profile multipliers
  • Persistence: Database-backed fleet reasoning results surviving across sessions
  • Web Interface: React-based fleet management dashboard with real-time visualization
  • Dual Deployment: Backend on Render, frontend on Vercel

Key Technologies

  • Backend: Python, FastAPI, LangGraph, SQLAlchemy
  • Database: CockroachDB with vector indexing
  • AI Models: Ollama (Qwen) with Groq fallback
  • Frontend: React, Vite, Tailwind CSS, Recharts
  • Simulation: OpenDC digital twin framework
  • Physics: CoolProp thermodynamic library
  • Deployment: Render (backend), Vercel (frontend)
  • AWS (s3,cloudwatch)

Challenges I ran into

One of the biggest challenges was discovering that real-time cooling-water telemetry from hyperscale AI data centers simply isn't publicly available. Cloud providers expose CPU/GPU utilization, but not facility-level cooling or water consumption. Instead of inventing arbitrary numbers, I built a thermodynamic estimation model from hardware utilization, estimated power draw, ambient temperature, humidity, and cooling efficiency.

The deepest challenge, though, was architectural: designing an AI system that genuinely learns from experience instead of just generating plausible-sounding responses. Getting from "retrieve a similar past event" to "know whether that past decision actually worked, and adjust confidence accordingly" was the real engineering problem , not the LLM call itself.

A more recent challenge involved resolving database schema conflicts between fleet reasoning results and RL training episodes. The system uses device-specific data isolation, which initially caused foreign key constraint violations when creating episodes for fleet racks. This was resolved by removing the foreign key constraint on episodes.rack_id, allowing fleet reasoning to work with virtual rack IDs while maintaining referential integrity for physical infrastructure.

Accomplishments that I'm proud of

I designed a complete hybrid architecture that combines real hardware telemetry, digital twin simulation, engineering-based water estimation, persistent operational memory, and retrieval-augmented multi-agent reasoning — running end to end on free-tier infrastructure.

I'm especially proud that memory isn't an afterthought or a pure RAG lookup. Every observation, recommendation, and operational outcome becomes a scored, reusable experience: strategies that work get more confident over time, and strategies that failed are flagged and avoided the next time a similar situation recurs — a real closed learning loop, not just retrieval.

I'm also proud of showing that a single laptop can become a realistic multi-rack AI data center through digital twin technology, making this kind of sustainability research accessible without expensive hardware.

What I learned

This project pulled from a lot of different disciplines: AI infrastructure monitoring, digital twins, thermodynamic cooling models, persistent memory architectures, distributed vector databases, retrieval-augmented generation, closed-loop feedback systems, and explainable AI decision-making.

The biggest lesson, though, is that sustainable AI isn't only about better models — it's about making the infrastructure behind those models smarter, more accountable, and more responsible with natural resources.

What's next for Aqua Rack

  • Deeper reinforcement learning on top of the existing episode/reward system — moving from Beta-distribution confidence toward full policy optimization
  • Multi-region AI data center support, using CockroachDB's distributed locality features across sites
  • Carbon footprint optimization alongside water conservation
  • Integration with real Building Management Systems (BMS) and DCIM platforms
  • Predictive maintenance for cooling infrastructure
  • Support for liquid cooling and immersion cooling technologies
  • Real-time anomaly detection using streaming AI
  • A long-term sustainability dashboard with historical water-saving analytics
  • Enterprise deployment across hybrid and multi-cloud environments

Aqua Rack — Where AI Learns to Save Water.

Built With

  • amazon-web-services
  • cockroachdb
  • coolprop
  • docker
  • fastapi
  • framer-motion
  • github
  • groq
  • langgraph
  • llm-orchestration
  • lucide-react
  • multi-agent-systems
  • ollama
  • opendc
  • openmatio
  • python
  • rag
  • react
  • react-router
  • recharts
  • tailwind-css
  • thermodynamic-modeling
  • time-series-data
  • vector-database
  • vector-search
Share this project:

Updates

Submission history