Inspiration

Track 02 asked for an intelligent automation system that manages energy on an isolated island from live sensor data — but most submissions to a prompt like that turn into a dashboard with some sliders. We wanted to build the thing underneath the dashboard: the actual decision engine a real island microgrid would need when a storm knocks generation down and there isn't enough power for everyone.

The core question that shaped everything: who gets electricity first, and why? Real islanded microgrids don't treat every building equally during a shortage — hospitals and water treatment are non-negotiable, everything else is a matter of degree. That tiered-triage idea, not a chatbot wrapper, became the actual project.

What we learned

Realism is a research problem, not a guess. Early on, every number — battery size, solar output, how much power a hospital actually draws — was a plausible-sounding placeholder. We went back and grounded every one of them: a small rural hospital's real load, reverse-osmosis desalination's real energy intensity,

$$E_{\text{desal}} = \dot{V}_{\text{water}} \times \epsilon, \qquad \epsilon \approx 3\text{–}4\ \frac{\text{kWh}}{\text{m}^3}$$

EIA residential consumption data, and real turbine power curves — output rises roughly with the cube of wind speed between cut-in and rated speed,

$$P_{\text{wind}}(v) \propto v^3 \quad \text{for } v_{\text{cut-in}} \lesssim v \lesssim v_{\text{rated}}$$

and collapses to zero above a safety cut-out speed. Stationary battery storage follows a C-rate convention we hadn't used correctly at first — max power is a fraction of capacity per hour, $P_{\max} = C \cdot E_{\text{capacity}}$, not an arbitrary constant.

A "smart-sounding" system needs a real algorithm underneath the AI. The LLM (gpt-oss-20b over NVIDIA NIM) narrates why the grid did something in plain English — but it never decides anything. The dispatch itself is deterministic, auditable code: a floor pass that guarantees Tier 1 before anyone else sees a watt, then a bounded market (Resilience Credits, not currency) for whatever true surplus is left. Every allocation is exact math we can defend line by line, and the LLM is a best-effort narration layer on top that a live demo can never depend on.

Testing catches the bugs that matter, not the cosmetic ones. Several real defects only surfaced under sustained or adversarial conditions: a battery that could discharge for Tier 1 during a storm but never recharge was mathematically guaranteed to hit zero given enough real time — it just took a multi-minute live demo to prove it. A trickle constant sized too small meant Households (Tier 3) were mathematically unreachable, not just under-served, regardless of how the simulation ran.

How we built it

  • Backend: FastAPI + a native WebSocket, no Socket.IO — an asyncio loop re-clears the grid every 2.2s and broadcasts to every connected client.
  • Engine: a two-phase allocator. Phase 1 splits supply tier-by-tier; if a tier can't be fully covered, the shortfall splits proportionally: $$\text{served}_i = \text{available} \times \frac{\text{floor}_i}{\sum_j \text{floor}_j}$$ Phase 2 clears any real surplus above every floor as a sealed-bid Resilience Credit market — highest priority per kWh first.
  • Weather + curtailment: six real weather profiles drive solar/wind multipliers; when the battery's full and nothing wants the surplus, generation is curtailed and reported as such — the same real-world response as a turbine feathering its blades or an inverter throttling back, not silently "wasted" power.
  • Frontend: a hand-built Three.js digital twin of the island — procedural terrain, live camera orbit, buildings that visibly change state — driven entirely by the same WebSocket tick, so nothing on screen is decorative.
  • Testing: 32 backend unit tests plus a full Playwright gauntlet (interaction, visual composition, pixel-level color verification, concurrency/stress) — well over 100 automated checks re-run after every change.

Challenges we faced

The hardest bugs weren't in the code we wrote — they were in the scale of the code we wrote. A trickle cap that looked reasonable on paper turned out to make an entire tier unreachable by construction. A recharge mechanism that worked for a 30-second demo silently failed over a 2-hour one. A UI panel that fit perfectly at one browser window size vanished — header and all — at another. Each one required tracing the actual arithmetic through, not just re-running until it looked fine, and each one made the final system meaningfully more honest about what it claims to do.

The last discipline we held onto: knowing what not to claim. Reef is a decision-support prototype with real, auditable dispatch logic — not a system licensed to run an actual electrical grid. That distinction matters as much as anything we built.

Built With

Share this project:

Updates

Submission history