-
-
CapacityOS shows the flexible grid absorbing the full 6.4 MW overload through coordinated batteries and buildings.
-
CapacityOS exposes scenarios where available flexibility cannot resolve the overload, before the program is deployed.
-
Drag data centres, housing developments, or EV depots onto the grid to test their effect on local capacity.
-
Explore an Ontario Southwest grid sandbox built from real demand patterns, modeled device clusters, and configurable new loads.
-
Adding a new load triggers an overload, prompting participating asset owners to offer flexibility before dispatch.
-
Adjust batteries, EV fleets, flexible buildings, enrollment, incentives, and reserve requirements to test different VPP designs.
-
1. Test the same VPP against each season’s historical worst day to reveal when the design holds—and when it fails.
-
Review owner offers, physics validation, dispatch decisions, costs, and the reasoning behind the final result.
-
Compare verified options such as changing program constraints or adding grid capacity when flexibility alone is insufficient.
Inspiration
Canada is about to build more than it has in a generation — more housing, more AI data centres, more electrified transit and industry. Every single one of those things needs power. And the honest answer to "where's it going to come from?" isn't always "build more wires." Sometimes it's "use the wires you already have, better."
That idea — that flexibility is infrastructure too — is not new. Utilities have talked about virtual power plants for years. But every conversation we had about it kept landing on the same wall: nobody can tell you, with any confidence, whether a specific flexibility program will actually work for a specific piece of new demand, in a specific place, under real conditions. It's either a hand-wavy slide deck or a six-month engineering study. There's nothing in between. We wanted to build the thing in between.
What it does
CapacityOS is a sandbox where you can drop a hypothetical data centre, housing development, or EV depot onto a modeled electrical zone — built on a real IESO Southwest demand profile — and find out, honestly, whether it can be accommodated.
It's not a toy simulation with made-up numbers. The demand curve underneath it is real historical grid data. The devices around it — batteries, EV fleets, flexible buildings — are a synthetic population with explicit device and owner constraints, not a black box. The demo world contains 234 synthetic device clusters grouped into 18 modeled owner portfolios. Battery offers are bounded by power, energy, state of charge, and reserve requirements; EV offers must preserve charging deadlines; building offers must stay inside modeled comfort limits.
When you push the zone over its capacity limit, CapacityOS sends a real flexibility request to every device owner in the zone, and each one — independently, using an LLM-backed agent reasoning over its own constraints — decides whether to offer flexibility, at what price, or decline. Every offer is checked against deterministic device constraints before it counts; an invalid offer gets exactly one bounded revision attempt. Only validated offers reach an OR-Tools optimizer, which first minimizes the remaining overload, then minimizes modeled dispatch cost among the remaining feasible options.
Then it tells you the truth: did it fully hold, partly hold, or not hold at all. And if it didn't fully hold, it doesn't stop there — it goes and tests fixes. More participation, a different incentive, more battery capacity, even a straight infrastructure upgrade as a fallback. Twenty to thirty real reruns of the simulation, not guesses — each pathway shows the actual, verified reduction in the remaining gap. Then it stress-tests that same program across different seasons of real historical data, because a plan that only works on one lucky day isn't a plan.
How we built it
We split the problem exactly where it should be split. The frontend is Next.js and React — a living, interactive world you can drop buildings onto and watch respond in real time. The backend is Python: FastAPI serving the API, Pandas and NumPy handling the real IESO demand data underneath it, and Google OR-Tools doing the one thing that actually has to be provably correct — the dispatch optimization. We deployed the frontend to Vercel and the backend to Render as a long-lived Python service, because OR-Tools and an LLM-backed owner-offer process don't belong on a serverless function.
The hardest architectural decision was keeping two very different kinds of "agent" honest and separate. There's the physical device — a battery has a state of charge, an EV fleet has a departure deadline, a building has a comfort limit, and none of that is negotiable. Then there's the owner — the economic decision-maker behind that device, and that's where we let an LLM (OpenAI) actually reason: should I offer flexibility, at what price, or say no? The rule we never broke: AI proposes. Physics constrains. Optimization decides. An LLM can talk its way into an offer. It can never talk its way past a battery's actual limits.
Challenges we ran into
The first real challenge was philosophical before it was technical: how do you build something that talks about flexibility and demand without silently pretending modeled data is measured fact? We ended up building a provenance system into the product itself — every number on screen is labeled Observed, Modeled, or Hypothetical, because the moment you blur that line, you've built a system that lies convincingly, which is worse than one that doesn't work at all.
The second was making AI agents genuinely independent without letting them touch anything that mattered. It would have been easy to let the LLM just output a dispatch number. We refused to. Every offer an owner agent makes gets checked against deterministic battery, EV, and building constraints before an optimizer ever sees it — which meant building a validator that could reject an LLM's proposal and ask it to revise, exactly once, under a hard bound, so a stubborn model can't stall the whole system.
Even producing the demo video became its own small engineering problem — AI video generation turned out to reliably hallucinate garbled on-screen text, which we solved by generating clean visuals and burning real titles in separately.
Accomplishments that we're proud of
We're proud that the "partly holds" result in our own demo video is real. Nobody manufactured a scenario to make the story land — we ran the actual simulation, got an actual partial outcome, and built the story around what really happened, because that's the whole point of the product.
We're proud that an LLM never once got to invent a number that mattered. Every dispatch decision in this product traces back to a deterministic, constraint-bounded optimizer.
And we're proud of the honesty layer itself — the fact that CapacityOS will tell you, on screen, in plain language, when it didn't fully hold, and then go find out what would actually change that, instead of quietly rounding a partial result up to a full one.
What we learned
We learned that the hardest part of building something trustworthy isn't the algorithm — it's the discipline to keep saying "not yet" to features that would look good but aren't verified. Every claim in this product, and in the story we're telling about it, had to survive being checked against the actual code, not just the pitch.
We learned that AI is an extraordinary proposer and a poor arbiter. The moment we stopped asking the LLM to decide anything and started asking it to only ever propose something a deterministic system could check, the whole product got both more capable and more honest at the same time.
What's next for CapacityOS
Next is a calibrated real-world pilot: a real utility's local constraints, a real device inventory instead of a modeled population, benchmarked against an actual planning case a utility has already done the hard way — so we can show, side by side, what this approach would have told them sooner.
Because the question was never just "how much should we build." It's "how much can we unlock from what's already standing" — and that's the question CapacityOS exists to keep answering, honestly, one zone at a time.
Log in or sign up for Devpost to join the conversation.