Inspiration

In April 2024, the electricity in Ouagadougou went out for an average of 14 hours a day, across the capital and every inland city. The cuts land hardest in the evening which is exactly when people come home and want water.

That collision is the whole project. No electricity means the pumps stop. When the pumps stop, the tanks empty and the taps run dry at the worst possible hour. Meanwhile the city carries a 57,000 m³ per day shortfall in drinking water, stated by the Minister of State for Water in May 2026.

Two things made us think a program could help. First: the sun is free at 1 p.m. and gone by 7 p.m., and the tanks are already built so the storage problem might already be solved, if only something decided when to fill them. Second: in this part of the world, every household has jerrycans. That means a message can move water.

What it does

WIIGA decides, every hour, which pump runs and on what power using sunshine, city electricity, or a diesel generator. Sunshine is free. The generator is expensive and dirty, and there is only a little of it. Its main trick: it fills the tanks at midday on free sunshine, when nobody is thirsty yet. A full tank at midday is a battery for the evening, and it costs nothing to build because the tank is already there. Its second trick: it can warn people to fill their jerrycans before the power goes. But if it warns and nothing happens, people stop believing it so it has to learn to speak only when it is sure. And when it has nothing useful to add at three in the morning, say it hands the station back to the operator and says so.

Measured over 365 simulated days against the best hand-written rulebook we could tune: 67 % fewer dry hours on the worst-served district, 48 % less CO2 than current practice, and 148,455 person-days a year kept above the WHO survival threshold of 20 litres, for 22,000 people.

How we built it

About 6,700 lines of Python. A custom Gymnasium environment models three district tanks, three pumps, three energy sources, a grid that fails, and demand that follows heat and holidays calibrated on three years of measured daily records from Open-Meteo, with no API key and no account.

A PPO agent (Stable-Baselines3) operates that twin. CPU only, eight parallel environments, about twenty minutes to train on a laptop. No GPU, no cloud API, no LLM. A utility in Ouagadougou can run this on the hardware it already has.

Two design decisions carry the project. The reward reads the worst-served district, not the average min, not mean because a policy that keeps two districts full and drains a third looks excellent on a dashboard and is unacceptable in a city. And credibility is a depletable resource: a correct warning buys +0.04 of the city's trust, a false one costs −0.15, and that asymmetry alone fixes a break-even accuracy of 78.9 % that the agent has to learn to stay above.

The demo is a single static HTML file with its data written inside it no server, no serverless function, no network request. It cannot break in front of a judge because an API changed its mind.

Challenges we ran into

Nine mechanisms were built, measured, and thrown away. Each is documented in the code at the exact line where it failed. The three worst:

Clamping trust at zero made lying free. Once at the floor, the subtraction was clipped, so a false alarm cost nothing while household response was already zero. The agent fell into that absorbing state and stayed, broadcasting seven times a day at 57 % accuracy. Below zero there is not an absence of trust there is active disbelief. The textbook form of reward shaping charged rent. γ·Φ(s') − Φ(s) costs (1−γ)·Φ every step merely for holding a reputation. A day with no warning at all cost −10.6, and a trusted utility paid twice what a discredited one paid. It rewarded destroying your own reputation. We caught ourselves measuring a pipe instead of a policy. Our demand-shock stress test swept up to ×4 and reported an agent that "loses" there. At ×4 the district asks for 82 m³/h from a pump that delivers 41. The sweep now computes its own hydraulic ceiling from the replayed year and refuses to run past it.

And one that cost us two days: a CVaR variant we trained never appeared in any table. The cause was a filename the command opened cvar_0.zip while the model was cvar_essai.zip, found nothing, and printed a table with the row silently missing. It now names what it cannot find.

Accomplishments that we're proud of

We beat the method the literature would actually use. Every other opponent in this repository is written by us, which is the objection a control engineer makes: nobody writes an if, they pose a receding-horizon model predictive controller and solve it. So we did a linear program re-solved every hour on the same information and the same max-min objective. The agent serves 82 % better. The controller runs the station for 48 % less money. Both halves are published.

We published everywhere it loses. Above 366 litres of diesel a day, a twenty-line rulebook is enough. It is worse than the rulebook on the market district. It loses in 2 of the 5 cities we transferred it to. Our CVaR experiment collapsed outright 48.2 dry hours a day against 0.156 and that failure is in the README, not in a drawer.

We tuned our own opponent against ourselves. The rulebook's constants were swept over a 6×5 grid; ours turned out to be under-tuned by 13 %, and the agent still beats the best of the family by 62 %.

Nineteen property tests, three training seeds, four ablations, a reward-hacking audit, and a literature review that narrowed our own claim rather than inflating it.

What we learned

The same fact defeats two different opponents, and we did not see it coming. Outage risk is bimodal: 59 % of hours sit below 0.1, 35 % above 0.6, and only 0.5 % in between. That is why moving a rule's threshold by a factor of four changes almost nothing. And it is why the MPC loses too a deterministic controller plans against the expectation of grid availability, so it acts as though half a grid were available on a risky evening. Half a grid does not exist.

Max-min fairness does not improve everything it flattens. The rulebook's districts run from 0.01 to 0.31 dry hours; ours run from 0.02 to 0.10. We are better on the worst and worse on the best. That is the trade, and anyone who would rather have one district suffer badly and two suffer nothing should not use this objective.

The binding limit turned out not to be the agent's judgement it was the pipe. A 30 % unannounced surge costs the agent about one minute of dry time a day. The dispensary's pump saturates at +36 % on the hottest day of the year. That number is worth more to a utility director than any of our percentages.

What's next for WIIGA

Twenty minutes on the phone with someone who runs a station. This is simulated, entirely and by construction, and no amount of extra compute fixes that. One conversation with an operator would be worth more than everything above.

Then, in order: a held-out year train on nine months, measure on three, so the agent is scored on days it has never seen. A scenario-based stochastic MPC, which would close part of the remaining gap honestly. And the idea we deliberately did not attempt under deadline: optimising the queue at the standpipe hours spent waiting rather than cubic metres delivered because that wait is carried overwhelmingly by women and children, and it would change what the agent optimises rather than how well it optimises it.

The object was never really water. It is operating an essential service on infrastructure that cannot be relied on a vaccine cold chain, a rural health battery, a motorbike charging network. About one billion people receive water from intermittent networks. Water is simply where we could measure it.

**Who this helps

Directly: the operators of a small urban water utility in a city with an unreliable grid Ouagadougou, Bobo-Dioulasso, and the several hundred cities across the Sahel and South Asia under the same constraint. Because the agent can hand any hour back, it can be adopted gradually rather than trusted all at once.

Through them: the 22,000 people in the three modelled districts.

**Impact Statement

About one billion people receive water from intermittent piped networks pipes that carry water for fewer than 24 hours a day. That is the constraint this project is about: not water scarcity in general, but water that arrives on a schedule somebody has to decide.

On the one station we model, measured over 365 simulated days against the best hand-written rulebook we could tune: 67 % fewer dry hours on the worst-served district, 48 % less CO2 than current practice, and 148,455 person-days a year kept above the WHO survival threshold of 20 litres 6.75 days per resident per year that would have been spent below it and were not.

The people who benefit are the households at the end of the pipe, and the ones who benefit most are those in whichever district is worst served, because that is the only district the objective looks at.

And this is where we stop. Multiplying 6.75 days by a billion people would make a headline and it would be dishonest. Everything here is simulated, and the repository publishes three measured boundaries where the agent stops helping including one where a twenty-line rulebook does the job better.

Built With

Share this project:

Updates

Submission history