Inspiration
I take the train and the same thing happens every time. You wait forever, then three of the same train show up together. That's bunching, and once a line starts doing it, it does it for the whole rush.
The MTA runs on a schedule written ahead of time. Dispatchers patch it live with holds and skip-stops, but each call is local. One dispatcher, one train, one platform. Nobody's steering the whole line at once.
That's a control problem, and control problems are what RL is actually good at. So I wanted to see what happens if every train is an agent and they all learn to space themselves out together.
What it does
HEADWAY is a multi-agent RL subway simulator. Autopilot for the MTA.
Every train on the line is its own agent, and at each station it makes one decision: depart, or hold one more tick. That's the whole action space. Two options, made hundreds of times per rush hour, across every train at once.
The interesting part is the reward. Each agent is scored on the number of people waiting across the entire line, not on its own punctuality. Holding costs a train something (people onboard are stuck), stranding someone costs it a lot more, and crucially the same number goes to every train. That's what stops the obvious failure: if you pay a train for its own schedule, it races ahead, skims the platform, and shoves the queue onto the train behind. Sharing the reward makes that a losing move.
Demand is real. Station arrival rates come from MTA hourly ridership on the state open data portal, AM rush, and they're wildly uneven the way the real line is. Trains don't fill evenly, which is what makes bunching emerge instead of being something I hardcoded.
Results, measured against a fixed timetable on the same demand:
| line | headway CV reduction |
|---|---|
| G | 88.1% |
| L | 80.0% |
| 7 | 79.1% |
| 1 | 74.0% |
| 6 | 68.6% |
Headway CV is the coefficient of variation of the gaps between trains. It's the direct measure of bunching: low CV means trains arrive evenly, high CV means you wait ten minutes and then three show up. Trained on the L, then run without retraining on four lines it had never seen.
How I built it
The environment. Hand-rolled PettingZoo ParallelEnv. Stations are nodes
with their own Poisson arrival processes scaled from real ridership, a Gaussian
rush profile on top (2.5x at peak), and dwell time driven by how many people are
actually on the platform. That last piece is what makes bunching self-inflicting:
a late train meets a bigger crowd, dwells longer, falls further behind, and the
train behind it hits an empty platform and closes the gap.
The agents. PPO from stable-baselines3, one shared MlpPolicy across every train via supersuit parameter sharing, so the policy learns a general rule for how a train should behave rather than memorizing positions. Each train sees 12 floats, all local: its position, direction, gap to the train ahead, gap to the train behind, its own load, its dwell timer, the queue at the next four stations, time through the episode, and a local crowding index. No agent ever sees the whole line. A dispatcher doesn't either, and a policy that needs global state is a policy you can't deploy.
Training. 8 parallel environments, 900-tick episodes, checkpoints at 1%, 3%, 6%, 12% and 25% of a 20M-step anneal horizon. The shipped policy is the 12% checkpoint: 2,400,048 timesteps. Later checkpoints existed and weren't better, which was its own small lesson.
The data. MTA Subway Stations (39hk-dx4f) for station geometry, MTA Subway
Hourly Ridership (wujg-7c2s) for demand, both from the state open data portal,
filtered to 7-9am across a normal October week and grouped by station complex.
Cached to disk so a training run is reproducible and doesn't depend on the API
being up.
The frontend. React and Vite with a three.js scene. It computes nothing at runtime. Every run in the UI is a replay trace the training pass already produced, 30 of them, about 21 MB committed to the repo: baseline plus five checkpoints, across five lines. Baloo 2 for display type, JetBrains Mono on every number. Nothing can fail on stage because nothing is running.
Challenges I ran into
Credit assignment. When one train holds, the payoff shows up two trains back, several minutes later, and the train that paid the cost never sees the benefit. The fix was giving every agent the same line-wide reward instead of a personal one, which sounds obvious written down and was not obvious at 2am.
Dependency hell, honestly. PettingZoo, supersuit and stable-baselines3 break against each other constantly, and Python 3.14 has no torch wheels. I ended up pinning every transitive dependency in one frozen block with a note telling future me not to bump a single line without re-running the smoke test.
And I didn't want live inference on stage. Precomputing every replay took longer to build but it can't break in front of a judge.
What I learned
That finding a problem which actually affects a lot of people is harder than I thought. It's easy to build something clever for a problem nobody has. Most of this hackathon was spent throwing ideas out until I landed on one where I could point at real people and say, this happens to them every single morning.
Also that the honest version of a result is more interesting than the impressive one. Passenger destinations in my sim are uniform random, which isn't how anyone travels, and knowing exactly where the model is thin is what tells me what to build next.
What's next for HEADWAY
Real origin-destination data instead of uniform random destinations, so trains fill and empty the way they actually do. Then multiple lines and transfer stations, where holding a train has consequences on a route it isn't even on.
But the version I actually want isn't autopilot. It's a dispatcher tool: here's the line right now, here are the three holds worth making, here's what each one buys you.

Log in or sign up for Devpost to join the conversation.