Inspiration
Midtown Manhattan runs on a few avenues, and a single double-parked delivery truck can take one of them down to fewer lanes for minutes at a time. The queue it creates spills back through the next signal, and then the one after that. NYC DOT's Traffic Management Center has hundreds of public cameras pointed at exactly these streets, but a limited number of operators watching them. Blockages tend to get noticed after the queue has already formed, and fixed signal plans keep feeding traffic into it.
We wanted to know: can the cameras the city already has catch a blocked lane on their own, and can we check whether a signal-timing change would actually help before anyone touches a light?
What it does
ellipsis watches NYC DOT's public traffic cameras in Midtown and flags vehicles that are blocking traffic: double parked, stopped in a travel lane, or blocking the box. For each alert it:
- shows the operator the camera frame, the vehicle, the lane, and how long it has been stopped,
- runs a SUMO simulation of the surrounding corridor comparing the current signal timing against candidate retimings, and
- recommends a change only if the simulation says it helps.
A person approves every change. ellipsis never touches a real signal.
How we built it
We started with a written plan (requirements, scope, non-goals) and split the system into 35 staged GitHub issues with clear owners, then built the subsystems in parallel and integrated them stage by stage.
Ingest. A poller pulls frames from NYC DOT's public camera feeds and checks feed health: a frozen picture, or a camera that has panned away from its usual view, pauses that camera so it can't raise false alerts.
Detection and tracking. Ultralytics YOLO11s detects vehicles in each frame, and an IoU tracker keeps one track_id per vehicle across frames and measures how long it has stayed still.
Lane context and dwell rules. A bounding box is not an incident. For each camera we hand-drew lane masks (curb-adjacent lanes, travel lanes, the intersection box) with a small mask editor. An event opens only when a tracked vehicle stays still inside a zone for longer than that zone's dwell threshold. For double parking:
$$ \text{alert} \iff \text{zone}(v) = \texttt{curb_adjacent} \;\wedge\; t_{\text{stationary}}(v) \ge 60\,\text{s} $$
Simulation. We built a SUMO network of Midtown (9th to 5th Ave, 29th to 39th St) from OpenStreetMap, with weekday midday demand from published 15 Penn Plaza FEIS counts scaled by 0.89 for congestion pricing. Each alert becomes a stopped vehicle in the same lane. After a 300 s warm-up, the stop is held for as long as the real one lasted (up to 5 minutes), followed by 3 minutes of recovery, over 3 random seeds. Three candidate plans are scored against the current one: shorten the upstream green, lengthen the blocked approach's green, or both. The cycle stays at 90 s, no green drops below 8 s, and no phase moves more than 20%. Delay is counted for every trip through the blocked block or the retimed signals, cross streets included, so a plan can't "win" by starving a side street.
Dashboard. A FastAPI backend streams incidents over a WebSocket to a React + MapLibre dashboard: a ranked incident queue, the evidence frame, traffic impact, a side-by-side Base vs. Sim view, and an Accept / Reject / False-positive decision bar. Gemini writes a short incident note from the detection and simulation data. A replay mode runs recorded footage through the live pipeline, and a public demo site at ellipsisnyc.tech shows real recorded incidents.
Results
We hand-tagged 15 real blockages across 9 cameras and scored the engine against them (scripts/evaluate.py):
$$ \text{precision} = \frac{TP}{TP + FP} = \frac{14}{16} = 88\%, \qquad \text{recall} = \frac{TP}{TP + FN} = \frac{14}{15} = 93\% $$
The median time to alert is 63 s, most of which is the deliberate 60 s dwell rule, and incident-type accuracy is 100% on the caught blockages. This is a small test set, not production-scale validation.
On the simulation side, for a box truck double parked on 7th Ave at 36th St, shortening the upstream green by 6.8 s cut average delay from 124.4 s to 119.4 s per vehicle (about 4%) and the queue from 9 to 8 vehicles. The gains are modest, and we say so. More on that below.
Challenges we ran into
- Tiny, noisy frames. The public feeds are 352×240 stills taken a few seconds apart. YOLO would box a yellow cab as a car, a truck and a bus at once, which the tracker counted as three vehicles. Class-agnostic NMS fixed that.
- Cameras that move. DOT operators pan and zoom the cameras, which instantly misaligns the lane masks. We compare each frame against a reference view and pause the camera until it comes back.
- Parked vs. just stuck in traffic. A car sitting at a red light looks exactly like a double-parked car. We hold alerts while too few vehicles are passing, so a stopped queue doesn't light up the whole avenue.
- Honest simulation. Our first recommendations showed no gain on real incidents. Working through why taught us that retiming limits spillback but can't move the truck. With a 3-minute stop, the same truck gains nothing from any plan, and at 3 seeds a few percent is within the noise. When the default timing wins, the dashboard says so instead of pushing a change.
- Signal plans aren't public. NYC doesn't publish its controller timing, so we built a pre-timed plan from published sources and kept a register of what's verified (the 90 s cycle, pedestrian intervals, a Barnes dance) and what's assumed (the clearance times and green splits). Adaptive Midtown in Motion control isn't modeled.
- Free-tier limits. Gemini's free tier allows about 20 requests a day per model, and each alert uses one, so we had to plan the demo around it.
What we learned
- Detection is the easy part. Persistence, lane geometry, and dwell time are what turn "there is a truck" into an incident an operator can trust.
- Measure against ground truth early. Hand-tagging footage and re-running the evaluator after every rule change kept us from fooling ourselves.
- Simulation is decision support, not proof. The most useful thing the simulator does is sometimes say don't change anything.
- Treat a hackathon like a real project. Requirements first, then issues with owners, then integration, made the parallel work actually fit together.
What's next
- Validate on night and bad-weather footage.
- Assisted lane calibration, so adding a camera doesn't mean drawing masks by hand.
- Real controller timing data from DOT in place of our modeled signal plans.
- Human approval on every change stays. That's a design choice, not a limitation.
Built With
- elevenlabs
- fastapi
- gemini
- github
- httpx
- maplibre-gl
- numpy
- opencv
- openfreemap
- openstreetmap
- pillow
- pydantic
- pytest
- python
- react
- sqlite
- sqlmodel
- sumi
- traci
- typescript
- ultralytics
- uvicorn
- vite
- websocket
- yolo


Log in or sign up for Devpost to join the conversation.