Inspiration
For someone who uses a wheelchair, a walker or a stroller, one car parked across a sidewalk can end a trip. They turn back, or roll into traffic. Cities are legally responsible for keeping pedestrian routes accessible, but they mostly learn about barriers from manual surveys and complaints. That's slow and expensive, and neighbourhoods that complain less get fixed less.
Meanwhile, cars with cameras already drive every street, and everyone carries a phone. We asked: can the street photos that already exist audit a city's accessibility automatically, and can residents report a barrier in 20 seconds without downloading anything?
That became SAGA: Street Accessibility and Guidance Audit. One rule shaped it: SAGA audits places, never people. Its output is a ranked list of locations for a city to inspect and fix, not a ticket for a driver, and every face and licence plate is blurred before a photo is stored.
What it does
Sees the street. A segmentation model labels every pixel of a photo into 19 classes: road, sidewalk, person, car, bicycle, and so on. Checks what everything is standing on. For each object, SAGA reads the strip of ground directly beneath it and beside it, then raises three flags: 🚧 Sidewalk obstruction: a car, truck, motorcycle or bike standing on the sidewalk 🚶 Pedestrian in roadway: a person forced to walk in the street ❌ No visible sidewalk: a built-up street with a visible curb but no sidewalk Ranks repairs with a transparent priority score score
Maps it. A city map of 747 audited US locations plus every resident report, with a top-25 repair priority list. Lets residents report from their phone. The SAGA iPhone web app opens the camera. The photo is audited on our server, the answer comes back as plain text first (VoiceOver-friendly) with a colour overlay, and the resident taps 👍 Correct / 👎 Wrong. That tap becomes a real-world label: SAGA measures its own precision on residents' photos and learns which photos to train on next. How we built it Data. We used Berkeley DeepDrive BDD100K (dashcam video from New York and the SF Bay Area), with fixed splits:
6,500 images for training; 500 for picking the best checkpoint; 1,000 official validation images as a test set never touched during training; 747 GPS-tagged test frames for the map. Model. SegFormer-B2 (27 M parameters), pre-trained on Cityscapes, which uses the same 19 classes. Run unchanged, it's an honest baseline.
We fine-tuned it on BDD100K for 8,000 steps with bf16 mixed precision on an NVIDIA RTX PRO 6000 Blackwell. It trained in about 8 minutes, against a projected 3.5 hours on a laptop. Segmentation quality is measured per class with intersection-over-union:
Test set (1,000 images) Cityscapes baseline SAGA mIoU (19 classes) 50.4% 63.4% Sidewalk IoU 51.4% 65.0% Person IoU 65.0% 74.5% Bus / Motorcycle IoU 28.0 / 50.7 77.4 / 65.3 Audit accuracy. Each test image is audited twice: once from the model's labels and once from the human annotators' labels. We report how well they agree:
Server: Python FastAPI serves the model, the audit engine, an anonymiser (blurs people and the plate band of vehicles using the model's own segmentation) and a SQLite report store. Frontend: plain HTML/JS with a Leaflet + OpenStreetMap city map, plus an installable iPhone web app (PWA) that shrinks photos on the phone before upload and reads location from the browser. Deployment. Live on a Vultr Ubuntu server at sagatech.us:
Caddy provides automatic Let's Encrypt HTTPS; systemd services restart on crash and on reboot; a firewall leaves only 80/443 open, and the app itself listens only on localhost; a login gate (four yes/no questions + a shared password, verified server-side against salted scrypt hashes) protects every page, photo and API call; per-report edit keys, rate limits and a 5-attempt lockout limit abuse. Tests. 24 Python unit tests covering the audit rules, login (user and 15-minute demo), temporary demo data, anonymisation and the report store.es
Challenges we ran into
Perspective lies. Our first rule flagged cars on a cross street as "parked on the sidewalk", because the near-side sidewalk sits between the camera and the car. We fixed it with a side check (sidewalk must also be beside the object), a distance limit and a frame-edge rule. That brought false alarms on the human-labelled set from 70 down to 14. "No sidewalk" vs. "sidewalk hidden by parked cars". In New York, parked cars hide the sidewalk in most frames. We only call a sidewalk missing when the curb zone is actually visible, which cut false flags from 183 to 32. Honest evaluation. It's easy to report a pretty number. We kept a test set the model never saw, reported 95% confidence intervals, and separately reviewed the rules by eye (8/14 obstruction flags plausible, 12/12 pedestrian flags correct). We call out our weakest flag openly: obstruction, with only 14 real test cases. Training stalled on a laptop. The Mac run hung after its first evaluation because extra data-loading processes were started mid-training. We fixed it and moved training to a rented GPU in an isolated workspace. Shipping, not just training. We hit a dead package repository, a 1 GB server that couldn't hold the model, HEIC photos from iPhones, and iOS Safari stripping GPS from uploads. Each needed a real fix: the Ubuntu-packaged Caddy, an in-place server upgrade, pillow-heif, and browser geolocation. Privacy by construction. Raw photos never touch disk, and phone numbers are stored only as salted hashes.
Accomplishments that we're proud of
+13 mIoU and +13.6 sidewalk IoU over the baseline, with the repair ranking reaching ρ = 0.85 against human-label audits. A live, standalone deployment: our first real resident report came from downtown Atlanta. The phone photo was audited on the server ("no visible sidewalk", priority 40), pinned by GPS, and confirmed with 👍. An accessibility tool that is itself accessible: text-first answers, large touch targets, no account needed.
What we learned
The hard part of "AI for accessibility" isn't the model; it's turning pixels into a claim a city can act on, and being honest about how often that claim is right. Geometry beats more data for rule failures: three simple perspective checks removed more false alarms than extra training would have. Residents' feedback is the cheapest, most relevant label source there is, so we built the product around collecting it.
What's next for SAGA
Curb ramps and crosswalks using Project Sidewalk's open accessibility labels. Aggregating many photos per block, so one blocked frame doesn't define a street. Retraining from residents' 👎 photos (active learning) to fix the obstruction flag. Individual accounts for city staff. Faster replies on dedicated CPU or GPU hardware: about 10 s now, about 3 s on 4 dedicated vCPUs, under 1 s on a GPU.
Log in or sign up for Devpost to join the conversation.