Inspiration

This semester at SFU I started writing a research paper on why Canada still does not use modern traffic surveillance systems, and the deeper I went into the data, the less sense it made. In 2023 alone, motor vehicle collisions in Canada caused 1,964 deaths and over 129,000 injuries, and the number of fatalities has been on a plateau for a decade, with no hint of decline. On top of that, more than 105,000 cars were stolen in 2022, and the Insurance Bureau of Canada called car theft a "national crisis", with insurers paying out more than C$1.5bn.

Then I looked at how other cities deal with the same problem. Moscow covers its roads with around 169 cameras per 10,000 people, which works, but turns into mass surveillance and a real privacy risk. Dubai, with more than 21 cameras per 10,000 residents, uses AI cameras to catch violations and respond faster, which I see as the balance. Vancouver has 221 traffic cameras for 710,000 people, only 140 of them for red-light and speed enforcement, and the feeds are rarely even recorded. Metro Vancouver is the largest of the three areas, with a fraction of the cameras.

But I do not think the answer is simply "more cameras". The real problem is that today's systems are built around licence plates, and plates are the first thing a thief swaps and the one thing a hit-and-run witness never catches. What people actually remember is the car: "It was a blue pickup truck, maybe a Toyota, it kept going east on Kingsway." So I asked myself a simple question: what if a city could search its existing cameras by exactly that? That is how ReTrace started, and StormHacks was the perfect chance to finally build it instead of only writing about it.

What it does

ReTrace turns the cameras a city already has into a searchable memory of every vehicle, and it indexes vehicles, not faces or plates.

  • Every feed is processed: each vehicle is detected, tracked, and recognised by make, model and generation (9,630 classes), body type, colour, and a 2048-dimensional appearance fingerprint.
  • Search the way people remember:
    • type a description;
    • say it out loud (a witness statement is transcribed and the colour, body type and make are pulled out of the words);
    • or upload a photo or a short video of the car.
  • Trace one car across the city: pick any vehicle and ReTrace finds it again on other cameras, draws its route on a map of Vancouver, and follows it with a camera that keeps the car centred.

The demo runs on 18 camera feeds with 697 indexed vehicles, built from public traffic datasets placed at illustrative Vancouver intersections, in day, dusk, night, rain and snow, on everything from 800×600 pole cameras to 4K overpass footage.

How I built it

1. Detection and tracking. I used YOLOv12n fine-tuned to a single vehicle class on 43,654 images (Roboflow 100 Vehicles, Stanford Cars and Udacity Self-Driving). On validation it reaches 88.8% precision, 86.1% recall and 92.4% AP50. ByteTrack keeps one ID per vehicle through occlusions, and every feed is tracked at its native resolution.

2. Choosing the right crops. This part turned out to matter more than I expected. The biggest box of a track is often the worst one: the car is cut off by the frame edge, hidden behind another car, or blurred. So every detection gets a quality score that combines truncation, overlap, aspect ratio, size, confidence and sharpness, and the best $k=4$ crops, spread over time, are kept. Their embeddings are averaged into one fingerprint:

$$ \mathbf{e} = \frac{\sum_{i=1}^{k} \hat{\mathbf{e}}i}{\left\lVert \sum{i=1}^{k} \hat{\mathbf{e}}_i \right\rVert}, \qquad \hat{\mathbf{e}}_i = \frac{f(x_i)}{\lVert f(x_i) \rVert} $$

3. Re-identification. Two EfficientNetV2-M models (300×300 input, 2048-D embedding, cross-entropy plus contrastive and circle loss) on GigaFlexhicle and Google Images of newer North American models, over 400,000 images:

  • a classifier with 9,630 make/model/generation classes;
  • an embeddings model on a US/Canada-focused subset (5,445 classes), with parameters tuned for embedding quality.

Make and model come from the classifier; matching comes from the embeddings model. On CityFlowV2 (36 cars, 5 cameras, zero-shot) the embeddings model reaches 73.3% Rank-1 and 67.4% mAP, and combining both models gives 75.0% Rank-1 and 73.3% mAP.

4. Journeys. Two sightings $i$ and $j$ on different cameras are linked only if they are each other's best match and pass a body-type and colour check:

$$ j = \arg\max_{m}\, s(i,m), \quad i = \arg\max_{m}\, s(j,m), \quad s(i,j) = \mathbf{e}_i^{\top}\mathbf{e}_j \ge \tau $$

I picked $\tau$ by F1 against the ground truth of the multi-camera dataset.

5. Search. LAION CLIP ViT-B/32 ranks every vehicle crop against a text description, Whisper transcribes voice, and the query words are parsed into colour, body type and make, which act as a filter on top of CLIP. For photo and video search, the same detector, crop selection and embeddings model run on the upload, so a picture from a phone is compared with every camera in the same space.

6. The product. The backend is FastAPI on a single RTX 3090. The website is React, TypeScript and Vite, with scroll-scrubbed video and bounding boxes painted on a canvas inside the video frame callback, so they never drift from the footage. The map of Vancouver is built from OpenStreetMap.

Cost. I measured everything on one RTX 3090: 68 fps for detection plus tracking at 1080p, 427 vehicles fingerprinted per second, 10 ms to match one car against a million vehicles, and 36 ms from a typed description to ranked results. At 10 analysed frames per second per camera:

$$ \text{cameras per GPU} = \left\lfloor \frac{68}{10} \right\rfloor = 6, \qquad \text{GPUs for Vancouver} = \left\lceil \frac{221}{6} \right\rceil = 37 $$

That is roughly US$37,000 of one-time hardware, or about US$167 per camera, with no new cameras installed.

All the training code, dataset splits, logs and weights are in the repository, together with a step-by-step guide to rebuild the datasets and reproduce every model.

Challenges I faced

Labels that look confident but are wrong. The first version of the classifier fell back to the same few classes on weak crops; one luxury SUV was the top-1 label for about a quarter of all vehicles. I clearly remember looking at the live map and seeing that SUV everywhere. I banned it, capped any single label at 3%, added a CLIP body-type cross-check, and a reliability gate that is stricter at night, in rain and in snow. Now only 259 of 697 vehicles get a make and model, and the rest honestly show colour and body type instead of a guess. I would rather show less than show something wrong.

False journeys. Plain nearest-neighbour matching linked cars with only 43% precision. Adding the body and colour check and mutual best match raised it to 66% precision at 63% recall, and the blue Tacoma in the demo is matched on 4 out of 4 cameras, confirmed by ground truth.

Boxes lagging behind the video. Re-rendering React on every frame made the boxes visibly drift. Moving all drawing into a canvas inside requestVideoFrameCallback fixed it completely.

Time. Building the website, the live map, journeys, voice and photo search in one weekend meant constantly choosing what matters most for the person using it. As always in hackathons, stress made me faster, and step by step, bug after bug, everything came together.

What I learned

  • Crop quality matters as much as model quality. A great model on a bad crop is still a bad answer.
  • Simple post-processing (attribute checks, mutual matching, label caps) often beats raw model confidence, and it is much easier to explain.
  • Witnesses describe cars, not plates, so search should start from what people actually remember.
  • Honesty is a feature: hiding unreliable labels and marking look-alikes as "not confirmed" makes the system more trustworthy, not less.

What's next

I want to see ReTrace piloted by a municipality on its existing camera network, with retention and access rules set by the city, an audit log for every search, and batched inference on smaller edge GPUs to bring the cost per camera even lower. The goal was never more cameras; it is getting more out of the cameras a city already has.

Built With

Share this project:

Updates

Submission history