Problem statement

Fixed mountaintop cameras, such as the HPWREN and ALERTCalifornia networks in Southern California, watch large areas of wildland because a fire is usually easiest to contain while it is still small. Nobody can watch every feed all the time, so automated smoke detectors are used to flag frames for a person to check. A detector score on its own doesn't tell that person what the model responded to, whether the alert lasted longer than one noisy frame, or how the detector behaves on fires and cameras it has never seen.

Research detectors also tend to be scored on benchmark clips whose construction can reward shortcuts, such as how much the picture has changed since the clip began.

Social impact: smoke alerts only help if the people receiving them can judge when to trust them. PlumeCheck shows its misses and false alarms next to its successes, which is what someone checking an alert needs before relying on a detector like this.

Solution overview

PlumeCheck is a research prototype that links a trained smoke detector to the evidence behind each of its decisions.

For every camera frame, it builds a background from the same camera's view 5-20 minutes earlier. A ResNet-18 network looks at both the current frame and what changed, and scores a 16 x 24 grid of image cells. That grid gives a heat-map of where the model responded. It is a model response, not a verified smoke location.

A persistent alarm needs two consecutive frames, at most 3 minutes apart, above a threshold chosen on validation data, after at least 20 minutes of camera history. A strict detection is a persistent alarm whose run starts after the annotated first visible plume. The evaluation code and the browser replay use the same versioned alarm rule, and tests check that they agree.

The demo is a static replay of saved model outputs for 24 selected held-out camera clips. It runs in any browser from a static web host, and nothing is inferred live.

Key features

  • Replay a camera sequence the model never trained on, frame by frame, with the model score, heat-map, alarm status and a timeline around the annotated first visible plume.
  • Filter clips by outcome (strict detection after the plume, no strict detection, pre-plume alarm, excluded), so a miss or an early false alarm is as easy to open as a success.
  • An Evidence tab with every result table, its definitions, the threshold policy, the seed and fold counts and the limitations, all regenerated from saved predictions.
  • Evaluation splits that hold out later fires and whole camera stations.
  • Preliminary shortcut checks. Global image statistics based on change since the clip began predict benchmark labels (AUROC 0.75-0.79, all frames) and still score 0.70-0.75 when retrained on an invented fire start 20 minutes before the plume. On the full-history frames PlumeCheck is scored on, the matched invented-onset control (10 minutes before the plume) gives 0.57-0.61 for those statistics, while a freshly trained PlumeCheck model scores 0.489 (later fires) and 0.496 (SmokeyNet test clips), close to chance. These are single-seed controls and don't show what the deployed model relies on.
  • Raw alert counts on 435.7 scored daylight camera-hours of ordinary archive footage from stations the scoring model never trained on.

Technologies used

Python, PyTorch (trained on Apple-silicon MPS), timm (ResNet-18), NumPy, pandas, scikit-learn and matplotlib for the model and evaluation. FastAPI for the research server, and plain JavaScript, HTML and CSS for the static demo. Tests use pytest and Node. The data is HPWREN's FIgLib fire-ignition image sequences and the HPWREN camera archive.

We trained the detector ourselves; it doesn't call an external AI API. AI coding assistants were used during development (see Data and attribution).

Target users

Camera-network operators and fire-lookout or dispatch staff who receive automated smoke alerts and need to judge them quickly, and developers and researchers evaluating camera-based smoke detectors. These are the intended users. We haven't tested PlumeCheck with them yet.

Results (exploratory development results)

Evaluation outputs informed later design choices, so these results are not an untouched final confirmation. Counts are camera clips, not distinct fires.

Held-out evaluation Frame AUROC Camera clips with a strict persistent detection after the plume Camera clips with a pre-plume alarm
Later fires (Aug 2021 - Sep 2026), 188 clips, 2 training seeds 0.858 (mean +/- s.d. 0.007) 77.9% 18.9%
Unseen camera stations, 459 clips, 5 folds, 1 seed 0.886 (mean of fold AUROCs) 80.6% 5.7%

Because a strict detection must start after the annotated first visible plume, "no strict detection" covers true misses and also clips whose alarm was already running before the plume.

Among detected clips, the median delay was 5.0 minutes for later fires and 6.0 minutes for unseen stations, measured from the benchmark's annotated first visible plume, not from ignition. AUROC measures ranking, not accuracy, and the detection and pre-plume columns can overlap.

The threshold is the 98th percentile of validation pre-plume scores. Its achieved validation false-positive rate is unverified because validation predictions were not saved. On later fires, the pre-plume frame false-positive rate was 8.2% and 12.0% for the two seeds. There is no head-to-head comparison with SmokeyNet.

The station folds include the 80 development clips used for model-family choices, so those results may be mildly optimistic, and a few possibly co-located station codes (for example smer/sm) were not merged. The later-fire split has neither issue.

Ordinary camera footage (raw alert counts)

On 435.7 scored daylight camera-hours (16 cameras x 2 days), the base model raised 65 persistent alerts and a variant trained with extra background camera frames raised 45. Those background frames were meant as negatives but aren't verified smoke-free. Each model is a single training run, and a camera-paired test of the difference is not significant (p ~ 0.12-0.15). Also, 8 of the 32 camera-days share a calendar date with the variant's background training footage from another station. So we don't attribute the drop to the extra background frames, and fewer alerts don't show fewer false alarms or preserved sensitivity. A blinded human review of these alerts is prepared but hasn't been done yet.

Limitations

  • Research prototype, not a validated fire-safety or decision-making system. No field deployment or user testing.
  • Southern California HPWREN cameras only. The imagery is mostly daylight: some benchmark clips are at night or in low light but weren't evaluated separately, and the footage evaluation used daylight frames only.
  • Most protocols use one training seed; later fires use two.
  • The shortcut checks are preliminary.
  • The heat-map is a model response, not a smoke location.

Data and attribution

Camera imagery comes from HPWREN (High Performance Wireless Research & Education Network), University of California San Diego, with support from the National Science Foundation (http://hpwren.ucsd.edu). FIgLib and the HPWREN camera archive are licensed CC BY-NC-ND 4.0 (https://www.hpwren.ucsd.edu/cc.html). Demo frames are resized for the web and keep their original HPWREN overlays; PlumeCheck draws the heat-map on top.

SmokeyNet split lists, used for one extra evaluation: Dewangan et al., Remote Sensing 14(4):1007 (2022). Using the lists is not a reproduction of the SmokeyNet model.

AI coding assistants were used for implementation, experiment orchestration, literature search support and documentation. They also produced preliminary labels for the archive alerts. Those labels are unverified, and we don't present them as evidence.

Team

Aditya Raut, Sagar Raut, Neeraj Movva and Sathvik Loke. All four of us contributed equally across research, implementation, testing, documentation and presentation.

Built With

Share this project:

Updates

Submission history