Guardian Angel
Track: Physical AI
AI that spots pedestrians crossing and alerts the vehicle in milliseconds.
For Physical AI, a correct answer can still arrive too late. Guardian Angel takes a decoded forward-camera RGB frame and decides whether a pedestrian is actively crossing. It returns one of two states:
CLEAR: no pedestrian is actively crossing in the current frameCROSSING_ALERT: a pedestrian is actively crossing in the current frame
I built Guardian Angel around one question:
How much can I reduce sensor-to-decision latency on Arm CPUs without changing the model's held-out decisions?
The result
I froze the dataset, sequence-level split, model checkpoint, preprocessing, threshold, quality rules, and test set before comparing runtimes.
| Arm platform | PyTorch eager FP32 | ExecuTorch/XNNPACK FP32 | p95 speedup | Decision agreement |
|---|---|---|---|---|
| Apple M1 Max | 11.650 ms | 2.240 ms | 5.201x | 24,759 / 24,759 |
AWS Graviton4 c8g.xlarge |
8.134 ms | 2.150 ms | 3.784x | 24,759 / 24,759 |
On Graviton4, decoded-RGB pipeline throughput increased from 127.104 to 479.012 FPS. The optimized runtime produced zero violations at the experimental 5 ms, 10 ms, and 20 ms latency budgets across 74,277 timing observations.
These are controlled comparisons within each platform. I am not presenting M1 Max and Graviton4 as a hardware race. Timing starts from an already decoded RGB frame and includes preprocessing, synchronous inference, and the deterministic alert decision. JPEG decode, file I/O, model loading, runtime conversion, and warmup are excluded.
What I built
To make the comparison reproducible, I built five pieces:
- A verified DVS-PedX CARLA RGB data pipeline
- A compact MobileNetV3 Small crossing classifier
- Interchangeable PyTorch, TorchScript, ONNX Runtime, and ExecuTorch runtime adapters
- A benchmark harness that preserves raw per-frame timing and prediction evidence
- An offline verifier that recomputes the published results from hashed evidence
The inference path is:
decoded forward-camera RGB frame
|
aspect-preserving resize and letterbox to 320 x 128
|
ImageNet-normalized FP32 tensor
|
MobileNetV3 Small crossing classifier
|
sigmoid plus frozen validation-selected threshold
|
CLEAR or CROSSING_ALERT
Guardian Angel is a present-frame perception component that could feed a larger ADAS or mobility stack. It does not control a vehicle, estimate collision risk, predict trajectories, or claim safety certification.
Why the optimization is credible
The main technical change is the CPU execution path.
The reference implementation uses PyTorch eager FP32. The selected implementation exports the same frozen model to ExecuTorch and delegates the graph to XNNPACK. I verify that the artifact contains an XNNPACK delegate call, verify the runtime thread count before model loading, and fail closed if the requested backend is unavailable. There is no silent fallback to eager execution.
I reproduced the result on two independent Arm64 systems:
- Apple M1 Max on macOS, using two verified XNNPACK threads
- AWS Graviton4 on Amazon Linux 2023, using four verified XNNPACK threads
The same dataset identity, split identity, checkpoint identity, preprocessing, threshold, and ExecuTorch artifact identity were checked before the Graviton run. Linux perf stat evidence records task-clock, cycles, instructions, cache behavior, branch misses, and context switches where available.
XNNPACK is portable, not Arm-exclusive. I describe the winning path as a verified Arm CPU deployment, not as custom NEON work. The same portable artifact produced a large relative gain on two different Arm microarchitectures.
Faster was not automatically better
Two INT8 candidates were faster than the selected FP32 runtime:
- ExecuTorch/XNNPACK PT2E INT8
- ONNX Runtime static INT8
Before ranking runtimes by speed, I set four acceptance limits on the validation split. Crossing recall and F1 could drop by no more than one percentage point, AUPRC could drop by no more than one percent relative, and event recall could not drop at all. Both INT8 versions crossed every limit, so I excluded them from the final result even though they were faster.
The selected FP32 XNNPACK runtime preserved every eager decision within each platform. Held-out Graviton quality was:
- Crossing recall: 0.717035
- F1: 0.795644
- Event recall: 19 / 20
- Confusion counts: TP 3,763, FP 448, TN 19,063, FN 1,485
The main limitation is also explicit. Recall was 0.920244 on standard scenarios and 0.327778 on adverse-weather scenarios. Guardian Angel is a systems optimization result with a real perception workload, not a claim of production-ready automotive accuracy.
Dataset discipline
I used the full synthetic CARLA RGB component of DVS-PedX under CC BY 4.0.
The acquired archive contains:
- 198 sequences
- 163,614 RGB frames
- 31,665 positive crossing frames
- 130 contiguous crossing events
- 117 standard sequences
- 81 adverse-weather sequences
I split by complete sequence, never by adjacent frame. The frozen split contains 139 training sequences, 29 validation sequences, and 30 held-out test sequences. Test feedback was not used to choose the model, threshold, runtime, or thread count.
The source paper describes 178,200 frames, while the verified distributed archive yielded 163,614. I documented the 14,586-frame discrepancy rather than fabricating or duplicating data.
Verifiable evidence
Every headline number can be recomputed from the raw evidence in the repository. The evidence package includes timing and prediction streams for M1 Max and Graviton4, platform records, Linux performance counters, sanitized AWS lifecycle provenance, benchmark configurations, and SHA-256 identities for the dataset, split, checkpoint, and runtime artifact.
Run:
git clone https://github.com/Shiv-aurora/Guardian-Angel
cd Guardian-Angel
python -m venv .venv
source .venv/bin/activate
pip install -e .
guardian-angel evidence verify --quick
The verifier independently recomputes:
- p50, p95, and p99 latency
- throughput and deadline violations
- confusion matrices and event metrics
- eager versus XNNPACK decision agreement
- the 5.201x M1 Max speedup
- the 3.784x Graviton4 speedup
- the rejection of the faster INT8 candidates
It runs offline after the AWS instance has been terminated.
Developer value
Other developers can reuse this workflow for latency-sensitive perception systems without letting the benchmark drift:
- Freeze task semantics and quality limits first
- Split temporally correlated data by sequence
- Bind data, model, preprocessing, threshold, runtime, and environment by identity
- Measure raw stage-level latency instead of one convenient average
- Verify that the requested Arm runtime is actually executing
- Reject candidates that improve speed but damage behavior
- Reproduce the relative improvement on another Arm platform
- Publish evidence that others can recompute
The same process can be applied to robot perception, industrial inspection, smart cameras, drones, and other sensor-to-decision systems where a late correct answer has little operational value.
What makes Guardian Angel stand out
The demo is simple: the same held-out camera frame produces the same alert under both runtimes, while the latency falls sharply. Behind that demo are reproducible data preparation, runtime adapters, fixed quality limits, cross-platform deployment scripts, raw evidence, performance-counter records, a one-command verifier, a comprehensive test suite, and documented failures.
The measured result I am submitting is:
On the same held-out crossing task, the optimized ExecuTorch/XNNPACK FP32 deployment reduced decoded-RGB-to-decision p95 latency by 80.77% on Apple M1 Max and 73.57% on AWS Graviton4, while preserving every eager decision within each platform.
Built with
Python, PyTorch, torchvision, MobileNetV3 Small, ExecuTorch, XNNPACK, ONNX Runtime, NumPy, OpenCV, pytest, Ruff, AWS EC2, AWS Graviton4, Amazon Linux 2023, Linux perf, DVS-PedX, and CARLA.
Links
- Source and reproduction: https://github.com/Shiv-aurora/Guardian-Angel
- Evidence quick start: https://github.com/Shiv-aurora/Guardian-Angel/tree/main/guardian-angel-evidence
- Cross-platform technical report: https://github.com/Shiv-aurora/Guardian-Angel/blob/main/reports/ARM_CROSS_PLATFORM_VALIDATION.md

Log in or sign up for Devpost to join the conversation.