Inspiration

A kestrel hovers nearly motionless, spending almost nothing, then strikes only when something moves, and only where it moved. Embedded vision does the opposite. A battery camera, a trail camera, a porch camera, an occupancy sensor: each runs a full neural network on every frame of a scene that is static most of the time, and pays for it in battery size, solar panel size, heat, and how long it can be left alone. I wanted to build the kestrel version: a detector that asks "should inference run at all, and on which pixels?" before spending anything on it.

What it does

Kestrel is a person-detecting camera that runs as a complete sense, decide, act loop on two microcontrollers:

Level Runs on Cost The question it answers
1. PIR sensor RP2350 (Cortex-M33), main board asleep 92 mA for the whole watcher board Is anything happening at all?
2. Frame-difference gate STM32H750 (Cortex-M7), just woken 173 microseconds per frame Did the scene actually change, and where?
3. INT8 person detector STM32H750 180 ms per run What is it?

Each level is orders of magnitude cheaper than the one it protects. Unchanged scene: the main board goes back toward sleep having spent a fraction of a millisecond. Changed scene: the gate returns the bounding box of what changed, and the detector runs on that region only. The model input is a fixed 192 by 192 pixels, so cropping does not make inference cheaper; it makes it better. The moving object fills the model's field of view instead of being shrunk with the whole frame: a distant person that would occupy about 8 pixels in a full-frame downscale occupies 40 or more in the crop, at identical cost. A free digital zoom.

Isn't this what a video doorbell already does? Not quite. A commercial PIR-wake camera wakes on any warm thing and then either streams video to a cloud model or runs the full network on every frame it captures. Kestrel adds a second gate that costs 173 microseconds, runs the network only on the pixels that actually moved, and does all of it on a microcontroller: no cloud, no Wi-Fi, no video leaving the device.

When the detector confirms a person, the event crosses a UART link to the RP2350, which drives a servo through a patrol sweep for as long as the person stays in frame. In a trail or remote-site deployment that same output pans a deterrent light or points a higher-resolution camera at the target; the detection already carries the target's position, so the servo is the hook, not the point.

Headline numbers, all measured on the hardware with a cycle counter and an inline power meter, none projected:

Measurement Value
Frames that skip inference on real indoor scenes (filmed: 20 min daylight, 12.5 min night) 98.2% / 99.1%
Average per-frame compute reduction about 56x daylight, about 115x night
Gate decision vs the inference it may avoid 173 microseconds vs 178 to 181 ms, roughly 1000x cheaper
Throughput, conventional vs gated 5 FPS to 15 FPS
Main board power, conventional vs gated and awake 1.24 W to 0.95 W
Main board asleep (STOP mode, panel off, camera hard powered down) 243 mA to 82 mA, 2.96x lower
Die temperature, 10 min conventional vs 10 min gated 46 C vs 36 C
Whole armed cascade idling, both boards 174 mA, below one board running the conventional way

No loss on moving subjects: every frame with motion still gets full inference, and the savings come entirely from frames where nothing happened. A subject that stops moving closes the gate; the last result is held (box for 3 seconds, servo for 2.5 seconds) and the subject is re-confirmed on its next motion.

What the currents mean in battery hours (arithmetic from the measured currents on a 2,000 mAh pack at the meter's 5 V rail, converter losses ignored; not a measured runtime):

State Current Hours on a 2,000 mAh pack
Conventional always-on, one board 243 mA about 8
Armed cascade idling, both boards 174 mA about 11.5
Main board asleep, this build 82 mA about 24
The same board in its maker's standby test 0.9 mA about 90 days

The last row is why the 82 mA floor is a development-board number, not an architecture number: in this build the board's regulators (a buck converter at light load feeding an LDO chain), the power LED, the camera module's supply rail and the display in its sleep state all stay powered. The compute and thermal savings above do not depend on that floor.

How I built it

Hardware: a WeAct STM32H750VBT6 board (Cortex-M7 at 480 MHz, 1 MB SRAM, 8 MB QSPI flash) with its onboard OV2640 camera over DCMI and DMA (the chip's camera port, frames copied to memory without the CPU) and a 0.96 inch ST7735 display as the live HUD; an RP2350-USB Mini (dual Cortex-M33) running the PIR pre-screen on core 0 and UART event parsing plus servo control on core 1; an HC-SR501 PIR sensor; an SG90 servo. Four wires connect the boards: UART both ways, a GPIO wake line, and ground. The PIR and the servo hang off the RP2350.

Firmware: bare-metal C, no operating system, in STM32CubeIDE, executing in place from QSPI flash (running straight out of the external flash chip) behind a USB bootloader, so no debug probe is needed to flash it. The detector is an INT8 person model from the ST Model Zoo (st_yololcv1, 192 by 192 input, 328 KiB flash, 160 KiB RAM, 62 million multiply-accumulates, 34.7% AP on COCO person) on the X-CUBE-AI 10.2.1 runtime (ST's neural-network runtime). The motion gate is my own module: about 240 lines of dependency-free C99, no vendor driver layer, no malloc, no floating point, resolution independent, with an optional Cortex-M7 SIMD path (USADA8, one Arm instruction that sums four pixel differences at once) that measured 3.7x faster than the scalar version (99 vs 364 microseconds in a warm benchmark loop) and is bit-exact with it. It has host-runnable tests and CI; anyone with a Cortex-M camera pipeline can drop it in.

Power engineering went down to solder-bridge level. The camera's power-down pin is routed through a solder bridge on the board (SB1) to one CPU pin (PA7); bridging it and driving it from deep sleep is what takes the board from 243 mA to 82 mA, together with the display's sleep-in command and the MCU's low-power-regulator STOP mode (its deep sleep).

Measurement: every timing number is from the Cortex-M7 DWT cycle counter (the CPU's built-in clock-cycle counter); every power number is from a FNIRSI FNB-C2 inline USB meter (20-bit ADC, 1 microamp resolution, published accuracy 0.05% plus 2 counts). The skip rates come from the firmware's own counters, streamed as CSV every 16 frames and filmed on the device. Raw evidence, the harness code, and a methodology section with threats to validity are in the repository.

Timeline and disclosure: everything here, hardware bring-up, firmware, benchmarks, and documentation, was designed, built, and measured between July 10 and August 14, 2026, inside the VoltHacks window; the public commit history is the timeline. The same build was also entered in the Arm AI Optimization Challenge, which closed August 15. Nothing in this submission predates the VoltHacks period.

Challenges I ran into

The biggest one was making the camera sleep and wake correctly: its software standby left the video frozen after wake, so power-down had to be done in hardware through a solder bridge on the board, and the camera then woke with wedged calibration (dark, purple video) until a full re-initialization on wake was added. The rest of the honest list, each written up in the repository's troubleshooting guide so nobody has to rediscover it:

  • a USB bootloader that rejects textbook linker scripts;
  • CubeMX (ST's code generator) silently reverting the voltage scaling that 480 MHz requires, which bricks the board on the next build;
  • a ceiling on how fast code can run straight out of the external flash chip: above a 120 MHz bus clock the board hard-faults;
  • one memory-protection attribute (a cache setting) that quietly made inference 5x slower;
  • a compiler pragma that blocked helper-function inlining and more than doubled the SIMD gate's cost (109 to 249 microseconds); the benchmark harness caught it on its first run and the change was reverted;
  • an RP2350 silicon gotcha: the hardware interpolator's blend mode takes its alpha from lane 1's masked result, and the reset-state mask truncates it to 1 bit, silently degrading bilinear to nearest-neighbor while looking plausible.

Accomplishments that I'm proud of

  • A measured 98 to 99% reduction in inference compute on real scenes, from a decision that costs a thousandth of what it saves, with no loss on moving subjects.
  • The entire sleeping cascade, both boards armed, drawing less than one board running the conventional way.
  • Publishing my own negative result. I benchmarked the RP2350's hardware interpolator for resizing model input, hoping for a speedup, and measured 0.90x: software at -O2 is faster (1554 vs 1736 microseconds). It is in the repository with full data and the silicon trap the benchmark uncovered. A loss you publish is how a reader knows the wins are real.
  • Reproducibility: every number traces to committed harness code, a CSV, or filmed on-device counters, and the benchmark report has a threats-to-validity section.
  • A reusable artifact. Motion-gated inference has been in research papers since 2017; I could not find a bare-metal Cortex-M implementation anywhere public as of July 2026, so I wrote one and shipped it as a standalone, host-tested C99 module.

What I learned

System-level optimization beats model-level optimization when the workload is bursty. Quantizing, pruning, and shrinking a model make one inference cheaper; none of them reach a 1000x saving. Refusing to run the model does. And measure before believing, including your own intuitions: two of mine (a software camera standby and a hardware-accelerated resize) were overruled by the hardware, and both failures produced documentation worth more than the features would have been.

What's next for Kestrel: Zero Wasted Inference

  • A resolution ladder: 192, 128, and 96 pixel variants of the detector dispatched by motion-region size, so inference cost genuinely scales with how much moved.
  • RP2350 dormant mode to push the watcher board below its current 92 mA, and hardware windowing inside the OV2640 so the crop saves camera bandwidth as well.
  • Off the development boards. The 82 mA sleep floor belongs to the dev boards (regulator chain at light load, power LED, camera supply rail, sleeping display), not to the architecture; the same board reaches 0.9 mA in its maker's standby test. A purpose-built board with a proper low-power supply and a radio is where this becomes a deployable trail camera.

Built With

  • arduino
  • arm-cortex-m33
  • arm-cortex-m7
  • c
  • cmsis
  • computer-vision
  • edge-ai
  • embedded-systems
  • low-power
  • ov2640
  • pico-sdk
  • pir-sensor
  • python
  • rp2350
  • servo
  • stm32
  • stm32cubeide
  • stm32h750
  • x-cube-ai
Share this project:

Updates

Submission history