Inspiration
A kestrel hovers nearly motionless, spending almost nothing, then strikes only when something moves, only where it moved. Embedded vision does the opposite: it runs a full neural network on every frame of a scene that is static 95% of the time. I wanted to build the kestrel version: a detector that asks "should inference run at all, and on which pixels?" before spending a single millijoule on it.
What it does
Kestrel is a complete sense-decide-act loop on two Arm microcontrollers:
- Level 1: a PIR sensor on an RP2350 (dual Cortex-M33) watches the room while the main board sleeps. On motion it pulses a wake line.
- Level 2: an STM32H750 (Cortex-M7, 480 MHz) wakes and runs a sub-millisecond SIMD frame-difference gate. Unchanged scene: back toward sleep. Changed: it localizes WHERE.
- Level 3: only then does the INT8 person detector run (X-CUBE-AI, st_yololcv1 192x192), and only on the motion region: the gate's ROI is cropped into the model's fixed-size input, acting as a digital zoom that gives distant subjects 5x more pixels at identical inference cost.
- Act: a confirmed person crosses a UART link to the RP2350, which drives the servo through a continuous patrol sweep for as long as the person stays in frame. The whole chain runs live on a desk.
Headline numbers, all measured on hardware (DWT cycle counter, FNIRSI FNB-C2 ammeter), none projected:
- 98.2% (daylight) / 99.1% (night) of frames skip inference on real scenes, filmed 20 and 12.5 minute runs, with zero accuracy loss: active frames always get full inference.
- The gate decision costs 173 us against the 178-181 ms inference it may avoid: roughly 1000x cheaper.
- Gating makes the system faster AND cooler at once: 5 to 15 FPS, 1.24 to 0.95 W, and a measured 10 C die-temperature drop.
- Deep sleep with the camera hard powered down: 243 to 82 mA, a 2.96x idle reduction. The entire armed two-board cascade idles at 174 mA: less than the single main board running always-on.
How I built it
WeAct STM32H750VBT6 (480 MHz core, app executing in place from QSPI flash behind a USB bootloader), OV2640 camera over DCMI+DMA, ST7735 LCD as the live HUD and debug console, X-CUBE-AI 10.2.1 runtime, and an RP2350-USB Mini running both cores: PIR pre-screen on core 0, UART event parsing and servo control on core 1. The motion gate is a dependency-free C99 module (SIMD USADA8 path on M7/M33) with host-runnable tests and CI. Power engineering went down to solder-bridge level: the camera's PWDN line is driven through the board's SB1 bridge for true hardware power-down in STOP mode, plus panel sleep-in and a low-power-regulator STOP for the MCU. All of it, hardware bring-up, firmware, benchmarks, and documentation, was designed, built, and measured during the challenge period; the repository's public commit history is the timeline.
Challenges I ran into
The honest list is long and it is all documented in the repo's troubleshooting guide so nobody has to rediscover it: a bootloader that rejects textbook linker scripts, CubeMX silently reverting the voltage scaling that 480 MHz requires, a QSPI execute-in-place ceiling that hard-faults above 120 MHz bus clock, an MPU attribute that quietly made inference 5x slower, a camera that resumes streaming after power-down but with wedged black-level calibration (dark, purple video) until a full re-init, and an RP2350 silicon gotcha: INTERP blend mode takes its alpha from lane 1's masked result, and the reset-state mask silently truncates it to 1 bit, degrading bilinear to nearest-neighbor while looking plausible.
Accomplishments I'm proud of
- A measured 98-99% compute reduction with zero accuracy loss, from a 173 us decision.
- The whole sleeping cascade drawing less than one always-on board.
- Publishing my own negative result: I benchmarked the RP2350 hardware interpolator for ML input resize hoping for a win and measured 0.90x (software at -O2 is faster). It is in the repo with full data, along with the silicon configuration trap the benchmark uncovered. I believe the measurement-driven honesty is worth more than a fabricated speedup. Every number earned on silicon; every result published, wins and losses alike.
- Reproducibility: every number traces to committed harness code, CSV or filmed on-device counters, with methodology and threats-to-validity sections in the benchmark report.
What I learned
System-level optimization beats model-level optimization when the workload is bursty: no amount of quantization reaches a 1000x saving, but refusing to run the model does. Also: measure before believing, including your own intuitions; two of mine (a software camera standby and a hardware-accelerated resize) were overruled by the hardware, and both failures produced documentation more valuable than the features.
What's next for Kestrel
A pan-tracking mount (the detection box already carries the target's position; the servo could carry the camera), a resolution ladder to make inference cost scale with motion area, and RP2350 dormant-mode work to push the watcher below its current 92 mA.


Log in or sign up for Devpost to join the conversation.