Inspiration

AI systems deployed on robots and edge devices usually run one fixed model configuration even though their real operating conditions constantly change. CPU load rises, thermal headroom falls, queues build up, and latency budgets tighten — but the inference stack typically keeps running exactly the same way.

PerceptShift was built around a simple idea: an Arm-powered physical AI system should be able to use multiple pre-certified inference profiles and adapt between them without sacrificing correctness or silently violating quality constraints.

What it does

PerceptShift is a local-first, deadline-aware adaptive AI inference platform for Arm64 robotics and edge systems.

It has two core components:

PerceptShift Forge profiles and certifies ONNX Runtime execution profiles against representative data. Rather than benchmarking inference in isolation, Forge measures the complete production execution path: preprocessing, inference, and postprocessing. Candidates are evaluated for latency, memory usage, model compatibility, output equivalence, quality constraints, provider behavior, and bundle integrity.

PerceptShift Runtime loads certified profiles ahead of time and dynamically selects between them according to current runtime conditions and application constraints. It tracks queue pressure, deadlines, profile health, failures, dwell time, cooldowns, and policy state. If no certified execution profile remains feasible, the runtime fails closed and requests a downstream control hold rather than silently using an unsafe or uncertified configuration.

PerceptShift integrates directly with ROS 2 Jazzy. A ROS image can flow through the real production pipeline:

sensor input → Arm preprocessing → ONNX Runtime → task adapter → normalized ROS output → runtime telemetry

The current release supports:

  • raw tensor inference
  • image classification
  • YOLOv8-style object detection

Arm optimization

PerceptShift includes native Arm64-specific engineering rather than simply running an existing AI application on Arm.

Its preprocessing pipeline has both a scalar reference implementation and an optimized AArch64 NEON implementation with runtime feature dispatch. The same versioned preprocessing contract is used for calibration, benchmarking, certification, and deployed inference so that the path Forge measures matches the path the runtime actually executes.

Execution profiles can also control ONNX Runtime CPU/XNNPACK behavior, thread configuration, CPU affinity, preprocessing strategy, and model variant.

Forge records provider evidence, complete executor latency, quality measurements, peak memory, environment information, model hashes, preprocessing hashes, and other evidence into reproducible profile bundles.

How we built it

PerceptShift combines C++20 and Python.

The latency-critical runtime is implemented in C++ using ONNX Runtime, native Arm preprocessing, bounded queues, deterministic profile switching, cryptographic bundle validation, and ROS 2 Jazzy integration.

The Forge optimization and certification layer is implemented in Python and orchestrates candidate generation, calibration, benchmarking, quality evaluation, Pareto selection, evidence generation, and profile-bundle construction.

The product also includes:

  • a local FastAPI service
  • a React operational console
  • ROS 2 lifecycle nodes and custom interfaces
  • Ed25519 bundle verification
  • Debian packaging for Ubuntu 24.04 Arm64
  • OCI container builds
  • systemd integration
  • reproducible clean-room validation tooling

Challenges we ran into

The hardest problem was ensuring that optimization evidence actually represented the production execution path.

It is easy to report a faster model while accidentally excluding preprocessing, postprocessing, provider fallback, queueing behavior, or different input transformations from the benchmark.

PerceptShift therefore uses a shared native execution path for offline benchmarking and online execution. We also built explicit quality and equivalence gates so that a candidate cannot become deployable simply because it is faster.

Arm64 validation also required careful handling of architecture-specific preprocessing, NEON equivalence, ONNX Runtime execution providers, ROS 2 integration, packaging, and clean-room Ubuntu Arm64 environments.

Accomplishments

The current PerceptShift release passes a 26-tier software verification suite covering:

  • native C++ and Python correctness
  • schema and contract testing
  • scalar/NEON preprocessing equivalence
  • ONNX Runtime executor tests
  • Forge end-to-end certification
  • quantization/calibration equivalence
  • quality degradation gates
  • adaptive controller behavior
  • cryptographic bundle integrity
  • ROS 2 Jazzy builds and real runtime inference
  • API → ROS → runtime integration
  • browser → API → ROS → ONNX Runtime end-to-end inference
  • Arm64 Debian package installation and removal
  • Arm64 OCI execution
  • ASan/UBSan
  • ThreadSanitizer
  • fuzz testing
  • coverage enforcement
  • clean-room Ubuntu 24.04 Arm64 reconstruction

We intentionally distinguish Arm64 software-correctness testing under virtualization/emulation from physical-device performance measurements. PerceptShift does not fabricate hardware performance, power, or thermal results.

What we learned

The biggest lesson was that AI optimization is not just model optimization.

For physical AI, the useful metric is the behavior of the complete execution path. Preprocessing, memory allocation, runtime provider selection, task-specific postprocessing, queue age, latency tails, and failure recovery can matter just as much as raw model inference time.

We also learned that optimization tooling becomes significantly more useful when its output is reusable. PerceptShift therefore produces portable, integrity-protected execution profiles and evidence rather than a one-off benchmark report.

What's next

PerceptShift is designed to become a reusable optimization and adaptive-execution layer for Arm-powered robotics, smart cameras, autonomous systems, and industrial edge AI.

Future work includes characterization on additional physical Arm platforms, broader execution-provider support, additional profile-search strategies, and certification against larger production workloads while keeping the same evidence-first architecture.

Built With

Share this project:

Updates

Submission history