Inspiration

AI inference on Arm64 is not simply a question of choosing a faster runtime. The same workload can behave very differently depending on how requests are scheduled, how runtime state is reused, and which execution backend is used.

We wanted to answer a simple question:

What if an optimization system could measure different execution strategies under the same workload and choose based on evidence instead of assumptions?

That became the Arm AI Optimization Harness.

What it does

The harness runs the same workload through different execution strategies and runtime adapters, measures latency, throughput, and wall-clock time, verifies the resulting evidence, and feeds those measurements into a scheduler.

The core comparison is:

Sequential execution — fresh execution context per request.

vs.

Dataflow execution — reusable workers / parallel execution across requests.

The important part is that the workload and output contract remain comparable while the execution environment changes.

The system then produces machine-verifiable evidence rather than relying on a manually selected benchmark result.

How we built it

The project is intentionally split into small, replaceable layers:

  • Workload layer — deterministic inference workloads.
  • Runtime adapters — Ollama, llama-server, and JSONL-compatible runtimes.
  • Execution harness — runs equivalent workloads under different execution strategies.
  • Evidence layer — records inputs, outputs, platform information, metrics, and SHA256 hashes.
  • Scheduler — compares measured configurations and selects a strategy from the evidence.
  • Cross-engine implementations — Python and native Rust implementations share the same inference contract.
  • Arm64 CI — benchmarks run on a real aarch64 environment rather than being labeled as Arm64 from a development machine.

The result is a portable benchmark harness rather than a benchmark tied to one runtime.

What we found

The most important result was not that every optimization worked.

Several configurations did not improve performance.

On the Arm64 CI environment, Ollama configurations were measured and ruled out when they failed to provide a meaningful speedup. A llama-server configuration using parallel request slots produced approximately 1.39× wall-time speedup under the tested workload.

The harness did not assume that result beforehand. It measured the alternatives and selected from the resulting evidence.

That distinction is the central idea of the project:

Don't assume an optimization is better. Measure it.

Challenges

The hardest part was making the comparison trustworthy.

It is easy to produce a benchmark that compares different runtimes. It is much harder to ensure that the workload, output contract, platform information, execution modes, and measurement data remain comparable.

We also had to prevent synthetic/demo results from being confused with real Arm64 measurements. The final system separates demo evidence from real Arm64 evidence and verifies the architecture recorded by the benchmark environment.

Another challenge was testing the harness itself against unusual execution engines. We used Rust for native conformance and even subjected the harness to pathological Malbolge workloads to verify that self-modifying code, illegal execution, and non-termination do not break the measurement layer.

What we learned

The biggest lesson was that optimization is contextual.

A configuration that looks promising in one environment or metric may be worse in another. Wall time, per-request latency, throughput, runtime behavior, and hardware architecture can tell different stories.

We also learned that reproducibility is part of the optimization problem itself. A performance claim is much more useful when another person can inspect the workload, platform, measurements, hashes, and decision path that produced it.

What's next

The next step is to expand the harness across more Arm64 hardware, larger models, longer workloads, and additional runtimes.

We also want the scheduler to become increasingly useful as more evidence accumulates, while keeping the core principle unchanged:

the system should choose based on measured execution behavior, not on assumptions about which runtime or strategy should be faster.

Built With

  • ai-inference
  • arm64
  • benchmarking
  • cli
  • cobol
  • dataflow
  • json
  • linux
  • llm
  • open-source
  • performance
  • performance-optimization
  • python
  • reproducibility
  • runtime-adapters
  • rust
  • systems
  • wasm
Share this project:

Updates

Submission history