ArmFlux Studio — AI Optimization for Arm

Don't just run AI on Arm. Make AI run better on Arm.

Track: Cloud AI

Project Overview

AI models are becoming increasingly capable, but deploying them efficiently on Arm64 infrastructure can require developers to manually experiment with precision, threading, graph optimization, runtimes, memory usage, and quality trade-offs.

ArmFlux Studio turns that process into a measurable, reproducible optimization workflow.

Instead of applying one optimization and assuming it is better, ArmFlux:

AI Model
   ↓
Analyze
   ↓
Establish FP32 Baseline
   ↓
Generate Valid Candidates
   ↓
Benchmark Each Candidate
   ↓
Apply Quality Constraints
   ↓
Compare Performance
   ↓
Pareto Analysis
   ↓
Select Best Configuration
   ↓
Explain Result
   ↓
Export & Reproduce

The goal is simple:

Automatically discover the best-performing configuration for an AI workload on an Arm64 target while preserving the required quality.

ArmFlux is designed as a reusable developer tool rather than a single-purpose AI application.


Why Cloud AI?

ArmFlux fits the Cloud AI track because its optimization workflow targets CPU-based Arm64 inference environments, including server-class Arm platforms such as Arm cloud infrastructure.

The same optimization engine can evaluate a model against different Arm64 target profiles and produce a reproducible configuration for deployment.

The project focuses on the exact optimization problems that matter when running AI on CPU-based infrastructure:

  • Inference latency
  • Throughput
  • Memory usage
  • Model size
  • CPU utilization
  • Model quality
  • Thread configuration
  • Graph optimization
  • Quantization
  • Developer workflow and reproducibility

The current prototype uses ONNX Runtime as its verified execution path and provides an Arm64 cloud target profile. It also detects whether the system is actually running on Arm hardware.

When running on a non-Arm development machine, ArmFlux explicitly enters Demo Mode rather than pretending the measurements are Arm measurements.


Inspiration

The inspiration came from a simple gap:

"AI works on Arm" is not the same as "AI is optimized for Arm."

Developers often have to manually test different quantization levels, thread counts, graph optimization settings, and runtimes to determine what actually performs best.

We wanted to turn that manual experimentation into an automated, evidence-driven process.

The key idea became:

What if an optimization tool could test the available configurations itself, measure the results, enforce a quality requirement, and tell the developer why one configuration won?

That became ArmFlux Studio.


What ArmFlux Studio Does

ArmFlux accepts a supported ONNX model and creates a complete optimization experiment.

1. Model Analysis

ArmFlux analyzes the model and reports:

  • Model architecture
  • Parameter count
  • Model size
  • Input/output information
  • Operator distribution
  • Potential computational bottlenecks

2. Baseline Measurement

Every optimization starts with a real FP32 baseline.

The baseline establishes the reference for:

  • Latency
  • Throughput
  • Memory
  • Model size
  • Accuracy where an evaluation set is provided

This prevents optimization claims from being made without a reference point.

3. Candidate Generation

ArmFlux generates hardware-aware candidate configurations.

The current verified prototype evaluates combinations involving:

  • FP32
  • INT8
  • Thread count
  • Graph optimization level
  • ONNX Runtime execution

Unsupported configurations are not presented as working optimizations.

4. Real Benchmarking

Every viable candidate is actually executed and measured.

ArmFlux records metrics such as:

  • Mean latency
  • p50 latency
  • p95 latency
  • p99 latency
  • Throughput
  • CPU utilization
  • Memory usage
  • Model size
  • Accuracy when an evaluation set is available

5. Accuracy-Aware Optimization

Performance is not allowed to win at any cost.

Developers can define a minimum acceptable accuracy threshold.

For example:

Minimum accuracy: 90%

Candidate A
Latency: 50 ms
Accuracy: 91.2%
→ Accepted

Candidate B
Latency: 38 ms
Accuracy: 86.7%
→ Rejected

This turns optimization into a constrained performance problem rather than simply selecting the fastest model.

6. Multi-Objective Optimization

Developers can optimize for:

  • Fastest inference
  • Highest throughput
  • Lowest memory
  • Smallest model
  • Balanced performance
  • Custom weighted objectives

The system evaluates the measured results according to the selected objective.

7. Pareto Analysis

ArmFlux identifies non-dominated configurations and visualizes the trade-off between performance and quality.

This allows developers to see configurations that provide different balances between:

  • Speed
  • Accuracy
  • Resource usage

Instead of hiding trade-offs, ArmFlux makes them visible.

8. Explainable Selection

After testing the candidates, ArmFlux explains why the winning configuration was selected.

For example:

INT8 + 4 threads + extended graph optimization was selected because it provided the best measured latency/throughput trade-off while remaining above the required accuracy threshold.

The explanation is generated from the actual benchmark results.

9. Reproducible Output

Every optimization run receives a unique Run ID.

ArmFlux stores:

  • Model hash
  • Target
  • Hardware snapshot
  • Optimization objective
  • Candidate configurations
  • Benchmark results
  • Accuracy threshold
  • Selected configuration
  • Optimization timeline

The winning configuration can be exported as a YAML artifact and reproduced through the CLI.


What We Optimized

The core optimization work focuses on real inference configuration choices, rather than simply converting a model.

Model-level optimization

The current verified prototype supports:

  • FP32 baseline
  • INT8 dynamic quantization through ONNX Runtime
  • Model-size comparison

Runtime-level optimization

ArmFlux evaluates:

  • ONNX Runtime execution
  • Graph optimization levels
  • Runtime configuration

CPU-level optimization

The candidate search evaluates different:

  • Intra-op thread counts
  • Parallel execution configurations

This matters because more threads do not always mean lower latency. Depending on the workload and hardware, additional parallelism can introduce overhead or contention.

Quality optimization

The system can enforce a minimum accuracy requirement and reject configurations that fall below it.

Developer workflow optimization

ArmFlux also optimizes the human side of AI deployment by turning manual experimentation into:

Profile → Search → Benchmark → Compare → Select → Export

This makes optimization experiments repeatable instead of dependent on undocumented trial and error.


Arm-Specific Design

ArmFlux is designed around Arm64-aware execution, rather than simply labeling a generic AI application as "Arm compatible."

The platform includes hardware detection for:

  • CPU architecture
  • CPU model
  • Physical/logical cores
  • RAM
  • Operating system
  • Detectable Arm CPU features

When running on Arm hardware, the system identifies the environment as:

ARM-NATIVE

When running elsewhere, it explicitly identifies:

DEMO MODE

This distinction is important because benchmark numbers from an x86 development machine should never be presented as Arm performance.

The project also provides Arm64 target profiles including:

  • Generic ARM64
  • Cortex-A76 class
  • Raspberry Pi-class Arm64
  • Android ARM64 artifact preparation
  • Arm Cloud Server

The primary Cloud AI target is the Arm64 cloud/server profile.


Benchmark Integrity

A major design principle of ArmFlux is:

Never fabricate an optimization result.

Every displayed benchmark result is generated by the benchmark engine.

The system records:

  • Warm-up runs
  • Timed iterations
  • Latency statistics
  • Resource measurements
  • Model size
  • Accuracy when an evaluation dataset is available
  • Hardware/environment information

The bundled demonstration workload is a small self-contained CNN trained on a synthetic-but-real task. Its measurements are genuine measurements of that workload, but they should not be interpreted as representative of production ImageNet or LLM performance.

For real Arm performance validation, the same workflow can be executed on an Arm64 environment.

ArmFlux also provides an export workflow for validating a winning configuration with Arm performance tooling such as Arm Performix. It does not claim a direct live Performix API integration in the current prototype.


How We Built It

Backend

  • Python
  • FastAPI
  • Pydantic
  • PyTorch
  • ONNX
  • ONNX Runtime
  • NumPy
  • psutil
  • PyYAML
  • Click

Frontend

  • React
  • TypeScript
  • Vite
  • Tailwind CSS
  • Recharts

Engineering

  • Docker
  • pytest
  • REST API
  • CLI
  • Modular optimization engine
  • JSON-based experiment storage

The backend separates the system into independent components:

Model Analyzer
      ↓
Profiler
      ↓
Candidate Generator
      ↓
Optimization Engine
      ↓
Benchmark Engine
      ↓
Evaluation Engine
      ↓
Pareto / Scoring
      ↓
Selection + Explanation
      ↓
Storage + Export

The frontend, REST API, and CLI all use the same core optimization pipeline, preventing different interfaces from producing different optimization behavior.


Setup Instructions

Requirements

  • Python 3.11+
  • Node.js 18+
  • npm
  • Linux/macOS/WSL recommended
  • For Arm-native validation: an Arm64 environment or Arm-powered device

1. Clone the repository

git clone <YOUR_GITHUB_REPOSITORY_URL>
cd armflux-studio

2. Run the setup script

bash scripts/setup.sh

This installs the backend/frontend dependencies, generates the bundled example model and evaluation set, and runs the test suite.

3. Start the backend

uvicorn backend.api.main:app --reload --port 8000

Open:

http://localhost:8000/docs

for the interactive API documentation.

4. Start the frontend

cd frontend
npm install
npm run dev

Open:

http://localhost:5173

5. Run the CLI

Check the detected hardware:

python3 cli/armflux.py hardware

Profile the bundled model:

python3 cli/armflux.py profile \
  examples/vision/armflux_demo_vision.onnx

Run an optimization experiment:

python3 cli/armflux.py optimize \
  examples/vision/armflux_demo_vision.onnx \
  --target arm_cloud \
  --objective balanced \
  --min-accuracy 0.90 \
  --eval-set examples/vision/eval_set.npz

List previous runs:

python3 cli/armflux.py list

Generate a report:

python3 cli/armflux.py report <RUN_ID>

Export the winning configuration:

python3 cli/armflux.py export <RUN_ID>

6. Docker

docker compose up --build

For Arm-native validation, build/run the deployment for an Arm64 environment:

docker build --platform linux/arm64 .

The benchmark should then be executed on the actual Arm64 environment so the resulting numbers represent the target hardware rather than the development machine.


Example Optimization Workflow

A typical ArmFlux experiment looks like:

1. Upload ONNX model
        ↓
2. Select ARM64 target
        ↓
3. Set accuracy threshold
        ↓
4. Select optimization objective
        ↓
5. Establish FP32 baseline
        ↓
6. Generate candidate configurations
        ↓
7. Quantize / configure candidates
        ↓
8. Benchmark candidates
        ↓
9. Reject candidates below quality threshold
        ↓
10. Compute Pareto frontier
        ↓
11. Select optimal configuration
        ↓
12. Explain the result
        ↓
13. Export reproducible configuration

The important part is that ArmFlux does not assume which optimization will win.

It measures the candidates and lets the data decide.


Challenges We Ran Into

The hardest part was discovering that optimization is not a single transformation.

INT8 quantization may reduce model size without necessarily producing the best latency on every workload.

Similarly, increasing the thread count can sometimes make performance worse instead of better.

This forced us to treat optimization as a search problem.

Hardware variability was another challenge.

Arm spans mobile processors, embedded systems, edge devices, and high-core-count cloud CPUs. A configuration that performs well on one Arm target cannot automatically be assumed to perform well on another.

We therefore built a target-aware configuration layer and explicitly distinguish between:

  • Actual detected hardware
  • User-selected target profiles
  • Demo/simulation mode

Benchmark methodology was another major challenge.

To make before/after comparisons meaningful, we needed controlled warm-up, repeated measurements, consistent inputs, and reproducible configurations.


Accomplishments We're Proud Of

We are particularly proud that ArmFlux is designed around evidence instead of assumptions.

The system can show:

What was tested
       ↓
What changed
       ↓
What improved
       ↓
What was rejected
       ↓
Why the winner was selected
       ↓
How to reproduce it

Key accomplishments include:

  • Real ONNX model analysis
  • Real FP32 baseline measurement
  • Real INT8 quantization
  • Target-aware candidate generation
  • Thread and graph optimization search
  • Real benchmark execution
  • Accuracy-aware rejection
  • Pareto frontier analysis
  • Composite ArmFlux performance scoring
  • Hardware detection
  • Reproducible experiment IDs
  • Exportable optimization configurations
  • REST API
  • CLI
  • Developer dashboard
  • Automated test suite
  • Explicit Demo Mode
  • Arm64 validation workflow

Most importantly:

ArmFlux makes AI optimization itself the product.


What We Learned

We learned that efficient AI is fundamentally a systems problem.

Model architecture matters, but so do:

  • Numerical precision
  • Memory behavior
  • Operator implementation
  • Runtime configuration
  • CPU parallelism
  • Hardware capabilities
  • Workload characteristics
  • Quality/performance trade-offs

We also learned that "faster" is not automatically "better."

An optimization that reduces latency but violates the required accuracy threshold is not a successful optimization.

ArmFlux therefore treats optimization as a constrained multi-objective problem:

$$ \text{Best Configuration} = \arg\min\left(\text{Latency}, \text{Memory}, \text{Model Size}\right) $$

subject to:

$$ \text{Quality} \geq \text{Required Threshold} $$

For throughput-oriented workloads, the optimization objective can instead prioritize maximizing throughput while respecting the same quality constraints.

This changed our approach from:

"How do we make this model faster?"

to:

"Which configuration gives this workload the best measurable trade-off on this target?"


What Makes ArmFlux Different

Most AI deployment workflows follow:

Model
 ↓
Convert
 ↓
Deploy

ArmFlux follows:

Model
 ↓
Profile
 ↓
Generate alternatives
 ↓
Measure
 ↓
Reject bad trade-offs
 ↓
Compare
 ↓
Find Pareto-optimal configurations
 ↓
Select
 ↓
Explain
 ↓
Reproduce

The developer does not have to guess which configuration is best.

ArmFlux provides evidence.


Why It Could Be Reused

ArmFlux is designed as a reusable optimization layer rather than a one-off benchmark.

Developers can extend it with:

  • New model architectures
  • New evaluation metrics
  • New Arm targets
  • New runtimes
  • New optimization strategies
  • New objective functions
  • New benchmark workloads

The optimization pipeline is intentionally separated from the UI, meaning the same engine can be used through:

  • Web dashboard
  • REST API
  • CLI
  • Automated CI/CD workflows

This makes the project useful beyond the competition.


Current Limitations

We intentionally avoid claiming capabilities that are not verified in the current prototype.

Current verified optimization path

  • ONNX
  • ONNX Runtime
  • FP32
  • INT8
  • Thread tuning
  • Graph optimization

Not currently implemented

  • FP16 execution
  • INT4 execution
  • In-process Android execution
  • In-process ExecuTorch execution
  • Direct XNNPACK execution
  • Direct Arm Performix API integration
  • Production multi-user database storage

Android artifact preparation is supported, but actual mobile execution requires a connected Arm-powered device or emulator.

The bundled demonstration was developed on an x86_64 host, so its benchmark numbers are not presented as Arm hardware measurements.


What's Next

The next stage for ArmFlux is to validate and expand the optimization engine on real Arm64 infrastructure.

Planned improvements include:

  • Native Arm64 benchmark runs
  • Arm cloud benchmarking
  • Additional Arm-optimized runtimes
  • More model architectures
  • Energy-aware optimization
  • More advanced optimization search
  • Automated hardware capability detection
  • Larger language-model workloads
  • TTFT and tokens/sec benchmarking
  • Mobile/edge deployment workflows
  • Continuous performance regression testing
  • Distributed Arm benchmarking
  • More deployment artifacts
  • Public benchmark datasets and recipes

The long-term vision is:

ArmFlux Studio becomes an AI performance engineering layer that automatically discovers how to run a workload most efficiently on its target Arm platform.


Final Takeaway

AI optimization should not be based on guesswork.

Developers should be able to ask:

"What is the best way to run this model on my Arm target?"

and receive:

A measured configuration, a performance comparison, an explanation, and a reproducible artifact.

That's ArmFlux Studio.

Profile. Optimize. Prove.

Built With

Share this project:

Updates

Submission history