ArmFlux Studio — AI Optimization for Arm
Don't just run AI on Arm. Make AI run better on Arm.
Track: Cloud AI
Project Overview
AI models are becoming increasingly capable, but deploying them efficiently on Arm64 infrastructure can require developers to manually experiment with precision, threading, graph optimization, runtimes, memory usage, and quality trade-offs.
ArmFlux Studio turns that process into a measurable, reproducible optimization workflow.
Instead of applying one optimization and assuming it is better, ArmFlux:
AI Model
↓
Analyze
↓
Establish FP32 Baseline
↓
Generate Valid Candidates
↓
Benchmark Each Candidate
↓
Apply Quality Constraints
↓
Compare Performance
↓
Pareto Analysis
↓
Select Best Configuration
↓
Explain Result
↓
Export & Reproduce
The goal is simple:
Automatically discover the best-performing configuration for an AI workload on an Arm64 target while preserving the required quality.
ArmFlux is designed as a reusable developer tool rather than a single-purpose AI application.
Why Cloud AI?
ArmFlux fits the Cloud AI track because its optimization workflow targets CPU-based Arm64 inference environments, including server-class Arm platforms such as Arm cloud infrastructure.
The same optimization engine can evaluate a model against different Arm64 target profiles and produce a reproducible configuration for deployment.
The project focuses on the exact optimization problems that matter when running AI on CPU-based infrastructure:
- Inference latency
- Throughput
- Memory usage
- Model size
- CPU utilization
- Model quality
- Thread configuration
- Graph optimization
- Quantization
- Developer workflow and reproducibility
The current prototype uses ONNX Runtime as its verified execution path and provides an Arm64 cloud target profile. It also detects whether the system is actually running on Arm hardware.
When running on a non-Arm development machine, ArmFlux explicitly enters Demo Mode rather than pretending the measurements are Arm measurements.
Inspiration
The inspiration came from a simple gap:
"AI works on Arm" is not the same as "AI is optimized for Arm."
Developers often have to manually test different quantization levels, thread counts, graph optimization settings, and runtimes to determine what actually performs best.
We wanted to turn that manual experimentation into an automated, evidence-driven process.
The key idea became:
What if an optimization tool could test the available configurations itself, measure the results, enforce a quality requirement, and tell the developer why one configuration won?
That became ArmFlux Studio.
What ArmFlux Studio Does
ArmFlux accepts a supported ONNX model and creates a complete optimization experiment.
1. Model Analysis
ArmFlux analyzes the model and reports:
- Model architecture
- Parameter count
- Model size
- Input/output information
- Operator distribution
- Potential computational bottlenecks
2. Baseline Measurement
Every optimization starts with a real FP32 baseline.
The baseline establishes the reference for:
- Latency
- Throughput
- Memory
- Model size
- Accuracy where an evaluation set is provided
This prevents optimization claims from being made without a reference point.
3. Candidate Generation
ArmFlux generates hardware-aware candidate configurations.
The current verified prototype evaluates combinations involving:
- FP32
- INT8
- Thread count
- Graph optimization level
- ONNX Runtime execution
Unsupported configurations are not presented as working optimizations.
4. Real Benchmarking
Every viable candidate is actually executed and measured.
ArmFlux records metrics such as:
- Mean latency
- p50 latency
- p95 latency
- p99 latency
- Throughput
- CPU utilization
- Memory usage
- Model size
- Accuracy when an evaluation set is available
5. Accuracy-Aware Optimization
Performance is not allowed to win at any cost.
Developers can define a minimum acceptable accuracy threshold.
For example:
Minimum accuracy: 90%
Candidate A
Latency: 50 ms
Accuracy: 91.2%
→ Accepted
Candidate B
Latency: 38 ms
Accuracy: 86.7%
→ Rejected
This turns optimization into a constrained performance problem rather than simply selecting the fastest model.
6. Multi-Objective Optimization
Developers can optimize for:
- Fastest inference
- Highest throughput
- Lowest memory
- Smallest model
- Balanced performance
- Custom weighted objectives
The system evaluates the measured results according to the selected objective.
7. Pareto Analysis
ArmFlux identifies non-dominated configurations and visualizes the trade-off between performance and quality.
This allows developers to see configurations that provide different balances between:
- Speed
- Accuracy
- Resource usage
Instead of hiding trade-offs, ArmFlux makes them visible.
8. Explainable Selection
After testing the candidates, ArmFlux explains why the winning configuration was selected.
For example:
INT8 + 4 threads + extended graph optimization was selected because it provided the best measured latency/throughput trade-off while remaining above the required accuracy threshold.
The explanation is generated from the actual benchmark results.
9. Reproducible Output
Every optimization run receives a unique Run ID.
ArmFlux stores:
- Model hash
- Target
- Hardware snapshot
- Optimization objective
- Candidate configurations
- Benchmark results
- Accuracy threshold
- Selected configuration
- Optimization timeline
The winning configuration can be exported as a YAML artifact and reproduced through the CLI.
What We Optimized
The core optimization work focuses on real inference configuration choices, rather than simply converting a model.
Model-level optimization
The current verified prototype supports:
- FP32 baseline
- INT8 dynamic quantization through ONNX Runtime
- Model-size comparison
Runtime-level optimization
ArmFlux evaluates:
- ONNX Runtime execution
- Graph optimization levels
- Runtime configuration
CPU-level optimization
The candidate search evaluates different:
- Intra-op thread counts
- Parallel execution configurations
This matters because more threads do not always mean lower latency. Depending on the workload and hardware, additional parallelism can introduce overhead or contention.
Quality optimization
The system can enforce a minimum accuracy requirement and reject configurations that fall below it.
Developer workflow optimization
ArmFlux also optimizes the human side of AI deployment by turning manual experimentation into:
Profile → Search → Benchmark → Compare → Select → Export
This makes optimization experiments repeatable instead of dependent on undocumented trial and error.
Arm-Specific Design
ArmFlux is designed around Arm64-aware execution, rather than simply labeling a generic AI application as "Arm compatible."
The platform includes hardware detection for:
- CPU architecture
- CPU model
- Physical/logical cores
- RAM
- Operating system
- Detectable Arm CPU features
When running on Arm hardware, the system identifies the environment as:
ARM-NATIVE
When running elsewhere, it explicitly identifies:
DEMO MODE
This distinction is important because benchmark numbers from an x86 development machine should never be presented as Arm performance.
The project also provides Arm64 target profiles including:
- Generic ARM64
- Cortex-A76 class
- Raspberry Pi-class Arm64
- Android ARM64 artifact preparation
- Arm Cloud Server
The primary Cloud AI target is the Arm64 cloud/server profile.
Benchmark Integrity
A major design principle of ArmFlux is:
Never fabricate an optimization result.
Every displayed benchmark result is generated by the benchmark engine.
The system records:
- Warm-up runs
- Timed iterations
- Latency statistics
- Resource measurements
- Model size
- Accuracy when an evaluation dataset is available
- Hardware/environment information
The bundled demonstration workload is a small self-contained CNN trained on a synthetic-but-real task. Its measurements are genuine measurements of that workload, but they should not be interpreted as representative of production ImageNet or LLM performance.
For real Arm performance validation, the same workflow can be executed on an Arm64 environment.
ArmFlux also provides an export workflow for validating a winning configuration with Arm performance tooling such as Arm Performix. It does not claim a direct live Performix API integration in the current prototype.
How We Built It
Backend
- Python
- FastAPI
- Pydantic
- PyTorch
- ONNX
- ONNX Runtime
- NumPy
- psutil
- PyYAML
- Click
Frontend
- React
- TypeScript
- Vite
- Tailwind CSS
- Recharts
Engineering
- Docker
- pytest
- REST API
- CLI
- Modular optimization engine
- JSON-based experiment storage
The backend separates the system into independent components:
Model Analyzer
↓
Profiler
↓
Candidate Generator
↓
Optimization Engine
↓
Benchmark Engine
↓
Evaluation Engine
↓
Pareto / Scoring
↓
Selection + Explanation
↓
Storage + Export
The frontend, REST API, and CLI all use the same core optimization pipeline, preventing different interfaces from producing different optimization behavior.
Setup Instructions
Requirements
- Python 3.11+
- Node.js 18+
- npm
- Linux/macOS/WSL recommended
- For Arm-native validation: an Arm64 environment or Arm-powered device
1. Clone the repository
git clone <YOUR_GITHUB_REPOSITORY_URL>
cd armflux-studio
2. Run the setup script
bash scripts/setup.sh
This installs the backend/frontend dependencies, generates the bundled example model and evaluation set, and runs the test suite.
3. Start the backend
uvicorn backend.api.main:app --reload --port 8000
Open:
http://localhost:8000/docs
for the interactive API documentation.
4. Start the frontend
cd frontend
npm install
npm run dev
Open:
http://localhost:5173
5. Run the CLI
Check the detected hardware:
python3 cli/armflux.py hardware
Profile the bundled model:
python3 cli/armflux.py profile \
examples/vision/armflux_demo_vision.onnx
Run an optimization experiment:
python3 cli/armflux.py optimize \
examples/vision/armflux_demo_vision.onnx \
--target arm_cloud \
--objective balanced \
--min-accuracy 0.90 \
--eval-set examples/vision/eval_set.npz
List previous runs:
python3 cli/armflux.py list
Generate a report:
python3 cli/armflux.py report <RUN_ID>
Export the winning configuration:
python3 cli/armflux.py export <RUN_ID>
6. Docker
docker compose up --build
For Arm-native validation, build/run the deployment for an Arm64 environment:
docker build --platform linux/arm64 .
The benchmark should then be executed on the actual Arm64 environment so the resulting numbers represent the target hardware rather than the development machine.
Example Optimization Workflow
A typical ArmFlux experiment looks like:
1. Upload ONNX model
↓
2. Select ARM64 target
↓
3. Set accuracy threshold
↓
4. Select optimization objective
↓
5. Establish FP32 baseline
↓
6. Generate candidate configurations
↓
7. Quantize / configure candidates
↓
8. Benchmark candidates
↓
9. Reject candidates below quality threshold
↓
10. Compute Pareto frontier
↓
11. Select optimal configuration
↓
12. Explain the result
↓
13. Export reproducible configuration
The important part is that ArmFlux does not assume which optimization will win.
It measures the candidates and lets the data decide.
Challenges We Ran Into
The hardest part was discovering that optimization is not a single transformation.
INT8 quantization may reduce model size without necessarily producing the best latency on every workload.
Similarly, increasing the thread count can sometimes make performance worse instead of better.
This forced us to treat optimization as a search problem.
Hardware variability was another challenge.
Arm spans mobile processors, embedded systems, edge devices, and high-core-count cloud CPUs. A configuration that performs well on one Arm target cannot automatically be assumed to perform well on another.
We therefore built a target-aware configuration layer and explicitly distinguish between:
- Actual detected hardware
- User-selected target profiles
- Demo/simulation mode
Benchmark methodology was another major challenge.
To make before/after comparisons meaningful, we needed controlled warm-up, repeated measurements, consistent inputs, and reproducible configurations.
Accomplishments We're Proud Of
We are particularly proud that ArmFlux is designed around evidence instead of assumptions.
The system can show:
What was tested
↓
What changed
↓
What improved
↓
What was rejected
↓
Why the winner was selected
↓
How to reproduce it
Key accomplishments include:
- Real ONNX model analysis
- Real FP32 baseline measurement
- Real INT8 quantization
- Target-aware candidate generation
- Thread and graph optimization search
- Real benchmark execution
- Accuracy-aware rejection
- Pareto frontier analysis
- Composite ArmFlux performance scoring
- Hardware detection
- Reproducible experiment IDs
- Exportable optimization configurations
- REST API
- CLI
- Developer dashboard
- Automated test suite
- Explicit Demo Mode
- Arm64 validation workflow
Most importantly:
ArmFlux makes AI optimization itself the product.
What We Learned
We learned that efficient AI is fundamentally a systems problem.
Model architecture matters, but so do:
- Numerical precision
- Memory behavior
- Operator implementation
- Runtime configuration
- CPU parallelism
- Hardware capabilities
- Workload characteristics
- Quality/performance trade-offs
We also learned that "faster" is not automatically "better."
An optimization that reduces latency but violates the required accuracy threshold is not a successful optimization.
ArmFlux therefore treats optimization as a constrained multi-objective problem:
$$ \text{Best Configuration} = \arg\min\left(\text{Latency}, \text{Memory}, \text{Model Size}\right) $$
subject to:
$$ \text{Quality} \geq \text{Required Threshold} $$
For throughput-oriented workloads, the optimization objective can instead prioritize maximizing throughput while respecting the same quality constraints.
This changed our approach from:
"How do we make this model faster?"
to:
"Which configuration gives this workload the best measurable trade-off on this target?"
What Makes ArmFlux Different
Most AI deployment workflows follow:
Model
↓
Convert
↓
Deploy
ArmFlux follows:
Model
↓
Profile
↓
Generate alternatives
↓
Measure
↓
Reject bad trade-offs
↓
Compare
↓
Find Pareto-optimal configurations
↓
Select
↓
Explain
↓
Reproduce
The developer does not have to guess which configuration is best.
ArmFlux provides evidence.
Why It Could Be Reused
ArmFlux is designed as a reusable optimization layer rather than a one-off benchmark.
Developers can extend it with:
- New model architectures
- New evaluation metrics
- New Arm targets
- New runtimes
- New optimization strategies
- New objective functions
- New benchmark workloads
The optimization pipeline is intentionally separated from the UI, meaning the same engine can be used through:
- Web dashboard
- REST API
- CLI
- Automated CI/CD workflows
This makes the project useful beyond the competition.
Current Limitations
We intentionally avoid claiming capabilities that are not verified in the current prototype.
Current verified optimization path
- ONNX
- ONNX Runtime
- FP32
- INT8
- Thread tuning
- Graph optimization
Not currently implemented
- FP16 execution
- INT4 execution
- In-process Android execution
- In-process ExecuTorch execution
- Direct XNNPACK execution
- Direct Arm Performix API integration
- Production multi-user database storage
Android artifact preparation is supported, but actual mobile execution requires a connected Arm-powered device or emulator.
The bundled demonstration was developed on an x86_64 host, so its benchmark numbers are not presented as Arm hardware measurements.
What's Next
The next stage for ArmFlux is to validate and expand the optimization engine on real Arm64 infrastructure.
Planned improvements include:
- Native Arm64 benchmark runs
- Arm cloud benchmarking
- Additional Arm-optimized runtimes
- More model architectures
- Energy-aware optimization
- More advanced optimization search
- Automated hardware capability detection
- Larger language-model workloads
- TTFT and tokens/sec benchmarking
- Mobile/edge deployment workflows
- Continuous performance regression testing
- Distributed Arm benchmarking
- More deployment artifacts
- Public benchmark datasets and recipes
The long-term vision is:
ArmFlux Studio becomes an AI performance engineering layer that automatically discovers how to run a workload most efficiently on its target Arm platform.
Final Takeaway
AI optimization should not be based on guesswork.
Developers should be able to ask:
"What is the best way to run this model on my Arm target?"
and receive:
A measured configuration, a performance comparison, an explanation, and a reproducible artifact.
That's ArmFlux Studio.
Profile. Optimize. Prove.
Built With
- ai-optimization
- arm
- arm64
- docker
- executorch
- fastapi
- hugging-face
- kleidiai
- machine-learning
- model-quantization
- numpy
- onnx
- onnx-runtime
- psutil
- pydantic
- pytest
- python
- pytorch
- react
- rest-api
- tailwind-css
- typescript
- vite
- xnnpack
Log in or sign up for Devpost to join the conversation.