-
-
Run AI inference with configurable runtime parameters and monitor generation performance
-
Generate and review detailed inference, benchmark, and optimization performance reports
-
AI-powered recommendations for selecting efficient inference configurations and deployment strategies
-
ArmPilot-AI: ARM64-first AI inference optimization and benchmarking platform
-
Unified dashboard for monitoring models, inference performance, benchmarks, and system insights
-
Configure runtime, inference, benchmark, and application settings for ArmPilot-AI
-
Benchmark LLM inference performance and compare key metrics across configurations
-
Track previous inference runs, benchmarks, configurations, and performance history
-
Optimize inference configurations to improve performance, efficiency, and resource utilization
Inspiration
As AI inference is shifting towards ARM64-based infrastructure due to its efficiency and versatility, optimizing an AI model for compute on ARM64 is not a trivial task
Developers often have to spend a lot of time experimenting with different configurations, number of CPU threads, batch sizes, context lengths, and other variables, and comparing their performance characteristics. Benchmarking in isolation often does not provide clear guidance on what changes to make next.
This prompted us to create ArmPilot-AI, an ARM64-first AI inference optimization and benchmarking platform, that transforms performance measurement into an optimization workflow.
What We Built
ArmPilot-AI is an end-to-end model management, inference, benchmarking, performance analysis, optimization, recommendation, reporting, and history platform for developers. At its core is the following optimization and benchmarking workflow:
Model → Configure → Benchmark → Analyze → Optimize → Recommend → Validate
By offering this kind of out-of-the-box model management and optimization, ArmPilot-AI aims to reduce the time-to-value for developers using various model types and configurations. It is focused on native ARM64 inference, but the design allows for easy integration of different inference engines such as llama.cpp/GGUF and ONNX Runtime, which is beneficial for both ARM64 and x86_64.
How We Built It
The interface was built using Next.js, React, Typescript, Tailwind CSS, and other modern frontend tools. The rest API backend was built using Python, FastAPI, Pydantic, and has a set of endpoints for auth, inference, benchmarks, optimization, recommendation, models, reports, history, and other features. The inference part uses a set of lightweigh runtimes and is focused on being performant on ARM64. The project uses config management, caching, JWT auth, Docker deployment, and a few other utilities.
The optimization is recommendation focused, meaning it analyzes candidate configurations and tries to recommend the ones that are most likely to be beneficial. It should be noted that such analysis is not a substitute for actual benchmarking.
Challenges We-faced
One of the main challenges we faced was getting the backend ready for cloud deployment while maintaining the focus on being lightweight.
The initial set of dependencies included a number of ML-related packages that were not actually used by the backend, which made the deployment build time and memory usage unnecessarily large. We had to do some cleanup and optimize the dependency list based on what the actual code was using. This became even more important for packages like llama.cpp that had to be compiled from source during deployment.
Another challenge was making sure that the recommendation and analysis parts could provide meaningful insight and drive the optimization workflow, and not simply be an exercise in number crunching.
What We Learned
Working on ArmPilot-AI taught us a number of lessons. First of all, it helped us realize that while model choice has a major impact on AI inference performance, the characteristics of the underlying hardware and supporting software are equally important.
Some of the key takeaways from the project include:
• The importance of measuring before optimizing
• The need to keep the recommendation logic explainable
• The value of separating analysis and observation from assumptions and optimization
• The importance of designing the deployment around actual needs and not dreams
• The need to avoid unsubstantiated performance claims
• The importance of designing the solution around the need to support multiple runtimes
Future Scope
The immediate future scope for ArmPilot-AI includes adding more support for different hardware on ARM64, automated configuration search, hardware-aware optimization, continuous benchmarking, monitoring, and edge/on-device inference optimization.
We hope to make optimizing AI model inference on ARM64 a non-trivial but ultimately straightforward process that follows a certain pattern rather than being trial-and-error.
Conclusion
ArmPilot-AI transforms the paradigm of ARM64 AI inference optimization from "how fast is my model" to "what should I change to make it faster".
Benchmark-driven model optimization is an iterative process that involves analyzing the characteristics of a particular setup, identifying potential improvements, implementing them, re-analyzing the results, and repeating this until further gains cannot be achieved. Benchmarks help identify weak points and guide the optimization process, but they rarely tell what exactly to do about them.
The key idea behind ArmPilot-AI is to make sure that every step in the optimization workflow is streamlined and repeatable.
It follows the optimization workflow below:
Benchmark → Analyze → Optimize → Validate
It helps developers move beyond simply asking "how fast is my model", and start asking more interesting questions, such as "what should I change to make my model faster".
Built With
- css
- docker
- fastapi
- gguf
- jwt
- llama.cpp
- next.js
- onnx
- pydantic
- python
- react
- tailwind
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.