Inspiration
Biological specimens can lose viability during transport because of temperature excursions, oxygen variation, vibration, delay, and handling stress. We wanted to explore whether a digital twin could act as a real-time rescue decision system rather than simply reporting that a sample had already degraded.
BioLoop Edge ARM Optimized was inspired by the idea of combining real biological reference data, simulation, control theory, and Arm-native optimization into one system. The goal is to model the specimen state, simulate possible future outcomes, evaluate rescue strategies, and execute those workloads efficiently on Google Axion.
What it does
BioLoop Edge creates a digital twin of a transported biospecimen.
The system starts from a real public CELLxGENE lung reference and combines it with transport-state variables such as temperature, oxygen, delay, vibration, integrity, and viability.
A user can:
load the real biological reference inject a simulated transport crisis observe the digital twin moving away from its reference state run Monte Carlo rescue simulations compare multiple controller strategies select a strong rescue strategy inspect Arm-specific optimization decisions view measured performance and Arm Performix profiling evidence
The Monte Carlo layer repeatedly evaluates candidate controllers such as PID, LQR, fixed, adaptive, and robust approaches across uncertain transport scenarios. Rather than optimizing against one deterministic future, BioLoop evaluates many possible outcomes and ranks the rescue strategies using control-performance and safety-related metrics.
The project also includes an ARM Optimization Autopilot. Instead of assuming that every optimization is beneficial, it benchmarks candidate execution paths and can reject regressions.
For example, INT8 quantization reduced the BioRisk model footprint by 45%, but generic INT8 inference was slower for the tested small MLP workload, so that latency path was rejected.
For the tested KleidiAI quantized matrix workload:
M = 512 N = 512 K = 1024 Block size = 32
we measured approximately: NEON DotProd : 1.721 ms NEON I8MM : 0.981 ms
How we built it
BioLoop Edge combines a web-based digital twin with an Arm-native backend running on Google Cloud Google Axion, aarch64.
The frontend contains the Jury Command Center, BioSpecimen Rescue Digital Twin, Risk AI Optimization Lab, and 3D Simulation Lab.
The backend is built with FastAPI and connects the digital-twin experience to biological-reference processing, risk calculation, Monte Carlo simulation, native controller execution, and optimization evidence.
For the biological layer, we use a public CELLxGENE reference containing 1,000 lung cells and 12 selected genes, with normalized expression summaries and PCA-based visualization.
For Arm optimization, we used:
Google Axion Arm NEON / Advanced SIMD I8MM, or 8-bit Integer Matrix Multiply KleidiAI Arm Performix ONNX Runtime native C++
We benchmarked the BioRisk model under FP32 and INT8 execution, evaluated KleidiAI matrix kernels, and implemented a native controller benchmark comparing scalar execution with NEON/FMA acceleration.
We also used Arm Performix to profile the actual KleidiAI execution. Performix collected 339,105 hotspot samples, with 98.89% attributed to the KleidiAI benchmark binary, and resolved the underlying kai_* matrix micro-kernel functions. We were able to observe both the I8MM and DotProd execution paths directly in the profiler.
This gave us two complementary forms of evidence: direct benchmark timing for speedup measurement and Performix profiling to confirm which Arm code paths were actually executing.
Challenges we ran into
One of the biggest lessons was that an optimization that looks good in theory does not necessarily improve the real workload.
Our first assumption was that INT8 inference would automatically be faster. Quantization successfully reduced the BioRisk model footprint by 45%, but repeated batch testing showed that generic ONNX Runtime INT8 execution was slower for this small network. Instead of hiding that result, we turned it into part of the architecture: the optimization layer can reject a slower execution path.
Profiling also required work. Our initial Performix CPU Microarchitecture captures successfully collected the workload, but symbolized function and call-stack information was empty.
We rebuilt the KleidiAI benchmark using:
-O2 -g -fno-omit-frame-pointer
This preserved optimization while exposing enough symbol and stack information for Performix Code Hotspots to resolve the underlying KleidiAI functions.
Another challenge was maintaining a clear scientific boundary. The CELLxGENE reference is real public biological data, but transport crises, recovery trajectories, gene perturbations, and rescue outcomes are simulated. We explicitly surface that distinction in the application and documentation.
Accomplishments that we're proud of
We are particularly proud that BioLoop Edge is not just a visual digital twin — it contains measurable Arm-native optimization work.
Key accomplishments include:
integrating real public CELLxGENE reference data into the digital twin building a Monte Carlo rescue engine that compares multiple controller strategies across uncertain scenarios implementing native ARM64 rescue-controller execution achieving 1.143x controller speedup using NEON/FMA reducing the BioRisk model footprint by 45% detecting that generic INT8 inference caused a latency regression and automatically rejecting that path building and benchmarking KleidiAI directly on Google Axion measuring an up to ~1.75x I8MM micro-kernel speedup over a comparable DotProd path for the tested workload measuring approximately 43% lower kernel time using Arm Performix to capture 339,105 profiling samples resolving real kai_* KleidiAI micro-kernel functions confirming that both NEON I8MM and DotProd execution paths were present during profiling creating an ARM Optimization Autopilot that presents these decisions directly in the application
The result is a system where the optimization evidence is visible, reproducible, and tied directly to the target Arm platform
What we learned
The biggest lesson was that hardware-aware AI optimization must be measurement-driven.
Quantization, vectorization, and specialized kernels should not be treated as automatic wins. The workload shape, model size, runtime, instruction support, and target CPU all matter.
We also learned the value of separating optimization layers. Model compression, inference latency, matrix-kernel performance, and native controller throughput are different problems and should be benchmarked independently.
Arm Performix reinforced this lesson. Benchmark timings tell us how fast a workload runs, while profiling shows where execution time is actually being spent and which optimized paths are active.
From the biological side, we learned that digital twins become far more meaningful when they are grounded in a real reference state and when simulated elements are clearly identified rather than presented as measured biological outcomes.
What's next for BioLoop Edge ARM Optimized
The next step is to evolve BioLoop from an optimization demonstration into a more autonomous Arm-native biospecimen rescue platform.
We want to extend the ARM Optimization Autopilot so that it can automatically profile candidate workloads, benchmark alternative execution paths, rank them, reject regressions, and deploy the best-performing configuration for the detected Arm hardware.
On the biology side, we plan to expand beyond the current 1,000-cell lung reference into larger CELLxGENE datasets, additional tissues, more cell types, and broader gene programs.
We also want to improve the Monte Carlo digital twin by incorporating richer uncertainty models, adaptive controllers, reinforcement-learning-based control policies, and more detailed transport environments.
Future versions could also explore edge deployment, where the digital twin and rescue controller run closer to the physical specimen rather than depending entirely on cloud infrastructure.
The long-term vision is:
A hardware-aware biological digital twin that continuously predicts specimen risk, explores possible futures, selects the best rescue action, and automatically chooses the most efficient Arm execution path for the hardware it is running on.
Built With
- arm-neon
- arm-performix
- c++
- cellxgene
- fastapi
- github
- google-axion
- i8mm
- javascript
- kleidiai
- monte-carlo-simulation
- onnx-runtime
- pca
- python
- vercel
Log in or sign up for Devpost to join the conversation.