💡 Inspiration
Deploying modern generative AI models and mathematical loops onto Arm CPUs is highly energy-efficient but technically challenging. Developers often struggle to manually write vector assembly, configure caching boundaries, or identify where bottlenecks reside. We wanted to build a visual, intelligent developer environment that acts as an "optimization copilot"—analyzing code, refactoring it for Arm SIMD hardware, and immediately demonstrating the speedup and carbon savings.
🧠 What it Does
- AST-Based Code Profiler: Parses C++ or Python code files to detect nested loops, naive array structures, and activation operations that are candidates for hardware acceleration.
- Agentic Refactoring Brain: Automatically rewrites naive loops into vectorized equivalents using Arm NEON SIMD vector intrinsics and links to Arm KleidiAI assembly microkernels.
- Target Compilation Preset Card: Dynamically generates target compiler commands (e.g.
g++ -O3 -march=armv9-a -lkai) tailored to processors like AWS Graviton, Google Axion, Snapdragon X Elite, and Raspberry Pi 5. - Sustainability & Benchmarking Simulator: Projects LATENCY, CACHE MISSES, ACTIVE POWER, INSTRUCTION COUNT, and calculates estimated CARBON OFFSET and CLOUD HOSTING COST savings.
🛠️ How We Built It
- Core backend: Node.js Express server executing token pattern mapping rules.
- Visual analytics: Chart.js rendering comparative latency, cache, bandwidth, instructions, and power diagrams.
- Immersive frontend: Modern dark theme styling with a custom scroll-synchronized line-number gutter, interactive particle emitter flow, and a 3D holographic wireframe register cube background rendered in pure HTML5 Canvas.
🚧 Challenges We Ran Into
- Developing a pure JS 3D perspective projection model without using heavy packages (to ensure Vercel's serverless cold starts remain instant).
- Mapping approximate polynomial Taylor expansions for vectorized exponentials (
vexpq_f32) inside Softmax loops.
🎉 Accomplishments We're Proud Of
- A zero-dependency, ultra-lightweight client-side 3D render loop running at 60 FPS.
- High-fidelity, direct mapping to the newly released Arm KleidiAI spec.
📚 What We Learned
- Armv9 instruction pipeline capabilities, register configurations, and compilation targets.
- The business value of Cloud AI code optimizations—where a minor cycle reduction cuts hosting costs by 60% and lowers global datacenter carbon footprints.
🔮 What's Next
- Expanding optimization libraries to support INT8/INT4 quantization pipelines.
- Support for tvm/halide IR compiler targets.

Log in or sign up for Devpost to join the conversation.