💡 Inspiration

Deploying modern generative AI models and mathematical loops onto Arm CPUs is highly energy-efficient but technically challenging. Developers often struggle to manually write vector assembly, configure caching boundaries, or identify where bottlenecks reside. We wanted to build a visual, intelligent developer environment that acts as an "optimization copilot"—analyzing code, refactoring it for Arm SIMD hardware, and immediately demonstrating the speedup and carbon savings.

🧠 What it Does

  • AST-Based Code Profiler: Parses C++ or Python code files to detect nested loops, naive array structures, and activation operations that are candidates for hardware acceleration.
  • Agentic Refactoring Brain: Automatically rewrites naive loops into vectorized equivalents using Arm NEON SIMD vector intrinsics and links to Arm KleidiAI assembly microkernels.
  • Target Compilation Preset Card: Dynamically generates target compiler commands (e.g. g++ -O3 -march=armv9-a -lkai) tailored to processors like AWS Graviton, Google Axion, Snapdragon X Elite, and Raspberry Pi 5.
  • Sustainability & Benchmarking Simulator: Projects LATENCY, CACHE MISSES, ACTIVE POWER, INSTRUCTION COUNT, and calculates estimated CARBON OFFSET and CLOUD HOSTING COST savings.

🛠️ How We Built It

  • Core backend: Node.js Express server executing token pattern mapping rules.
  • Visual analytics: Chart.js rendering comparative latency, cache, bandwidth, instructions, and power diagrams.
  • Immersive frontend: Modern dark theme styling with a custom scroll-synchronized line-number gutter, interactive particle emitter flow, and a 3D holographic wireframe register cube background rendered in pure HTML5 Canvas.

🚧 Challenges We Ran Into

  • Developing a pure JS 3D perspective projection model without using heavy packages (to ensure Vercel's serverless cold starts remain instant).
  • Mapping approximate polynomial Taylor expansions for vectorized exponentials (vexpq_f32) inside Softmax loops.

🎉 Accomplishments We're Proud Of

  • A zero-dependency, ultra-lightweight client-side 3D render loop running at 60 FPS.
  • High-fidelity, direct mapping to the newly released Arm KleidiAI spec.

📚 What We Learned

  • Armv9 instruction pipeline capabilities, register configurations, and compilation targets.
  • The business value of Cloud AI code optimizations—where a minor cycle reduction cuts hosting costs by 60% and lowers global datacenter carbon footprints.

🔮 What's Next

  • Expanding optimization libraries to support INT8/INT4 quantization pipelines.
  • Support for tvm/halide IR compiler targets.
Share this project:

Updates

Submission history