Vulkore: CUDA-Inspired GPU Compute for Every Android Phone
Inspiration
Modern smartphones contain incredibly powerful GPUs capable of running physics simulations, image processing pipelines, and even large language models. Yet unlike NVIDIA GPUs, Android has no CUDA.
Developers are forced to choose between writing hundreds of lines of low-level Vulkan boilerplate or using heavyweight frameworks that still expose much of Vulkan's complexity. Worse, GPU kernels are usually written in shading languages instead of general-purpose C, making compute development feel completely different from desktop GPU programming.
We wanted to change that.
Our goal was simple:
Make GPU programming on Android feel as easy as CUDA, while remaining completely portable across Vulkan devices.
That vision became Vulkore—a lightweight C++ framework that brings CUDA-inspired ergonomics to Android GPUs.
What it does
Vulkore is a pure C++20 Vulkan compute runtime that lets developers write GPU kernels in OpenCL C, compile them once with Google's clspv, and launch them using a CUDA-like API.
Instead of manually creating descriptor sets, pipelines, command buffers, synchronization primitives, and memory barriers, developers simply write:
Vulkore::launch(kernel, grid, args...);
Behind the scenes, Vulkore automatically:
- Loads SPIR-V kernels
- Parses clspv reflection metadata
- Binds kernel arguments automatically
- Creates and caches compute pipelines
- Records command buffers
- Handles synchronization
- Manages Vulkan memory
The result is a runtime that's only around 1.7 MB, yet capable of running workloads ranging from simple vector operations to LLM inference entirely on a phone GPU.
How we built it
The project consists of two major pieces.
Build-time kernel pipeline
GPU kernels are written in OpenCL C instead of GLSL.
During compilation, Google's clspv converts the kernels into Vulkan-compatible SPIR-V while generating reflection metadata describing every kernel argument.
This lets developers write real C with pointers, loops and structs while still targeting Vulkan.
Lightweight Vulkan runtime
The runtime is responsible for everything Vulkan normally requires developers to do manually:
- Device discovery
- Memory allocation (VMA)
- Descriptor management
- Pipeline creation
- Buffer uploads/downloads
- Automatic argument binding
- Command recording
- Synchronization
One of our biggest design goals was hiding Vulkan without sacrificing performance.
The runtime also introduces Batch, allowing hundreds of kernel launches to be recorded into a single command buffer and submitted once—a huge improvement for dispatch-heavy workloads like transformers.
Challenges we ran into
The hardest problems weren't writing Vulkan code—they were making the framework portable.
One major issue was non-coherent memory.
Desktop Vulkan drivers typically expose coherent memory, meaning CPU writes automatically become visible to the GPU. Many Android GPUs don't.
Early versions of Vulkore worked perfectly on desktop but failed on real phones because CPU cache flushing wasn't handled correctly. We redesigned the memory subsystem around memory visibility instead of assuming coherence, making the same runtime work correctly across desktop GPUs, integrated GPUs, and Android devices.
Another challenge was command submission overhead.
Running an LLM involves hundreds of GPU dispatches for every generated token. Submitting each dispatch individually wasted significant CPU time, so we introduced Batch, recording hundreds of dispatches into one command buffer and performing only a single queue submission.
We also learned that desktop testing wasn't enough. Some kernels passed validation, compiled successfully, and even launched without errors—but silently produced incorrect results on actual phone hardware. That led us to build an extensive on-device validation suite and continuously test on real Android devices rather than relying solely on desktop Vulkan implementations.
What we learned
Building Vulkore taught us far more than Vulkan programming.
We learned how modern GPU runtimes abstract low-level APIs, how SPIR-V and compiler reflection can remove huge amounts of boilerplate, and why mobile GPU drivers behave very differently from desktop drivers.
Perhaps the biggest lesson was that developer experience matters just as much as performance.
The fastest GPU runtime isn't very useful if writing a simple compute kernel requires hundreds of lines of setup code. Good abstractions don't hide power—they make it accessible.
Results
Vulkore isn't just a programming framework—it has been validated on real workloads.
Some highlights include:
- CUDA-inspired API for Vulkan compute
- OpenCL C kernels compiled offline with clspv
- Automatic descriptor binding using compiler reflection
- 91 automated tests across desktop and Android
- Validated on real Adreno hardware
- Batch execution for dispatch-heavy workloads
To demonstrate real-world performance, we ported Gemma 3 1B to Vulkore and ran it entirely on the GPU of a OnePlus 15 (Snapdragon 8 Elite, Adreno 840).
Our implementation achieves:
- 60–70 tokens/sec decode throughput
- 8,192-token context window
- ~1.9 second model load time
On the same phone, Google AI Edge Gallery running the same Gemma 3 1B model achieved approximately:
- ~45 tokens/sec decode throughput
- ~1,600-token context window
While the applications use different runtimes and optimization pipelines, the comparison demonstrates that Vulkore can deliver competitive—and in our testing, substantially higher—on-device LLM performance using a lightweight Vulkan compute runtime.
What's next
Vulkore is still evolving.
Our roadmap includes kernel fusion, persistent pipeline caches, timeline semaphores, and additional compiler optimizations to further reduce dispatch overhead.
Our long-term vision is ambitious but simple:
Make portable GPU programming on Android as approachable as CUDA is on desktop, enabling developers to build high-performance AI and compute applications without becoming Vulkan experts.
Built With
- c++
- clspv
- vulkan
Log in or sign up for Devpost to join the conversation.