Inspiration

Waiting for an AI response often means waiting for tokens to appear one at a time. I wanted to make that process faster by getting more out of the GPU.

What it does

FlashInference accelerates Qwen3-4B on a single NVIDIA H100. It improves both how quickly the model reads a prompt and how quickly it generates its response, while checking output quality against the original model.

How we built it

I profiled the model to find where time was going, then built specialized CUDA and Triton kernels. I combined operations, optimized how weights and cached information move through memory, and used CUDA graphs to reduce repeated scheduling overhead.

Challenges we ran into

Faster individual operations didn’t always make the whole model faster. Small numerical differences could also affect generated tokens, so every optimization needed careful performance and correctness testing.

Accomplishments that we're proud of

We reached 1,290.2 tokens per second on the official leaderboard, placing second at that point. We also built a testing pipeline that helped us distinguish real improvements from noisy measurements.

What we learned

Fast AI requires more than powerful hardware. Memory movement, scheduling, and small implementation details all matter, measuring the complete system is also essential.

What's next for FlashInference

I'm working on faster prompt processing and more efficient token generation across different batch sizes and context lengths. Our next goal is to turn those improvements into a stronger leaderboard result without sacrificing output quality.

Built With

Share this project:

Updates

Submission history