Hackathon project description
How the solution addresses the problem
The project preserves the supplied pre-normalization, causal Transformer-layer
formula while replacing expensive implementation details with GPU-oriented
paths. BaselineTransformer provides the original explicit PyTorch reference,
and UserOptimizedTransformer receives the same weights and inputs before its
output is checked against that reference. A result is accepted only when every
element satisfies the required absolute-error or relative-error limit.
The optimized implementation packs the Q/K/V projections, uses PyTorch scaled dot-product attention without materializing unnecessary attention matrices, fuses residual addition with LayerNorm, fuses the FFN input projection with GELU, caches inference weights, and replays eligible fixed-shape work with CUDA Graphs. Shape checks select measured specializations where one general kernel is not best: exact batch chunking for shape 6, grouped matrix kernels for shape 8, a selected attention backend for shape 11, and bounded streaming for the otherwise infeasible shape 14. Unsupported hardware and unknown shapes retain correct PyTorch fallbacks.
Each candidate optimization was evaluated on the target laptop for both latency and numerical accuracy. Faster candidates that exceeded the error limit or regressed full-model latency were rejected. This directly addresses the challenge's requirement to fuse or specialize kernels by known input shape without changing the mathematical layer being tested.
Development and AI tools
| Tool | Use in this project |
|---|---|
| Visual Studio Code | Source editing, integrated-terminal runs, and result inspection |
| Windows PowerShell | Virtual-environment setup and benchmark automation |
| Git and GitHub | Version control and reference research for GPU-kernel techniques |
NVIDIA nvidia-smi |
GPU, driver, clock, power, temperature, and memory checks |
| PyTorch Profiler | CUDA operator timing and identification of kernel bottlenecks |
| OpenAI Codex | Development-time code analysis, optimization ideas, regression investigation, test planning, and documentation assistance |
OpenAI Codex was used as an engineering assistant rather than as part of the submitted runtime. AI-proposed changes were reviewed through source inspection, accuracy tests, repeated CUDA-event timings, and memory checks before being retained. No Colab or Jupyter notebook is required to build or run the project.
APIs used
The implementation uses PyTorch's scaled dot-product-attention and backend selection APIs, CUDA Graph capture/replay, CUDA events, device-memory queries, and profiling APIs. Triton's JIT API is used to compile the custom fused normalization, FFN, and grouped matrix kernels. The benchmark itself does not call OpenAI or any other external web API at runtime.
Libraries and frameworks
- Python 3.11 and its standard library provide the command-line runners, subprocess isolation, statistics, configuration dataclasses, and file paths.
- PyTorch provides tensors, the reference model, CUDA integration, SDPA, correctness checks, CUDA Graphs, timing, and profiling.
- Triton provides the custom GPU-kernel language and JIT compiler. On the tested
Windows machine this is installed through
triton-windows. - The NVIDIA CUDA runtime and driver execute the generated PyTorch and Triton kernels on the RTX 3060 Laptop GPU.
Datasets and assets
No external dataset, pretrained model, manually labelled data, or media asset
is used. The official 14 input-shape configurations are the challenge-provided
benchmark specification. Inputs are deterministic synthetic tensors generated
with seeded torch.randn; model parameters are initialized once and copied
from the baseline to the optimized model so both implementations receive an
identical test. Consequently, the reported result measures kernel execution
rather than data loading, tokenization, or pretrained-model quality.
Log in or sign up for Devpost to join the conversation.