Hackathon project description

How the solution addresses the problem

The project preserves the supplied pre-normalization, causal Transformer-layer formula while replacing expensive implementation details with GPU-oriented paths. BaselineTransformer provides the original explicit PyTorch reference, and UserOptimizedTransformer receives the same weights and inputs before its output is checked against that reference. A result is accepted only when every element satisfies the required absolute-error or relative-error limit.

The optimized implementation packs the Q/K/V projections, uses PyTorch scaled dot-product attention without materializing unnecessary attention matrices, fuses residual addition with LayerNorm, fuses the FFN input projection with GELU, caches inference weights, and replays eligible fixed-shape work with CUDA Graphs. Shape checks select measured specializations where one general kernel is not best: exact batch chunking for shape 6, grouped matrix kernels for shape 8, a selected attention backend for shape 11, and bounded streaming for the otherwise infeasible shape 14. Unsupported hardware and unknown shapes retain correct PyTorch fallbacks.

Each candidate optimization was evaluated on the target laptop for both latency and numerical accuracy. Faster candidates that exceeded the error limit or regressed full-model latency were rejected. This directly addresses the challenge's requirement to fuse or specialize kernels by known input shape without changing the mathematical layer being tested.

Development and AI tools

Tool Use in this project
Visual Studio Code Source editing, integrated-terminal runs, and result inspection
Windows PowerShell Virtual-environment setup and benchmark automation
Git and GitHub Version control and reference research for GPU-kernel techniques
NVIDIA nvidia-smi GPU, driver, clock, power, temperature, and memory checks
PyTorch Profiler CUDA operator timing and identification of kernel bottlenecks
OpenAI Codex Development-time code analysis, optimization ideas, regression investigation, test planning, and documentation assistance

OpenAI Codex was used as an engineering assistant rather than as part of the submitted runtime. AI-proposed changes were reviewed through source inspection, accuracy tests, repeated CUDA-event timings, and memory checks before being retained. No Colab or Jupyter notebook is required to build or run the project.

APIs used

The implementation uses PyTorch's scaled dot-product-attention and backend selection APIs, CUDA Graph capture/replay, CUDA events, device-memory queries, and profiling APIs. Triton's JIT API is used to compile the custom fused normalization, FFN, and grouped matrix kernels. The benchmark itself does not call OpenAI or any other external web API at runtime.

Libraries and frameworks

  • Python 3.11 and its standard library provide the command-line runners, subprocess isolation, statistics, configuration dataclasses, and file paths.
  • PyTorch provides tensors, the reference model, CUDA integration, SDPA, correctness checks, CUDA Graphs, timing, and profiling.
  • Triton provides the custom GPU-kernel language and JIT compiler. On the tested Windows machine this is installed through triton-windows.
  • The NVIDIA CUDA runtime and driver execute the generated PyTorch and Triton kernels on the RTX 3060 Laptop GPU.

Datasets and assets

No external dataset, pretrained model, manually labelled data, or media asset is used. The official 14 input-shape configurations are the challenge-provided benchmark specification. Inputs are deterministic synthetic tensors generated with seeded torch.randn; model parameters are initialized once and copied from the baseline to the optimized model so both implementations receive an identical test. Consequently, the reported result measures kernel execution rather than data loading, tokenization, or pretrained-model quality.

Built With

Share this project:

Updates

Submission history