Optimizing to achieve higher speedups and lower runtimes
Providing support for systems that support CUDA and those that do not
What it does
Benchmarks the UserOptimizedTransformer against the Baseline Transformer, ensuring correctness within the given error limits and better performance with non-CUDA and CUDA support
How we built it
Non-CUDA implementation was done through manual prompting on ChatGPT web interface (Go subscription)
CUDA implementation was done through agentic capabilities of Codex (subscription provided by NUS)
Challenges we ran into
One of us didn't have access to an Nvidia GPU so we had to separate our development workflows then combine both solutions at the end
Understanding of the transformer model and its complexity within the given time frame
Sometimes Codex doesn't adhere to what was asked of it (e.g. asking for a CUDA implementation but it gives one that doesn't use CUDA)
Accomplishments that we're proud of
Providing a non-CUDA implementation and a CUDA implementation, both with substantial speedups
Being able to complete this in the short amount of time given and balancing school work at the same time
What we learned
Learnt more about CUDA and the various specific optimizations available
Learnt more about how to develop fast in a Linux environment
Log in or sign up for Devpost to join the conversation.