← Back to ProjectsSeptember 9, 2026. Posted by Peach
TikTokTechjam - GPU kernel for a Transformer layer

Didn't win TechJam 2026, but honestly this was fun HAHAH
Our track was writing GPU kernels for a Transformer layer : CUDA graphs, Triton, mixed precision, none of which any of us had really touched before. We had under a week. Most of it was: be confused, measure something, be slightly less confused, repeat.
The task gave us 14 different input shapes, and they were wildly different from each other. Batch sizes from 1 to 10,000. Sequence lengths from 32 to 100,000. One shape was so large that the original implementation would need a 20 terabyte intermediate tensor just to produce a single number ..it genuinely cannot run on any GPU that exists.
The obvious way to handle that spread is to write rules. "If the batch is bigger than 32, use half precision." We decided not to write any, because those numbers are really just fitted to whichever GPU you happened to test on, and nobody can tell from reading the code whether they're right.
So we let the code work it out at runtime instead ..try the options on the real input, time them, keep whichever wins. What it picked was not what we would have guessed: half precision at batch 1, full precision at batch 4 and 16, then half precision again at 64. No rule we could have written produces that.
Ended up around 10× faster than the reference across the suite, with every shape still numerically correct.
The part I'll actually remember isn't the number though. We'd written our own safety checks to decide when to switch optimizations on, and two of them turned out to be wrong, not making our results incorrect, just quietly making us slower. We only found them because we kept measuring instead of trusting what we'd written. Good lesson, slightly painful way to learn it.
Congrats to the winning teams, genuinely. A lot of sharp work in that room hehe ><
And thank you to my teammates for all our late nights! lets go to more hackathons togetherrr