Back to the notes

Training transformers in seconds, not hours

Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.

Processor Technology SOL 20 Computer
Swtpc6800 en:User:Swtpc6800 Michael Holley / Wikimedia Commons. Resized and converted to WebP. Public domain

A project called TurboGPT hit Hacker News today with a claim that sounds like benchmark fiction: train a transformer in 13 seconds. The model is 22 kilobytes. That is smaller than most PNG screenshots. The interesting bit is not that someone made a tiny model. You can shrink a transformer by cutting layers and embedding dimensions until it fits in a tweet. The interesting bit is the training speed. Thirteen seconds for a transformer that actually converges is not normal, even at toy scale. The repo does not hide the trick. They rewrote the training loop in raw CUDA and C#, bypassing PyTorch entirely. No autograd overhead. No framework abstractions. Just matrix multiplies and gradient calculations written as kernels that run directly on the GPU. This is the kind of thing you do when you need to train thousands of small models in a pipeline, not when you are prototyping one large one. Medical imaging models that run per-patient. Embedded models that retrain on device. Anything where the startup cost of PyTorch or JAX is larger than the compute you actually need. The trade-off is obvious. You lose the ecosystem. No Hugging Face Transformers. No pre-trained checkpoints. No wandb integration. You are writing CUDA by hand, which means you are also debugging CUDA by hand when the gradients explode. But if your bottleneck is iteration speed and you know the architecture will not change, this makes sense. Thirteen seconds means you can try a hundred hyperparameter combinations in half an hour. With PyTorch that same loop takes six hours. I would not rewrite production training this way unless the speed delta was paying for an engineer’s salary. But for research where the model is fixed and the data keeps changing, this is worth the file size of a decent README.


Source: Show HN: TurboGPT: train 22KiB transformer in 13s

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work