machine-learning A field note by Vikrant Sharma
Training transformers in seconds, not hours
Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.
A project called TurboGPT hit Hacker News today with a claim that sounds like benchmark fiction: train a transformer in 13 seconds. The model is 22 kilobytes. That is smaller than most PNG screenshots. The interesting bit is not that someone made a tiny model. You can shrink a transformer by cutting layers and embedding dimensions until it fits in a tweet. The interesting bit is the training speed. Thirteen seconds for a transformer that actually converges is not normal, even at toy scale. The repo does not hide the trick. They rewrote the training loop in raw CUDA and C#, bypassing PyTorch entirely. No autograd overhead. No framework abstractions. Just matrix multiplies and gradient calculations written as kernels that run directly on the GPU. This is the kind of thing you do when you need to train thousands of small models in a pipeline, not when you are prototyping one large one. Medical imaging models that run per-patient. Embedded models that retrain on device. Anything where the startup cost of PyTorch or JAX is larger than the compute you actually need. The trade-off is obvious. You lose the ecosystem. No Hugging Face Transformers. No pre-trained checkpoints. No wandb integration. You are writing CUDA by hand, which means you are also debugging CUDA by hand when the gradients explode. But if your bottleneck is iteration speed and you know the architecture will not change, this makes sense. Thirteen seconds means you can try a hundred hyperparameter combinations in half an hour. With PyTorch that same loop takes six hours. I would not rewrite production training this way unless the speed delta was paying for an engineer’s salary. But for research where the model is fixed and the data keeps changing, this is worth the file size of a decent README.