Training transformers in seconds, not hours
Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.
Read the note 1 min read
From the notebook
5 notes on this topic.
Someone built a 22KiB transformer that trains in 13 seconds. The trick is not the model size, it is how they squeezed PyTorch out of the picture.
A new transformer variant ditches the attention loop entirely. Faster inference, cleaner scaling, same performance.
A startup ditched everything GPUs do well and made a chip that runs one architecture 20 times faster.
Researchers found that transformer attention mechanisms lack the executive control functions that let human brains manage working memory. The models can retrieve information, but they cannot suppress irrelevant context.
A programmer rewrote PyTorch's transformer architecture using Rust and abstract algebra. The result is dense but fast.