Transformers without loops: the Inception architecture drops self-attention
A new transformer variant ditches the attention loop entirely. Faster inference, cleaner scaling, same performance.
Read the note 1 min read
From the notebook
1 note on this topic.