transformers A field note by Vikrant Sharma
Transformers without loops: the Inception architecture drops self-attention
A new transformer variant ditches the attention loop entirely. Faster inference, cleaner scaling, same performance.
Most transformer improvements fiddle with the attention mechanism. Make it sparse, make it linear, quantise the weights. This post from zartbot goes the other direction: throw out the loop. The Inception architecture replaces multi-head self-attention with a fixed projection trick borrowed from convolutional networks. No queries walking through keys. No quadratic memory scaling. Just matrix multiplications you can parallelise without the sequential bottleneck. The interesting claim is that you do not lose much. On language modelling benchmarks, Inception matches GPT-scale transformers at similar parameter counts. The win is inference speed: no iterative attention means fewer steps to cache, fewer steps to recompute during autoregressive generation.
Why this matters for deployment
Most production transformer costs come from repeated attention computations during text generation. You generate one token, recompute attention over the entire context, generate the next token, repeat. Inception collapses that into a single forward pass per layer. The author benchmarks a 1.3B parameter Inception model at 40% faster inference than a comparable GPT-2 setup. That is the kind of speedup that changes what you can run on a single GPU or what batch size fits in memory. The tradeoff is expressiveness. Self-attention learns which tokens matter for each position. Inception projects everything through a fixed mixing function. For tasks where context really is sparse or position-dependent, that might cost you accuracy. For tasks where the model just needs to aggregate features, it might not. I am curious whether this holds at GPT-4 scale. Most architectural shortcuts that work at 1B parameters break at 175B. But if it does scale, this is the kind of change that makes fine-tuning and serving cheaper without rewriting the entire training stack.
Source: On Next-Gen Transformer: Loops Are Not What You Need