Etched built an ASIC that only runs transformers
A startup ditched everything GPUs do well and made a chip that runs one architecture 20 times faster.
Read the note 1 min read
From the notebook
4 notes on this topic.
A startup ditched everything GPUs do well and made a chip that runs one architecture 20 times faster.
Three years after first commit, the project that made local LLMs possible ships a stable release.
Iroh built a system that splits LLM inference across volunteer nodes. The networking stack handles dropouts mid-inference. Wild.
An enthusiast loaded a 1T-parameter model into 768GB of Intel Optane DIMMs and got 4 tokens per second on a single GPU. Slow, but it worked.