Vikrant
Someone ran a trillion-parameter LLM on consumer hardware with Optane memory
An enthusiast loaded a 1T-parameter model into 768GB of Intel Optane DIMMs and got 4 tokens per second on a single GPU. Slow, but it worked.
Read post
1 post tagged.