Mercury 2.5 LLM clocks 770 tokens per second
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
Read the note 1 min read
From the notebook
4 notes on this topic.
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
The Python library you call from a Jupyter notebook might be Rust underneath. PyO3 makes that translation invisible.
A network engineer pushed Go to saturate a 100 Gbps link using AF_XDP. The surprising bit is not the speed, it is that Go got there at all.
Manifest deprecated their LLM router after six months. The reason is not what I expected.