llm A field note by Vikrant Sharma
Mercury 2.5 LLM clocks 770 tokens per second
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
Mercury 2.5 is generating 770 tokens per second on Artificial Analysis benchmarks. That is faster than I can read. It is also faster than most web interfaces can render streaming responses without the text flickering or skipping lines. For context, GPT-4 sits around 40 to 60 tokens per second on the same benchmarks. Claude 3.5 Sonnet hits 80 to 100. Gemini Pro pushes 120. Mercury 2.5 is seven to ten times faster than the best production models most people use daily. The interesting bit is not the raw speed. It is what happens when generation speed stops being the constraint. If the model can spit out a full A4 page of text in under two seconds, the slow part becomes parsing that text, validating it, or rendering it in a UI that does not break. The optimisation work shifts from the inference layer to everything downstream. This also changes how you think about retry logic. Right now, if a prompt fails, you retry and wait another five seconds. If retries take half a second, you can brute-force your way through edge cases by running three variations of the same prompt in parallel and picking the best one. Speed makes redundancy cheap. The Hacker News thread is full of people asking whether Mercury 2.5 sacrifices quality for speed. That is the right question. Artificial Analysis does not publish quality benchmarks yet, so we do not know if this model is fast because it is small or fast because the inference stack is optimised. If it is the latter, this is a preview of what 2027 looks like. If it is the former, it is still useful for tasks where speed beats nuance. Either way, 770 tokens per second is no longer a model speed. It is a rendering problem.