Mercury 2.5 LLM clocks 770 tokens per second
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
Read the note 1 min read
From the notebook
3 notes on this topic.
A new language model is generating text faster than most tools can render it. The bottleneck just moved from the API to the browser.
Someone built a tool to make AMD GPUs run local LLMs without the usual driver nightmares. The benchmarks show Vulkan beating HIP by 20 percent.
Anthropic open-sourced a framework for testing how well AI models find security bugs. It includes 32 real CVEs and a scoring system. Time to feed it some of my old projects.