Back to the notes

Someone built an LLM leaderboard that knows what you can afford

A live benchmark that ranks models by performance per dollar spent, not just raw scores. Turns out the best model depends on your budget.

A detailed electronic circuit board lit in blue and orange
Umberto / Unsplash Unsplash License

Most LLM benchmarks rank models by raw performance. MMLU scores, coding ability, reasoning benchmarks. The model at the top is the one you should use, right? Best Model For Your Budget flips that. It ranks models by performance per dollar spent. Updated daily. You pick your budget, it tells you which model gives you the most capability for that spend. The interesting bit is how much the answer changes with budget. At 10 cents per million tokens, you are looking at completely different models than at 10 dollars per million tokens. The cheapest option is not always a scaled-down version of the best one. Sometimes it is a completely different architecture that happens to be efficient at the low end. This matters if you are building anything that calls an LLM in a loop. A RAG system that processes a thousand documents per query. A monitoring agent that reads logs every minute. A classification job that runs on every support ticket. The model that wins on pure capability might cost you 50 times more than the second-best option. If that second option scores 85 percent of the top model’s performance, you just found your answer. I have been treating model selection as a quality decision, not a cost-benefit decision. This site suggests that is backwards. The best model is the one that fits your budget and still does the job. Raw benchmarks do not tell you that. The other thing this makes obvious: pricing changes constantly. A model that was cost-effective last week might be beaten by a new release today. If you hardcoded a model name into your production config six months ago, you are probably overpaying now. Treat model selection like dependency updates, not architecture decisions.


Source: Best LLM for every budget, updated daily

Back to all notes

Behind the notes

Vikrant
Sharma.

Artificial Intelligence Engineer intern at Voxon Photonics in Adelaide. Studying a Master of Information and Communications Technology at UniSC, with a focus on data, machine learning and security.

Meet the person behind the work