Someone built an LLM leaderboard that knows what you can afford
A live benchmark that ranks models by performance per dollar spent, not just raw scores. Turns out the best model depends on your budget.
Read the note 1 min read
From the notebook
2 notes on this topic.
A live benchmark that ranks models by performance per dollar spent, not just raw scores. Turns out the best model depends on your budget.
Dan Luu measured the same coding task twenty times with the same prompt. The variance in output quality was higher than the difference between model versions.