Lovelaice framing

Model benchmark

Definition

The public leaderboards — MMLU, GPQA, SWE-bench, LMArena. Standardized tests that rank models on general tasks. Useful for narrowing a shortlist, useless for picking a winner: a model topping the chart can still be the worst option for your specific task.

Model benchmarks are the public leaderboards used to rank foundation models: MMLU for general knowledge, GPQA for graduate-level science, SWE-bench for software engineering, LMArena for head-to-head preference. They're the industry's shared shorthand for 'which model is best' and the reason a new model release comes with a chart.

Why it matters

Benchmarks measure general capability on problems nobody on your team has. They're also gamed, and the test data leaks into training sets over time. A model topping the leaderboard can still be the worst option for your specific task — the only benchmark that decides anything is your own.