Domain-specific benchmark
Definition
The same idea as a model benchmark, run on your data, your constraints, your users' language: your golden dataset, scored across candidate models and prompts. The number that should decide anything — and it routinely disagrees with the public leaderboards.
A domain-specific benchmark is what you get when you score a set of candidate models and prompts against your own golden dataset. It's the same idea as a model benchmark — comparable numbers across systems — but the questions come from your domain, your users, your edge cases. It's the only benchmark that maps to the decision you're actually making.
Why it matters
An MMLU score doesn't tell you whether a model can read a German insurance policy. Your own benchmark does, and it routinely disagrees with the leaderboard. This is where the decision to switch models — up or down the size ladder — should live.
Related terms