Benchmark
A fixed set of tasks with a score, used to compare models — and only as good as the tasks resemble your work.
A score is a number about one set of tasks, not a property of the model. Two models a point apart on a public leaderboard can be far apart on your codebase, and the harness around the model — how many attempts, what tools, what prompt — often moves the number more than the model does.
Read the method before the number: how many runs, whether the tasks leaked into training, who paid for the evaluation. A benchmark whose answers are already in the training data measures memory, not ability.