A benchmark creates a repeatable evaluation target by specifying inputs, tools, model settings, success criteria, scoring, and aggregation. Different benchmarks can focus on coding, computer use, scientific research, reasoning, safety, or other bounded capabilities.
A score is evidence about performance under that test, not proof that one model is best for every workflow. Task mix, prompts, tools, effort settings, judges, versions, and cost can change both the result and its practical meaning.
ELI5
A benchmark is a fixed set of tasks and scoring rules used to compare AI systems. It creates a common test so results from different models or versions can be discussed using the same measurements.
For example, a coding benchmark may give every model the same broken programs and count how many are repaired correctly. One benchmark cannot prove that a model is best for every real job, especially if the test is narrow or its answers have leaked into training data.









