What is a benchmark?

Definition

An artificial intelligence benchmark defines a repeatable evaluation target, such as coding tasks, reasoning questions, multimodal problems, or serving performance. A score is meaningful only with the exact dataset, method, model settings, and grading rules.

Benchmarks can be contaminated, optimized against, selectively reported, or poorly matched to real use. Strong conclusions compare several evaluations and disclose uncertainty rather than treating one exceptional result as definitive.

Acronyms and aliases

AI benchmark synonymmodel benchmark synonymartificial intelligence benchmark variant

Frequently asked questions

What makes an artificial intelligence benchmark useful?

A useful benchmark has clear tasks, representative data, transparent procedures, reliable scoring, and enough discrimination to compare relevant systems.

Why can benchmark results disagree?

Results can differ because of datasets, prompts, sampling, model versions, tool access, scoring, test contamination, and reporting choices.

Videos explaining benchmark

  1. The Week Open Models Closed the Gap
    AI Search43:531 VIEW
  2. Why OpenAI Is Cutting Cursor Model Access
    Theo27:051 VIEW