What is benchmark cherry-picking?

Definition

Selective reporting can make progress appear smoother or more general than it is. A claim may emphasize tasks where a model improved, ignore regressions, choose a convenient baseline or assemble historical milestones that fit a preferred trend.

A stronger evaluation defines selection rules in advance and reports representative successes, failures and uncertainty. Independent replication and complete versioned results make it harder to hide contradictory evidence.

ELI5

Benchmark cherry-picking happens when someone shows only the AI tests where a model looks strong and leaves out weaker or conflicting results. A narrow success can then appear to prove general superiority.

For example, a provider might publish the best two scores from a large evaluation while omitting tasks where the model performed poorly. A fair comparison shows representative tests, settings, repeated runs and uncertainty so readers can see the complete pattern.

Acronyms and aliases

selective benchmark reporting synonymAI benchmark cherry-picking variant

Frequently asked questions

How can benchmark cherry-picking distort an AI claim?

It can exaggerate progress by selecting favorable tasks, baselines, time periods or model versions and hiding contradictory results.

How can evaluators reduce cherry-picking?

They can predefine tests, publish complete versioned results, include uncertainty and support independent replication.

Videos explaining benchmark cherry-picking