Selective reporting can make progress appear smoother or more general than it is. A claim may emphasize tasks where a model improved, ignore regressions, choose a convenient baseline or assemble historical milestones that fit a preferred trend.
A stronger evaluation defines selection rules in advance and reports representative successes, failures and uncertainty. Independent replication and complete versioned results make it harder to hide contradictory evidence.
ELI5
Benchmark cherry-picking happens when someone shows only the AI tests where a model looks strong and leaves out weaker or conflicting results. A narrow success can then appear to prove general superiority.
For example, a provider might publish the best two scores from a large evaluation while omitting tasks where the model performed poorly. A fair comparison shows representative tests, settings, repeated runs and uncertainty so readers can see the complete pattern.

