Artificial intelligence benchmark cherry-picking presents a model through the evaluations where it performs best without showing the broader evidence. Selection can involve tasks, datasets, metrics, prompts, models or repeated runs.
The practice can make a narrow strength look like general superiority. Readers should look for complete methodology, representative workloads, uncertainty and independent replication before treating selected scores as conclusive.