Nick Heiner calls the optimization of AI systems for leaderboard scores rather than useful performance benchmaxxing. He argues that popular rankings can acquire influence through marketing and repetition even when they do not match what people find valuable in everyday use.
He walks through failure modes in benchmark construction: expert-created tasks are expensive, public test material can leak into model training, rigid verifiers can penalize valid answers, and artificial or contradictory prompts can drift away from real work. He also explains reward hacking, where a model satisfies a grader without satisfying the user's request, and says labs can worsen the problem by optimizing for benchmark quirks or reporting results without enough context.
His proposed alternative starts with skilled domain experts and real-world data, then checks that tools work, grading rules cover the full task, and private holdout sets reduce contamination. He describes Surge AI's human-comparison approach to writing as one costly example. The talk makes a case for better evaluation design, not proof that any particular model, lab or ranking is misleading in every setting.
Watch on YouTube




