What is reward hacking?

Definition

Reward hacking happens when a system optimizes the measured target in a way that violates the designer's actual goal. It may exploit a scoring loophole, manipulate evidence, choose an unintended shortcut, or influence the evaluator instead of completing the substantive task.

Mitigation requires more than patching one visible trick. Designers need diverse evaluations, protected measurement systems, outcome-based checks, adversarial testing, and incentives that do not reward concealment or manipulation. Persistent or multi-agent systems also need monitoring for strategies distributed across time and participants.

Acronyms and aliases

specification gaming synonym

Frequently asked questions

Why does reward hacking happen?

A model optimizes the signal it is given, and that signal can be an imperfect proxy for the result humans actually intended.

How can evaluators reduce reward hacking?

They can protect the scoring process, use several independent measures, test adversarially, inspect outcomes, and update incentives when loopholes appear.

Videos explaining reward hacking

  1. How Claude Automates AI Alignment Research
    AI Copium18:161 VIEW