Capability measurement uses defined tasks, environments, and success criteria to estimate what an AI system can do. It can cover reasoning, tool use, planning, coding, scientific work, persuasion, autonomy, or other abilities relevant to deployment and safety decisions.
A single benchmark is not enough. Measurement should include realistic and adversarial conditions, repeated trials, uncertainty, prompt sensitivity, hidden limitations, and changes caused by tools or surrounding systems.
ELI5
Capability measurement tests what an AI can actually do instead of relying on claims or impressive demonstrations. It defines tasks and success conditions, then checks performance across enough cases to understand strengths and limits.
For example, before giving an agent powerful research tools, evaluators can test whether it plans long tasks, follows constraints, and recovers from errors. Repeated and difficult cases reveal more than one carefully chosen success.
