What is capability measurement?

Definition

Capability measurement uses defined tasks, environments, and success criteria to estimate what an AI system can do. It can cover reasoning, tool use, planning, coding, scientific work, persuasion, autonomy, or other abilities relevant to deployment and safety decisions.

A single benchmark is not enough. Measurement should include realistic and adversarial conditions, repeated trials, uncertainty, prompt sensitivity, hidden limitations, and changes caused by tools or surrounding systems.

ELI5

Capability measurement tests what an AI can actually do instead of relying on claims or impressive demonstrations. It defines tasks and success conditions, then checks performance across enough cases to understand strengths and limits.

For example, before giving an agent powerful research tools, evaluators can test whether it plans long tasks, follows constraints, and recovers from errors. Repeated and difficult cases reveal more than one carefully chosen success.

Frequently asked questions

Why is capability measurement important for safety?

Controls and deployment decisions need evidence about the actions a system can perform and the conditions under which those abilities appear.

Can benchmark scores fully measure AI capability?

No. Real capability also depends on prompts, tools, context, environments, repeated reliability, and tasks the benchmark did not include.

Videos explaining capability measurement