What is an evaluation rubric?

Definition

An evaluation rubric defines the qualities an evaluator should inspect and how different results should be scored. For an AI output, criteria might cover factual support, completeness, safety, relevance and the preservation of domain-specific decisions.

A rubric improves consistency but can become outdated or reward superficial compliance. Teams should test it against real failures, record expert reasoning and revise criteria when models, standards or the consequences of errors change.

Acronyms and aliases

scoring rubric synonymevaluation criteria rubric variant

Frequently asked questions

What should an AI evaluation rubric include?

It should include clear criteria, examples, severity guidance, scoring rules and instructions for ambiguous or high-risk cases.

Why can a static evaluation rubric fail?

It can miss new failure modes and changing standards, or reward outputs that match wording without satisfying the real task.

Videos explaining evaluation rubric