An evaluation rubric defines the qualities an evaluator should inspect and how different results should be scored. For an AI output, criteria might cover factual support, completeness, safety, relevance and the preservation of domain-specific decisions.
A rubric improves consistency but can become outdated or reward superficial compliance. Teams should test it against real failures, record expert reasoning and revise criteria when models, standards or the consequences of errors change.