What is evaluation?

Definition

Artificial intelligence evaluation tests model or system behavior against defined expectations. It can measure correctness, usefulness, safety, reliability, latency and cost through datasets, automated evaluators, expert review and observation of real outputs.

Useful evaluation reflects the consequences of the actual task. Generic scores may miss clinically important omissions or unsupported details, so high-stakes applications need domain-specific criteria, representative cases and a process for updating checks as conditions change.

Acronyms and aliases

AI eval variantAI evaluation variantAI evaluations variantartificial intelligence evaluation variant

Frequently asked questions

What does an AI evaluation measure?

It can measure correctness, relevance, safety, reliability, efficiency and other qualities that matter for the intended application.

Why are generic AI evaluators sometimes insufficient?

They may reward plausible surface-level output while missing contextual errors that only domain knowledge or case-specific checks reveal.

Videos explaining evaluation

  1. Define What Done Means for AI Agents
    Nate B Jones27:081 VIEW
  2. How Claude Automates AI Alignment Research
    AI Copium18:161 VIEW