What is data annotation?

Definition

Data annotation converts raw examples into evidence with explicit meaning. Domain experts or trained reviewers may label correctness, preference, error type, safety, relevance or another quality dimension needed by an AI system.

Annotation quality depends on clear guidelines and reviewer calibration. Self-serve tools can reduce operational cost, but disagreements, uncertainty and ownership should remain visible rather than being hidden behind a single label.

ELI5

Data annotation adds useful labels or judgments to raw examples so they can support AI training or evaluation. The labels explain what an example contains or how it should be interpreted.

For example, reviewers might mark customer messages as billing questions, technical problems, or complaints. Clear instructions and checks are needed because inconsistent labels teach the system an unclear or misleading pattern.

Acronyms and aliases

data labeling synonym

Frequently asked questions

Who should annotate AI evaluation data?

Reviewers need enough domain knowledge to make the quality decision, with guidelines and escalation for ambiguous or high-impact cases.

How can teams improve annotation quality?

They can use clear rubrics, examples, reviewer calibration, disagreement tracking, audits and tools that preserve the source context.

Videos explaining data annotation