Human evaluation can use ratings, rankings, pairwise choices, expert review or task completion studies. It is especially important when quality depends on context or subjective judgment that an automatic metric cannot fully represent.
Human judgments also vary with the evaluator, instructions and examples. A sound study uses clear rubrics, representative participants, blinded comparisons where practical and enough independent judgments to measure disagreement rather than hiding it.
ELI5
Human evaluation means asking people to judge how well an AI result works. People can notice usefulness, beauty, intent and awkward mistakes that a numeric test may miss.
For example, several creators could compare two versions of a generated video edit and choose which one better preserves the intended mood. Their opinions may differ, so the evaluation needs clear questions and more than one person's view.
