An artificial intelligence evaluation metric converts an aspect of system behavior into a value that can be compared across outputs, models or versions. Metrics can measure accuracy, omission rates, unsupported claims, task completion, safety violations, latency, cost and many other properties.
No single metric captures every important quality. A high aggregate accuracy score can hide rare but consequential failures, so teams often combine several metrics with case-level inspection and expert judgment that reflects the real cost of different errors.