A calibrated system is uncertain at the right times. Across many cases, claims assigned similar confidence should be correct at approximately the corresponding frequency, rather than confidence merely sounding persuasive.
Calibration helps an agent decide whether to proceed, ask a question, retrieve more evidence, or escalate. It should be evaluated separately for facts, inferred preferences, recommendations, and actions because one overall score can conceal important differences.
ELI5
Confidence calibration checks whether an AI system knows how sure it should be. A system that says it is highly confident should be right much more often than when it says it is uncertain.
For example, if an agent labels one hundred product matches as 80 percent confident, roughly eighty should prove correct over time. If only half are correct, its confidence is too high and needs adjustment.

