A model is calibrated when predictions assigned a particular confidence level succeed at approximately that rate across comparable cases. Calibration is usually assessed over groups of predictions with tools such as reliability diagrams, calibration error metrics, and held-out evaluation data.
Calibration can change across populations, tasks, time, and unfamiliar inputs. A globally calibrated system may still be unreliable in a high-risk subgroup, so teams should evaluate relevant slices, preserve uncertainty around estimates, and combine confidence with decision costs and human review.
Acronyms and aliases
AI confidence calibration acronymmodel calibration synonymartificial intelligence confidence calibration variant
Related terms
Frequently asked questions
What does it mean for artificial intelligence confidence to be calibrated?
It means that predictions given a stated confidence level are correct at roughly that frequency across an appropriate set of comparable cases.
Can artificial intelligence confidence calibration drift?
Yes. New users, changed data, unfamiliar conditions, and model updates can make earlier calibration measurements inaccurate.