Artificial intelligence confidence estimation can use predictive probabilities, ensembles, repeated sampling, distance from familiar data, auxiliary models, or domain-specific measures. The estimate may apply to a complete prediction, a region, a token, a structure, or another defined part of the output.
A confidence score becomes useful when its meaning is explicit and empirically tested. Different scores are not automatically comparable, and a high value may still be wrong on unfamiliar data. Calibration checks whether the estimate corresponds to observed accuracy or another relevant success criterion.