Large language model judge calibration compares automated judgments with labels in a golden dataset. Teams revise the judge prompt, rubric or scoring method, then inspect whether agreement improves across representative and difficult cases.
Calibration is not permanent. Model changes, prompt edits and new product behavior can alter scores, so the process should be versioned, repeated and monitored with visible examples and reviewer ownership.