Variation can come from model sampling, task selection, scoring judgment, service changes and the small number of cases tested. A narrow subset may give very different rankings when only a few results change.
Repeated runs, confidence intervals and larger representative sets help quantify uncertainty. A score should not be treated as a precise capability measurement when its expected variation is large.
ELI5
Benchmark variance describes how much a model's test score might move if the test were repeated or used a different small sample. High variance means the reported number is less stable.
For example, on a ten-task set, one changed result moves the score by ten percentage points. Repeating the test and adding more representative tasks gives a clearer view of whether the apparent lead is dependable.
