Task generalization goes beyond reproducing training examples. A model must identify the structure that matters and adapt its behavior when objects, initial states, wording, tools, or environmental details change.
Evaluation uses held-out tasks and controlled variations to measure this ability. Strong results on one test set still need broader validation because real deployments can introduce conditions absent from the original evaluation.
