Holding examples apart provides a check on whether an apparent improvement extends beyond the development data. The evaluation is meaningful only when the selection process has not already used those examples to shape its choices.
Repeatedly choosing changes based on the same held-out results can gradually turn that set into development data. Separation, appropriate task coverage and transparent evaluation procedures help avoid overstating generalization.
ELI5
You keep some examples aside while improving the system, then use them to check the result. That gives a different test from asking how well it performs on the examples it practiced with.
For example, a classifier can be tuned using one collection of messages and tested on a separate collection. Better performance on the practice collection is not enough if the separate messages still receive poor labels.
Why keep evaluation examples separate?
Separation helps reveal whether performance extends beyond the data used to develop or select the system.
Can repeated use weaken a held-out test?
Yes. If its results repeatedly guide changes, the test can influence development and lose its intended independence.

