An optimization process proposes prompt variants, evaluates their outcomes and uses feedback to select or generate further candidates. Feedback can include task scores, execution traces and explanations of errors rather than a single aggregate number.
The evaluation must reflect the intended goal and include cases not used to select the prompt. Better results on a fixed development set can arise from overfitting or scoring flaws, so changes need checks on representative held-out work.
ELI5
This means trying different ways to explain a job to the AI and keeping the versions that work better under a clear test. The model itself does not have to be retrained.
For example, an assistant may miss an important rule in a support task. A revised prompt can state that rule more clearly, then be tested on fresh questions to see whether the improvement is real.
Does prompt optimization change model weights?
Not necessarily. It can improve the supplied instructions while leaving the underlying trained model unchanged.
Can an optimized prompt work poorly on new tasks?
Yes. It may overfit the examples used during the search, so representative held-out evaluation matters.


