Artificial intelligence prompt sensitivity occurs when wording, order, examples or surrounding context substantially affect model behavior. Casual prompts may understate capability, while highly tuned prompts can overstate how well the model handles ordinary use.
Evaluation should use several realistic prompt forms and preserve the harness and settings. Consistent performance across variations provides stronger evidence than one carefully selected prompt that produces an impressive answer.