Marc Klingen explains how production traces become the evidence for improving an AI application. Instead of changing prompts by intuition alone, teams collect failures and feedback, curate representative datasets, compare experiments and deploy only after checking whether the proposed change improves the behavior they actually want.
The demonstration uses a changelog agent to show how a coding agent can propose prompt and evaluation changes and test them against earlier cases. Klingen emphasizes human ownership of goals, dataset quality and deployment boundaries: an automatically improved score can otherwise reflect overfitting or reward hacking rather than a better product.
Watch on YouTube




