Self-Improving Agents with Traces, Datasets and Evals

AI Engineer17m 7s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Marc Klingen explains how production traces become the evidence for improving an AI application. Instead of changing prompts by intuition alone, teams collect failures and feedback, curate representative datasets, compare experiments and deploy only after checking whether the proposed change improves the behavior they actually want.

    The demonstration uses a changelog agent to show how a coding agent can propose prompt and evaluation changes and test them against earlier cases. Klingen emphasizes human ownership of goals, dataset quality and deployment boundaries: an automatically improved score can otherwise reflect overfitting or reward hacking rather than a better product.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Marc Klingen against a black background beside the blue and white headline “AGENT IMPROVEMENT”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 6 October 2026 and duration 17m 7s.

    Marc Klingen presents an open-source agent improvement loop that turns production evidence into datasets, evaluations and experiments, while keeping goals and deployment decisions under human control.