The talk frames a software factory as a workflow that moves work from issue triage toward production. Its first improvement loop has an agent inspect a triage agent's runs and human feedback, then propose revisions to the triage skill. The example keeps those changes in a pull request for human review, so an erroneous revision can be caught before the skill changes.
A second loop extracts reusable facts and outcomes from past runs into an agent-scoped memory store. The example concerns recurring software issues: stored root-cause context can help a later run avoid rediscovering the same information. The speaker also describes versioning, source tracing and human review of memories.
Finally, the talk argues for assigning different model classes to different tasks instead of using one expensive model everywhere. It describes task-specific routing rules and an evaluation approach that compares models on a team's own workflows. The evaluation and customer-facing routing capabilities are presented partly as current internal practice and partly as future work, not as a measured general performance result.
Watch on YouTube




