Two Anthropic Applied AI presenters trace a progression from text-only model APIs to agent SDKs and managed runtimes. They identify the production concerns that appear once agents run for many users, including hosting, session management, credentials, code execution, isolation and observability.
The presenters argue that an agent harness must change as model capabilities change. They describe one case where a context-reset workaround became overhead after a later model improved, then explain their design choice to separate the agent loop from tool execution and persist session events so work can resume after failures.
A site reliability incident demo illustrates agent definitions, restricted execution environments, uploaded evidence and durable sessions. The presenters also discuss credential vaults, parallel container setup, context recovery from session logs, private tool connections, periodic memory updates and a separate grader that checks task outcomes against a user-defined rubric. Their performance and product claims are presented as their own reported experience.
Watch on YouTube




