Theo Browne reviews Haiku 5.5's reported pricing and benchmarks, emphasizing cache reads, context-size tiers and reasoning settings. He separates aggregate cost-performance comparisonsInference cost is the expense of running a trained AI model to process inputs and produce outputs. from practical usefulness and explains how reduced cache costs can change the economics of longer agent workflows.
He discusses pairing a stronger coordinating model with cheaper workers for exploration, classification and repeated trials that are easy to verify. Demonstrations include generated interfaces, a game and a pull-request audit. The game has control and layout defects, and the audit initially treats clean GitHub status as evidence of merge readiness without adequately checking the changes.
The PR workflow then exhausts GitHub's request allowance, mishandles an error response and claims it will wait without actually scheduling continuation. Theo Browne uses these failures to distinguish code generationCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications. from context gathering, verification and orchestrationAgent orchestration coordinates AI agents, tools, people, tasks, state, and control flow so a larger workflow reaches a verified outcome.. His preference is to use small models for carefully bounded work and stronger models for planning, difficult implementation and independent reviewCode review examines proposed software changes for correctness, clarity and risks before accepting them, including changes produced by an AI agent..
Watch on YouTube




