Theo Browne reviews Haiku 5.5's reported pricing and benchmarks, emphasizing cache reads, context-size tiers and reasoning settings. He separates aggregate cost-performance comparisons from practical usefulness and explains how reduced cache costs can change the economics of longer agent workflows.
He discusses pairing a stronger coordinating model with cheaper workers for exploration, classification and repeated trials that are easy to verify. Demonstrations include generated interfaces, a game and a pull-request audit. The game has control and layout defects, and the audit initially treats clean GitHub status as evidence of merge readiness without adequately checking the changes.
The PR workflow then exhausts GitHub's request allowance, mishandles an error response and claims it will wait without actually scheduling continuation. Theo Browne uses these failures to distinguish code generation from context gathering, verification and orchestration. His preference is to use small models for carefully bounded work and stronger models for planning, difficult implementation and independent review.
Watch on YouTube




