OpenAI should be scared of this one

Theo31m 46s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Theo Browne reviews Sonnet 5.5 after switching much of his own coding work to Opus 5.5. He contrasts benchmark results with practical behavior, arguing that model choice should account for the task, reasoning setting and total token usage rather than headline token prices alone.

    Theo Browne explains why cached-input costs can become a larger share of an agent's bill as other token prices fall. In his comparisons, Sonnet 5.5's extra token consumption sometimes eliminates the expected savings against Opus 5.5. He strongly criticizes maximum reasoning settings, reporting much higher token use without reliably better results on his tests.

    The review distinguishes output speed from end-to-end completion time and discusses weaknesses in generated frontend designs. Theo Browne reports that his fish-game task took longer with Sonnet 5.5 than with Opus 5.5 despite faster token generation. He also notes that the Opus run used a Fable subagent, limiting the fairness of that comparison.

    Theo Browne finds a clearer advantage in an architectural audit of a large code change: Sonnet 5.5 delivers useful analysis at substantially lower cost and latency in his custom evaluation. He suggests that stronger models could delegate well-scoped investigations to Sonnet 5.5 rather than using it for every coding task.

    Theo Browne demonstrates a generated 3D game and praises its movement, mechanics and recognizable assets. His conclusion is task-specific: Sonnet 5.5 can be a valuable analytical tool within a larger workflow, while Opus 5.5 remains his preferred general coding model. The video presents personal evaluations, not universal rankings.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Theo Browne in a blue top beside the blue and white headline 'SONNET'S REAL VALUE' on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 29 September 2026 and duration 31m 46s.

    Theo Browne finds Sonnet 5.5 most compelling for delegated codebase investigation, while warning that token consumption and reasoning settings complicate its apparent price advantage.