Matthew Berman compares Anthropic's reported Opus 5 results with Fable 5 and earlier Opus models. Coding, computer use and interactive-game scores improve in the presented charts, while some legal and health measures decline.
His central point is to compare successful task cost, not just token pricing. Reasoning settings can change both cost and accuracy, and the most expensive setting does not always produce the strongest result. He highlights a reported ARC-AGI 3 jump but does not present a comprehensive independent test.
Matthew Berman also discusses launch pricing and automatic safety-triggered model fallbacks. His explanation of reduced cybersecurity capability and possible government review is explicitly speculative, rather than a verified account of training or approval.
Watch on YouTube




