Matthew Berman begins with Opus 5.5 release claims, benchmark comparisons and the relationship between token prices, reasoning effort and total task cost. The discussion presents vendor and third-party evaluation results as specific measurements rather than a universal ranking of which model is best.
Matthew Berman interviews Thariq Shihipar about agent tools, testing, deployment, memory and model-assisted development. Their conversation emphasizes the interaction between model capabilities and the surrounding harness, including how evaluation and prompting practices change as models improve.
Matthew Berman is joined by Forward Future editors Alex and Future Brian for game and animation demonstrations created with early model access. Alex discusses iterative game development, while Future Brian shows coded animation and puzzle experiments. These projects illustrate possible creative workflows, not controlled comparisons or evidence that every result required no iteration.
Matthew Berman then turns to GPT-6 Sol and Luna, comparing reported performance, availability, reasoning settings and compute costs. The closing discussion separates efficient workhorse behavior from frontier capability and notes that different benchmarks, budgets and tasks can support different model choices.
Watch on YouTube




