The presenter tests Claude Opus 5.5 through Claude Code on browser puzzles and a browser-based physics simulation. The model completes several interactive challenges, but a parking puzzle takes more than ten minutes, which the presenter counts as a practical failure. Its first physics implementation scores poorly in a self-critique loop and freezes in the browser before further revision produces a usable simulation.
Further demonstrations ask the model to perform a piano solo, build a mathematical explainer, reconstruct a property in Blender, create an Unreal Engine game and compose music through a desktop audio workstation. The presenter reports some impressive results, but the runs take substantial time; the property reconstruction is not fully faithful, and the final music track falls short of professional mastering. These are individual demonstrations described in the transcript, not controlled cross-model tests.
The review also identifies visual-recognition errors in the model's chat interface and surveys published benchmark and pricing claims. The presenter distinguishes those vendor or leaderboard results from his own trials. Together, the examples suggest substantial capability with meaningful limits in speed, fidelity and accuracy.
Watch on YouTube



