Bijan Bowen reviews Opus 5.5's stated efficiency and benchmark improvements before trying several hands-on coding tasksCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications. at different reasoning effortsReasoning effort is the amount of internal computational work an AI model applies before producing an answer or action.. He explores a browser operating-system simulation with working applications and shared windows, then tests a self-contained skateboarding game and a watch product page.
Bijan Bowen reports strong creative details in game graphics, animations, sound and interfaces, but also identifies invisible walls, missing views and controls that need follow-up. A physical robot-arm taskEmbodied AI perceives and acts through a physical or simulated body while interacting with an environment. uses a simulated plan yet fails to pick up the car; careful reasoning alone does not complete the real-world action.
Bijan Bowen adds an unfamiliar guitar-store game prompt to explore performance beyond recurring demonstration tasks. The result creates hundreds of instrument models but suffers from lag and incomplete gameplay. These are qualitative prototypes and personal observations, not a controlled benchmarkA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions. proving broad reliability or generalization beyond training data.
Watch on YouTube




