Bijan Bowen tests Claude Sonnet 5.5 with a browser-based desktop, a driving game, a C++ skateboarding game and a subway shooter. The generated environmentsGame generation uses AI to create or assemble playable assets, scenes, rules, audio and interactions. impress him with their detail and interactivity, but several maximum-effort tasks take close to two hours. Switching effort settingsReasoning effort is the amount of internal computational work an AI model applies before producing an answer or action. becomes part of the experiment, and missing sound or incomplete details remain visible limitations in his account.
Bijan Bowen reports a fast completion in his robot-arm challenge but explicitly treats it as an accidental success rather than competent control. The arm pushes against the table and almost drops the object, illustrating how a favourable completion time can conceal a poor execution strategyAI trajectory evaluation assesses the sequence of reasoning-relevant states, tool calls, decisions, and side effects that led to an agent's final result.. This contrasts sharply with the model's stronger performance on visual software tasks.
Bijan Bowen adds an unfamiliar backyard pool-party game prompt to probe whether the model performs beyond his repeated tests. The resulting Blender and Godot project, followed by a RuneScape-style recreation, strengthens his positive assessment of its visual coding abilities. His comparison with GPT-6 Sol is limited to these hands-on tasks, not a controlled general-purpose benchmark, and his reported subscription usage is specific to this session.
Watch on YouTube




