Bijan Bowen gives GPT-6.1 Sol and Claude Sonnet 5.5 four practical tasks: a destruction-derby game, a desktop robot's branding and voice software, radio-controlled car parking, and music-video production. Both produce working elements, but the game demonstrations retain fidelity and implementation defects rather than meeting the reference exactly.
Bijan Bowen prefers Sonnet's desktop-robot presentation and finds that its voice integration handles commands that fail in the competing demonstration. In the parking task, both agents use camera observations and movement models, but neither completes parking within the allotted session. Sonnet tests combined controls and requests clarification about the available gap; its later debrief request is initially refused.
Bijan Bowen runs the creative test on higher reasoning settings and prefers Sonnet's audio changes and video effects, while acknowledging a transition where GPT may do better. His overall preference concerns these specific tasks and subjective output quality, not a controlled universal benchmark or a cost-normalized comparison. His explanation that GPT's capability is being restricted remains speculation.
Watch on YouTube




