Bijan Bowen Tests GPT-6.1 Sol and Sonnet 5.5 on Four Tasks

Bijan Bowen1h 11m
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Bijan Bowen gives GPT-6.1 Sol and Claude Sonnet 5.5 four practical tasks: a destruction-derby game, a desktop robot's branding and voice software, radio-controlled car parking, and music-video production. Both produce working elements, but the game demonstrations retain fidelity and implementation defects rather than meeting the reference exactly.

    Bijan Bowen prefers Sonnet's desktop-robot presentation and finds that its voice integration handles commands that fail in the competing demonstration. In the parking task, both agents use camera observations and movement models, but neither completes parking within the allotted session. Sonnet tests combined controls and requests clarification about the available gap; its later debrief request is initially refused.

    Bijan Bowen runs the creative test on higher reasoning settings and prefers Sonnet's audio changes and video effects, while acknowledging a transition where GPT may do better. His overall preference concerns these specific tasks and subjective output quality, not a controlled universal benchmark or a cost-normalized comparison. His explanation that GPT's capability is being restricted remains speculation.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Bijan Bowen beside the blue and white headline FOUR AGENT TASKS on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 1 October 2026 and duration 1h 11m.

    Bijan Bowen compares two coding agents on four practical tasks, preferring Sonnet 5.5 overall while showing imperfect outputs and no successful parking result.