Pat Simmons gives Claude Opus 5.5, Claude Fable 5.1 and GPT-6 Astra similar high-effort build prompts in their respective agent environmentsAn AI agent environment is the bounded world in which an agent observes state, uses tools, interacts with others, receives feedback, and produces consequences.. The three tasks are different apps: a code-drawn Moby Dick storybookCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications., a product-launch film and website, and a Tony Hawk's Pro Skater-inspired browser gameGame generation uses AI to create or assemble playable assets, scenes, rules, audio and interactions.. This is a narrated hands-on judgment of these examples, not a controlled benchmark of all coding workA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions..
For the storybook, Simmons favors Opus 5.5's detailed scenes over Fable 5.1's interactive chapters and Astra's faster but less complete result. For the launch task, the models generate imagery for a short film, assemble video and sound with code, and build a matching websiteFront-end code generation uses AI to create or modify the user-interface code of websites and applications from instructions or examples.. Simmons again prefers Opus's visual pacing and closer website interpretation, while criticizing sound artifacts across the outputs.
The final game test compares rendered characters, ramps, controls and skating physics. Simmons judges Opus the strongest of the three builds, but says Astra remains a personal choice for everyday large refactors and site architecture. Reported API-equivalent costs come from agents estimating their own session logs, and Simmons explicitly questions their accuracy; the first task also includes a usage-limit interruption.
Watch on YouTube



