Bijan Bowen tests Cognition SWE-2, which he says was post-trained from Kimi K3, across a varied set of coding and creative tasks. The demonstrations cover a browser-style interface, a C++ skateboard simulation, game and 3D assets, a subway game and a watch-related site.
The results are mixed. Some generated applicationsCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications. are usable or visually promising, while a robot-arm task fails in ways Bowen considers unsafe for real-world deploymentReal-world AI evaluation measures model or agent behavior in genuine deployment conditions where actions encounter actual users, systems, constraints, and consequences.. He discusses how the model behaves during iterations rather than relying on one benchmark numberA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions..
Bowen's comparisons, speed impressions and pricing remarks are his observations in this session, not independent controlled measurements. The video is most useful for seeing specific success and failure modesModel reliability is the degree to which an AI model produces dependable behavior across repeated, varied and operationally realistic use. before deciding whether SWE-2 fits a workflow.
Watch on YouTube




