Bijan Bowen tests the two models in their respective coding harnesses with the same prompts. His goal is to replace broad claims that an open-weights model has caught a frontier closed model with outputs viewers can inspect side by side.
The tasks cover browser interfaces, a playable game and Blender work, which exposes differences that aggregate benchmark scores can hide. Bowen evaluates whether each model follows the specification, produces coherent structure and completes the requested interaction rather than rewarding a visually impressive first frame alone.
The comparison shows that model choice depends on the kind of work and the failure modes a developer can tolerate. GLM 5.2 is a credible option for some agentic coding tasks, while Claude's strengths remain visible in consistency and execution across the broader test set.
Watch on YouTube



