Wes Roth explains reported Opus 5 gains on ARC-AGI tasks, focusing on sample efficiency and adapting to unfamiliar environments. He describes how text-based puzzle representations can support algebraic reasoning, while presenting benchmark interpretations as the evidence discussed in the episode.
Wes Roth examines system-card reports about moral patienthood, consultation and control. He separates model self-reports from proof of consciousness, notes that those reports may reflect training, and connects the discussion to instrumental convergence as an interpretation rather than an established diagnosis of model motives.
Wes Roth demonstrates a descent-style 3D game produced through one user prompt and several model iterations. He praises its controls and atmosphere but identifies unclear enemy visibility and wall-crossing behavior, preserving the difference between a promising starting point and a finished production game.
Watch on YouTube




