Why ARC-AGI-3 Breaks LLM Agents

0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Tim Scarfe joins Benjamin Crouzier, Jeroen Cottaar, Dries Smit, Stefano Viel and Michal Tešnar to examine ARC-AGI-3, an interactive benchmark in which an agent must infer the rules and goal of an unfamiliar game from raw frames and a limited action budget. Humans bring strong priors about players, mazes, keys and exits, while a model can identify many of the same concepts yet still act inefficiently or settle on the wrong objective.

    The team describes a harness that combines language-model reasoning with multiple representations of the game state, code execution, structured tools and reward shaping. Their experiments show that models benefit from human concepts encoded in language and color, but long-context reasoning remains brittle: once an agent forms a bad hypothesis, it can keep interpreting every later observation through that frame instead of revising its plan.

    The discussion separates competence from performance and asks what a high benchmark score would actually prove. Action efficiency limits brute force, while unseen games require general strategies rather than memorized solutions. The team expects progress to combine detailed harness engineering, reinforcement learning and stronger base models, with the most valuable result being reusable ways to acquire abstractions rather than a system tuned to a fixed set of games.

    Original YouTube thumbnailWatch on YouTube