Elon promised this one would be good...

Theo27m 53s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Theo Browne examines the gap between Grok 4.7's release claims, benchmark scoresA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions. and his own development experience. He distinguishes tokens per API request from total cost per taskCost per completed task measures the total AI, tool, infrastructure, retry, and repair expense for each verified useful outcome., arguing that additional tool calls and prolonged verification can make unchanged token pricesToken pricing is the rate an AI provider charges for processing input tokens, generating output tokens, or reading cached tokens. misleading as a guide to practical expense.

    Theo Browne describes a codebase-review experimentCode review examines proposed software changes for correctness, clarity and risks before accepting them, including changes produced by an AI agent. in which models propose and validate improvements to T3 Code. Grok 4.7 performs useful investigation in that limited test, while his frontend examples and fish-game generation are substantially weaker. He also discusses noisy effort-dependent benchmark scores and unusual looping behavior.

    Theo Browne's conclusion is mixed: he enjoys aspects of the model's persistence and curiosity but finds its speed, cost and subscription tradeoffs hard to justify against alternatives. The rankings are his observations from specific runs, including model-assisted judgingA large language model as a judge is an evaluation method in which a language model scores, compares or critiques another system's output using stated criteria., rather than a controlled proof of universal model quality.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Theo Browne in blue lightly touches his chin beside blue “GROK 4.7” and white “USEFUL BUT COSTLY” on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 22 September 2026 and duration 27m 53s.

    Theo Browne finds useful codebase investigation in Grok 4.7 but questions whether its inconsistent results and higher task costs justify the upgrade.