Matthew Berman reviews Grok 4.7's launch claims and Elon Musk's predictions, comparing reported results with competing frontier models. He questions selective comparisons and emphasizes that parity on particular tests does not establish overall leadership.
Matthew Berman separates token pricesToken pricing is the rate an AI provider charges for processing input tokens, generating output tokens, or reading cached tokens. from total task costsCost per completed task measures the total AI, tool, infrastructure, retry, and repair expense for each verified useful outcome.: a model that reasons longer can consume more tokens despite a low unit price. He discusses effort settingsReasoning effort is the amount of internal computational work an AI model applies before producing an answer or action., benchmark omissionsA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions. and the difficulty of judging value from headline scores alone.
Matthew Berman reports stronger results in some knowledge-work tasks alongside less convincing coding and terminal performance. He examines independent rankings and context-window limitsA context window is the maximum amount of tokenized information an AI model can consider during one processing session., while keeping outside demos distinct from controlled testing. This is an initial assessment of reported evidence, not a comprehensive hands-on evaluation.
Watch on YouTube




