Matthew Berman reviews Grok 4.6 as an iterative upgrade rather than a new base-model training run. The model improves substantially over Grok 4.5 on knowledge-work, terminal, legal and coding benchmarks, although it still trails the best coding models on some developer-focused evaluations.
The comparison emphasizes cost per completed task instead of benchmark quality alone. Grok 4.6 becomes more expensive than its predecessor because it uses more reasoning, but it remains cheaper than several similarly capable frontier models. A simple profile-card test also shows that its visual coding output is usable but less polished than the strongest result in the comparison.
Berman attributes xAI's progress to combining large-scale compute with Cursor's coding data and product feedback. Grok 4.5 helped regenerate supervised fine-tuning trajectories for Grok 4.6 across software engineering, science and knowledge work, illustrating how current models already contribute to training their successors.
Grokbot packages the same capabilities for a broader audience by hiding model selection and code while returning documents, presentations and other work products. Berman argues that this pairing gives xAI a clearer product strategy and increases competition with OpenAI and Anthropic, which could improve models and lower prices for users. Newsletter cross-promotion is omitted.
Watch on YouTube


