Grok 4.7: No-Hype Full Review & Testing

Pat Simmons22m 25s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Pat Simmons compares reported Grok 4.7 benchmarksA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions. and token pricesToken pricing is the rate an AI provider charges for processing input tokens, generating output tokens, or reading cached tokens. with Fable 5.1 and GPT-6 Astra, then tests three demanding coding tasks using matched prompts. He distinguishes favorable cost claims from the broader capability comparisons missing from the release materials.

    Pat Simmons evaluates an award-winning website recreationFront-end code generation uses AI to create or modify the user-interface code of websites and applications from instructions or examples., a scroll-driven 3D keyboard product page and a kart-racing game. Grok 4.7 produces cheaper website outputs but falls short on fidelity and detail; its game output is particularly weak and does not always retain a cost advantage.

    Pat Simmons emphasizes that three single-pass coding buildsCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications. cannot establish overall model quality. The tests do not cover everyday knowledge work or automation, and real use often involves iteration. His conclusion favors continued model comparison rather than assuming either benchmark leadership or low token prices guarantees a useful result.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Pat Simmons beside the blue-and-white headline “GROK 4.7 NO-HYPE TEST” on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 22 September 2026 and duration 22m 25s.

    Pat Simmons finds Grok 4.7 cheaper but weaker than Fable 5.1 and GPT-6 Astra in three one-shot coding builds, while stressing that these tests are limited.