Mistral Large 4 Tested on Eight Real Coding Tasks

AICodeKing9m 14s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    AICodeKing connects the Mistral Large 4 preview to OpenCode through OpenRouter and tests eight repeatable prompts in fresh folders. The review grades delivered artifacts rather than the agent's explanation, checking actual browser rendering and interaction. An elevator simulation, folding table and panda SVG work well, although smaller visual and interaction defects prevent perfect scores.

    Other tasks expose reliability problems: the first contact-lens case and counting attempts produce no usable answer, the archery game has tiny targets and layout issues, and a 3D watch initially fails because of missing module paths. The presenter keeps separate repair attempts outside the main total, but counts a successful local Gemma fine-tuning rerun after correcting a permissions problem caused by the test setup.

    The resulting first-pass score is 42 out of 80. The presenter treats that as evidence from one particular harness and prompt set, not a universal model ranking, and recommends keeping browser testing close when using the model for longer work.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Blue and white “MISTRAL LARGE 4 CODING TEST” headline above blue code brackets around a white check on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 7 October 2026 and duration 9m 14s.

    AICodeKing's eight-task test finds Mistral Large 4 capable of useful visual prototypes and local fine-tuning, but inconsistent at completing tasks without repairs.