Bijan Bowen introduces Mistral Large 4 as a preview model and separates reported benchmark results from his own hands-on tests. He explains that its eventual open-weight release does not make a trillion-parameter model practical for ordinary local hardware, and evaluates the preview through Mistral's Vibe coding harness.
Bijan Bowen's browser operating-system test produces working utilities and a notably creative mechanism for installing custom apps. Skateboarding, pool-party and subway games demonstrate useful coding capability but also show inverted text, broken interactions and modest visual quality. An image-guided revision initially fails because the harness is not actually providing native image access, illustrating how tooling can confound a model evaluation.
Bijan Bowen finds further weaknesses in watch-site scale, apartment geometry and reference-image replication, while a multi-era city captures some period details despite spatial errors. His roughly $47 test bill and mixed outputs lead to a qualified assessment: a significant improvement over earlier Mistral models, not proof of superiority over competitors or a finished production system.
Watch on YouTube




