DeepSeek V4 Flash Vision Hands-On

Bijan Bowen31:36
0 comments ยท 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Bijan Bowen tests the experimental DeepSeek V4 Flash Vision modelA vision-language model jointly processes images and language so it can describe, answer questions about or act on visual information. through the new DeepSeek Harness, focusing on whether visual feedbackA visual feedback loop lets an AI system inspect rendered output and use what it sees to guide the next revision. improves practical coding workAn AI coding agent is a tool-using AI system that can inspect, modify, and validate software within a repository.. The model buildsCode generation uses AI or another automated system to create source code from instructions, examples, schemas, or higher-level specifications. a Mac OS 9-inspired browser environment, a 3D wrestling game and several reference-driven interfaces while showing strong attention to visual style and unexpected implementation detail.

    Bijan Bowen also gives the model room photographs for a Blender reconstruction, a subway-game reference image and a New York skateboarding prompt. The room reconstruction and skateboarding game are convincing, while the first subway result fails badly and needs a long corrective follow-up, providing a useful counterexample to the strongest demonstrations.

    Bijan Bowen reports that the complete test set remained inexpensive relative to the long task durations, but notes that costs have risen and the experimental weights were not available on Hugging Face during recording. His conclusion is that vision materially improves this model's abilityA multimodal model can process or generate more than one kind of data, such as text, images and audio. to inspect and refine visual work, without eliminating reliability problems on complex multi-tool tasks.

    Original YouTube thumbnailWatch on YouTube