
Just now, the first multimodal model in DeepSeek V4, DeepSeek-V4-Flash-Vision-Exp, became available on Hugging Face under the MIT License.

The publicly released materials include model files, the Tokenizer, a reference implementation for Prompt Encoding, and a minimal PyTorch inference implementation covering core modules such as the vision encoder, Aligner, DFlash Attention, MoE, Hyper-Connections, and DSpark.
ITHome learned that DeepSeek-V4-Flash-Vision-Exp is the first visual model in the DeepSeek V4 series to natively support image input. It is officially labeled an “Exp” experimental version and was launched on the DeepSeek API platform on August 21, 2026.
In addition to text, the model supports image input, enabling it to describe images, recognize text in screenshots, analyze charts, and more. Supported image formats include JPEG, PNG, GIF, and WebP.
DeepSeek officials stated that V4-Flash-Vision-Exp demonstrates strong multimodal adaptability across various Agent frameworks, effectively supporting users in unlocking more practical work scenarios with diverse Agent tools. It matches the full V4-Flash version on text-only tasks such as Agent, reasoning, and world knowledge, fully retaining its existing strengths. On Agent benchmarks requiring visual understanding, it significantly outperforms V4-Flash, with multimodal Agent capabilities now approaching Opus-4.8.
