DeepSeek's First Image-Capable V4 Model Matches Claude Opus 4.8
The V4 line now covers image input, and the open model matches Claude Opus 4.8 on vision-inclusive benchmarks, putting a closed-model-class capability under an MIT license.
Reporting from 1 source: GIGAZINE.
DeepSeek released DeepSeek-V4-Flash-Vision-Exp as an open model, the first in the V4 series to accept image input. The 305-billion-parameter multimodal model builds on the V4-Flash architecture and adds vision through post-training. On benchmarks that include image recognition, it scores on par with Claude Opus 4.8 while keeping text performance at or above the earlier V4-Flash-0731. It is available from Hugging Face and ModelScope under an MIT license, and through the API.
In benchmark comparisons against DeepSeek-V4-Flash-0731 and Claude Opus 4.8, the model holds text performance at or above the earlier V4-Flash release while adding vision, and it scores on par with Claude Opus 4.8 on tests that include image recognition.
Weights are published on Hugging Face and ModelScope under an MIT license, and the model is reachable through the API. The API accepts JPEG, PNG, GIF, and WebP, with a maximum input resolution of 8192 pixels on the long side, dropping to 4096 when more than 15 images are sent. Large images are resized to an 800 by 800 pixel equivalent, capping tokens at about 384 per image. Peak-hour pricing per million tokens is $0.44 for input, $0.014 for cached input, and $1.32 for output, with off-peak rates at half that.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.