DeepSeek Releases Experimental Image-Capable Model
DeepSeek has released an experimental multimodal model called "V4-Flash-Vision-Exp" that can process both text and images. In the company's agent benchmark evaluation, it demonstrated performance comparable to Anthropic's Claude Opus 4.8, with some items surpassing it. The release currently remains at the experimental stage.

DeepSeek, a Chinese artificial intelligence company, has released an experimental multimodal model called "V4-Flash-Vision-Exp" that can handle both text and images. This adds image understanding capabilities to the company's existing text-focused model "V4-Flash," and is currently positioned as being in the experimental stage.
Multimodal refers to the ability to process not just text, but also images, videos, and other media together. While most artificial intelligence models to date have specialized in text-based interactions, there has been growing demand for use cases such as "showing an image and having its content explained" or "reading diagrams and providing answers." Major development companies are competing in multimodal support. V4-Flash-Vision-Exp is positioned as DeepSeek's new move in line with this trend.
When DeepSeek tested the model with its own agent benchmark—an evaluation standard that measures an AI's ability to autonomously complete tasks—V4-Flash-Vision-Exp demonstrated performance close to "Claude Opus 4.8," a model provided by Anthropic, and in some areas exceeded those results. However, it should be noted that this evaluation was conducted by DeepSeek itself and is not an independent verification by a third party.
The reason this deserves attention is that DeepSeek has a track record of repeatedly releasing models that achieve high performance at relatively low cost, making a strong impact on the artificial intelligence industry. By releasing this new model with multimodal capabilities that come close to Anthropic's top-tier models, even in experimental form, DeepSeek once again demonstrates the speed of its development pace.
On the other hand, as the name "Experimental" suggests, this model is not yet an official release, and its stability in actual operating environments and its ability to handle a wide range of tasks will require further verification. Even if the results on agent benchmarks are excellent, whether that translates to actual user experiences remains a separate question, and future third-party evaluations and real-world usage reports will be important indicators for assessment.
The applications of artificial intelligence have already expanded beyond text generation to include document reading, image analysis, and autonomous task support that combines multiple tools—more complex use cases. DeepSeek's experimental release of an image-capable model demonstrates that the competition to meet diverse needs is expanding globally, including among Chinese artificial intelligence companies. Going forward, the transition to an official release and the results of independent benchmark evaluations will be key milestones in determining the true value of this model.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.