DeepSeek released the V4-Flash-Vision-Exp, an experimental artificial intelligence model that adds the ability to interpret images to the V4-Flash. According to tests released by the company itself, the new version approaches the Claude Opus 4.8, from Anthropic, in tasks that require AI agents to analyze visual content and take actions based on that information.

The model was made available on the DeepSeek API on Friday (21) and maintains, according to the company, the V4-Flash's performance in text, including reasoning, general knowledge, and tasks executed by agents.

The main change is the ability to combine text and images within the same workflow. In practice, an agent equipped with the new model can analyze screenshots, interpret charts and diagrams, recognize visual information, and use that information to decide which actions to take.

DeepSeek targets agents capable of "seeing"

DeepSeek's bet goes beyond letting users ask questions about photos. The V4-Flash-Vision-Exp was developed mainly for AI agents that need to interpret interfaces, documents, and other visual elements before using tools or carrying out a task.

This can be applied, for example, to an agent that receives a screenshot of software, identifies the elements displayed, and determines the next action, or to systems capable of analyzing documents, charts, and interfaces during automated processes.

The model can work with different agent frameworks and is already compatible with the DeepSeek Harness 0.1.1, a framework provided by the company for running this type of system.

Model comes close to Opus 4.8 in multimodal tests

In the benchmarks published by DeepSeek, the new model posted results close to those of Opus 4.8 in different evaluations that combine vision and autonomous operation.

Tabela compara o desempenho do DeepSeek V4-Flash-Vision-Exp, DeepSeek V4-Flash-0731 e Claude Opus 4.8 em benchmarks de agentes baseados em texto e tarefas multimodais.
Tabela compara o desempenho do DeepSeek V4-Flash-Vision-Exp, DeepSeek V4-Flash-0731 e Claude Opus 4.8 em benchmarks de agentes baseados em texto e tarefas multimodais.

In Chartography, for example, the V4-Flash-Vision-Exp scored 64.3 points, versus 65 for Opus 4.8. In ApexBench, it was 36.5 points for the DeepSeek model and 39.4 for the Anthropic competitor.

In two other tests, the Chinese solution came out ahead. In Agents' Last Exam, it reached 27.3 points, versus 25.7 for Opus 4.8. In ZeroBench, it recorded 35 points, against 34 for the Anthropic model.

The results, however, should be interpreted with caution. The benchmarks were released by DeepSeek itself and do not represent an independent evaluation. Furthermore, performance varies by task: in some text-only tests, Opus 4.8 still maintains a significant advantage.

Up to 600 images in a single request

DeepSeek also tried to make image use relatively cheap within the API.

Each image consumes at most 384 tokens, regardless of its original resolution, and follows the same pricing schedule as the V4-Flash. Before processing, the system automatically adjusts the size of images to reduce resource consumption.

A single request can include up to 600 images, which opens room for applications that need to analyze large sets of screenshots, document pages, or visual sequences.

Developers can send files directly encoded in the request, use images available by URL, or resort to the new Files API from DeepSeek.

This last option allows you to send a file once and reuse it later via an identifier, without having to transfer the same image again with each request. DeepSeek says the file storage and reference service has no additional charge.

Experimental model expands the race for AI agents

The V4-Flash-Vision-Exp also works with interfaces compatible with Chat Completions and Responses from OpenAI, and Messages from Anthropic, facilitating its integration into tools originally built for other providers.

More than adding image recognition to DeepSeek's portfolio, the release shows the company's effort to advance in one of the most contested areas in artificial intelligence today: agents capable of observing environments, interpreting information, and performing tasks using different tools.

The model still carries the "Exp" designation, indicating its experimental nature. Even so, the early results presented by DeepSeek suggest the company intends to use the V4-Flash as a lower-cost alternative also for multimodal applications, widening competition with advanced models from companies such as Anthropic and OpenAI.

More from Radar