DeepSeek Releases V4.1 Flash With Native Visual Understanding

Sep 15, 2026
DeepSeek has released its V4.1 Flash model with native visual understanding, a new architecture, lower cache requirements and reduced API prices.

DeepSeek released DeepSeek-V4.1-Flash on 10 September, the smallest model in a new architecture family, with native visual understanding and a new Causal Encoder Decoder architecture. The 552 billion parameter mixture of experts model activates 8 billion parameters for input and 16 billion for output.

DeepSeek says new pre training methods and larger scale reinforcement learning put its benchmark results ahead of flagship models including DeepSeek-V4-Pro, and that tests by multiple parties place it ahead of V4-Pro on performance, cost, speed and total runtime. Its key value cache requires one quarter of the high bandwidth memory and one eighth of the SSD storage used by the previous generation.

DeepSeek-V4.1-Flash is available through the DeepSeek API under the deepseek-flash identifier. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, while DeepSeek-V4-Pro requests began routing to the new model on 14 September at V4.1 Flash rates, an arrangement DeepSeek says will last until V4.1-Pro launches. DeepSeek also cut API prices from 10 September, with off peak rates at half of peak rates. The model and a technical report are available on Hugging Face.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.