DeepSeek unveils V4.1-Flash, its smallest AI model with native visual understanding

DeepSeek unveils V4.1-Flash, its smallest AI model with native visual understanding

Chinese AI lab DeepSeek announced the launch of DeepSeek-V4.1-Flash on Wednesday, describing the release as “smarter, faster, more efficient”. According to the company, the model is the smallest member of its newly introduced architecture family and the first in that family to ship with native visual understanding, multimodal capabilities built directly into the model rather than bolted on. DeepSeek said the architecture was designed from the ground up for greater capability, faster inference and higher throughput, and to serve as a foundation that scales up to larger sibling models.

At the heart of the release is what DeepSeek calls an asymmetric architecture: a 552-billion-parameter Mixture-of-Experts model built on a novel Causal Encoder–Decoder design, with just 8 billion active parameters dedicated to processing input and 16 billion for generating output. The company says new pre-training methods combined with larger-scale reinforcement learning post-training push benchmark results ahead of the previous generation. Efficiency gains extend beyond compute as well, V4.1-Flash's key-value cache requires only one quarter of the HBM and one eighth of the SSD storage compared with its predecessor, a compression DeepSeek argues translates into major savings, since cache-hit charges often make up a large share of agentic AI workloads.

V4.1-Flash is live now on the DeepSeek API under the model name “deepseek-flash”, with native multimodal support; the older V4-Flash and V4-Flash-Vision-Exp models have been retired, though legacy endpoints temporarily route to the new release for compatibility. Reflecting the reduced serving costs, DeepSeek is cutting API prices and continuing its peak/off-peak pricing scheme, in which off-peak rates run at 50 percent of peak rates, an incentive for developers to schedule flexible workloads during quieter hours. In keeping with its open-source tradition, the company has published the model on Hugging Face and says it will work closely with the open-source community on inference support, inviting operators planning large-scale deployments of 2,000 GPUs or more to get in touch.

by Paul

Add as a preferred source on Google
DeepSeek iconDeepSeek
  207
  • ...

DeepSeek is an AI chatbot utilizing advanced natural language processing and machine learning to deliver intelligent, conversational support for a wide range of tasks. It excels in answering questions, generating creative content, and solving complex problems. Key features include Dark Mode and an ad-free experience. Rated 3.1, DeepSeek offers alternatives for users seeking different solutions.

No comments so far, maybe you want to be first?
Gu