Hangzhou-based AI startup DeepSeek has released DeepSeek-V4.1-Flash, a compact but powerful AI model that surpasses its larger V4-Pro version in multiple benchmarks while offering significantly reduced operating costs, marking a notable advancement in AI model efficiency and deployment.

  • V4.1-Flash integrates image understanding and advanced cache compression.
  • Performance benchmarks show it surpasses DeepSeek V4-Pro and rivals US competitors.
  • Significant cost reduction with API rerouting and open-weight release under MIT license.

What happened

Chinese AI startup Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd. has launched DeepSeek-V4.1-Flash, the smallest and latest iteration in their new model architecture family. This model contains 552 billion parameters, almost doubling the previous 284 billion in V4-Flash. It uses a new causal encoder-decoder design, activating just 8 billion parameters during input processing and 16 billion during output generation, optimizing resource use.

The new model replaces both the experimental vision model and the earlier V4-Flash version released in August, carrying forward integrated image understanding capabilities. DeepSeek has rerouted API requests from the older V4-Pro model to the new V4.1-Flash starting September 14, billing users at the smaller model’s significantly lower rates until a V4.1-Pro version is released. The company also made the model weights available on Hugging Face under an MIT license.

Why it matters

V4.1-Flash shows a substantial leap in performance and efficiency. Its innovative approach to key-value cache storage using a four-bit floating-point format reduces memory footprint to about a quarter compared to the earlier model, and persistent SSD cache storage requirements fall to approximately one-eighth. This optimization improves speed and runtime costs, offering developers a roughly 70% reduction in output token API pricing compared to V4-Pro.

Benchmark comparisons reveal V4.1-Flash narrowly outscores notable US models Anthropic’s Claude Opus 5 and OpenAI’s GPT-5.6 Sol on Terminal-Bench 2.1 and performs strongly on DeepSWE v1.1 software engineering tests. Although US models maintain an edge on certain science reasoning tests, DeepSeek’s advancements signal growing global competition in AI model development and demonstrate China’s increasing technical capability in this space.

What to watch next

DeepSeek plans to collaborate with the open-source community to extend inference support and explore broader deployment opportunities. The upcoming release of a V4.1-Pro model is anticipated to potentially improve performance further. Monitoring how these developments affect DeepSeek’s market traction and its positioning against leading AI companies will be important for industry observers.

Additionally, Anthropic recently named DeepSeek as one of seven Chinese labs involved in distillation campaigns targeting its Claude AI model, highlighting ongoing competitive and geopolitical dynamics in AI research. DeepSeek’s significant funding backing and valuation indicate strong financial support to sustain aggressive innovation and expansion in AI capabilities.

Source assisted: This briefing began from a discovered source item from SiliconANGLE. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings