DeepSeek V4.1-Flash is smaller, faster and already outperforms the Pro model
DeepSeek continues its aggressive development pace. They have introduced the DeepSeek-V4.1-Flash model, the smallest in their new architectural family, which not only focuses on lower price, but also on higher speed and better performance.
DeepSeek says the new architecture is designed for faster model execution, higher throughput, and easier scaling to much larger systems. V4.1-Flash is a first look at the technology that will likely underpin the company's future larger models.
Even more interesting is the comparison with the existing DeepSeek-V4 Pro. According to their data, V4.1-Flash has surpassed the Pro version in terms of performance, speed, cost and total execution time. Until the arrival of V4.1 Pro, requests intended for the current Pro model will therefore be redirected to V4.1-Flash and billed at the lower Flash version tariff.
The model also brings native support for multimodal content, meaning the new architecture is no longer focused solely on text processing. This is an important step for DeepSeek, which aims to compete more directly with systems from OpenAI, Google, and Anthropic on more demanding agent-based and visual tasks.
V4.1-Flash also comes with new API pricing. During off-peak hours, a million input tokens without caching will cost approximately 1 yuan (€0.13), while a million output tokens will cost 4 yuan (€0.42). During peak hours, prices will double.
DeepSeek is also preparing for an initial public offering of shares on Shanghai's STAR Market.



















