DeepSeek V4.1 Flash Launches With Lower API Prices and New Architecture
Table of Contents
Chinese AI company DeepSeek has launched DeepSeek V4.1 Flash, describing it as the smallest model in its new architecture family. The company says the model is designed to deliver faster inference, higher throughput and lower costs while adding native visual understanding.
The launch comes with lower API prices and a planned transition away from the existing V4 Pro API tier.
A 552B-Parameter Model With an Asymmetric Architecture
According to DeepSeek, V4.1 Flash is a 552-billion-parameter mixture-of-experts model built around a new Causal Encoder-Decoder architecture.
The company says the model activates only 8 billion parameters for input processing and 16 billion for output generation. DeepSeek says this asymmetric design is intended to deliver greater efficiency than models of a similar overall size.
DeepSeek also says V4.1 Flash uses one-quarter of the previous generation’s HBM requirement for its KV cache and one-eighth of the SSD storage requirement. These figures are claims from DeepSeek’s technical announcement rather than independently verified measurements.
Native Visual Understanding
One of the main additions in DeepSeek V4.1 Flash is native multimodal visual understanding.
DeepSeek says the new model can process visual information directly through its API. This replaces the previous arrangement in which the company offered a separate experimental V4 Flash Vision model.
The model is being positioned for agentic workloads, coding and applications where inference speed and operating costs are important.
DeepSeek Reports Strong Benchmark Results
DeepSeek says V4.1 Flash performed ahead of several flagship models in its benchmark testing, including its own V4 Pro.
The company’s announcement says the model’s new pretraining and reinforcement-learning methods contributed to improved benchmark performance. DeepSeek also says multiple tests placed V4.1 Flash ahead of V4 Pro in performance, cost, speed and total runtime.
However, these should be treated as company-reported results. Independent benchmark validation for the newly released model was not yet available when it launched.
That distinction is important when comparing V4.1 Flash with competing AI models because benchmark outcomes can depend on the datasets, prompts and evaluation methods used.
Lower API Prices
DeepSeek is also reducing API prices with the new model.
The company says the new pricing took effect on September 10 and continues its peak/off-peak structure, with off-peak rates set at half the peak rates.
| API Category | V4.1 Flash Off-Peak Price |
| Uncached input | $0.15 per million tokens |
| Output | $0.60 per million tokens |
| Cached input | $0.003 per million tokens |
The lower pricing is designed to make the model cheaper to operate at scale, particularly for applications that make frequent API calls.
V4 Pro Requests Will Move to Flash
DeepSeek is also planning to phase out its V4 Pro API model.
Starting at 04:00 UTC on September 14, 2026, requests made to deepseek-v4-pro are scheduled to be automatically routed to V4.1 Flash and billed at V4.1 Flash rates. DeepSeek says this arrangement will remain in place until V4.1 Pro launches.
This means V4 Pro has not simply been discontinued immediately. Instead, DeepSeek has announced a transition in which existing V4 Pro API requests will be redirected to the newer Flash model.
DeepSeek’s IPO Plans Add to the Momentum
The DeepSeek V4.1 Flash launch comes as the company prepares for a potential IPO in China.
Reuters reported that DeepSeek has engaged CITIC Securities to prepare for an initial public offering on Shanghai’s STAR Market, with the company aiming to begin the process this year. The final timing, IPO size and valuation have not been determined.
Reuters also reported that DeepSeek’s latest funding round could value the company at around 500 billion yuan, or approximately $75 billion. That figure represents a potential valuation in the financing round rather than a finalized IPO valuation.
What Makes V4.1 Flash Different?
| Feature | DeepSeek V4.1 Flash |
| Total parameters | 552B |
| Active input parameters | 8B |
| Active output parameters | 16B |
| Architecture | Causal Encoder-Decoder + MoE |
| Visual understanding | Native |
| Main focus | Speed, efficiency and agentic workloads |
| Off-peak input | $0.15 per million tokens |
| Off-peak output | $0.60 per million tokens |
| V4 Pro transition | Planned routing from Sept. 14 |
What the Launch Means for AI Developers
The release reflects a broader trend in AI development: using large mixture-of-experts models while activating only a portion of their parameters for individual tasks.
DeepSeek’s approach is designed to reduce the resources required to run the model while maintaining high performance. If those efficiency and performance claims are supported by independent testing, the architecture could make advanced AI models more economical to deploy.
The lower API prices could also increase competitive pressure on other AI providers, especially in coding, agentic applications and other workloads that consume large numbers of tokens.
Conclusion
DeepSeek V4.1 Flash introduces a new architecture, native visual understanding and lower API pricing while maintaining a large 552-billion-parameter model with a smaller number of active parameters.
DeepSeek is also preparing to move V4 Pro API requests to the new Flash model from September 14, positioning V4.1 Flash as the company’s more efficient option for developers.
The model’s benchmark results are promising, but independent testing will be important before drawing firm conclusions about how it compares with leading AI systems. For developers, the combination of multimodal capabilities, lower pricing and a focus on inference efficiency makes V4.1 Flash a notable new release in the competitive AI model market.