DeepSeek V4.1 Flash Reaches Sixth on OpenRouter Days After Launch
Table of Contents
DeepSeek V4.1 Flash has quickly gained traction following its September 10, 2026 release, reaching sixth place in OpenRouter’s model token-consumption ranking within its first three days, according to data reported by Gelonghui.
The model processed approximately 4.94 trillion tokens during that period, highlighting strong early demand on the OpenRouter platform. The rapid rise came alongside the launch of DeepSeek’s new multimodal Mixture-of-Experts architecture, which combines 552 billion backbone parameters with selective parameter activation.
Flash Reaches Sixth Place Within Days
The early adoption of V4.1 Flash is one of the most notable developments following its launch.
According to OpenRouter data reported by Gelonghui, V4.1 Flash reached sixth place in token consumption within three days of becoming available, recording around 4.94 trillion tokens.
The ranking measures model usage on OpenRouter rather than the entire global AI market. Therefore, “sixth globally” should be understood specifically as sixth in OpenRouter’s reported model token-consumption ranking.
The wider OpenRouter market also recorded approximately 127 trillion tokens during the week of September 7–13, representing growth of more than 10% from the previous week.
Chinese AI models accounted for approximately 61.17 trillion tokens, remaining ahead of U.S. models for the 20th consecutive week, according to the same data.
Xiaomi’s MiMo-V2.5 also recorded a sharp increase, with reported token consumption rising 230% week over week to reach fifth place.
A 552B Mixture-of-Experts Model
V4.1 Flash was released on September 10 as the smallest model in DeepSeek’s new architecture family.
It contains 552 billion total backbone parameters, but its Mixture-of-Experts design activates only a portion of them during inference. According to DeepSeek, the model activates 8 billion parameters when processing input and 16 billion when generating output.
The previous V4 Flash contained 284 billion parameters and activated approximately 13 billion parameters per token.
| Feature | DeepSeek V4.1 Flash |
| Release date | September 10, 2026 |
| Total parameters | 552 billion |
| Active input parameters | 8 billion |
| Active output parameters | 16 billion |
| Architecture | Causal Encoder-Decoder |
| Context window | Up to 1 million tokens |
| KV cache | About 890 bytes/token |
| Visual understanding | Native |
| Licence | MIT |
This architecture allows DeepSeek to increase overall model capacity without activating the entire parameter set for every token.
Lower Memory Requirements
DeepSeek has also redesigned the model’s KV cache, which stores information from previous tokens during inference.
The company reports a cache requirement of approximately 890 bytes per token, or roughly one-quarter of the memory footprint of V4 Flash’s corresponding cache.
The change could be important for long-context and high-volume AI workloads. However, V4.1 Flash should not be considered a lightweight model simply because only a portion of its parameters are activated during individual operations.
DeepSeek has not published comprehensive VRAM requirements, inference-speed benchmarks or batch-size guidance, leaving developers to evaluate deployment requirements for their specific environments.
V4 Flash and V4 Flash Vision Transition
V4.1 Flash also replaces DeepSeek’s earlier V4 Flash and V4 Flash Vision offerings.
DeepSeek’s API documentation says requests using the legacy V4 Flash and V4 Flash Vision model names are automatically served by V4.1 Flash and billed at the Flash rate.
The model weights are publicly available under an MIT licence, allowing developers to access and experiment with the model outside DeepSeek’s hosted API.
V4 Pro Remains Available
DeepSeek initially announced that V4 Pro API requests would be redirected to V4.1 Flash from September 14 and billed at the lower Flash rate until V4.1 Pro was released.
However, the company later reversed that plan following user demand. DeepSeek’s current API documentation says V4 Pro will continue to be available, with its billing method unchanged.
This means V4.1 Flash should not be described as a complete replacement for V4 Pro.
Multimodal and Benchmark Performance
V4.1 Flash offers native visual understanding and supports context windows of up to one million tokens.
DeepSeek reports strong results across coding, agentic and multimodal benchmarks:
| Benchmark | DeepSeek-reported score |
| DeepSWE v1.1 | 74.2 |
| Terminal-Bench 2.1 | 90.6 |
| CyberGym | 88.1 |
| DocVQA | 95.6 |
| MMMU-Pro | 56.5 |
| LongBench-V2 | 45.2 |
These figures come from DeepSeek’s own evaluations and should not be treated as independent verification.
V4.1 Flash also trails V4 Pro on some evaluations. For example, DeepSeek reports 45.2 on LongBench-V2 for V4.1 Flash compared with 51.5 for V4 Pro.
Alibaba Cloud Adds Hosted Access
The model’s early adoption also coincided with wider availability through third-party platforms.
According to the reported industry data, Alibaba Cloud launched hosted DeepSeek V4.1 Flash access on September 13, offering API access through its cloud platform and a Token Plan connected with tools including Qoder and Codex.
This expands the ways developers can access V4.1 Flash beyond DeepSeek’s own services.
Lower API Pricing
V4.1 Flash also comes with relatively low API pricing, particularly during off-peak periods.
| API usage | Off-peak price per million tokens |
| Uncached input | $0.15 |
| Cached input | $0.003 |
| Output | $0.60 |
Actual costs vary according to token usage, caching and peak or off-peak pricing.
Why the Sixth-Place Ranking Matters
V4.1 Flash’s rapid move into sixth place on OpenRouter within days of launch provides an early indication of developer interest.
The 4.94 trillion-token figure is particularly notable because it was accumulated during only the model’s first three days on the platform. However, token consumption alone does not establish that the model is technically superior to every model above or below it.
Usage can be influenced by pricing, availability, routing, developer experimentation and the types of applications using a particular model.
Conclusion
DeepSeek V4.1 Flash reached sixth place in OpenRouter’s model token-consumption ranking within three days of its launch, processing approximately 4.94 trillion tokens according to data reported by Gelonghui.
Its rapid adoption comes alongside a new 552-billion-parameter MoE architecture, selective 8B/16B activation, an approximately 890-byte-per-token KV cache, native visual understanding and a one-million-token context window.
The model also offers low API pricing and has been made available through additional platforms such as Alibaba Cloud.
While the early OpenRouter numbers suggest strong developer interest, they represent activity on one platform rather than the entire AI market. Similarly, DeepSeek’s benchmark results require independent evaluation before broader conclusions can be drawn.
For now, V4.1 Flash’s combination of rapid early adoption, large-scale architecture and lower inference pricing makes it one of the more notable AI model releases of September 2026.