Skip to content
Break Read Break Read Break Read
Break Read Break Read Break Read
  • Blog
  • Contact
  • Blog
  • Contact
Close

Search

Home/AI/DeepSeek V4.1 Flash Launches With Lower API Prices and New Architecture
DeepSeek V4.1 Flash
AI

DeepSeek V4.1 Flash Launches With Lower API Prices and New Architecture

September 12, 2026 4 Min Read

Table of Contents

A 552B-Parameter Model With an Asymmetric Architecture
Native Visual Understanding
DeepSeek Reports Strong Benchmark Results
Lower API Prices
V4 Pro Requests Will Move to Flash
DeepSeek’s IPO Plans Add to the Momentum
What Makes V4.1 Flash Different?
What the Launch Means for AI Developers
Conclusion

Chinese AI company DeepSeek has launched DeepSeek V4.1 Flash, describing it as the smallest model in its new architecture family. The company says the model is designed to deliver faster inference, higher throughput and lower costs while adding native visual understanding.

The launch comes with lower API prices and a planned transition away from the existing V4 Pro API tier.

A 552B-Parameter Model With an Asymmetric Architecture

According to DeepSeek, V4.1 Flash is a 552-billion-parameter mixture-of-experts model built around a new Causal Encoder-Decoder architecture.

The company says the model activates only 8 billion parameters for input processing and 16 billion for output generation. DeepSeek says this asymmetric design is intended to deliver greater efficiency than models of a similar overall size.

DeepSeek also says V4.1 Flash uses one-quarter of the previous generation’s HBM requirement for its KV cache and one-eighth of the SSD storage requirement. These figures are claims from DeepSeek’s technical announcement rather than independently verified measurements.

Native Visual Understanding

One of the main additions in DeepSeek V4.1 Flash is native multimodal visual understanding.

DeepSeek says the new model can process visual information directly through its API. This replaces the previous arrangement in which the company offered a separate experimental V4 Flash Vision model.

The model is being positioned for agentic workloads, coding and applications where inference speed and operating costs are important.

DeepSeek Reports Strong Benchmark Results

DeepSeek says V4.1 Flash performed ahead of several flagship models in its benchmark testing, including its own V4 Pro.

The company’s announcement says the model’s new pretraining and reinforcement-learning methods contributed to improved benchmark performance. DeepSeek also says multiple tests placed V4.1 Flash ahead of V4 Pro in performance, cost, speed and total runtime.

However, these should be treated as company-reported results. Independent benchmark validation for the newly released model was not yet available when it launched.

That distinction is important when comparing V4.1 Flash with competing AI models because benchmark outcomes can depend on the datasets, prompts and evaluation methods used.

Lower API Prices

DeepSeek is also reducing API prices with the new model.

The company says the new pricing took effect on September 10 and continues its peak/off-peak structure, with off-peak rates set at half the peak rates.

API CategoryV4.1 Flash Off-Peak Price
Uncached input$0.15 per million tokens
Output$0.60 per million tokens
Cached input$0.003 per million tokens

The lower pricing is designed to make the model cheaper to operate at scale, particularly for applications that make frequent API calls.

V4 Pro Requests Will Move to Flash

DeepSeek is also planning to phase out its V4 Pro API model.

Starting at 04:00 UTC on September 14, 2026, requests made to deepseek-v4-pro are scheduled to be automatically routed to V4.1 Flash and billed at V4.1 Flash rates. DeepSeek says this arrangement will remain in place until V4.1 Pro launches.

This means V4 Pro has not simply been discontinued immediately. Instead, DeepSeek has announced a transition in which existing V4 Pro API requests will be redirected to the newer Flash model.

DeepSeek’s IPO Plans Add to the Momentum

The DeepSeek V4.1 Flash launch comes as the company prepares for a potential IPO in China.

Reuters reported that DeepSeek has engaged CITIC Securities to prepare for an initial public offering on Shanghai’s STAR Market, with the company aiming to begin the process this year. The final timing, IPO size and valuation have not been determined.

Reuters also reported that DeepSeek’s latest funding round could value the company at around 500 billion yuan, or approximately $75 billion. That figure represents a potential valuation in the financing round rather than a finalized IPO valuation.

What Makes V4.1 Flash Different?

FeatureDeepSeek V4.1 Flash
Total parameters552B
Active input parameters8B
Active output parameters16B
ArchitectureCausal Encoder-Decoder + MoE
Visual understandingNative
Main focusSpeed, efficiency and agentic workloads
Off-peak input$0.15 per million tokens
Off-peak output$0.60 per million tokens
V4 Pro transitionPlanned routing from Sept. 14

What the Launch Means for AI Developers

The release reflects a broader trend in AI development: using large mixture-of-experts models while activating only a portion of their parameters for individual tasks.

DeepSeek’s approach is designed to reduce the resources required to run the model while maintaining high performance. If those efficiency and performance claims are supported by independent testing, the architecture could make advanced AI models more economical to deploy.

The lower API prices could also increase competitive pressure on other AI providers, especially in coding, agentic applications and other workloads that consume large numbers of tokens.

Conclusion

DeepSeek V4.1 Flash introduces a new architecture, native visual understanding and lower API pricing while maintaining a large 552-billion-parameter model with a smaller number of active parameters.

DeepSeek is also preparing to move V4 Pro API requests to the new Flash model from September 14, positioning V4.1 Flash as the company’s more efficient option for developers.

The model’s benchmark results are promising, but independent testing will be important before drawing firm conclusions about how it compares with leading AI systems. For developers, the combination of multimodal capabilities, lower pricing and a focus on inference efficiency makes V4.1 Flash a notable new release in the competitive AI model market.

Share this article on
  • Facebook
  • Pinterest
  • Twitter
  • Linkedin
  • Whatsapp
Author

Lalith Raj

Follow Me
Other Articles
iPhone price hike
Previous

Apple Raises iPhone Prices in India by Up to 41%

GPT-Live-1
Next

OpenAI Launches GPT-Live-1 API With $0.05-Per-Minute Voice Layer

Search...

Recent Posts

  • Asian Games 2026
    Asian Games 2026 Begin in Aichi-Nagoya With 43 Sports
    by Lalith Raj
    September 16, 2026
  • Snapchat just brought AI powered conversational ads to its app. 2
    Snapchat Launches Sponsored Interactive AI Ads Inside Chat
    by Nithin
    March 1, 2026
  • Lovable just launched its vibe coding app on iOS and Android
    Lovable Mobile App Launches Vibe Coding Experience on iOS and Android
    by Nithin
    March 4, 2026
  • Apple just introduced a cheaper option for App Store subscriptions
    Apple introduces a new subscription model: Monthly Plans with 12-Month Commitment
    by Nithin
    March 8, 2026

Categories

  • AI
  • Business
  • Cars
  • Entertainment
  • Finance
  • Music
  • News
  • Science
  • SEO
  • Sports
  • Technology
  • Trending

Break Read

Stay ahead in the fast-moving world of technology with expert articles, industry updates, and practical insights.

Latest Posts

  • Asian Games 2026 Begin in Aichi-Nagoya With 43 SportsSeptember 16, 2026
  • Google Launches Gemini 3.8 Live With Real-Time Voice and Background ReasoningSeptember 16, 2026
  • Chandrayaan-1 Data Reveal Evidence of Ancient Australe Basin on MoonSeptember 16, 2026

Pages

  • Contact
  • Terms and Conditions
  • Privacy Policy
  • Refund Policy
Copyright 2026 — Break Read. All rights reserved.
Go to mobile version