DeepSeek V4.1 Flash Officially Released: 552B Parameters Punching Above Its Weight, V4 Pro Fully Retired

9.10 On September 10, DeepSeek officially released the V4.1 Flash model, the smallest in its new architecture series, yet one that comprehensively surpasses its previous flagship V4 Pro. Alongside the launch, DeepSeek announced an orderly retirement of V4 Pro: after 12:00 on September 14, all V4 Pro requests will be routed to V4.1 Flash and billed at the new rates.

On September 10, DeepSeek officially released the V4.1 Flash model, the smallest in its new architecture series, yet one that comprehensively surpasses its previous flagship V4 Pro. Alongside the launch, DeepSeek announced an orderly retirement of V4 Pro: after 12:00 on September 14, all V4 Pro requests will be routed to V4.1 Flash and billed at the new rates.

Asymmetric Architecture: 8B Input, 16B Output

V4.1 Flash has 552B total parameters in a MoE architecture, but its true innovation lies in the new Causal-Encoder-Decoder structure.

The architecture’s defining feature is asymmetric input-output activation: processing input activates only 8B parameters, while generating output activates 16B. This design significantly reduces input-side computation for tasks like code repository analysis and long-document processing, where large amounts of information are read to produce relatively limited judgments.

Another efficiency gain comes from major KV Cache compression. Compared to the previous generation, V4.1 Flash reduces HBM requirements to 1/4 and SSD requirements to 1/8; against the original architecture, cumulative compression is roughly 437x. Since cache-hit costs dominate many agent scenarios, this directly lowers long-context task costs.

agentic-benchmark

Performance Surpasses V4 Pro Across Key Benchmarks

Official comparison data shows V4.1 Flash surpassing V4 Pro and several competitors on core benchmarks:

BenchmarkV4.1 FlashV4 ProGLM 5.3Kimi K3GPT 5.6-Sol
GPQA Diamond90.992.488.192.994.1
Codeforces34713348———
Terminal-Bench 2.190.687.988.288.388.8
DeepSWE v1.174.262.766.967.573.0
CyberGym88.183.384.580.084.5
Automation-Bench54.843.248.846.745.8

In competitive programming (Codeforces), V4.1 Flash achieved a rating of 3471, exceeding V4 Pro’s 3348. It also leads in terminal task execution and long-horizon coding, and tops the cybersecurity benchmark CyberGym at 88.1.

Native Multimodal: Vision Integrated into the Main Model

V4.1 Flash is DeepSeek’s first mainline model with native multimodal visual understanding, no longer relying on a bolt-on vision module. Previously requiring separate calls to V4 Flash (text) and V4 Flash Vision Exp (images), the two paths are now unified into one model, reducing maintenance for developers.

Price Reduction: Cache-Hit at ¥0.02 per Million Tokens

API pricing has been lowered, still with peak/off-peak tiers:

PeriodInput (Cache Hit)Input (Cache Miss)Output
Off-Peak¥0.02¥1.0¥4.0
Peak¥0.04¥2.0¥8.0

Peak hours are weekdays 9:00–12:00 and 14:00–18:00; all other times (including weekends) are off-peak. New pricing took effect at 12:00 on September 10.

API Changes and Ecosystem Integration

On the API side, the model name deepseek-flash now calls V4.1 Flash. The older V4 Flash and V4 Flash Vision Exp have been retired, temporarily routed to the new model for compatibility. After 12:00 on September 14, deepseek-v4-pro requests will also be routed to V4.1 Flash and billed at new rates.

Tencent’s WorkBuddy and CodeBuddy, along with OpenCode, have fully integrated as official partners.

V4.1 Flash marks a “flagship handover” at DeepSeek—replacing the larger previous-generation Pro with the smallest Flash model. The asymmetric architecture and KV Cache compression turn “small cost, big tasks” from slogan into verifiable engineering practice. The deeper implication: when efficiency gains drive costs down, AI adoption across office, coding, customer service, and other scenarios is likely to accelerate. For developers, getting stronger capabilities at lower prices is the most direct benefit.

https://www.deepseek.com/news/deepseek-v4-1-flash/ 

DeepSeek V4.1 Flash Officially Released: 552B Parameters Punching Above Its Weight, V4 Pro Fully Retired | AIYXL