
Just after completing its first major external funding round, DeepSeek—the AI company renowned for its "extreme cost-performance"—has released another major technical breakthrough. On June 27, DeepSeek, in collaboration with Peking University, officially released and open-sourced the DSpark inference acceleration framework, along with the full-stack speculative decoding training toolkit DeepSpec.
This is not an iteration of model capabilities, but an engineering revolution in inference efficiency. In real production traffic on DeepSeek-V4, DSpark has increased single-user generation speed by 60% to 85%, achieving 661% throughput advantage in high-concurrency scenarios.
Why DSpark? — The "Autoregressive Curse" of LLM Inference
LLM text generation is fundamentally a "token-by-token guessing game": each new token requires a complete forward pass. Generating 100 tokens means the model must "re-digest" what it has written 99 times.
The industry has long had a classic solution—Speculative Decoding: use a lightweight "draft model" to quickly generate candidate tokens, then have the full model verify them in a single batch, accepting the correct guesses.
But existing approaches each have flaws:
Autoregressive draft models (e.g., Eagle3): Generate token-by-token—high quality but slow; only short candidate blocks are practical in deployment
Parallel draft models (e.g., DFlash): Generate entire blocks in one pass—fast, but each position cannot depend on previous tokens, causing "suffix decay": the longer the candidate, the lower the acceptance rate
More critically, in high-concurrency scenarios, fixed-length verification forces the target model to waste precious batch processing power on tokens that are highly likely to be rejected—causing overall throughput to decline rather than improve. This is the fundamental reason DeepSeek used to "spin" and even crash during peak evening hours.
DSpark's Two "Knives": Semi-Autoregressive Generation + Confidence-Scheduled Verification
DSpark's core innovation lies in two complementary mechanisms:
Knife 1: Semi-Autoregressive Generation Architecture
DSpark retains the high throughput of parallel draft models while adding a lightweight sequential module (Markov Head or RNN Head) that injects "what was the previous token" information during sampling, so subsequent tokens are no longer "blind guesses".
The results are impressive: DSpark with just 2 layers of Transformer outperforms DFlash with 5 layers across all test domains, effectively suppressing suffix decay. On Qwen3-4B as the target model, DSpark's average acceptance length exceeds Eagle3 by 30.9% and DFlash by 16.3%.
Knife 2: Confidence-Scheduled Verification
DSpark trains an additional confidence estimation head to predict each draft token's "prefix survival probability," using Sequential Temperature Scaling (STS) to calibrate prediction error from 3%-8% down to approximately 1%.
A hardware-aware prefix scheduler then dynamically determines how many tokens to verify per request based on current system load and each token's survival probability—aggressively verifying more under light load, and pruning low-confidence suffixes under high concurrency.
The paper formally guarantees: "the acceptance rule preserves the target distribution exactly, speculative decoding accelerates generation without any quality loss."
Production Results: Speed Up 85%, Throughput Surges 661%
DSpark has been deployed in the preview service engines of DeepSeek-V4-Flash and V4-Pro, and completely replaced the previous single-token baseline MTP-1 two weeks after the V4 preview release.
Production metrics:
| Engine | SLA Requirement | Throughput Gain | Key Takeaway |
|---|---|---|---|
| V4-Flash | 80 tok/s/user | +51% | Stable improvement under moderate load |
| V4-Flash | 120 tok/s/user | +661% | MTP-1 near collapse; DSpark maintains normal operation |
| V4-Pro | 35 tok/s/user | +52% | Stable improvement under low SLA |
| V4-Pro | 50 tok/s/user | +406% | Significant advantage under strict SLA |
At equivalent throughput levels, DSpark increases single-user generation speed by 57% to 85%. A response that previously took 10 seconds now arrives in 5-6 seconds.
More critically, DSpark prevents the system from "crashing" under high concurrency. When traffic spikes hit, the scheduler automatically shortens verification length to avoid consuming critical batch processing capacity—handling sudden traffic surges without scaling up.
DeepSpec Open Source: Lowering the Industry's Inference Cost Floor
Alongside DSpark, DeepSeek open-sourced DeepSpec, a full-stack codebase for training and evaluating speculative decoding draft models.
DeepSpec breaks the workflow into three stages: data preparation, training, and evaluation. It currently includes three draft model implementations (DSpark, DFlash, Eagle3) and supports Qwen3 and Gemma as target model families.
This means developers can use these tools to train dedicated inference acceleration draft models for their own models, without building infrastructure from scratch. As TMTpost commented: "This pulls the industry's inference cost benchmark down another notch."
Editor's Note
DSpark's true value isn't measured by benchmark scores—it's about solving a real-world engineering problem: keeping the system running when users flood in.
In 2024-2025, DeepSeek made its name on "extreme cost-performance." By 2026, with compute shortages escalating from "startup problems" to "battles between giants," inference efficiency competition has become a core battleground.
Liang Wenfeng's move to bet on inference optimization right after securing funding sends a clear signal: accelerate both model iteration and productization, while seizing the high ground in compute efficiency competition. In an era where compute is more precious than gold, efficiency itself is the deepest moat.
This article is based on the DeepSeek official paper and reports from Machine Heart, Zhidx, ifeng Tech, Huanqiu Tech, TMTpost, and other media outlets.