July 16, On the eve of the World Artificial Intelligence Conference (WAIC), Moonshot AI dropped a bombshell — the official release of its next-generation foundation model Kimi K3, with a total parameter count of 2.8 trillion, making it the largest open-source large language model in the world.

A New High in the Parameter Race
2.8 trillion. The number itself is a declaration. Previously, DeepSeek V4 had 1.6 trillion parameters, and Wenxin 5.0 had 2.4 trillion. Kimi K3 pushes the parameter ceiling for open-source models to nearly 3 trillion in one leap.
Why does parameter count matter so much? A Moonshot AI representative used a vivid analogy: "Parameters are like neural connections in the human brain. Nearly 3 trillion parameters mean this model can pack more knowledge and patterns into its 'brain'—understanding more, thinking deeper, and answering more accurately."
But size isn't everything. Moonshot AI emphasizes that K3's development focus isn't simply about expanding parameter count, but about translating scale advantages into capability improvements through architectural innovation. This is what sets K3 apart from the "parameter-stacking" approach.
Architectural Innovation: Smoother Information Flow
Kimi K3's technical foundation rests on two proprietary innovations: the KDA (Kimi Delta Attention) hybrid linear attention mechanism and Attention Residuals technology. These innovations aim to solve the problem of information loss in ultra-long sequences, allowing information to flow more smoothly through longer sequences and deeper models.
The 1 million-token context window and native visual understanding capabilities enable K3 to handle complex tasks requiring a "panoramic perspective"—whether analyzing massive codebases or conducting deep research on lengthy documents.
On efficiency, through coordinated optimization of model structure, training methods, and data recipes, Kimi K3 achieves a approximately 2.5x improvement in overall scaling efficiency compared to the previous generation.
Benchmark Domination: Frontend Programming Surpasses Claude Fable 5
Parameter count is the "hardware"; benchmarks are the "real-world performance."
On Arena.AI's Frontend Code Arena, Kimi K3 topped the leaderboard with 1679 points, surpassing Anthropic's Claude Fable 5. Across 7 sub-categories in the frontend domain, K3 ranked first in 6.
In broader benchmark tests, K3 performed equally strongly:
| Benchmark | Score | Ranking |
|---|---|---|
| GDPval-AA v2 (44 occupations/9 industries) | 1687 | 3rd (behind Claude Fable 5 Max and GPT-5.6 Sol Max) |
| AA-Briefcase (private agent benchmark) | 1527 | 2nd (ahead of GPT-5.6 Sol Max) |
| BrowseComp (long-context retrieval) | 91.2 | Highest score |
Across the full evaluation suite, Kimi K3's overall intelligence level ranks just behind Claude Fable 5 and GPT-5.6 Sol, but consistently surpasses other models like GPT-5.5 and Opus-4.8.
From Benchmarks to "Real Work": 48-Hour Autonomous Chip Design
Even more impressive than the benchmarks is K3's performance on real-world tasks.
Moonshot AI demonstrated a proof-of-concept: Kimi K3, over 48 hours of continuous autonomous operation, independently used open-source electronic design automation tools to complete the full design flow for a 4mm² chip—from architecture design to optimization and verification—achieving 100MHz timing closure with simulation decoding speeds exceeding 8,700 tokens per second.
"What does this mean? Large models can not only write code but also build their own hardware," commented one developer.
In software engineering scenarios, K3 can combine source code, rendered results, execution logs, and test feedback to determine the next modification direction—applicable to frontend engineering, game development, infrastructure optimization, and research programming.
In knowledge work, K3 can process massive amounts of material to generate consulting-grade industry research reports and dynamic graphic videos. For digital content creation, the model can combine 3D reasoning, programming, and visual capabilities to transform concepts, images, or videos into interactive experiences.
Pricing Strategy: Not a "Price War," But Still Competitive
Unlike some domestic models that continue the "ultra-low-price" approach, Kimi K3 has adopted a differentiated pricing strategy.
According to Kimi's Open Platform, K3 uses pay-per-use pricing: input at RMB 2 per million tokens (cache hit) / RMB 20 (cache miss), output at RMB 100 per million tokens. The API is compatible with OpenAI's SDK, with international pricing at $3 per million input tokens and $15 per million output tokens.
Axios notes that K3's usage cost of approximately $12 per million tokens does not follow the "ultra-low-price" approach of some Chinese models, positioning it closer to Anthropic's mid-range offerings, but still below the pricing of several U.S. premium models it is challenging.
Notably, thanks to Mooncake's disaggregated inference architecture, Kimi's official API achieves a cache rate exceeding 90% in coding scenarios, effectively reducing actual input costs to one-quarter of the standard price.
Open Source and "Going Global": Silicon Valley's Alert
The full model weights for Kimi K3 will be publicly released by July 27, 2026. This means global developers will soon be able to freely download and deploy this 2.8 trillion-parameter open-source model.
The news has sparked intense interest in Silicon Valley. Elon Musk responded to a review of Kimi K3 on social media with a single word: "Impressive."
Axios commented that Kimi K3's strong performance in multiple early tests shows some capabilities now rival top U.S. models, and "America's lead in advanced AI is shrinking month by month."
Mozilla CTO Rafik Kerikorian put it bluntly: "Now, this is a U.S.-China issue. American AI companies are clearly worried."
The release of Kimi K3 sends three notable signals:
First, the "open-source parameter race" has entered a new magnitude. A 2.8 trillion-parameter open-source model further narrows the capability gap between open and closed-source models. When an open-source model already tops the frontend programming charts, how deep is the closed-source "moat" really?
Second, the real battleground is shifting from benchmarks to "real work." A 48-hour autonomous chip design is far more compelling than any benchmark number. The AI model competition is moving from "who answers questions better" to "who can complete complex real-world tasks more effectively."
Third, China AI's "cost-performance approach" is triggering global ripple effects. Although K3 hasn't adopted a "rock-bottom" pricing strategy, its prices remain below comparable U.S. models. When Chinese AI simultaneously offers "approaching capabilities" and "cost advantages," Silicon Valley's vigilance is not unfounded.
WAIC has just begun, and Kimi K3 is just one appetizer at this AI feast.