SuperCLUE 2025 Annual Report Released: Domestic Large Models Accelerate Catch-Up, Open-Source Ecosystem Leads Globally

2.24 Recently, SuperCLUE released the 2025 Annual Chinese Large Model Benchmark Evaluation Report, comprehensively reviewing the development landscape and core breakthroughs in the global large model industry over the past year. The report indicates that 2025 is a pivotal year for large models transitioning from "technological explosion" to "agentic deployment." Domestic large models have demonstrated strong momentum in catching up across reasoning capabilities, code generation, and open-source

ScreenShot_2026-02-24_163620_104


Recently, SuperCLUE released the 2025 Annual Chinese Large Model Benchmark Evaluation Report, comprehensively reviewing the development landscape and core breakthroughs in the global large model industry over the past year. The report indicates that 2025 is a pivotal year for large models transitioning from "technological explosion" to "agentic deployment." Domestic large models have demonstrated strong momentum in catching up across reasoning capabilities, code generation, and open-source ecosystems, with some tasks already achieving global leadership.

I. Three Key Advancements in Large Models in 2025: Rise of Agents, Reasoning Breakthroughs, and Open-Source Dominance

The report divides the development of large models in 2025 into three key stages

  • The Era of Hundreds of Models and Multimodal Emergence (2023–2024): ChatGPT ignited global attention, with domestic players such as Baidu, Alibaba, and iFLYTEK rapidly following suit. Open-source models like Baichuan and ChatGLM2 drove technological democratization.

  • Multimodal Explosion and Reasoning Breakthroughs (2024–2025): OpenAI released Sora and GPT-4o, advancing video generation and real-time interaction. Domestic video models such as Kling AI and Vidu achieved breakthroughs overseas. Reasoning models like DeepSeek-R1 and k0-math emerged en masse.

  • Rise of Agents and Ecosystem Restructuring (2025–present): DeepSeek-R1 took the world by storm with its exceptional cost-performance ratio. Chinese open-source models now account for half of the global market. Agent products such as Manus and AutoGLM have been deployed, with programming agents emerging as a standout.


ScreenShot_2026-02-24_163707_696

 

II. Global Leaderboard: Overseas Still Lead, Domestic Models Accelerate to Keep Pace

In the 2025 annual SuperCLUE general benchmark evaluation, Claude-Opus-4.5-Reasoning ranked first globally with a score of 68.25, followed by Gemini-3-Pro-Preview (65.59) and GPT-5.2(high) (64.32)

Domestic models delivered impressive performances: Kimi-K2.5-Thinking from Moonshot AI (61.50) ranked fourth globally, while Alibaba's Qwen3-Max-Thinking (60.61) ranked sixth, both breaking into the global top ten. The report notes that domestic large models are accelerating their evolution from "following" to "running alongside," and are already globally competitive in tasks such as code generation and mathematical reasoning

 

ScreenShot_2026-02-24_165300_723 

 

III. Open-Source Ecosystem: Domestic Models Lead, Surpassing Overseas Counterparts

In the open-source model track, domestic models are overwhelmingly dominant. Kimi-K2.5-Thinking topped the open-source leaderboard with a score of 61.50, followed by DeepSeek-V3.2-Thinking and GLM-4.7 in second and third place respectively, significantly outperforming the best overseas open-source model, gpt-oss-120b
. The report states: "Chinese open-source models have come to dominate the global open-source ecosystem"

 

ScreenShot_2026-02-24_165418_474 

 

IV. Deep Dive into Six Tasks: Domestic Models Top Code Generation, Instruction Following Remains a Weakness

This evaluation covered six tasks: mathematical reasoning, scientific reasoning, code generation, agent (task planning), precise instruction following, and hallucination control. Key findings include:

  • Code Generation: Kimi-K2.5-Thinking ranked first globally with 53.33 points, surpassing top overseas models such as Grok-4 and Claude, with particularly outstanding performance in WebCoding subtasks.

  • Mathematical Reasoning: Qwen3-Max-Thinking tied with Gemini-3-Pro-Preview for first place globally (80.87 points).

  • Agent (Task Planning): Qwen3-Max-Thinking ranked third globally (70.13), with Kimi-K2.5-Thinking close behind.

  • Precise Instruction Following: Domestic models trail overall, with a gap of over 13 points from overseas leaders—a clear weakness for domestic models.

  • Hallucination Control: GLM-4.7 broke into the global top three, narrowing the gap with the overseas first tier to within 5 points.

 

ScreenShot_2026-02-24_165848_538 

 

V. Closed-Source vs. Open-Source Comparison: Closed-Source Leads Overall, Open-Source Achieves Niche Breakthroughs

In the comparison between closed-source and open-source models, closed-source models maintain a lead across all six tasks, particularly in agent, instruction following, and hallucination control. However, open-source models achieved a niche breakthrough in code generation, with Kimi-K2.5-Thinking claiming the top global spot, demonstrating the strong potential of the open-source camp in vertical domains.


VI. Cost-Performance and Efficiency: Domestic Models Offer Significant Cost-Performance Advantages, with Room for Improvement in Inference Efficiency

In terms of cost-performance, domestic models achieve near-top-tier international performance at lower prices. Kimi-K2.5-Thinking, Qwen3-Max-Thinking, and Doubao-Seed-1.8 all fall within the high cost-performance range, while overseas models with comparable performance are typically priced more than three times higher than their domestic counterparts.

In inference efficiency, overseas models still occupy the high-efficiency zone, but domestic models are accelerating optimization. Kimi-K2.5-Thinking achieved an approximately 14% improvement in reasoning capability and a nearly 3x increase in inference speed over its predecessor Kimi-K2-Thinking, demonstrating initial progress toward the synergistic optimization of "high performance + high efficiency".


VII. Human Consistency Validation: SuperCLUE Highly Correlated with LMArena

The report also conducted human consistency validation, correlating SuperCLUE scores with LMArena's public anonymous voting results. The Pearson correlation coefficient reached 0.8239, and the Spearman correlation coefficient reached 0.8321, indicating that SuperCLUE benchmark results are highly consistent with human preferences and carry strong credibility.


VIII. Specialized Evaluation Matrix Fully Upgraded: Covering Multimodal, Agent, Programming, and Other Frontier Areas

Beyond the general benchmark, SuperCLUE 2025 also launched multiple specialized evaluation modules covering:

  • Programming Arena: Claude-Opus-4.5-Reasoning leads, with Kimi-K2.5-Thinking in second place.

  • Image/Video Arena: International models such as Gemini-3-Pro-Image-Preview and veo-3.0 lead, with domestic models like Seedream and Kling closely following.

  • Agent Specialization: Covering CUA, general agents, software engineering (SWE), deep research, and other directions, comprehensively evaluating model capabilities in real-world scenarios.

  • Hallucination and Instruction Following Specialization: DeepSeek-R1 leads in factual hallucination tasks, while the GPT-5 series performs best in precise instruction following.

 

ScreenShot_2026-02-24_170058_810 

 

IX. 2026 Outlook: Domestic Models Poised for Full Catch-Up

The report concludes that 2025 has been a pivotal year for domestic large models transitioning from "following" to "running alongside." In areas such as code generation, mathematical reasoning, and open-source ecosystems, domestic models have achieved global competitiveness; however, there is still room for improvement in tasks such as agent and instruction following. With the release of next-generation models such as Kimi-K2.5-Thinking and Qwen3-Max-Thinking, domestic large models are accelerating their progress toward top-tier international standards.

SuperCLUE stated that it will continue to deepen its evaluation framework across multimodal, agent, and reasoning directions, providing more scientific and comprehensive reference for the industry to support the high-quality development of China's large model sector.


About SuperCLUE

SuperCLUE is the continuation of the Chinese Language Understanding Evaluation (CLUE) benchmark in the era of large models, focusing on comprehensive evaluations of general-purpose large models. It is committed to providing objective, neutral third-party assessment services for industry, academia, and research institutions.

For the full report, please visit:Full text of the report