On August 4, 2025, independent third-party evaluator SuperCLUE released the H1 2025 Comprehensive Chinese Large Model Benchmark Report, conducting a comprehensive assessment of 45 mainstream large models from both domestic and international sources. The report shows that overseas models maintain a significant advantage in reasoning tasks, while domestic models excel in agent capabilities and hallucination control. In the open-source model space, China has already achieved leadership.
I. Overall Ranking: Overseas Models Take Top Three
This evaluation covered six tasks: mathematical reasoning, scientific reasoning, code generation, agent capabilities, hallucination control, and precise instruction following, with a total of 1,288 original questions.
1st Place: OpenAI o3 (73.78 points)
2nd Place: OpenAI o4-mini(high) (73.32 points)
3rd Place: Google Gemini-2.5-Pro (68.98 points)
Top Domestic Performer: ByteDance Doubao-Seed-1.6-thinking-250715 (68.04 points, 4th globally)

The gap between top domestic and international models narrowed from 10.42% in May to 7.78%, demonstrating that domestic models are catching up rapidly.

II. Domestic Models Showcase Notable Strengths
1. Global Leadership in Agent Tasks
Doubao-Seed-1.6-thinking-250715 ranked first globally with a score of 90.67. GLM-4.5 and SenseNova V6 Reasoner tied for second domestically with 83.58 points.

2. Excellent Performance in Hallucination Control
Doubao-Seed, ERNIE-X1-Turbo-32K-Preview, and Hunyuan-T1-20250711 secured the top three domestic spots, with the gap to top overseas models being less than 1 point.
3. Significant Lead in Open-Source Models
DeepSeek-R1-0528 (66.15), Qwen3-235B-A22B-Thinking-2507 (64.34), and GLM-4.5 (63.25) ranked top three among open-source models. The best overseas open-source model scored only 46.37, trailing domestic models by nearly 20 points.
III. Overseas Models Still Dominate Reasoning Tasks
o3 scored 75.02 in reasoning tasks, leading the best domestic model DeepSeek-R1-0528 (65.74) by nearly 10 points. Overseas models also occupied four of the top five spots in code generation, with closed-source models averaging 76.16, significantly outperforming open-source models at 54.94.

IV. Small-Parameter Models Shine
Among models with under 10 billion parameters:
Qwen3-8B(Thinking) topped the list with 48.38 points, surpassing overseas models of comparable scale. In the sub-5B edge-side category, domestic models secured the top two spots, demonstrating strong deployment potential.

V. Cost-Performance and Efficiency Analysis
Top overseas models (e.g., o3) are expensive, offering lower cost-performance compared to domestic models. Among domestic models, Hunyuan-T1, GLM-4.5, and Doubao-Seed combine high performance with reasonable cost. In inference efficiency, overseas models respond faster, with only SenseNova V6 Reasoner from the domestic side entering the high-efficiency zone.

VI. Specialized Capability Maturity
According to the SC Maturity Index, domestic models perform as follows:
Medium Maturity (0.5–0.8) : Mathematical reasoning, agent, scientific reasoning, code generation
Low Maturity (0.1–0.5) : Hallucination control, precise instruction following
VII. Representative Model Analysis
Doubao-Seed-1.6-thinking-250715: Strong in agent, code generation, and hallucination control.
DeepSeek-R1-0528: Excels in complex reasoning and precise instruction following; top open-source model.
GLM-4.5: Top domestic in agent tasks, leading in scientific reasoning.
kimi-k2-0711-preview: Third domestically in code generation, with strong reasoning capabilities.
VIII. Specialized Benchmark Extensions
The report also introduces several specialized evaluation benchmarks:
Agent Series: Including DeepResearch, AgentCLUE-General
Multimodal Series: Covering visual reasoning, text-to-video, image-to-video, text-to-image, etc.
Text and Reasoning Series: Focusing on hallucination control, factuality, instruction following, etc.
Performance Series: Including third-party platform integration stability and search capability evaluations.
Conclusion
In the first half of 2025, Chinese large models have made significant progress in general capabilities, open-source ecosystems, and agent applications, demonstrating strong competitiveness especially in Chinese-language scenarios. However, in core capabilities such as complex reasoning and code generation, overseas models still maintain their lead. Going forward, domestic models need to continue breaking through in technical depth, efficiency optimization, and multimodal integration.
Source: SuperCLUE Team
Full Report