H1 2025 Chinese Large Model Evaluation Report Released: Domestic Models Rise, Overseas Still Lead in Reasoning

9.2 On August 4, 2025, independent third-party evaluator SuperCLUE released the H1 2025 Comprehensive Chinese Large Model Benchmark Report, conducting a comprehensive assessment of 45 mainstream large models from both domestic and international sources. The report shows that overseas models maintain a significant advantage in reasoning tasks, while domestic models excel in agent capabilities and hallucination control. In the open-source model space, China has already achieved leadership.

On August 4, 2025, independent third-party evaluator SuperCLUE released the H1 2025 Comprehensive Chinese Large Model Benchmark Report, conducting a comprehensive assessment of 45 mainstream large models from both domestic and international sources. The report shows that overseas models maintain a significant advantage in reasoning tasks, while domestic models excel in agent capabilities and hallucination control. In the open-source model space, China has already achieved leadership.

I. Overall Ranking: Overseas Models Take Top Three

This evaluation covered six tasks: mathematical reasoning, scientific reasoning, code generation, agent capabilities, hallucination control, and precise instruction following, with a total of 1,288 original questions.

  • 1st Place: OpenAI o3 (73.78 points)

  • 2nd Place: OpenAI o4-mini(high) (73.32 points)

  • 3rd Place: Google Gemini-2.5-Pro (68.98 points)

Top Domestic Performer: ByteDance Doubao-Seed-1.6-thinking-250715 (68.04 points, 4th globally)


250603

 

The gap between top domestic and international models narrowed from 10.42% in May to 7.78%, demonstrating that domestic models are catching up rapidly.


250609


II. Domestic Models Showcase Notable Strengths

1. Global Leadership in Agent Tasks

Doubao-Seed-1.6-thinking-250715 ranked first globally with a score of 90.67. GLM-4.5 and SenseNova V6 Reasoner tied for second domestically with 83.58 points.


250626

 

2. Excellent Performance in Hallucination Control

Doubao-Seed, ERNIE-X1-Turbo-32K-Preview, and Hunyuan-T1-20250711 secured the top three domestic spots, with the gap to top overseas models being less than 1 point.

3. Significant Lead in Open-Source Models

DeepSeek-R1-0528 (66.15), Qwen3-235B-A22B-Thinking-2507 (64.34), and GLM-4.5 (63.25) ranked top three among open-source models. The best overseas open-source model scored only 46.37, trailing domestic models by nearly 20 points.

III. Overseas Models Still Dominate Reasoning Tasks

o3 scored 75.02 in reasoning tasks, leading the best domestic model DeepSeek-R1-0528 (65.74) by nearly 10 points. Overseas models also occupied four of the top five spots in code generation, with closed-source models averaging 76.16, significantly outperforming open-source models at 54.94.


250632

 

IV. Small-Parameter Models Shine

Among models with under 10 billion parameters:

Qwen3-8B(Thinking) topped the list with 48.38 points, surpassing overseas models of comparable scale. In the sub-5B edge-side category, domestic models secured the top two spots, demonstrating strong deployment potential.


250633

 

V. Cost-Performance and Efficiency Analysis

Top overseas models (e.g., o3) are expensive, offering lower cost-performance compared to domestic models. Among domestic models, Hunyuan-T1, GLM-4.5, and Doubao-Seed combine high performance with reasonable cost. In inference efficiency, overseas models respond faster, with only SenseNova V6 Reasoner from the domestic side entering the high-efficiency zone.


250635


VI. Specialized Capability Maturity

According to the SC Maturity Index, domestic models perform as follows:

  • Medium Maturity (0.5–0.8) : Mathematical reasoning, agent, scientific reasoning, code generation

  • Low Maturity (0.1–0.5) : Hallucination control, precise instruction following

VII. Representative Model Analysis

  1. Doubao-Seed-1.6-thinking-250715: Strong in agent, code generation, and hallucination control.

  2. DeepSeek-R1-0528: Excels in complex reasoning and precise instruction following; top open-source model.

  3. GLM-4.5: Top domestic in agent tasks, leading in scientific reasoning.

  4. kimi-k2-0711-preview: Third domestically in code generation, with strong reasoning capabilities.

VIII. Specialized Benchmark Extensions

The report also introduces several specialized evaluation benchmarks:

  • Agent Series: Including DeepResearch, AgentCLUE-General

  • Multimodal Series: Covering visual reasoning, text-to-video, image-to-video, text-to-image, etc.

  • Text and Reasoning Series: Focusing on hallucination control, factuality, instruction following, etc.

  • Performance Series: Including third-party platform integration stability and search capability evaluations.

Conclusion

In the first half of 2025, Chinese large models have made significant progress in general capabilities, open-source ecosystems, and agent applications, demonstrating strong competitiveness especially in Chinese-language scenarios. However, in core capabilities such as complex reasoning and code generation, overseas models still maintain their lead. Going forward, domestic models need to continue breaking through in technical depth, efficiency optimization, and multimodal integration.

Source: SuperCLUE Team
Full Report

https://pan.baidu.com/s/1p2gjbT3C22KzUYE2iNC40A?pwd=b8ia