On May 28, the SuperCLUE team released the May 2025 Comprehensive Chinese Large Model Benchmark Report, conducting a comprehensive evaluation of 43 large models from both domestic and international sources. The report shows that OpenAI's o4-mini(high) topped the overall rankings with a total score of 70.51, demonstrating exceptional capabilities across reasoning, code generation, agent capabilities, and other dimensions, leading the best domestic model by 7.35 points. While domestic models still lag in overall capabilities, they excel in niche areas such as text creation and code generation, with some metrics even surpassing overseas models.

Domestic and International Model Gap Narrows, but Instruction Logic Remains a Weakness
The report notes that over the past 25 months, the general capability gap between domestic and international large models in the Chinese language domain has gradually narrowed. However, domestic models perform relatively weakly in instruction logic tasks, with the highest-scoring Hunyuan-T1-20250403 (36.97) trailing o4-mini(high) by 31.1 points—indicating significant room for improvement.
Domestic Models Shine in Multiple Areas
ByteDance's Doubao-1.5-thinking-pro-250415 led the text understanding and creation task with a score of 81.04. Alibaba's Qwen3-235B-A22B achieved 90.53 in code generation, with a negligible gap from top overseas models. Additionally, small-parameter models such as the Qwen3 series demonstrated remarkable potential, with the 4B, 8B, and 14B versions all scoring over 50 in reasoning tasks—surpassing some closed-source large models.

Open-Source Models and Cost-Performance Advantages
Domestic open-source models delivered impressive performances, with DeepSeek-R1 from DeepSeek and Alibaba's Qwen3 series leading the global open-source ecosystem in Chinese-language scenarios. Domestic models also hold significant cost-performance advantages—for example, DeepSeek-V3-0324 delivers high-quality output at a low price, while overseas models like Gemini 2.5 Pro, though high-performing, come at a steep cost.

Maturity Analysis: Text Understanding Strongest, Scientific Reasoning Needs Improvement
Among domestic large models, text understanding and creation tasks show the highest maturity (SC Index 0.91), while scientific reasoning (0.26) and mathematical reasoning (0.38) remain in the low-maturity range, requiring further optimization.

Future Outlook
The SuperCLUE team stated that as AI technology rapidly evolves, domestic models have demonstrated the ability to compete globally in specific domains. However, continued breakthroughs are needed in directions such as instruction logic and multimodal interaction. The report also called for industry attention to the extreme cost-performance potential of small-parameter models to drive widespread adoption of edge-side applications.
This evaluation covered six tasks—mathematical reasoning, scientific reasoning, code generation, and others—with all questions being original and the question bank updated every two months to ensure objectivity and timeliness. As an independent third-party evaluator, SuperCLUE is committed to providing professional references for the industry and supporting the technological development of Chinese large models.
Click: Full version of the Chinese Large Model Benchmark Evaluation May 2025 Report