On September 28, SuperCLUE, the comprehensive evaluation benchmark for Chinese general-purpose large models, released its September 2025 ranking. The data shows that while top-tier international models continue to maintain their lead, domestic large models are closing the gap at an astonishing pace, demonstrating strong competitiveness across multiple key capabilities.

Domestic Large Models Make Rapid Strides, SuperCLUE September Ranking Reveals New AI Landscape
International Giants Temporarily Lead, OpenAI Continues to Dominate
In this evaluation, OpenAI's GPT-S(high) topped the overall ranking with a comprehensive score of 69.37, excelling in numerical reasoning (73.64) and hallucination control (78.34). Another OpenAI model, o3(high), achieved an impressive 77.27 in numerical reasoning, demonstrating the company's deep technical accumulation in logical reasoning.
Anthropic's Claude-Opus-4.1-Reasoning delivered the best performance in hallucination control with an outstanding score of 86.23. Google's Gemini-2.5-Pro led other models in agent capabilities with 83.06. The top five positions were all captured by international vendors, forming the current first tier in the large model space.
Domestic Players Rise with Strength, Achieving Breakthroughs Across Multiple Domains
Notably, domestic large models are showing strong momentum in catching up. DeepSeek-V3.1-Terminus-Thinking leads the domestic contingent with 61.44 points, narrowing the gap with the top spot to less than 8 points. ByteDance's Doubao-Seed-1.6-thinking follows closely with 60.96, delivering standout performances in hallucination control (77.78) and agent capabilities (82.50).
Baidu's ERNIE-XL1 and Alibaba's Qwen3-Max scored 60.33 and 59.68 respectively, securing spots in the highly competitive domestic first tier. Huawei's openPangu-Ultra-MoE-7188 demonstrated solid capability with a score of 58.87.

Diversified Technical Approaches, Chain-of-Thought Proves Highly Effective
From a technical perspective, models with "Thinking" (chain-of-thought) capabilities generally performed better. Thinking versions from DeepSeek, ByteDance, Alibaba, and others showed significant improvements over their base versions, indicating that chain-of-thought technology is highly effective in enhancing model reasoning capabilities.
In specific capability dimensions, domestic models demonstrated differentiated strengths. Alibaba's Qwen3-Max led domestic models in code generation with a score of 67.18; ByteDance's Doubao achieved an impressive 82.50 in agent capabilities; and DeepSeek showed balanced development in scientific reasoning and agent capabilities.-
Industrial Ecosystem Maturing, Broad Application Prospects
Currently, the large model industry has formed a healthy, multi-layered, and multi-technical-route development landscape. From major internet companies to specialized AI firms, players are actively positioning themselves along both open-source and closed-source technical paths. This diversified development provides rich options for different application scenarios and will drive the prosperity of the entire AI industrial ecosystem.
Industry experts suggest that the rapid progress domestic large models have achieved in just a few years reflects China's strong innovation capabilities and industrial strength in the AI field. As technologies continue to mature and application scenarios expand, large models will play an increasingly important role in the digital economy, injecting new momentum into industrial upgrading and innovative development.
Looking ahead, domestic large models are expected to further leverage their advantages in Chinese language understanding and localized application scenarios, while continuing to break through in foundational capability building, advancing the high-quality development of China's AI industry.
Links: https://superclueai.com/
Full report:
https://pan.baidu.com/s/1V1e2NmhWuukBGnmuCMZ7-Q?pwd=xga3