Since 2023, AI large models have set off the largest wave of artificial intelligence in history on a global scale. Entering 2024, the competitive landscape of global large models has become increasingly intense. With the releases of GPT-4o, Claude 3.5, Gemini 1.5 Pro, and Llama 3, domestic large models likewise engaged in a magnificent race to catch up in the first half of 2024. The Chinese large model benchmark SuperCLUE has continuously tracked in real time the development trends and comprehensive performance of large models both domestically and internationally.


I. Key Progress of Domestic Large Models
1. Key Progress of Large Models in 2023 and the Panorama of Chinese Large Models
Domestic academia and industry have also made substantial breakthroughs over the past year and a half. This can be roughly divided into three stages: the preparation period (after the release of ChatGPT, domestic industry, academia, and research quickly reached a consensus on large models), the growth period (the quantity and quality of domestic large models began to gradually increase), and the explosion period (open-source and closed-source large models emerged endlessly across all walks of life, forming a competitive landscape of a hundred models contending).

2. Panorama of Chinese Large Models Worth Watching in 2024
As of now, more than a hundred open-source, closed-source general large models and industry-specific large models have been released domestically. SuperCLUE has compiled a panorama of large models worth watching in 2024.

3. Domestic and International Large Model Technology Development Trends from 2023 to 2024
Since May 2023, the capabilities of large models both domestically and internationally have continued to develop. Among them, the best overseas models represented by the GPT series have undergone multiple iterations and upgrades from GPT-3.5, GPT-4, GPT-4 Turbo, to GPT-4o. Domestic models have also gone through a magnificent 14-month iteration cycle, during which the top-ranked model changed hands 8 times, continuously raising the strongest combat power of domestic models.
In terms of the overall trend, the gap in general capabilities in the Chinese language domain between first-tier large models domestically and internationally has continued to narrow, shrinking from 30.12% in May 2023 to 4.94% in June 2024.

II. SuperCLUE General Capability Evaluation
1. Introduction to the Chinese Large Model Benchmark SuperCLUE
The Chinese Language Understanding Evaluation (CLUE) benchmark is a language model evaluation benchmark dedicated to scientific, objective, and neutral assessment, initiated in 2019. It has successively launched widely cited evaluation benchmarks such as CLUE, FewCLUE, KgCLUE, and DataCLUE.
SuperCLUE is the development and continuation of the CLUE benchmark in the era of large models. It focuses on comprehensive evaluation of general large models. Based on years of evaluation experience and the widespread application of general large models in academia, industry, and on the user side, SuperCLUE has constructed a multi-level, multi-dimensional comprehensive evaluation benchmark.
Differences Between Traditional Evaluation and SuperCLUE

Three Major Characteristics of SuperCLUE
1) Independent Third-Party Evaluation, Not Led by Large Model Developers
As competition among large models both domestically and internationally becomes increasingly fierce, evaluations led by model developers may carry the risk of favoring their own products. In sharp contrast, SuperCLUE, as a completely independent third-party evaluation organization, is committed to providing unbiased and objective evaluation results. SuperCLUE adopts advanced automated evaluation technology to effectively eliminate the uncertainty caused by human factors, ensuring that every evaluation is fair and impartial.
2) Evaluation Methods Consistent with Real User Experience Goals
Unlike traditional evaluations that use multiple-choice questions, SuperCLUE aims to remain consistent with real user experience goals, and therefore incorporates evaluations of open-ended subjective questions. Through a multi-dimensional, multi-perspective, and multi-level evaluation system as well as a dialogue format, it simulates the application scenarios of large models and genuinely and effectively examines the generation capabilities of models.
3) "Live" Updates, with Evaluation Systems/Methods Advancing with the Times
Unlike evaluations in traditional academic fields, SuperCLUE continuously upgrades and iterates its evaluation system, evaluation dimensions, and methods in accordance with global large model technology development trends, so as to quantify the degree of technological evolution of large models as accurately as possible.
2. SuperCLUE Evaluation System and Dataset Description


To more authentically reflect the capabilities of large models, this semi-annual evaluation adopts a multi-dimensional, multi-level comprehensive evaluation scheme, consisting of three major dimensions: Science, Liberal Arts, and Hard.
[Science Tasks] are divided into calculation, logical reasoning, and code evaluation sets;
[Liberal Arts Tasks] are divided into seven evaluation sets: encyclopedic knowledge, language understanding, long text, role-playing, generation and creation, safety, and tool use;
[Hard Tasks] For the first time, this evaluation incorporates an precise instruction-following evaluation set. In addition, complex multi-step reasoning and high-difficulty problem-solving Hard evaluation sets will be launched successively.

3. List of Evaluated Models
The data for this evaluation is drawn from the SuperCLUE June evaluation results, and the models selected are the June versions of 33 representative large models from both domestic and international sources.

Click: Chinese Large Model Benchmark Evaluation 2024 First Half Report (Complete Edition)