Artificial Analysis AI Year in Review — 2024 Highlights

Open-source models are approaching top labs. Represented by open-weight models developed by companies such as Meta, Mistral, and Alibaba Cloud, they have approached and surpassed the intelligence level of GPT-4.

Frontier Models

In 2024, multiple labs caught up with OpenAI's GPT-4, and the first models surpassing GPT-4's intelligence level emerged.

Frontier Language Model Intelligence, Evolution Over Time

bg01

 Key Trends:

Competing labs catching up with OpenAI's GPT-4: After OpenAI launched GPT-3.5 (ChatGPT) in November 2022, it kicked off the language model race; over the following 18 months, competing labs worked hard to catch up.

Open-source models approaching top labs: Represented by open-weight models developed by companies such as Meta, Mistral, and Alibaba Cloud, they have approached and surpassed the intelligence level of GPT-4.

Sparks of intelligence beyond GPT-4: The final months of 2024 witnessed the first major intelligence leap beyond GPT-4, led by OpenAI's o1. Themes such as inference-time compute scaling, data quality, and new reinforcement learning techniques, together with pre-training compute scaling, became the main levers for improving models.

Country of Origin of Language Models

The United States dominates the intelligence frontier; China appears to be second, with only a few other countries demonstrating frontier-level training capabilities.

bg02

Open-Source Language Model Intelligence

Driven by models from companies such as Meta, Mistral, and Alibaba Cloud, the performance gap between open-source and proprietary models has narrowed significantly.

bg03

Language Model Inference Pricing

In 2024, inference pricing for intelligent language models at all levels dropped significantly; GPT-4o mini approached GPT-4's intelligence level while being 100 times cheaper.

bg04

Language Model Size

Smaller models reaching intelligence levels previously achievable only by larger models is a key factor behind falling inference pricing and improved speed.

bg05

Key Takeaways:

Over the past 12 months, small models have improved significantly. This rate of improvement has greatly exceeded that of large models, leading to a substantial narrowing of the performance gap.

What Is Driving the Improvement?

Compute Scaling

• "Chinchilla-optimal" is a thing of the past: Language models are now used at scale in real-world applications, and increasing training compute to reduce inference is a clear win.

• Open-weight models are now frequently trained on over 10T tokens — which was almost unheard of two years ago.

Knowledge Distillation

• The intelligence gains of larger models have been distilled into smaller versions.

• Reinforcement learning techniques such as Direct Preference Optimization have driven the latest generation of small models, including Llama 3.1/2.

High-Quality Data

• Leveraging high-quality data is particularly beneficial for small models, especially Microsoft's Phi series of models.

Language Model Context Window

The context window has significantly increased, with 128k becoming the new standard; long-context reasoning enables models to process more data at once.

bg06

Key Takeaways:

Since Q3 2023, the median context length of frontier models has increased 32-fold. The expansion was initially led by proprietary models, but by Q3 2024, open-source models had caught up. Recent advances have also increased the context length of certain models (Gemini, Nova) to 2 million tokens.

What Is Driving the Improvement?

New Technologies

• Techniques for increasing context length range from hardware-aware distributed attention implementations to attention approximation and length extrapolation methods such as RoPE.

Developer Demand

• 70% of developers in the Artificial Analysis developer survey said that context window is important to them when choosing a model.

• Labs are responding to a clear industry demand signal.

Implications

Reduced Complexity

• Manage smaller context windows in production.

• Longer contexts reduce the need to build retrieval, summarization, and truncation strategies.

New Methods and Applications

• Agent tools quickly exhaust the context window.

• Many applications do not rely on text alone. Larger context windows support multimodal inputs including images, video, and more.

Inference Strategies

• More compute is being used per task, not only in generating additional reasoning tokens but also in agentic workflows. Longer contexts enable more flexibility.

Language Models: Landscape of Major Players

Participants in the AI value chain vary in their degree of vertical integration; Google, from TPU accelerators to Gemini, is the most vertically integrated player.

bg07_20250102_17358055435963600

 Insights from Language Model Developers

Demand for AI models is concentrated on releases from top AI labs; model inference quality and price are the main decision factors when choosing a model.

bg08


Companies are adopting a variety of technical approaches to using LLMs, but no single approach dominates; most LLM users intend to use multimodality.

bg09

Most AI model users intend to use multiple models in their applications; ~3/4 of users access models through hosted serverless endpoints.

bg10

Image Generation Quality

Image generation quality improved rapidly in 2024, with significant leaps in photorealism, prompt adherence, and text rendering.

bg11

Image Generation Landscape

Progress and competition in image models continued to accelerate in 2024; the top 5 models in the Artificial Analysis Image Arena were all launched after Q3 2024.

bg12

Video Generation Landscape

When OpenAI previewed Sora in February 2024, there was little competition, but by the time it launched in December 2024, the competitive landscape had become much more intense.

bg13

Text-to-Speech Landscape

The latest generation of transformer-based text-to-speech models achieved new quality milestones in 2024, surpassing long-standing Hyperscaler offerings.

bg14

 

语音转文本现状


OpenAI 在 2022 年底开源 Whisper 时改变了 AI 转录的格局,使云推理服务商能够进入市场并在和价格方面展开竞争。


bg15

Click:Artificial Analysis 2024 China Artificial Intelligence State Report (Complete Edition)