Frontier Models
In 2024, multiple labs caught up with OpenAI's GPT-4, and the first models surpassing GPT-4's intelligence level emerged.
Frontier Language Model Intelligence, Evolution Over Time

Key Trends:
Competing labs catching up with OpenAI's GPT-4: After OpenAI launched GPT-3.5 (ChatGPT) in November 2022, it kicked off the language model race; over the following 18 months, competing labs worked hard to catch up.
Open-source models approaching top labs: Represented by open-weight models developed by companies such as Meta, Mistral, and Alibaba Cloud, they have approached and surpassed the intelligence level of GPT-4.
Sparks of intelligence beyond GPT-4: The final months of 2024 witnessed the first major intelligence leap beyond GPT-4, led by OpenAI's o1. Themes such as inference-time compute scaling, data quality, and new reinforcement learning techniques, together with pre-training compute scaling, became the main levers for improving models.
Country of Origin of Language Models
The United States dominates the intelligence frontier; China appears to be second, with only a few other countries demonstrating frontier-level training capabilities.

Open-Source Language Model Intelligence
Driven by models from companies such as Meta, Mistral, and Alibaba Cloud, the performance gap between open-source and proprietary models has narrowed significantly.

Language Model Inference Pricing
In 2024, inference pricing for intelligent language models at all levels dropped significantly; GPT-4o mini approached GPT-4's intelligence level while being 100 times cheaper.

Language Model Size
Smaller models reaching intelligence levels previously achievable only by larger models is a key factor behind falling inference pricing and improved speed.

Key Takeaways:
Over the past 12 months, small models have improved significantly. This rate of improvement has greatly exceeded that of large models, leading to a substantial narrowing of the performance gap.
What Is Driving the Improvement?
Compute Scaling
• "Chinchilla-optimal" is a thing of the past: Language models are now used at scale in real-world applications, and increasing training compute to reduce inference is a clear win.
• Open-weight models are now frequently trained on over 10T tokens — which was almost unheard of two years ago.
Knowledge Distillation
• The intelligence gains of larger models have been distilled into smaller versions.
• Reinforcement learning techniques such as Direct Preference Optimization have driven the latest generation of small models, including Llama 3.1/2.
High-Quality Data
• Leveraging high-quality data is particularly beneficial for small models, especially Microsoft's Phi series of models.
Language Model Context Window
The context window has significantly increased, with 128k becoming the new standard; long-context reasoning enables models to process more data at once.

Key Takeaways:
Since Q3 2023, the median context length of frontier models has increased 32-fold. The expansion was initially led by proprietary models, but by Q3 2024, open-source models had caught up. Recent advances have also increased the context length of certain models (Gemini, Nova) to 2 million tokens.
What Is Driving the Improvement?
New Technologies
• Techniques for increasing context length range from hardware-aware distributed attention implementations to attention approximation and length extrapolation methods such as RoPE.
Developer Demand
• 70% of developers in the Artificial Analysis developer survey said that context window is important to them when choosing a model.
• Labs are responding to a clear industry demand signal.
Implications
Reduced Complexity
• Manage smaller context windows in production.
• Longer contexts reduce the need to build retrieval, summarization, and truncation strategies.
New Methods and Applications
• Agent tools quickly exhaust the context window.
• Many applications do not rely on text alone. Larger context windows support multimodal inputs including images, video, and more.
Inference Strategies
• More compute is being used per task, not only in generating additional reasoning tokens but also in agentic workflows. Longer contexts enable more flexibility.
Language Models: Landscape of Major Players
Participants in the AI value chain vary in their degree of vertical integration; Google, from TPU accelerators to Gemini, is the most vertically integrated player.

Insights from Language Model Developers
Demand for AI models is concentrated on releases from top AI labs; model inference quality and price are the main decision factors when choosing a model.

Companies are adopting a variety of technical approaches to using LLMs, but no single approach dominates; most LLM users intend to use multimodality.

Most AI model users intend to use multiple models in their applications; ~3/4 of users access models through hosted serverless endpoints.

Image Generation Quality
Image generation quality improved rapidly in 2024, with significant leaps in photorealism, prompt adherence, and text rendering.

Image Generation Landscape
Progress and competition in image models continued to accelerate in 2024; the top 5 models in the Artificial Analysis Image Arena were all launched after Q3 2024.

Video Generation Landscape
When OpenAI previewed Sora in February 2024, there was little competition, but by the time it launched in December 2024, the competitive landscape had become much more intense.

Text-to-Speech Landscape
The latest generation of transformer-based text-to-speech models achieved new quality milestones in 2024, surpassing long-standing Hyperscaler offerings.

语音转文本现状
OpenAI 在 2022 年底开源 Whisper 时改变了 AI 转录的格局,使云推理服务商能够进入市场并在和价格方面展开竞争。

Click:Artificial Analysis 2024 China Artificial Intelligence State Report (Complete Edition)