OpenAI Launches GPT-6.1 Sol Ultrafast: 6× the Price for 8× the Speed, Opening the Era of the Agent "Speed Premium"

Chinese

On October 8, OpenAI officially launched the Ultrafast version of GPT-6.1 Sol across the API, Codex, and ChatGPT Work. It runs at up to 8× the speed of the standard version (about 300 tokens/second) and is priced at 6× the standard version — $12 per million input tokens and $60 per million output tokens.

On October 8, OpenAI officially launched the Ultrafast version of GPT-6.1 Sol across the API, Codex, and ChatGPT Work. It runs at up to 8× the speed of the standard version (about 300 tokens/second) and is priced at 6× the standard version — $12 per million input tokens and $60 per million output tokens.

This is the landmark move in which OpenAI, following the release of Ultrafast mode at DevDay in late September, "hands down" that capability from the flagship model GPT-6 Astra to Sol, its value-for-money mainstay — and it is a key step in "speed" being formally productized and priced.

The Same Model, Two Speeds: Ultrafast Is Not a New Model

Ultrafast is essentially a service tier (service_tier), not a second model. Developers call `gpt-6.1-sol` by setting `service_tier: “ultrafast”`, and the model weights, context window (1.05 million tokens), and knowledge cutoff date (April 30, 2026) all remain unchanged.

What you pay extra for is not a smarter AI, but a shorter wait. OpenAI's official definition of Ultrafast is also deliberately conservative: "shortening the time between generated output tokens."

Price Tiers: Three Speed Levels, a Sixfold Premium

GPT-6.1 Sol is now divided into three tiers by speed:

Ultrafast's pricing is exactly 6× that of the standard version. Notably, the cached read price is $0.60 per million tokens, likewise scaled up by 6×.

Who Needs to Pay for Speed?

Ultrafast's value is highly scenario-dependent. The typical use cases OpenAI cites are real-time coding assistants, interactive agents, and customer-facing real-time experiences.

For continuously running coding agents, the advantage compounds: if a task requires 12 rounds of tool calls, and each round's generation speed is 8× faster, the total completion time is shortened by far more than "a little faster." One developer's testing found that the standard version of Sol outputs at about 42 tokens/second, while Ultrafast reaches 304 tokens/second — more than 7× faster; the wait from question to first character dropped from 7.1 seconds to 2.9 seconds, and the entire response went from 17 seconds to 4.3 seconds.

But for a single complex reasoning request, most of the time is spent "thinking" rather than "outputting," so paying 6× the price may just mean staring at the same loading spinner for a little less time.

Availability: A "Privilege" for Pro 500 and Enterprises

On the API side, Ultrafast is open to all API users, billed by usage.

But in Codex and ChatGPT Work, Ultrafast is available only to Pro 500 subscribers ($500 per month), eligible pay-as-you-go enterprise editions, and quota-billed education editions; enterprise users must have an administrator enable it manually.

Pro 500's compute quota is 25× that of Plus, but Ultrafast deducts quota at 8× the standard tier. This means that using the entire 25× quota on Ultrafast is equivalent to only a little over 3 Plus's worth of usage.

Data Residency and Compliance

GPT-6.1 Sol Ultrafast fully supports U.S. and EU data residency compliance services. Notably, the Ultrafast version of the flagship model GPT-6 Astra supports only U.S. data residency, while Sol covers the EU (EEA + Switzerland). For enterprises with EU data compliance requirements, this is a substantive differentiating advantage.

The launch of Ultrafast marks OpenAI's formal pricing of the "speed" dimension. As AI agents move from "question-answering tools" toward "continuously running productivity," latency is becoming a variable as important as intelligence. When AI writes code for you, operates software, and handles tickets, every second of waiting means a blocked workflow and lost attention.

Whether 6× the price for 8× the speed is a good deal depends on how high the cost of "humans waiting for AI" is in your business scenario. For high-frequency agent pipelines, Ultrafast may be a necessity; for individual users' everyday conversations, the standard version is already sufficient. OpenAI is not trying to get everyone to upgrade, but instead hands the choice to developers, using price signals to guide resource allocation.

From a broader perspective, Ultrafast is the inevitable result of "tiered inference compute." When model capabilities converge and compute cost becomes the focus of competition, using price to differentiate service levels is a mature model proven by the cloud services industry. OpenAI is turning AI inference into a finely tuned compute business.