On October 7, Anthropic officially released Claude Haiku 5.5, calling it the fastest, cheapest, and most capable small model in the Haiku series. The model's running costs are on average about 75% lower than those of the previous-generation Haiku 4.5, further strengthening Anthropic's competitiveness in the "small model price-performance" race.
A "Step-Change" Price Cut: Cached Reads Just $0.01
Haiku 5.5 uses tiered pricing based on prompt length, with different prices for requests of no more than 100,000 tokens and requests exceeding 100,000 tokens. Compared with the previous-generation Haiku 4.5, the reductions are significant:
Taking a request of no more than 100,000 tokens as an example, the input price is just $0.10 per million tokens and the output price is $0.50—a price level that is highly competitive in the small-model race.
Notably, the cached read price has dropped to $0.01 per million tokens, just one-tenth of the previous generation. For agent tasks that need to repeatedly call the same context, the cost savings from this adjustment are especially significant.
Sonnet 5.5 Cached Pricing Also Cut
Anthropic also announced a reduction in the cached read price for Sonnet 5.5, from $0.20 per million tokens to $0.10. Anthropic said this adjustment can reduce Sonnet 5.5's running costs in most agent tasks by about 20%.
This series of price adjustments sends a clear signal: cache costs are becoming a new focal point of competition among large-model vendors. For continuously running agents, the cache fees generated by repeatedly reading context often account for the bulk of costs, and lowering cache prices improves the economics of long-term operation more substantially than lowering the price of a single call.
Performance: Balanced Knowledge Work and Agent Capabilities
Despite being positioned as a small model, Haiku 5.5 performs in a balanced way across multiple benchmarks:
- GDPval-AA v2.1 (knowledge work): 1620 points
- AA-Briefcase v1.1 (agent benchmark): 1578 points
- OSWorld 2.1 (offline subset, computer operation): 72.4%
- Terminal-Bench 4.0 (terminal coding): 39.2%
The First Haiku Model to Support "Adjustable Performance"
Haiku 5.5 is the first Haiku model to support adjustable performance settings, allowing users to configure a choice between cost and intelligence according to task requirements. This design gives the small model greater scenario adaptability—simple tasks use a low-cost configuration, while complex tasks increase the level of intelligence.
Anthropic also positions Haiku 5.5 as the sub-agent model when Opus 5.5 and Sonnet 5.5 perform coding tasks. In multi-agent collaboration architectures, flagship models handle planning and decision-making, while small models such as Haiku 5.5 handle the execution of specific subtasks. This "large-and-small model collaboration" model is becoming the mainstream architecture for agent systems.
The release of Haiku 5.5 reflects a key shift in competition among large AI models—as the performance gap between flagship models gradually narrows, the focus of competition is shifting toward "the cost per unit of intelligence." By cutting the cached read price to $0.01 and reducing running costs by 75% versus the previous generation, Haiku 5.5 is essentially redesigning the cost structure for "continuously running agents," an emerging workload.
Even more noteworthy is its "adjustable performance" design—this marks the evolution of small models from "cheap substitutes with limited capabilities" into "tools for finely tuning the cost-intelligence balance." When Anthropic positions Haiku 5.5 as a sub-agent for Opus 5.5 and Sonnet 5.5, a multi-agent collaboration paradigm of "flagship planning, small-model execution" has clearly emerged. This is not merely a supplement to the product line, but a laying of groundwork for the infrastructure of the agent era.