Anthropic launches Claude Opus 4.8: Focus on agent reliability and knowledge work accuracy

5.29 As an incremental upgrade to Opus 4.7, the new model focuses on strengthening the reliability of agent programming, the stability of multi-step reasoning and the accuracy of knowledge work while keeping the price unchanged, and significantly reduces the incidence of hallucinations and unfounded conclusions. rate

On May 29, AI security pioneer Anthropic officially released its flagship model Claude Opus 4.8. As an incremental upgrade to Opus 4.7, the new model focuses on strengthening the reliability of agent programming, the stability of multi-step reasoning and the accuracy of knowledge work while keeping the price unchanged, and significantly reduces the incidence of "illusions" and unfounded conclusions.


More reliable and sharper: comprehensive optimization of agent behavior

Compared with previous generations,Although the update range of Opus 4.8 is small, it directly hits the core pain point of the current implementation of AI agents-unreliability. According to feedback from Anthropic officials and many early testers, Opus 4.8 is "more reliable and has sharper judgment" in complex, multi-step tasks.


Specifically:

Actively question unreasonable plans: When user instructions have logical loopholes or vague goals, the model will proactively ask questions or raise objections instead of blindly implementing them;

Identify and correct your own mistakes: You can trace the context in a long task chain, discover inconsistencies and correct themselves;

Increase in code defect annotation rate: the probability of allowing the code you write to have defects without explaining them is reduced to one-quarter of Opus 4.7;

Reduce unfounded conclusions: Be more willing to proactively flag uncertainties and avoid"Speak confidently."

This series of improvements stems from Anthropic continues to deepen its "alignment" technology. In terms of pro-social indicators, Opus 4.8 has hit new highs in dimensions such as "supporting user autonomy" and "acting in the best interests of users"; at the same time, the incidence of mismatch behaviors such as deception and manipulation has been further reduced, approaching the level of its cutting-edge research model Claude Mythos Preview.


2026052903_20260529_17800471614251740


Double optimization of performance and cost: fast mode speed up 2.5 Times, the cost is reduced to 1/3

Anthropic simultaneously announced performance and pricing changes for Opus 4.8:

Fast Mode (Fast Mode: Reasoning speeds up to 2.5 times faster for response time-sensitive scenarios;

Model cost: overall reduction to previous model 1/3, providing an economic basis for large-scale deployment.


Maintenance of a dual-track pricing system:

General Mode: Input $5 /million tokens, output $25 /million tokens (same as Opus 4.7);

Fast Mode: Input $10 /million tokens, output $50 /million tokens.

Furthermore,claude.ai has added an effort level control function that allows users to switch between high, extra (xhigh in Claude Code), max and other gears to flexibly balance quality and cost. The default high gear consumes similar token consumption in coding tasks to Opus 4.7, but the effect is better; higher gear increases computing investment in exchange for extreme output.

Benchmarking:SWE-Bench Pro reaches 69.2%, surpassing GPT-5.5 and Gemini 3.1 Pro


In authoritative evaluations,Opus 4.8 performed well:

SWE-Bench Pro (Software Engineering Benchmark): Score 69.2%, surpassing OpenAI GPT-5.5 and Google Gemini 3.1 Pro;

Multi-domain knowledge reasoning: In professional scenarios such as law, finance, and scientific research, factual consistency and logical rigour are significantly improved.

However, in terminal programming (In specific tasks such as Terminal-based coding, GPT-5.5 still maintains the lead, indicating that the differentiation of advantages between different models in segmented areas is becoming increasingly obvious.


2026052905


Industry background: AI has entered the "era of responsibility" and reliability has become the focus of competition

The release of Opus 4.8 comes at a critical juncture as the global AI industry shifts from a "ability race" to a "responsibility race." As GitHub Copilot fully shifts to usage billing and companies deeply embed AI into production processes, model reliability, interpretability and cost controllability have become the core criteria for selection.

Anthropic's focus on seemingly "conservative" improvements such as "reducing illusions" and "proactively labeling uncertainty" is actually an accurate response to high-responsibility scenarios (such as corporate automation, medical assistance, and financial analysis). In the process of AI changing from a "toy" to a "tool", stability is more important than surprise, and honesty is more precious than intelligence.