Claude Fable 5 developers measured-80.3% SWE-Bench Pro benchmark verification,"super long autonomous work" ability sparked heated discussions in the programming community

6.11 The ability to "long autonomous work" is particularly critical. It allows AI to transition from "assistive tools" that require human confirmation every time to "autonomous agents" that work independently for a long time.

Anthropic released the Claude Fable 5 model on June 9. After two days of developer testing, its 80.3% SWE-Bench Pro benchmark has been verified in actual coding tasks. Developers report significant improvements over Opus 4.8 in complex multi-file refactorings, autonomous debugging sessions, and large codebase analysis. The model's "long autonomous work" ability (which can continue to work for longer periods of time without human input) has become the biggest highlight.

 


Core technical features of Fable 5

 

CharacteristicsDescriptionDeveloper feedback
SWE-Bench Pro80.3%(highest in the industry)The accuracy of multi-file reconstruction has been significantly improved
Over-long autonomous workLonger continuous work than any predecessor Claude modelReduced frequency of long session interruptions
secure routingCybersecurity/biology queries automatically transfer to Opus 4.8The trigger rate was less than 5%, in line with expectations
free periodPro/Max/Team/Enterprise plan is free until June 22Developers actively try
pricingPoints are required from June 23Specific prices have not yet been announced



 

Key considerations for API developers

 

There are important changes in Fable 5's API behavior: When the model refuses to answer for security reasons, it returns an HTTP 200 status code (instead of an error code) with a field called 'stop_reason: "refuse". If no output is generated, no charge will be paid. Developers need to handle this "soft rejection" scenario gracefully.

 


"Programming Three Kingdoms Kill" with GPT-5.5 and Gemini 3.5 Pro

 

The release of Fable 5 brings the AI programming model competition into a new stage:


modelCompanySWE-Bench ProCore advantage
Claude Fable 5Anthropic80.30%Long-term independent work, corporate safety and compliance
GPT-5.5OpenAI~58.6%The largest developer ecosystem, Codex integration
Gemini 3.5 ProGoogleundisclosedDeep integration with Google Cloud for cost advantages


  

 

Claude Fable 5's 80.3% SWE-Bench Pro verification is a "new ceiling" for AI programming capabilities. While GPT-5.5 is still hovering around 58%, Anthropic has pushed its benchmark to 80%-which means AI can solve more than 4/5 of real software engineering problems autonomously. The ability to "long autonomous work" is particularly critical: it allows AI to transition from "assistive tools"(requiring human confirmation each time) to "autonomous agents"(working independently for a long time). But feedback from the developer community also reveals risks: Fable 5's "soft rejection" mechanism protects security, but also increases the processing complexity of developers; and the pricing after the free period ends on June 22 will determine whether it can transform from "technological leadership" to "commercial success." For Anthropic, Fable 5 is not only a model release, but also a key bargaining chip to showcase the "technology moat" to investors on the eve of its IPO.