Musk Strikes Back! SpaceXAI Releases Grok 4.6: Matches GPT-5.6 Sol in Composite Score at Half the Price

8.13 On August 12, SpaceXAI officially released its next-generation large model, Grok 4.6. In the Artificial Analysis Intelligence Index, Grok 4.6 scored 61, tying OpenAI's GPT-5.6 Sol Max for third place globally, behind only Anthropic's Claude Opus 5 and Fable 5 Max. This release marks a significant milestone, firmly establishing Musk's AI venture among the global frontier model leaders.

On August 12, SpaceXAI officially released its next-generation large model, Grok 4.6. In the Artificial Analysis Intelligence Index, Grok 4.6 scored 61, tying OpenAI's GPT-5.6 Sol Max for third place globally, behind only Anthropic's Claude Opus 5 and Fable 5 Max. This release marks a significant milestone, firmly establishing Musk's AI venture among the global frontier model leaders.

2026081301

I. Performance Leap: Agent and Coding Capabilities Lead the Charge

Compared to the July release of Grok 4.5, the new model shows substantial gains across key dimensions. Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, a 5-point increase from the previous 56, and tied GPT-5.6 Sol Max across 9 comprehensive benchmarks.

The model's strengths lie in long-horizon agent tasks and complex programming. On the DeepSWE v1.1 benchmark for software engineering, it scored 65.9% , significantly outperforming Grok 4.5's 54%; on Terminal-Bench v3.0, it achieved 26% , up from 15.7%.

In knowledge work, Grok 4.6 achieved an Elo score of 1753 on GDPval-AA v2, surpassing GPT-5.6 Sol Max (1728) and Claude Fable 5 (1741); on AA-Briefcase, it scored 1577 Elo, second only to Claude Opus 5.

2026081302

II. Targeted Optimization: From "Answering Questions" to "Completing Projects"

The core upgrade focuses on enabling the model to handle multi-step tasks rather than just answering questions. SpaceXAI stated Grok 4.6 underwent extended training with curated model-generated reasoning data and high-quality engineering data.

Training leveraged Grok 4.5 to generate SFT trajectories across reasoning intensities and domains like STEM and software engineering, with model-based checks to filter problematic trajectories. Extensive reinforcement learning covered general programming, knowledge work, kernel optimization, web development, and CAD.

III. Pricing Strategy: Flagship Performance at "Half the Price"

Pricing is another key differentiator. API pricing starts at $2 per million input tokens** and **$6 per million output tokens, less than half the price of GPT-5.6 Sol.

Investor Gavin Baker noted Grok 4.6's performance is roughly equivalent to Fable 5 Max, but costs 80% less for input and 88% less for output. However, actual task costs require comprehensive evaluation.

IV. Availability and Safety

Grok 4.6 is now available on Cursor and Grok Build, via API, and through partners like OpenRouter, Vercel, and Cloudflare. A first-week promotion doubles usage quotas for Cursor and Grok Build users.

Safety assessments, the largest conducted by SpaceXAI to date, cover pre-deployment capability testing, safety calibration, and third-party testing.

Grok 4.6's release signals SpaceXAI's strategy of competing on the performance-cost curve—excelling in specific tasks rather than pursuing a one-size-fits-all model. As capabilities converge, "cost per unit of intelligence" becomes key. Grok 4.6 proves that "good enough + cheap enough" remains effective, sending a clear signal to the developer market.