The competition in voice AI is shifting from "hearing and speaking" to "reasoning and taking action." On July 29, SpaceXAI under Elon Musk released the next-generation speech-to-speech model Grok Voice Think Fast 2.0, redefining the standard for real-time voice interaction with a 0.7-second first-response time and a top-ranked agentic performance.
Performance Leap: 82.9% Overall Score, #1 in Agentic Capability
According to speech-to-speech benchmark data from Artificial Analysis, Grok Voice Think Fast 2.0 achieved an AA Speech-to-Speech Quality Index score of 82.9% , a 7.2 percentage point improvement over the previous 1.0 version's 75.7%. On the Tau Voice benchmark measuring agentic performance, the model ranked first with 56.5% , surpassing competitors including OpenAI's GPT-Realtime-2.1 High (45.7%) and Google's Gemini 3.1 Flash (37.7%).
On key sub-metrics, the new model performed exceptionally well:
Speech Reasoning (Big Bench Audio) : 97.2%
Conversational Dynamics (Full Duplex Bench) : 95.1% (up from 77.8% in the previous generation)
Response Speed (Time to First Audio) : 0.70 seconds
This means that in the Full Duplex Bench test measuring real-time multi-turn voice interaction, the new model jumped from 77.8% to 95.1%, approaching the level of natural human conversation.

0.7-Second Response: 44% Faster, 60% Less Inference Tokens
Response speed is the most tangible upgrade. The average time to first audio is just 0.70 seconds , a 44% reduction from the previous version's 1.25 seconds. It is the only model under 1 second among the top 5 on the Artificial Analysis leaderboard, faster than GPT-Realtime-2 High (1.14s) and GPT-Realtime-2.1 High (1.21s).
SpaceXAI states the new model supports "reasoning while speaking" —it can perform inference simultaneously with speech generation, without sacrificing latency-1. Compared to its predecessor, median inference token consumption has been reduced by approximately 60% (to 40% of the 1.0 version's usage), allowing tool calls to typically execute before the agent finishes its first sentence.
Transcription: 24 Languages, 1.5-2x Accuracy Improvement
For transcription, SpaceXAI's internal evaluations across thousands of phrases in 24 languages claim that Grok Voice Think Fast 2.0 outperforms Deepgram Nova 3 and ElevenLabs Scribe v2 by 1.5 to 2.0 times , and improves 1.4 times over the 1.0 version.
In scenarios with significant background noise and telephony compression, xAI claims the gap can widen to approximately 10 times compared to specialized speech-to-text models, demonstrating robust performance in real-world challenging environments.
Pricing and Availability: $0.08 per Minute, Default Upgrade on August 5
The model is priced at **$0.08 per minute of audio** (approximately $4.80 per hour of input audio), which is higher than the previous version's $3.00/hour but only about 45% of the price of GPT-Realtime-2.1 High.
SpaceXAI plans to automatically upgrade the grok-voice-latest model from version 1.0 to 2.0 on August 5. Developers who still need to use the old version should lock the 1.0 version identifier in advance.
Developer Feedback: "A Leap Forward"
Several developers shared their hands-on experiences on X. Nick White, CEO of AI company Helionova AI, said he has completed the full-stack real-time voice upgrade on his product Tradecraft, calling it "absolutely a leap-forward progress." In tests, the model showed almost no noticeable pauses during multi-turn real-time phone conversations, advancing the narrative and providing real-time feedback.
SpaceXAI software engineer Parker Conrad noted: "The entire voice team poured a lot of effort, thought, and passion into this. The model's intelligence, accuracy, and capabilities feel almost surreal."
Grok Voice is currently undergoing A/B testing in Starlink's phone service, delivering measurable improvements in sales conversion and customer service response efficiency.
As voice models evolve from "hearing and speaking" to "reasoning and taking action," the competitive arena has shifted from speech recognition accuracy to the comprehensive capabilities of conversational agents. Grok Voice Think Fast 2.0, with its low-latency "reasoning-while-speaking" architecture and top-ranked agentic performance, is pushing voice AI toward a new phase where it can "converse like a human and work like a human." With the default model switch on August 5, this capability will officially enter developer and enterprise production environments.