Skip to main content
Product Updates

Elon Musk’s SpaceXAI’s Most Intelligent AI Voice Model, Grok Voice Think Fast 2.0 Debuts

Elon Musk’s SpaceXAI company released Grok Voice Think Fast 2.0 yesterday (July 29), its most powerful and intelligent speech-to-speech AI model to date.

Elon Musk’s SpaceXAI’s Most Intelligent AI Voice Model, Grok Voice Think Fast 2.0 Debuts

Elon Musk’s SpaceXAI company released Grok Voice Think Fast 2.0 yesterday (July 29), its most powerful and intelligent speech-to-speech AI model to date.

In terms of pricing, the model costs $0.08 per minute of audio (note: approximately 0.54 yuan at the current exchange rate). The company plans to upgrade the grok-voice-latest model on August 5, from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0.

The model focuses on improving speech reasoning, transcription accuracy, conversational capabilities, and tool-calling reliability. According to speech-to-speech benchmark data published by Artificial Analysis, the new model’s AA Speech-to-Speech Quality Index is 82.9%, an increase of 7.2 percentage points from Grok Voice Think Fast 1.0’s 75.7%.

Elon Musk’s SpaceXAI’s Most Intelligent AI Voice Model, Grok Voice Think Fast 2.0 Debuts
Elon Musk’s SpaceXAI’s Most Intelligent AI Voice Model, Grok Voice Think Fast 2.0 Debuts
Elon Musk’s SpaceXAI’s Most Intelligent AI Voice Model, Grok Voice Think Fast 2.0 Debuts

SpaceXAI said the model can perform reasoning simultaneously while speaking, without adding extra latency to gain reasoning capabilities. The company also said the new version reduces the number of reasoning tokens required for each response; in production environments, tool calls are typically completed before the agent finishes speaking its first sentence.

In an internal evaluation covering thousands of phrases in 24 languages, xAI said Grok Voice Think Fast 2.0’s transcription performance was 1.5–2.0 times better than that of Deepgram Nova 3 and ElevenLabs Scribe v2, and 1.4 times better than version 1.0. In scenarios with significant background noise and telephone compression, xAI said its performance gap over dedicated speech-to-text models could widen to approximately 10 times.