xAI launches Grok Voice Transcribe 2.0 with half the errors
xAI has launched Grok Voice Transcribe 2.0, a highly accurate speech-to-text model that slashes error rates at no extra cost, enabling developers to build more reliable voice applications.

Elon Musk's AI venture, xAI, has released Grok Voice Transcribe 2.0, a major upgrade to its speech-to-text model designed for both prerecorded and live audio. The new model immediately secured the top spot among 32 streaming speech-to-text models on the Artificial Analysis public leaderboard. Despite the performance boost, xAI is keeping its pricing unchanged at $0.10 per audio hour for batch processing and $0.20 per hour for streaming. This flat rate includes advanced features like speaker diarization, timestamps, and key-term biasing.
According to xAI's internal evaluations, Transcribe 2.0 makes roughly half as many errors as its predecessor across real-world production datasets, including customer-support calls, spoken credentials, and short voice commands. The most dramatic improvement occurred in multilingual short phrases, where the word error rate plummeted from 20.6% to 6.8%—a 67% relative reduction. This makes the model highly effective at processing brief inputs like vehicle commands, even when the audio contains background noise, compression artifacts, or low-quality 8 kHz call-center signals.
For developers and practitioners, the upgrade offers a seamless transition. Existing API integrations will automatically transition to the new model unless they explicitly pin their integration to the older grok-voice-transcribe-1.0 version. The API provides robust controls for real-time voice applications, supporting multichannel transcription for up to eight channels, turn detection, filler-word removal, and key-term biasing for up to 100 domain-specific terms per request.
Atlassian has signed on as the launch customer, using Transcribe 2.0 to transcribe Loom videos. In practice, Atlassian feeds these transcripts directly into the Cursor coding agent to create a record-to-code workflow, where high accuracy on technical terms and variable names is critical. By outperforming rival systems like ElevenLabs Scribe v2, Deepgram Nova-3, Google Chirp 3, Azure Speech to Text, AssemblyAI Universal-3.5, and OpenAI's live transcription, xAI's latest release establishes a highly competitive option for enterprise voice pipelines.
This is our own summary of reporting by AlphaSignal



