SpaceXAI shipped Grok Voice Transcribe 2.0 on 18 September 2026 with two claims stacked on top of each other: twice as accurate as version 1.0, and first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Both are defensible. Neither is the reason the release matters.

The reason it matters is the price line nobody put in a headline. Grok Voice Transcribe 2.0 costs $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming, unchanged from version 1.0. On the Artificial Analysis non-streaming board it posts 2.3% word error rate. That puts it in a two-model tie for the cheapest paid transcription on the board, alongside a model that is also more accurate than everything priced above it. The accuracy race in speech-to-text has been over for a while. The price race just ended too, and it ended at ten cents.

What SpaceXAI actually shipped

Grok Voice Transcribe 2.0 is a drop-in replacement. SpaceXAI says it is rolling into existing API integrations without code changes, version 1.0 will be deprecated in the coming weeks, and anyone who needs the old behaviour can pin grok-voice-transcribe-1.0 explicitly. The feature list is unusually complete for the price: word-level timestamps, speaker diarization at no additional cost, multichannel transcription across up to 8 channels, key term biasing at 100 terms per request, text formatting, filler word removal and smart turn detection. It handles dozens of languages with automatic detection, including mid-recording language switches.

The headline accuracy figure SpaceXAI chose to lead with is a word error rate on multilingual short phrases that fell from 20.6% to 6.8%. The company also says the model leads every tested system on telephony audio, and cites production use in customer support calls, video narration transcription and voice agents in Tesla vehicles. Pricing is confirmed independently on xAI's model documentation, which lists speech to text at $0.10 per hour for REST and $0.20 per hour for streaming. Worth noting for anyone wiring this today: that page carries the price but not yet a model ID for transcription, while the speech-to-speech entry grok-voice-think-fast-2.0 is listed in full.

The two numbers differ by nearly 3x, and both are real

SpaceXAI's headline number is 6.8%. The leaderboard number for the same model is 2.3%. Neither is wrong, and the gap between them is the most useful thing in the announcement.

Artificial Analysis computes AA-WER from a fixed blend: AA-AgentTalk at 50%, VoxPopuli-Cleaned-AA at 25% and Earnings22-Cleaned-AA at 25%, roughly 8 hours of audio spanning varied accents and domain language. Multilingual short phrases are not in that mix. SpaceXAI ran its own evaluation on the audio shape it cares most about, which is what voice agents actually receive, and reported the result honestly. The two figures describe the same model on different tests, and only one of them can be compared with anyone else's.

The "twice as accurate" framing sits between them. On the non-streaming Artificial Analysis board, version 1.0 scores 4.0% and version 2.0 scores 2.3%, which is a 1.7x reduction in errors. On SpaceXAI's own short-phrase test, 20.6% to 6.8% is a 3x reduction. Pick your benchmark and the same release is either a solid generational step or a transformation. This is the same measurement problem that showed up when four independent evaluations of one classifier were read side by side, where the spread on a single public benchmark reached 10.7 points.

Four bars showing 20.6 and 6.8 on one test and 4.0 and 2.3 on another for the same speech to text model
The same model, two tests: 20.6 to 6.8 on SpaceXAI's short-phrase evaluation, 4.0 to 2.3 on the Artificial Analysis blend.

Where Grok Voice Transcribe 2.0 actually lands

The first-place claim is specifically about the streaming board, which is a separate and smaller ranking of 32 models. On the non-streaming board, which compares 61 models in total, the picture is different. Prices below are the per-1,000-minute figures Artificial Analysis publishes, converted to cost per hour of audio.

ModelAA-WERPer 1,000 minPer hour
Fun-Realtime-ASR-preview1.7%$0.00$0.00
StepAudio 3 ASR1.7%$7.00$0.42
MAI-Transcribe-22.0%$1.67$0.10
Scribe v2 (ElevenLabs)2.2%$3.67$0.22
Grok Voice Transcribe 2.02.3%$1.67$0.10
Gemini 3.5 Transcribe2.6%$5.00$0.30
GPT Transcribe3.3%$4.50$0.27
Deepgram Nova-35.2%$4.30$0.26

Fifth on one board, first on another, and both statements were true on the same day. That is not a vendor being slippery. It is a category where streaming and batch are genuinely different problems, and the board you happen to read decides the winner. Microsoft ran into the same thing: when it launched MAI-Transcribe-2 on 3 September it said the model ranked second on the Artificial Analysis word error rate leaderboard. It now sits third at 2.0%, behind two entries tied at 1.7%. Nothing about the model changed. The board did.

Five bars showing word error rates of 1.7, 1.7, 2.0, 2.2 and 2.3 on the non-streaming leaderboard
Fifth of 61 on the non-streaming board at 2.3%, behind two entries tied at 1.7%.

Ten cents is now the accuracy leaders' price

Read that table by price instead of by rank and the useful pattern appears. The two cheapest paid models on the board, MAI-Transcribe-2 and Grok Voice Transcribe 2.0, are both $1.67 per 1,000 minutes, and they are third and fifth on accuracy. Every model priced above them is less accurate than at least one of them. Gemini 3.5 Transcribe costs 3x as much and scores 2.6%. GPT Transcribe costs 2.7x as much and scores 3.3%. Deepgram Nova-3 costs 2.6x as much and scores 5.2%, more than double the error rate of the models that cost a third as much.

There is one asterisk on the tie. Microsoft describes its $0.10 as a limited-time offer running until the end of the year, while SpaceXAI simply lists $0.10 as the price and points out it did not change between versions. Those are different commitments to the same number, and by January only one of them is guaranteed to still be there.

The move happened fast. Meta Muse Voice Transcribe claimed first place on the Artificial Analysis streaming board on 1 September at 3.1% word error rate, priced at $3 per 1,000 audio-minutes, roughly $0.18 an hour, and that looked aggressive at the time. Seventeen days later the streaming crown had moved and the batch floor had halved. For anyone transcribing at volume, a 100-hour month costs $10 at the floor against $30 on Gemini 3.5 Transcribe, and the cheaper option has the lower error rate.

Three stepped platforms engraved with $0.10, $0.22 and $0.30 per hour of audio
Cost per hour of audio: the accuracy leaders now sit on the lowest step.

Why utterance length decides who wins

The single biggest variable in these comparisons is not the model. It is how long the utterances in the test are, and almost nobody discloses it in a headline.

Wispr published the clearest evidence a day before this launch. Its Canto model scored 3.4% word error rate on an evaluation of real dictations, the best of the six models it tested. On short dictations, the one and two word utterances people fire off all day, the same model scores 21.4%. That is a 6x swing from utterance length alone, and two of the models Canto beat land on 21.4% as well, meaning the ranking that held on long audio collapses into a three-way tie on short audio.

This is exactly why SpaceXAI's 6.8% deserves more attention than its 2.3%. Short multilingual phrases are the hardest commonly-measured category in speech recognition, and they are precisely what a voice agent, a command interface or a live caption feed has to handle. A model at 6.8% on that category is doing something genuinely difficult. It just cannot be laid next to somebody else's 3.1% on clean long-form audio and called better.

Microsoft's numbers show the same effect from the other direction. MAI-Transcribe-2 posts 2.0% on the Artificial Analysis blend but 5.2% average on the FLEURS benchmark across 60 languages, because FLEURS spreads across languages where far less training data exists. Same model, two honest numbers, a 2.6x difference.

Two blocks showing 3.4 word error rate on long dictation against 21.4 on short utterances
One model, 3.4% on real dictation and 21.4% on one and two word utterances.

How to pick one this week

Run your own five-minute test before trusting any of these figures, because the only benchmark that matters is your audio.

  1. Cut a 10-minute sample from real material. Not clean studio narration. Use the messiest thing you actually process: an interview with crosstalk, a call recording, a voice memo with background noise, whatever dominates your pipeline.
  2. Include the short stuff deliberately. If any part of your workflow captures commands, names, prices or one-word answers, put a few minutes of that in the sample. This is the category that separates the models, and it is the one the leaderboards under-weight.
  3. Run the same file through three endpoints. Grok Voice Transcribe 2.0 and MAI-Transcribe-2 at $0.10 an hour, plus whichever model you use now. Ten minutes of audio costs under two cents on each, so there is no reason to test fewer than three.
  4. Diff the transcripts rather than reading them. Correct one by hand, then diff the others against it. Count proper nouns, numbers and technical terms separately from ordinary words; those are the errors that cost you editing time, and a single wrong price or name does more damage than ten wrong articles.
  5. Price the whole job, not the model. If you need diarization, check whether it is bundled. SpaceXAI includes speaker diarization at no extra cost and Meta reports a 17.5% diarization error rate on its own model, which matters more than the transcription figure when you are labelling a three-person interview.

What to check before you switch

Three things will decide whether the ten-cent tier survives contact with your pipeline. First, the deprecation clock: version 1.0 goes away in the coming weeks, so if you are already on SpaceXAI and depend on its exact output formatting, pin the old model now and diff before the switch is made for you. Second, the price commitment, which is open-ended at SpaceXAI and explicitly time-limited at Microsoft. Third, streaming versus batch, because the $0.20 streaming rate is double the batch rate at SpaceXAI, and a workflow that could tolerate a few minutes of latency is paying twice as much for nothing.

Frequently asked questions

Is Grok Voice Transcribe 2.0 actually the most accurate speech-to-text model?

It is first among 32 models on the Artificial Analysis streaming board according to SpaceXAI, and fifth of 61 on the non-streaming board at 2.3% word error rate. Both are accurate descriptions of different rankings. On the non-streaming board it trails Fun-Realtime-ASR-preview and StepAudio 3 ASR at 1.7%, MAI-Transcribe-2 at 2.0% and ElevenLabs Scribe v2 at 2.2%.

Why does SpaceXAI say 6.8% when the leaderboard says 2.3%?

They measure different audio. The 6.8% figure is SpaceXAI's own evaluation on multilingual short phrases, down from 20.6% on version 1.0. AA-WER is a fixed blend of AA-AgentTalk, VoxPopuli-Cleaned-AA and Earnings22-Cleaned-AA across roughly 8 hours, which contains no multilingual short-phrase set. Short utterances produce far higher error rates in every model, so the two numbers are not comparable and neither is inflated.

What does Grok Voice Transcribe 2.0 cost?

$0.10 per hour of audio for batch transcription and $0.20 per hour for streaming, the same as version 1.0. Speaker diarization, word-level timestamps, multichannel up to 8 channels and key term biasing are included rather than billed separately.

How does it compare with MAI-Transcribe-2 on price?

They are identical at $1.67 per 1,000 minutes, or $0.10 an hour. MAI-Transcribe-2 is slightly ahead on AA-WER at 2.0% against 2.3%, but Microsoft describes its price as a limited-time offer until the end of the year while SpaceXAI lists $0.10 as its standing price.

Do I need to change my code to use version 2.0?

No. SpaceXAI says the model is rolling into existing API integrations without code changes and will become the default, with version 1.0 deprecated in the coming weeks. If you need the old model, pin grok-voice-transcribe-1.0 before that happens.

Which model should I use for short commands and voice agents?

Test rather than trust a leaderboard. Short utterances are where rankings collapse: Wispr Canto leads at 3.4% on real dictation but scores 21.4% on one and two word utterances, tied with two models it beats everywhere else. SpaceXAI's 6.8% on multilingual short phrases is the only published figure in this group measured on that category specifically.