Microsoft AI has released MAI-Transcribe-2, a speech-to-text model it calls the fastest, most accurate, and cheapest in the world. Launched on September 3, 2026, it transcribes audio at $0.10 per hour, handles 60 languages, and ships with speaker diarization and word-level timestamps built in.

Try It: Caption a Video in Minutes

The fastest way to see what MAI-Transcribe-2 does is to run a clip through the MAI Playground. Upload an interview or a raw voiceover, choose "verbatim" or "clean" style, and you get back a transcript with per-word timestamps you can drop straight into a subtitle track. Because the model marks who is speaking, a two-person podcast comes back already split by speaker, which removes the tedious manual labeling step most captioning workflows still require.

Why It Matters for Creators

Transcription sits at the front of countless creator workflows: subtitles, show notes, searchable archives, and repurposing long video into clips. Price and speed have been the friction. At $0.10 an audio hour, a limited-time rate through the end of 2026, MAI-Transcribe-2 makes bulk transcription cheap enough to run across an entire back catalog, and its speed advantage means a feature-length recording comes back in minutes rather than a coffee break. Keyword biasing also lets you feed in names, product terms, or jargon so the model stops mangling the words that matter most to your niche. For creators comparing options, our guide to the best AI speech-to-text tools shows how the field has been shifting on both accuracy and cost.

Key Details

Accuracy: Ranks first on the FLEURS benchmark across 60 languages with an average word-error rate of 5.2 percent, and second on the Artificial Analysis WER leaderboard.

Speed: Microsoft reports it is 10 times faster than GPT-Transcribe, 7 times faster than ElevenLabs Scribe v2, and 5 times faster than Gemini 3.5 Transcribe.

Features: Speaker diarization, word-level timestamps, keyword biasing for domain terms, and code-switching for blended languages like Hinglish and Spanglish. See the full model page for language coverage.

What to Do Next

If you build transcription into an app or pipeline, MAI-Transcribe-2 is available through Microsoft Foundry and on OpenRouter for a provider-agnostic call. Benchmark it against your current tool on a representative sample of your own audio, including noisy and multilingual clips, before you switch a production workflow.