What Microsoft shipped, and where it landed

Microsoft released MAI-Transcribe-2 on September 3, 2026, and it now ranks first on FLEURS, the standard multilingual speech benchmark covering 60 languages, with a 5.2 percent average word-error rate. On the separate Artificial Analysis leaderboard it placed second.

The model adds speaker diarization, word-level timestamps, keyword biasing, configurable verbatim or clean transcription styles, mid-sentence code-switching, and automatic language identification. It is available through Microsoft Foundry, the MAI Playground, and Open Router.

The number that matters is the price, not the score

Microsoft is charging $0.10 per hour of audio processed, a limited-time offer running through the end of 2026. A Microsoft statement called it 'not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.'

A benchmark ranking is a claim any lab can dispute with its own numbers next quarter. A price this low, from a company with Microsoft's distribution, is harder to walk back once customers have built workflows around it.

The speed comparison, in Microsoft's own numbers

Microsoft's published speed claims, compared against the three transcription models it names directly, look like this:

ModelVendorSpeed vs MAI-Transcribe-2
MAI-Transcribe-2MicrosoftBaseline
GPT-TranscribeOpenAI10 times slower
Scribe v2ElevenLabs7 times slower
Gemini 3.5 TranscribeGoogle5 times slower

These are Microsoft's own benchmark comparisons rather than an independent third-party test, so a business evaluating a switch should still run its own sample audio through each option before committing.

Why this is a margin question, not just a model question

Plenty of European tools sell voice-to-text as a bundled seat feature inside a call-recording, meeting-notes, or customer-support product, priced well above what raw model access now costs. That gap used to be the cost of decent accuracy; increasingly, it is just the vendor's markup.

A business paying a monthly per-seat transcription fee has a straightforward number to test against now: multiply its actual monthly audio-hours by $0.10, and compare that to what it is billed. A large gap is worth a conversation with the vendor, not an assumption that the fee still reflects the underlying cost.

The catch: a promotional price, not a floor

Microsoft's own announcement frames $0.10 per audio-hour as available 'through the end of 2026,' not as a permanent list price. Pricing wars in AI infrastructure have moved this fast before, and vendors that cut prices to win share have raised them again once the share was won.

Any contract or internal budget built around this price this quarter should assume it may not hold into next year, and should include a re-pricing or exit clause rather than treating the promotional rate as the baseline for a multi-year plan.