What Microsoft shipped, and where it landed
Microsoft released MAI-Transcribe-2 on September 3, 2026, and it now ranks first on FLEURS, the standard multilingual speech benchmark covering 60 languages, with a 5.2 percent average word-error rate. On the separate Artificial Analysis leaderboard it placed second.
The model adds speaker diarization, word-level timestamps, keyword biasing, configurable verbatim or clean transcription styles, mid-sentence code-switching, and automatic language identification. It is available through Microsoft Foundry, the MAI Playground, and Open Router.
The number that matters is the price, not the score
Microsoft is charging $0.10 per hour of audio processed, a limited-time offer running through the end of 2026. A Microsoft statement called it 'not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.'
A benchmark ranking is a claim any lab can dispute with its own numbers next quarter. A price this low, from a company with Microsoft's distribution, is harder to walk back once customers have built workflows around it.
The speed comparison, in Microsoft's own numbers
Microsoft's published speed claims, compared against the three transcription models it names directly, look like this:
| Model | Vendor | Speed vs MAI-Transcribe-2 |
|---|---|---|
| MAI-Transcribe-2 | Microsoft | Baseline |
| GPT-Transcribe | OpenAI | 10 times slower |
| Scribe v2 | ElevenLabs | 7 times slower |
| Gemini 3.5 Transcribe | 5 times slower |
These are Microsoft's own benchmark comparisons rather than an independent third-party test, so a business evaluating a switch should still run its own sample audio through each option before committing.
Why this is a margin question, not just a model question
Plenty of European tools sell voice-to-text as a bundled seat feature inside a call-recording, meeting-notes, or customer-support product, priced well above what raw model access now costs. That gap used to be the cost of decent accuracy; increasingly, it is just the vendor's markup.
A business paying a monthly per-seat transcription fee has a straightforward number to test against now: multiply its actual monthly audio-hours by $0.10, and compare that to what it is billed. A large gap is worth a conversation with the vendor, not an assumption that the fee still reflects the underlying cost.
The catch: a promotional price, not a floor
Microsoft's own announcement frames $0.10 per audio-hour as available 'through the end of 2026,' not as a permanent list price. Pricing wars in AI infrastructure have moved this fast before, and vendors that cut prices to win share have raised them again once the share was won.
Any contract or internal budget built around this price this quarter should assume it may not hold into next year, and should include a re-pricing or exit clause rather than treating the promotional rate as the baseline for a multi-year plan.
Read next: Microsoft Bets 2.5 Billion That Tools Are Not Enough | Microsoft Now Caps What Its Own Engineers Spend on AI



