The Number That Matters

Meta Superintelligence Labs launched Muse Voice Transcribe on September 1, and it took the top spot on Artificial Analysis's independent streaming speech-to-text leaderboard the same day.

The model merges two jobs that most transcription pipelines have run as separate stages: speaker diarization, which tells you who is talking, and endpointing, which tells you when they've stopped. Meta trained it on more than 70 languages, with 25 validated at launch, and built it to handle 20 or more simultaneous speakers, recordings that run past an hour, and speakers who switch languages mid-conversation without triggering a language reset. Access runs through Meta's Model API at USD 3 per 1,000 audio-minutes, a rate that works out to USD 0.18 per hour of audio transcribed.

What Businesses Have Been Paying

Most companies running call transcription, meeting notes, or customer-support audio through an API today are paying for two things at once, even if the invoice doesn't spell it out: the transcript itself, and the extra processing that turns it into something a manager can actually use.

OpenAI's Whisper API has publicly listed transcription-only pricing around USD 0.006 per minute, which works out to roughly USD 0.36 per hour, with diarization typically handled as a separate step or a different tool entirely. Azure Speech and Google Speech-to-Text price their standard tiers in a broadly similar range, often landing near USD 1 per hour before speaker identification gets added as its own feature or its own pricing tier. Every one of those setups asks a business to pay twice and integrate twice for a result Meta now ships as one model at one price.

ServicePrice per 1,000 minutesPrice per hourDiarization included
Meta Muse Voice TranscribeUSD 3.00USD 0.18Yes, built in
OpenAI Whisper API (transcription only)~USD 6.00~USD 0.36No, separate step
Typical cloud speech-to-text standard tier~USD 15-24~USD 0.90-1.44Often a separate add-on

The EUR Math, Approximately

Convert Meta's per-hour rate into euros and the figure an EU or UK finance team can actually hold up against an invoice lands close to EUR 0.17 per hour.

That conversion uses a rough dollar-to-euro rate near 0.92 and it'll drift as currency markets move, so treat it as a ballpark for comparison, not a number to put on a purchase order. The same math puts the per-1,000-minute rate at roughly EUR 2.76, or about GBP 2.40 for teams billing in sterling. Any business currently paying meaningfully more than that for transcription, diarization, and endpointing combined has, for the first time, a public reference price to hold its own vendor contract against rather than just a sales rep's quote.

Buy, Build, or Switch

The decision facing any EU or UK team running audio pipelines is not whether Meta's model is technically impressive, it is whether the current vendor stack still earns its markup now that diarization and endpointing no longer require separate tools to stitch together.

The model is already load-bearing inside Meta's own products, powering Mac dictation features and the company's Muse Code tool, which is a signal that the system has been hardened in production rather than tuned for a leaderboard demo. For a team paying a premium to combine transcription, speaker labels, and turn-detection from three different vendors, that combination of a public price and a live production workload is the part worth acting on. The Mac dictation feature is simply where most people will notice the underlying model first.

Why We Do This

We do this for everyone trying to keep up with what technology is doing to our lives. The people who build it, and the people it happens to. The Servola Journal exists so that what we learn belongs to all of them.

Nobody pays us for this. No ads, no paywall, free to everyone. We just believe that understanding what's happening to all of us shouldn't depend on who can afford to pay for it.

If it gave you something today, tell us to keep going. Follow us, leave a like, or write a positive comment. We read every one, and they are what keeps us going.