The two lines in the price list

xAI's public pricing page currently carries both versions of its speech-to-speech model side by side, which is the clearest way to see what is about to happen. Grok Voice Think Fast 1.0 is listed at 0.05 dollars a minute of audio, or 3.00 dollars an hour, plus 0.004 dollars for text input. Grok Voice Think Fast 2.0 is listed at 0.08 dollars a minute, or 4.80 dollars an hour, with the same text input charge. Same product category, same vendor, same page, a 60 percent gap between two rows.

The release note dated 29 July is where the date lives. xAI records that grok-voice-think-fast-2.0 is now available with Speech to Speech and that grok-voice-latest will route to this model starting 5 August 2026. The company's own migration guidance is that no action is needed to upgrade, and that anyone wishing to stay on the older model should pin grok-voice-think-fast-1.0 before then. Read those two sentences together and the shape of the week becomes clear: the cheaper option still exists, and it is not the one you get by leaving your code alone.

What the upgrade actually improves

None of this is a case of a vendor charging more for less. On the independent Artificial Analysis speech-to-speech benchmark, Think Fast 2.0 scored 82.9 percent against 75.7 percent for version 1.0. Time to first audio, the pause a caller hears before the model begins speaking, fell from 1.25 seconds to 0.70. Median relative reasoning token use dropped to 0.4 times that of the predecessor, because the new model reasons in parallel with speech rather than thinking first and talking afterwards. Transcription accuracy improved as well. By every published measure this is a better model that costs its maker less to run.

That last clause is the interesting one. A model that burns 60 percent fewer reasoning tokens and answers in roughly half the time is cheaper to serve, and in the text market that fact reaches the customer within weeks. It reached OpenAI's customers on 30 July, when the company cut GPT-5.6 Luna by 80 percent and Terra by 20 percent and attributed the move to a fall in its own serving costs. Here the same class of improvement arrived attached to a 60 percent increase. Nothing about that is dishonest. It is a straightforward consequence of the unit the meter counts.

Why a better model costs more here

Text and voice are sold on different clocks, and the difference decides who keeps the efficiency. Text is billed per token, which is a measure of work done. When a model needs fewer tokens to reach the same answer, your invoice falls without anyone deciding that it should, because you are buying units of computation and you now need fewer of them. Voice is billed per minute of audio, which is a measure of time elapsed. A caller explaining a delivery problem takes the same ninety seconds whether the model behind the line is fast or slow, so the number of billable units is set by human speech, not by model efficiency.

The consequence is that voice is the one line in an AI budget that structurally cannot benefit from the efficiency race, and it is also the line that scales with your customers rather than with your engineering. Every improvement your vendor makes to a per-token product eventually shows up as a smaller bill. Every improvement to a per-minute product shows up as a better product at whatever price the vendor sets, and the saving stays where it was earned. An owner watching model prices collapse all year could reasonably assume that all AI is deflating. The meter, not the model, decides whether that is true for a given workload.

The arithmetic on your own call volume

Put real numbers against it, because the percentage understates how quickly this compounds. A four minute call costs 20 cents on version 1.0 and 32 cents on version 2.0. Ten thousand such calls a month is 40,000 minutes, which is 2,000 dollars against 3,200 dollars, a difference of 1,200 dollars a month or 14,400 dollars a year, arriving from a routing change rather than from any decision you made. For a European operator the rate is quoted in dollars while the revenue behind those calls is in euros or pounds, so the increase lands on a cost line that is already exposed to a currency you do not invoice in.

It is worth checking whether the speed gain pays for any of this, and it does not come close. Suppose the faster start somehow removed a full ten seconds from that four minute call, far more than the 0.55 second improvement in time to first audio would plausibly deliver. Ten seconds at the new rate is worth about 1.3 cents, against 12 cents of additional cost on the same call. The generous version of the argument returns roughly a ninth of the increase. Buying the better model is defensible on quality and on the caller experience. It is not a cost-saving move, and any business case that presents it as one has confused a faster model with a cheaper one.

Pin it or budget for it

There are only two honest positions before Wednesday and both require an action this week. If your voice agent is doing routine, scripted work where 75.7 percent on a benchmark was already sufficient, pin grok-voice-think-fast-1.0 in your configuration and keep the 3.00 dollar hourly rate. If callers regularly hit the limits of the older model, take 2.0 deliberately, reforecast the voice line at 4.80 dollars an hour of audio, and record in the budget note that the increase bought latency and accuracy rather than savings. What you cannot defend afterwards is discovering the change in a September invoice.

The wider habit worth forming is to read the unit before the price. Ask of every AI service in your stack what the meter counts, because that single fact predicts whether vendor efficiency will ever reach you: per token and per request bills fall as models improve, while per minute, per seat and per conversation bills do not. Then find the aliases. Any identifier ending in latest is a standing instruction to accept whatever the vendor ships next, including its price, and the only workloads where that is the right setting are the ones where you have decided in advance that you do not care what the next version costs.