Four Days From Launch to Sold Out
Moonshot AI released Kimi K3 on 16 July, a model of roughly 2.8 trillion parameters with a one-million-token context window. On 20 July the company published a notice to its users describing compute capacity constraints and a subscription suspension. Requests over the first 48 hours, it said, had surged far beyond its projections and brought its existing compute cluster close to maximum capacity.
The remedy was to stop selling. Moonshot said it had decided to immediately suspend new consumer subscriptions and would dedicate all available computing resources to current subscribers so that their service remained unaffected, while it expanded infrastructure and reopened places in batches. This is not an outage and not a price rise. It is a vendor declining new revenue because it cannot serve it.
The Rationing Rule Is the Part Worth Reading
When supply ran short, the tiebreaker was tenure, not willingness to pay. Existing subscribers kept their benefits intact. Prospective customers, including any European team that spent the weekend evaluating K3 and planned to commit on Monday, were simply told no. In a market where buyers are used to capacity being a solved problem behind a credit card, that is an unfamiliar shape of failure.
It is also a defensible choice, and worth crediting as such. Degrading service for paying users in order to onboard more of them would have been the worse decision. But the consequence for a buyer is concrete: an account opened early is not merely a cheaper account, it is a claim on scarce supply. That is a property of the purchase that no benchmark table records.
A Benchmark Assumes You Can Buy the Thing
Two days before the suspension, independent evaluation had placed K3 at the top of a widely watched frontend coding leaderboard, ahead of the best closed models. That result was real and it was useful. It also silently assumed something that turned out to be false, which is that a team persuaded by the number could then act on it.
The practical correction is small and cheap. When you evaluate a hosted model, record two more things beside price and score: what the provider says about capacity headroom, and what it has actually done when demand exceeded supply. Moonshot has now given a public answer to the second question, which is more than most vendors have. Treat that transparency as information, not as a mark against it.
What Changes on 27 July
Moonshot has said the full K3 weights are due to be published on 27 July. If that holds, the constraint moves. A model you can host is limited by the accelerators you can rent or own, not by another company's queue, and for a European operator that also moves the data-residency question onto ground you control.
The economics change rather than disappear. K3 has been priced at roughly 3 dollars per million input tokens and 15 dollars per million output, about 13 euros for the latter, and self-hosting a model in this size class is not a casual undertaking. The honest framing is that 27 July converts a queuing problem into a capital and engineering problem. For some workloads that is a better problem. For most it is simply a different one.
What to Do Before Your Next Model Commitment
Configure and periodically test a second provider for any workflow that a model now sits inside. The test matters more than the configuration, because a fallback nobody has exercised is a plan, not a capability. Keep the switch at the level of your own code rather than relying on a vendor abstraction you would also lose.
Then decide deliberately whether you are buying capability or capacity. If the model is genuinely load-bearing for revenue, early subscription and reserved throughput are worth paying for before you need them, on the evidence of this week. If it is exploratory, the queue costs you nothing and you should not pay to skip it.
Read next: Moonshot's Biggest Model Has No Thinking Dial | Alibaba Withheld the Number That Sets Your Cost



