The week a startup dropped its frontier vendor

Andy Fang, a co-founder of DoorDash, described the company's use of Moonshot AI for an experimental command-line tool in plain commercial terms: better quality at a cheaper cost. Cursor used Moonshot's Kimi in building its Composer 2 coding agent. The startup Lindy reportedly dropped Anthropic's tools altogether in favour of DeepSeek's V4 models. Airbnb and Siemens have both been experimenting with Alibaba and DeepSeek systems to hold down the cost of daily operations. None of these are ideological decisions. They are procurement decisions, taken one workload at a time.

The scale is visible in the routing data. On OpenRouter, the marketplace where a large share of model traffic is brokered, US companies sent more than 30 percent of their token consumption to Chinese models in every single week from 8 February onward, with a peak week at 46 percent. In the first half of 2025 the same figure was 4.5 percent. A Hugging Face study in March 2026 found Chinese open-source models accounted for 41 percent of downloads on that platform. The price gap driving it is not marginal: the Chinese systems typically run 60 to 90 percent cheaper than the leading American models, and at some tiers the difference is larger still.

What makes this a decision story rather than a market story is who captured the saving. The firms in that list did not switch because they discovered a cheaper vendor this quarter. They switched because they were already able to. Somebody at each of them had built the ability to run the same workload against a different model and see, in numbers, whether the output held up.

The option to switch is bought long before it pays

Most European companies standardised on a single frontier vendor in 2024 and 2025, and did so for defensible reasons: one contract, one security review, one integration, one set of prompts to maintain. The cost of that decision was invisible while prices were stable. It became visible the moment a 60 percent gap opened, because a firm with one vendor and no evaluation harness cannot tell whether a cheaper model is good enough for its own traffic, and no vendor benchmark answers that question for it.

This is the standard shape of an option you did not know you were selling. Standardising is a genuine efficiency, and it prices in a world where the alternatives are roughly equivalent. When the spread widens, the firm that kept a way to test alternatives converts the spread into margin, and the firm that did not simply watches it. The asymmetry is not about being clever. It is about who paid a small, boring cost in advance, and the boring cost here is a fixed evaluation set: a few hundred requests drawn from your real traffic, with graded expected outputs, that any candidate model can be run against in an afternoon.

Resist the reflex to route everything. The pattern in the reporting is selective, not wholesale. Teams are sending the high-volume, low-stakes work to the cheapest model that clears the bar and keeping the judgment-heavy work where it was. That is the correct shape, and the reason is base rates: a model that matches a frontier system on 95 percent of your requests can still be materially worse on the 5 percent that generate complaints, refunds or legal exposure. Segment by consequence of error before you segment by price.

Why the European version of this is a different trade

An American firm choosing a Chinese model is usually choosing an API. A European firm making the same nominal choice faces a question the price comparison does not contain: where does the request go, and under what transfer basis. Sending customer data to a model API hosted outside the EU is a transfer question with documentation obligations attached, and for regulated workloads the answer is often that the cheap API is simply unavailable to you.

The route that does work is the one the price tables do not describe. Most of the Chinese systems in this story are open weight, which means you can run them on infrastructure you control, inside your own boundary, with no request leaving your estate. Yasir Atalan of CSIS put the appeal plainly: open-source models give relief to those who want to keep their data. But that route converts a per-token bill into a different cost structure entirely, namely GPU capacity you have to buy or rent, an inference stack somebody has to run, and the staff to keep both healthy. That can still be much cheaper at volume. It is a fundamentally different decision from changing an API key, and it belongs in a capex conversation rather than a procurement one.

So the useful comparison is three-way, not two-way. Frontier API, cheaper API, and self-hosted open weights each carry a different cost curve, a different data posture and a different failure mode. Run all three against the same evaluation set on your own traffic, including a volume at which the self-hosted option breaks even, and you will have something a board can act on. Run none of them, and the 60 percent number in the headlines is simply a saving other companies are taking.