Shanghai answered the question it skipped in July

On 3 August Alibaba launched Qwen3.8-Max, the model it had previewed on 19 July at the World AI Conference in Shanghai without the numbers that mattered. The specification is now public: 2.4 trillion total parameters, 95 billion of them active at inference, built on the Qwen 3.5 architecture and handling text, images, video and documents in one system. Hong Kong shares rose about 6 percent to 124.00 Hong Kong dollars in early trading on the announcement.

The company also committed to the thing it would not date in July. It said the weights will be released openly next week, which would make this the first Max-class Qwen that anyone can download and run on hardware they control. For a European operator watching Chinese frontier models specifically because open weights permit self-hosting, inspection and exit, that is the part of the announcement with consequences attached.

Ninety-five billion active is the genuinely good news

Why it matters: the activation ratio is what turns a parameter count into a bill. Qwen3.8-Max engages roughly 4 percent of its network on any given query, so the compute each request consumes sits closer to a 95 billion parameter model than to a 2.4 trillion one. Alibaba says inference costs are lower than the previous generation, and for once the disclosed architecture supports the claim rather than obscuring it.

This is precisely the figure that was missing from the July preview, and it has arrived on the favourable side. For anyone buying tokens through an API, the practical read is that a model marketed on 2.4 trillion parameters should not be priced like one. The preview still runs at 10 percent of standard rates through Alibaba's Token Plan, Qoder and QoderWork, and the predecessor Qwen3.7-Max lists at 2.50 US dollars per million input tokens and 7.50 per million output, about 2.30 and 6.90 euros. Judge any offer against that list, never against the promotional rate.

Then count the memory, not the parameters

The number nobody put in a headline: holding 2.4 trillion parameters takes roughly 2.4 terabytes of memory at 8-bit precision, and around 1.2 terabytes if you quantise to 4-bit and accept the quality cost. An eight-GPU H200 server provides about 1.1 terabytes of high-bandwidth memory in total. Open weights at this size therefore do not mean one server. They mean at least two full nodes at 8-bit, or a tight fit on one node after aggressive quantisation, before you have allocated a single byte to the key-value cache that actually serves concurrent users.

That lands in the worst possible month. Memory is the scarcest input in the industry right now: roughly 70 percent of next year's supply is already committed, DRAM and HBM contract prices have climbed all quarter, and the squeeze reaches far enough down the market that Apple cannot keep an entry-level laptop in stock and is steering buyers to a dearer one. The freedom that open weights confer is real, and for most European operators it is currently theoretical. You are being handed the keys to a building whose floor space you cannot rent.

The benchmark claim is still Alibaba scoring Alibaba

Alibaba says the model is broadly competitive with the leading American systems, beating them on several coding, multimodal and engineering benchmarks while trailing on some general reasoning tests, and it has again placed itself close behind Anthropic's Fable 5. Every one of those placements comes from Alibaba's own launch material. No independent leaderboard has scored Qwen3.8-Max yet, exactly as none had scored the July preview.

Yes, but: the open-weight release changes the weight of that caveat rather than removing it. Once weights are public, independent evaluation stops depending on vendor goodwill or API access, and a ranking claim becomes falsifiable within days by anyone with the hardware to load it. A vendor that opens its weights one week after making a ranking claim is at minimum willing to be checked. The correct response is to wait for the check, which is now days away rather than indefinite.

What this changes for a European buyer this month

Nothing in your production stack, yet. The sequence worth following is narrow: let the weights land, let the independent scores arrive, then evaluate on your own workload at full list price rather than the preview rate. That is a September decision rather than an August one, and there is no penalty for being second to a model that will still be there in six weeks.

The bottom line: the strategic question this launch raises is not whether a Chinese model is good enough. It is whether open weights still buy independence when the hardware needed to exercise them is rationed. If self-hosting is your reason for tracking these releases at all, the binding constraint has quietly moved from licence terms to memory procurement, and that is where the planning effort now belongs.