At WAIC in Shanghai this month, across 1,100 exhibitors and more than 140 forums, the phrase that kept coming up was not AGI. It was Token工厂, token factory.
There is nothing subtle about the strategy behind it. Take a unit of AI output and standardize it. Industrialize the production. Drive the cost down until nobody else wants to compete, then export the surplus. It is the solar and battery playbook, applied to inference.
I run a cross-border business out of Shanghai, and before that spent seven years running a commerce agency here on a P&L that lived or died on Chinese platform economics. That makes me skeptical in two directions: of Western coverage that waves off Chinese infrastructure, and of Chinese coverage that treats a pilot as a finished export industry.
To be clear up front, none of this argues Chinese models are weak. Stanford's AI Index put the top US model 2.7 percent ahead of the top Chinese model in March. The capability question is close to settled. What follows is about infrastructure and economics.
Five assumptions hold the strategy up. One stands. Four are shaky. If you only want the practical part, skip to the last section.
Open a row to see what the available data does to it.
The export loop worksOne pilot
Shantou closed the loop in April: data in, inference domestic, answer out by API. No volume and no revenue have been published, and the city was picked for a cable landing station most of the country does not have.
Read the sectionThe price gap is realStands
Open-weight Chinese models run 60 to 90 percent below leading Western pricing, and developers have already routed accordingly. This is the strongest leg the strategy has.
Read the sectionDomestic silicon can carry itNot yet
Huawei’s 910C delivers roughly 60 percent of H100 inference performance, and most of DeepSeek’s tokens are still inferenced on Western hardware.
Read the sectionThe factories are runningMostly idle
Effective usage of 36.8 percent in 2025, rack utilization of 20 to 30 percent at many centers, against leading platforms running at 90 to 95 percent. A coordination problem, not a supply glut.
Read the sectionVolume converts to revenueInverted
Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue. China owns the commodity lane; the profit sits in the other one.
Read the section36Kr's WAIC recap called it the first year the token economy became a core conference topic.
Make inference a metered commodity
In March, China's National Data Administration fixed the official Chinese term for token as 词元, cí yuán, and called it a settlement unit linking technical supply to commercial demand. Beijing now tracks daily token volume and reports it at policy events the way it reports steel output.
The headline number is a thousandfold rise in two years. About 100 billion daily calls in early 2024, more than 140 trillion by March 2026.
State-reported. Unaudited. Counts consumption, not value.
Treat that with some care. It is state-reported, nobody audits it, and it counts consumption instead of value. That is not a China problem specifically. No government or vendor anywhere publishes an audited token number. An agent stuck in a bad loop overnight burns an absurd quantity and produces nothing.
By end-2025 China had built more than 100,000 high-quality datasets totaling over 890 petabytes, the supply side of the same policy push.
The mechanism is real, but it is one pilot
In April, Shantou in Guangdong closed a loop under a policy called 来数加工, inbound data processing. Overseas data enters a digital bonded zone. Domestic compute runs the inference. The answer goes back out by API. Power, compute and revenue all stay in China.
- Overseas data inEnters a digital bonded zone
- Domestic computeRuns the inference
- Answer out by APIOnly the output crosses back
Chinese state media summed it up with a line that traveled widely: the salt can travel the world, but the salt fields must stay home. Good line. I have caught myself repeating it. It is also promotional framing from the same system that publishes the consumption figures.
What has not been published is volume. One city, one pilot, no disclosed revenue. Shantou was picked for reasons that predate AI entirely. It holds one of three mainland submarine cable landing stations, and latency to Singapore runs about 32.7 milliseconds. Most of the country has neither. That part tends to get left out.
- 1 of 3 mainland submarine cable landing stations
- 32.7 ms latency to Singapore
- n/a export volume
- n/a disclosed revenue
Chinese coverage frames token exports as an impossible triangle of latency, data compliance and very cheap green power.
It is real, and developers have already voted
Nobody argues with this one. It is the strongest leg the strategy has.
DeepSeek V4 Flash costs $0.14 per million input tokens. GPT-5.5 costs $5.00. Open-weight Chinese models run 60 to 90 percent below leading Western pricing. On OpenRouter, which routes calls to whichever provider wins on price, Chinese models went from under 10 percent of usage in early 2025 to a record 58 percent of tokens processed by US firms in July.
Open weights run 60 to 90 percent below leading Western pricing.
Of tokens processed by US firms.
In agent workloads the gap compounds, because one task can fan out across dozens of parallel sub-agent calls.
OpenRouter weekly volume grew from roughly 5 trillion tokens in April 2025 to over 20 trillion a year later.
Where it breaks first: the compute underneath
This is the gap most coverage understates, and it is the one that bothers me. You cannot industrialize output on equipment you cannot get enough of.
Huawei's 910C delivers roughly 60 percent of H100 inference performance, by DeepSeek's own published assessment. The Council on Foreign Relations puts the best US chips at about five times more powerful than Huawei's best today, widening toward 2027. Then there is the SemiAnalysis detail, which matters more than any chip spec: most of DeepSeek's tokens are still inferenced on Western hardware. The flagship of Chinese open-weight AI is not running the Chinese stack at scale.
of H100 inference performance
more powerful today, widening toward 2027
the power, for about 1.7× the compute
The flagship of Chinese open-weight AI is not running the Chinese stack at scale.
The energy argument, which usually gets made next, has the same problem. China generates more than twice the electricity of the United States and added 543 gigawatts in 2024 alone. But a four-fold power penalty on domestic accelerators eats much of a two-fold price advantage. Cheap power ends up subsidizing inefficient silicon rather than producing cheaper tokens.
2× the electricity of the United States, plus 543 gigawatts added in 2024 alone.
4× the power draw on domestic accelerators, against a two-fold price advantage.
Chinese firms are not sitting still. Zhipu trains and serves on domestic accelerators. Alibaba is building data centers across South Korea, Malaysia, Thailand and Mexico, renting foreign capacity to sidestep the silicon problem. Both are rational moves, and neither closes the gap this year.
Huawei's CloudMatrix 384 draws roughly 4.1 times the power of Nvidia's NVL72 for about 1.7 times the compute.
A lot of the capacity is sitting idle
The factory framing implies the plants are running. Chinese domestic reporting suggests many are not.
Tencent Cloud puts average GPU utilization at China's intelligent computing centers below 30 percent. CAICT data cited by Huxiu shows an effective usage rate of 36.8 percent in 2025, meaning roughly 70 yuan of every 100 invested is idling. Securities Times reported rack utilization of 20 to 30 percent at many centers.
Calling this oversupply misses it. The market has split, and I have watched the same pattern in Chinese logistics: healthy national totals sitting on stranded regional capacity. Leading platforms run at 90 to 95 percent with order books into 2028. The high-quality compute is genuinely short, while the low-quality stuff cannot find a customer at any price. That is a coordination problem, and those take years to clear.
Leading platforms run at 90 to 95 percent with order books into 2028. Genuinely short.
order books into 2028The low-quality stuff cannot find a customer at any price. A coordination problem, and those take years to clear.
no customer at any priceA thousand-card computing center costs over 30 million yuan a year to operate, which is why low utilization turns into losses fast.
Volume is not revenue, and that is the gap that matters
If you read one number here, make it this one. Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue.
Anthropic, on OpenRouter. The lanes are priced differently.
The market is splitting into a commodity lane and a premium lane. China owns the commodity lane. The profit sits in the other one.
The economics are moving the wrong way for everyone. Huxiu documents Alibaba Cloud and Baidu AI Cloud raising prices by up to 34 percent this year. Agents consume 100 to 1,000 times what a chatbot does. Unit prices keep falling while invoices keep growing, and nobody has an accepted way to link tokens consumed to work completed.
Inference now accounts for roughly 85 percent of enterprise AI budgets, against near-zero marginal cost in traditional software.
The one no price cut touches
Under China's 2017 National Intelligence Law, companies must cooperate with state intelligence work. That is a jurisdictional fact. It says nothing about how any particular company behaves, and it is still enough to keep regulated finance, healthcare and government workloads with Western providers.
Washington is escalating separately. The Treasury Secretary raised the prospect of sanctions in July over alleged distillation of US models, days ahead of a September bilateral AI dialogue. If part of the Chinese cost advantage comes from not paying frontier training costs, it is less durable than it looks. Nobody has proven that case. It is not trivial either.
Large enterprises continue buying mostly from Anthropic, OpenAI, Azure and Google Cloud.
What would change my mind
I am not predicting failure. The export story is running ahead of the evidence, which is different. Three signals would move me: Shantou publishing real export volume and revenue, Chinese computing center utilization crossing 50 percent, and a regulated Western enterprise putting a production workload on a China-hosted endpoint.
Shantou publishes real export volume and revenue.
Chinese computing center utilization crosses 50 percent.
A regulated Western enterprise puts a production workload on a China-hosted endpoint.
Any of those would tell you more than another month of routing charts. If all three move, I have this wrong, and I would want to know.
What to actually do about it
Watching signals is for analysts. If you run a budget, the question is operational. Almost nobody hassorted their workloads by whether they need frontier capability. It gets discussed, then it sits on the list.
A rough method: pull last month's token spend by application. Flag anything doing classification, extraction, summarization, translation, first drafts or routine code completion. That bucket is usually 60 to 80 percent of volume and rarely needs a frontier model. Run a week of parallel traffic against a cheap open model, measure task completion rather than benchmark scores, and keep the expensive model for what actually fails.
Usually 60 to 80 percent of volume, and rarely a frontier job.
At a 35-to-1 input price gap the arithmetic is blunt. A team spending $40,000 a month with 70 percent of volume in that bucket carries about $28,000 where the cheap option runs closer to $1,000. Allow for higher retry rates and some workloads bouncing back, and the saving is still not marginal.
Coding tool vendors, support platforms and translation shops moved first, because their margins are directly exposed to inference cost. Regulated industries have not, and probably will not.
One thing to know before you route anything: because the weights are open, Western hosts serve the same Chinese models with no data touching China, at roughly double first-party pricing and still far below frontier rates. For most companies that is the sensible version of this trade.
Western hosts serving open-weight Chinese models charge roughly double the first-party price, with no data touching China.
