Skip to content

Insights/East-West model notes

China wants to industrialize the token.

Five assumptions hold the export story up. The price gap survives. The chips, the idle racks and the revenue do not.

140Tdaily tokens, state-reported
36.8%effective compute usage
46%of revenue, 12% of volume
July 25, 20269 min read
A technician walks a service aisle in a computing hall, half the racks unlit, network cabling running loose along a scuffed concrete floor under one warm service lamp

At WAIC in Shanghai this month, across 1,100 exhibitors and more than 140 forums, the phrase that kept coming up was not AGI. It was Token工厂, token factory.

There is nothing subtle about the strategy behind it. Take a unit of AI output and standardize it. Industrialize the production. Drive the cost down until nobody else wants to compete, then export the surplus. It is the solar and battery playbook, applied to inference.

I run a cross-border business out of Shanghai, and before that spent seven years running a commerce agency here on a P&L that lived or died on Chinese platform economics. That makes me skeptical in two directions: of Western coverage that waves off Chinese infrastructure, and of Chinese coverage that treats a pilot as a finished export industry.

To be clear up front, none of this argues Chinese models are weak. Stanford's AI Index put the top US model 2.7 percent ahead of the top Chinese model in March. The capability question is close to settled. What follows is about infrastructure and economics.

Five assumptions hold the strategy up. One stands. Four are shaky. If you only want the practical part, skip to the last section.

The assumption ledger

Open a row to see what the available data does to it.

1/5stands
The export loop worksOne pilot

Shantou closed the loop in April: data in, inference domestic, answer out by API. No volume and no revenue have been published, and the city was picked for a cable landing station most of the country does not have.

Read the section
The price gap is realStands

Open-weight Chinese models run 60 to 90 percent below leading Western pricing, and developers have already routed accordingly. This is the strongest leg the strategy has.

Read the section
Domestic silicon can carry itNot yet

Huawei’s 910C delivers roughly 60 percent of H100 inference performance, and most of DeepSeek’s tokens are still inferenced on Western hardware.

Read the section
The factories are runningMostly idle

Effective usage of 36.8 percent in 2025, rack utilization of 20 to 30 percent at many centers, against leading platforms running at 90 to 95 percent. A coordination problem, not a supply glut.

Read the section
Volume converts to revenueInverted

Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue. China owns the commodity lane; the profit sits in the other one.

Read the section
Survives the dataRunning ahead of the evidence
36Kr's WAIC recap called it the first year the token economy became a core conference topic.
36氪July 2026
The move

Make inference a metered commodity

In March, China's National Data Administration fixed the official Chinese term for token as 词元, cí yuán, and called it a settlement unit linking technical supply to commercial demand. Beijing now tracks daily token volume and reports it at policy events the way it reports steel output.

The headline number is a thousandfold rise in two years. About 100 billion daily calls in early 2024, more than 140 trillion by March 2026.

100Bdaily calls, early 2024
140Tdaily, March 2026

State-reported. Unaudited. Counts consumption, not value.

Treat that with some care. It is state-reported, nobody audits it, and it counts consumption instead of value. That is not a China problem specifically. No government or vendor anywhere publishes an audited token number. An agent stuck in a bad loop overnight burns an absurd quantity and produces nothing.

By end-2025 China had built more than 100,000 high-quality datasets totaling over 890 petabytes, the supply side of the same policy push.
新华社March 24, 2026
The export

The mechanism is real, but it is one pilot

In April, Shantou in Guangdong closed a loop under a policy called 来数加工, inbound data processing. Overseas data enters a digital bonded zone. Domestic compute runs the inference. The answer goes back out by API. Power, compute and revenue all stay in China.

  1. Overseas data inEnters a digital bonded zone
  2. Domestic computeRuns the inference
  3. Answer out by APIOnly the output crosses back
Power staysCompute staysRevenue stays

Chinese state media summed it up with a line that traveled widely: the salt can travel the world, but the salt fields must stay home. Good line. I have caught myself repeating it. It is also promotional framing from the same system that publishes the consumption figures.

What has not been published is volume. One city, one pilot, no disclosed revenue. Shantou was picked for reasons that predate AI entirely. It holds one of three mainland submarine cable landing stations, and latency to Singapore runs about 32.7 milliseconds. Most of the country has neither. That part tends to get left out.

On the record
  • 1 of 3 mainland submarine cable landing stations
  • 32.7 ms latency to Singapore
Not published
  • n/a export volume
  • n/a disclosed revenue
Chinese coverage frames token exports as an impossible triangle of latency, data compliance and very cheap green power.
澎湃新闻 and 同花顺财经2026
The price gap

It is real, and developers have already voted

Nobody argues with this one. It is the strongest leg the strategy has.

DeepSeek V4 Flash costs $0.14 per million input tokens. GPT-5.5 costs $5.00. Open-weight Chinese models run 60 to 90 percent below leading Western pricing. On OpenRouter, which routes calls to whichever provider wins on price, Chinese models went from under 10 percent of usage in early 2025 to a record 58 percent of tokens processed by US firms in July.

Input, per 1M tokens
DeepSeek V4 Flash
$0.14
GPT-5.5
$5.00

Open weights run 60 to 90 percent below leading Western pricing.

Chinese model share on OpenRouter
under 10%early 2025
58%July 2026, a record

Of tokens processed by US firms.

In agent workloads the gap compounds, because one task can fan out across dozens of parallel sub-agent calls.

OpenRouter weekly volume grew from roughly 5 trillion tokens in April 2025 to over 20 trillion a year later.
CNBCJuly 2026
The compute

Where it breaks first: the compute underneath

This is the gap most coverage understates, and it is the one that bothers me. You cannot industrialize output on equipment you cannot get enough of.

Huawei's 910C delivers roughly 60 percent of H100 inference performance, by DeepSeek's own published assessment. The Council on Foreign Relations puts the best US chips at about five times more powerful than Huawei's best today, widening toward 2027. Then there is the SemiAnalysis detail, which matters more than any chip spec: most of DeepSeek's tokens are still inferenced on Western hardware. The flagship of Chinese open-weight AI is not running the Chinese stack at scale.

Huawei 910C vs H10060%

of H100 inference performance

Best US chip vs best Huawei

more powerful today, widening toward 2027

CloudMatrix 384 vs NVL724.1×

the power, for about 1.7× the compute

The flagship of Chinese open-weight AI is not running the Chinese stack at scale.

The energy argument, which usually gets made next, has the same problem. China generates more than twice the electricity of the United States and added 543 gigawatts in 2024 alone. But a four-fold power penalty on domestic accelerators eats much of a two-fold price advantage. Cheap power ends up subsidizing inefficient silicon rather than producing cheaper tokens.

The advantage

the electricity of the United States, plus 543 gigawatts added in 2024 alone.

The penalty

the power draw on domestic accelerators, against a two-fold price advantage.

Chinese firms are not sitting still. Zhipu trains and serves on domestic accelerators. Alibaba is building data centers across South Korea, Malaysia, Thailand and Mexico, renting foreign capacity to sidestep the silicon problem. Both are rational moves, and neither closes the gap this year.

Huawei's CloudMatrix 384 draws roughly 4.1 times the power of Nvidia's NVL72 for about 1.7 times the compute.
SemiconductorX2026
The idle racks

A lot of the capacity is sitting idle

The factory framing implies the plants are running. Chinese domestic reporting suggests many are not.

Tencent Cloud puts average GPU utilization at China's intelligent computing centers below 30 percent. CAICT data cited by Huxiu shows an effective usage rate of 36.8 percent in 2025, meaning roughly 70 yuan of every 100 invested is idling. Securities Times reported rack utilization of 20 to 30 percent at many centers.

Average GPU utilization at intelligent computing centersTencent Cloud
Effective usage rate in 2025CAICT, via Huxiu
Rack utilization at many centersSecurities Times

Calling this oversupply misses it. The market has split, and I have watched the same pattern in Chinese logistics: healthy national totals sitting on stranded regional capacity. Leading platforms run at 90 to 95 percent with order books into 2028. The high-quality compute is genuinely short, while the low-quality stuff cannot find a customer at any price. That is a coordination problem, and those take years to clear.

High-quality compute

Leading platforms run at 90 to 95 percent with order books into 2028. Genuinely short.

order books into 2028
Stranded capacity

The low-quality stuff cannot find a customer at any price. A coordination problem, and those take years to clear.

no customer at any price
A thousand-card computing center costs over 30 million yuan a year to operate, which is why low utilization turns into losses fast.
虎嗅 and 观察者网July 2026
The revenue

Volume is not revenue, and that is the gap that matters

If you read one number here, make it this one. Anthropic holds around 12 percent of token volume on OpenRouter and captures close to 46 percent of the revenue.

Share of token volume
12%
Share of revenue
46%

Anthropic, on OpenRouter. The lanes are priced differently.

The market is splitting into a commodity lane and a premium lane. China owns the commodity lane. The profit sits in the other one.

The economics are moving the wrong way for everyone. Huxiu documents Alibaba Cloud and Baidu AI Cloud raising prices by up to 34 percent this year. Agents consume 100 to 1,000 times what a chatbot does. Unit prices keep falling while invoices keep growing, and nobody has an accepted way to link tokens consumed to work completed.

up to 34%price rises at Alibaba Cloud and Baidu AI Cloud this year
100 to 1,000×what an agent consumes against a chatbot
85%of enterprise AI budgets now going to inference
Inference now accounts for roughly 85 percent of enterprise AI budgets, against near-zero marginal cost in traditional software.
腾讯新闻June 2026
The constraint

The one no price cut touches

Under China's 2017 National Intelligence Law, companies must cooperate with state intelligence work. That is a jurisdictional fact. It says nothing about how any particular company behaves, and it is still enough to keep regulated finance, healthcare and government workloads with Western providers.

Washington is escalating separately. The Treasury Secretary raised the prospect of sanctions in July over alleged distillation of US models, days ahead of a September bilateral AI dialogue. If part of the Chinese cost advantage comes from not paying frontier training costs, it is less durable than it looks. Nobody has proven that case. It is not trivial either.

Large enterprises continue buying mostly from Anthropic, OpenAI, Azure and Google Cloud.
BaiguanJuly 2026
The test

What would change my mind

I am not predicting failure. The export story is running ahead of the evidence, which is different. Three signals would move me: Shantou publishing real export volume and revenue, Chinese computing center utilization crossing 50 percent, and a regulated Western enterprise putting a production workload on a China-hosted endpoint.

Shantou publishes real export volume and revenue.

Chinese computing center utilization crosses 50 percent.

A regulated Western enterprise puts a production workload on a China-hosted endpoint.

Any of those would tell you more than another month of routing charts. If all three move, I have this wrong, and I would want to know.

Practice

What to actually do about it

Watching signals is for analysts. If you run a budget, the question is operational. Almost nobody hassorted their workloads by whether they need frontier capability. It gets discussed, then it sits on the list.

A rough method: pull last month's token spend by application. Flag anything doing classification, extraction, summarization, translation, first drafts or routine code completion. That bucket is usually 60 to 80 percent of volume and rarely needs a frontier model. Run a week of parallel traffic against a cheap open model, measure task completion rather than benchmark scores, and keep the expensive model for what actually fails.

Flag these first
ClassificationExtractionSummarizationTranslationFirst draftsRoutine code completion

Usually 60 to 80 percent of volume, and rarely a frontier job.

At a 35-to-1 input price gap the arithmetic is blunt. A team spending $40,000 a month with 70 percent of volume in that bucket carries about $28,000 where the cheap option runs closer to $1,000. Allow for higher retry rates and some workloads bouncing back, and the saving is still not marginal.

Run it on your own billAt a 35-to-1 input price gap.
$28,000carried in that bucket today
$1,000the cheap option, rounded up for retries
$27,000a month, before workloads bounce back

Coding tool vendors, support platforms and translation shops moved first, because their margins are directly exposed to inference cost. Regulated industries have not, and probably will not.

One thing to know before you route anything: because the weights are open, Western hosts serve the same Chinese models with no data touching China, at roughly double first-party pricing and still far below frontier rates. For most companies that is the sensible version of this trade.

Western hosts serving open-weight Chinese models charge roughly double the first-party price, with no data touching China.
OpenRouterJune 2026

Talk to us

Talk to us about AI.

A conversation with the senior team about your markets, your data, and where AI would actually pay back for you. No slides, no obligation, and if the honest answer is that your current setup is fine, you will hear that too.

Talk to us about AI