AI Model Costs 2026: US Frontier vs Chinese Open Source

Living tracker. Last updated: 25 July 2026. Bookmark it — prices in this market have a half-life of about a quarter.

How much do AI models cost in 2026? On raw API rates, the US–China gap is 5–10x. OpenAI’s flagship GPT-5.6 Sol charges $5.00 per million input tokens and $30.00 per million output tokens, and Anthropic’s Claude Opus 4.8 sits at $5.00/$25.00, as of July 2026 (TLDL pricing index; OpenRouter). China’s DeepSeek V4 Pro charges $0.435/$0.87 for the same million tokens — roughly one-tenth of Sol’s input rate and about 3% of its output rate (BenchLM, synced 23 July 2026). Alibaba’s Qwen3.5 Plus is cheaper still at $0.40/$2.40, and Zhipu’s open-weight GLM-5.2 lands at $1.40/$4.40 — about one-sixth of GPT-5.5-class pricing (Fello AI, July 2026).

That headline gap is real, but it is not the whole bill. Chinese models routinely burn more tokens per task — Artificial Analysis found GPT-5.6 finishing coding-agent tasks on roughly one-ninth of the output tokens DeepSeek V4 consumed (TNW, July 2026) — and the training-cost claims underneath the cheap prices remain fiercely contested. This tracker follows all of it: API rates, cost-per-task reality, training economics, Asia’s budget subscription tiers, and the $700 billion capex question. It is the economics companion to our US vs China AI capability tracker.

What do frontier APIs cost per million tokens?

All prices below are standard published API rates in USD per 1 million tokens, verified against provider documentation and third-party pricing indices as of 25 July 2026. Cached-input discounts (typically 90%) and batch rates are excluded for comparability.

ModelOriginInput / 1MOutput / 1MNotes (as of July 2026)
GPT-5.6 SolOpenAI (US)$5.00$30.00Flagship tier
GPT-5.6 TerraOpenAI (US)$2.50$15.00Mid tier
GPT-5.6 LunaOpenAI (US)$1.00$6.00Fast/cheap tier
Claude Opus 4.8Anthropic (US)$5.00$25.001M-token context; released May 2026
Gemini 3.1 ProGoogle (US)$2.00$12.00Prompts ≤200K tokens; higher above
Gemini 3.6 FlashGoogle (US)$1.50$7.50Launched 21 July 2026
DeepSeek V4 ProDeepSeek (China)$0.435$0.87Off-peak; 2x during Beijing peak hours
DeepSeek V4 FlashDeepSeek (China)$0.14$0.28Off-peak; 2x during Beijing peak hours
Qwen3.5 PlusAlibaba (China)$0.40$2.401M context; rises above 256K input
Qwen3.5 FlashAlibaba (China)$0.10$0.40Cheapest 1M-context tier tracked
GLM-5.2Zhipu / Z.ai (China)$1.40$4.40Open weights
Kimi K3Moonshot AI (China)$3.00$15.002.8T params; open weights due 27 July

Sources: TLDL OpenAI index (verified 17 July 2026), OpenRouter, Fello AI Gemini guide, BenchLM DeepSeek, BenchLM Qwen (synced 22 July 2026), Fello AI GLM guide, Tom’s Hardware.

Two things jump out of the table. First, Qwen3.5 Flash at $0.10/$0.40 means a million-token-context model now costs less per token than most US “nano” tiers — GPT-5.4 nano is $0.20/$1.25 (TLDL). Second, Kimi K3 at $3.00/$15.00 breaks the “Chinese equals cheap” pattern: Moonshot AI is pricing its 2.8-trillion-parameter model like a Western mid-tier, a 5x jump on input versus Kimi K2.6, its immediate predecessor (Tom’s Hardware, July 2026). Chinese labs are no longer uniformly racing to zero.

DeepSeek now charges by the clock

The most interesting pricing innovation of 2026 came from Hangzhou. On 30 June 2026, DeepSeek announced peak-and-off-peak API pricing — the first major lab to price tokens by time of day — launching alongside V4 in mid-July (Servola, July 2026). During peak windows of 09:00–12:00 and 14:00–18:00 Beijing time, rates double: V4 Pro output goes from roughly ¥6 (~$0.87) to ¥12 (~$1.74) per million tokens, and V4 Flash from ¥2 to ¥4. Off-peak rates are unchanged from prior pricing. For European and American users the peak windows fall largely overnight, which effectively hands Western developers a permanent discount — and tells you DeepSeek’s constraint is inference capacity, not demand.

The gap is 5–10x on price — but tokens aren’t equal

Per-token prices flatter the Chinese models, because Chinese models talk more. Benchmarking firm Artificial Analysis found GPT-5.6 completed coding-agent tasks using roughly one-ninth of the output tokens DeepSeek V4 consumed, while scoring higher — meaning a model that looks 10x cheaper per token can converge on price parity per finished task (TNW, July 2026). It is why some Chinese developers now pay for GPT-5.6 despite the sticker shock.

GLM-5.2 is the clearest case. Artificial Analysis clocks it at roughly 43,000 output tokens per task on its Intelligence Index — about 37,000 of them reasoning tokens — roughly double what comparable GPT-5.5-class runs emit, against ~24,000 for MiniMax-M3 and ~35,000 for Kimi K2.6 (Artificial Analysis data, June 2026). Even so, GLM-5.2’s cheap rates keep it at roughly $0.46 per Index task — still on the cost-intelligence Pareto frontier, just not one-sixth of the US bill once verbosity is priced in. The practical rule for buyers, as of July 2026: multiply Chinese per-token quotes by 2–3x for agentic workloads before comparing, and remember output tokens also cost latency, not just money.

US enterprises are already routing tokens east

Whatever the caveats, the traffic has moved. Chinese-origin models peaked at 46.4% of weekly token usage on routing platform OpenRouter in mid-2026 (CNBC), up from an 11% average over the prior 12 months and just 4.5% in the first half of 2025, according to CNBC’s analysis of OpenRouter data (Yahoo Finance, July 2026). In the week of 9–15 February 2026, Chinese models surpassed US models on the platform for the first time, at 4.12 trillion weekly tokens. DeepSeek alone carries 17.6% of OpenRouter tokens and Qwen 13.9%, with Chinese models running 60–90% cheaper than leading Anthropic and OpenAI offerings — and Ramp’s chief economist calling cost “the deciding factor” for many procurement teams.

The pattern rhymes with what we documented when the first shock hit: cheap, good-enough open weights reset the floor price for the entire industry. Read the origin story in our analysis of how DeepSeek broke global AI economics — the difference in 2026 is that the discount is no longer hypothetical; it is showing up in US enterprise routing data.

What did these models really cost to train?

The famous number is DeepSeek’s ~$5.6 million (rounded in headlines to $6 million) for V3’s pre-training run — the claim that vaporised nearly $600 billion of Nvidia market cap in January 2025. The rebuttal is equally famous: SemiAnalysis estimated DeepSeek’s total server capex at roughly $1.6 billion, with GPU investments above $500 million across a fleet of ~50,000 Hopper-class chips, and noted the $6 million covered only the GPU time of the final pre-training run — excluding R&D, failed experiments, salaries and infrastructure (SemiAnalysis, 2025). Both numbers are true; they measure different things. The marginal run was cheap. The organisation was not.

The narrative playbook has since become standard issue in China. Baidu released ERNIE 5.1 in May 2026 claiming a pre-training cost of just 6% of the industry average, alongside a top-Chinese-model ranking on LMArena (CnTechPost, 9 May 2026) — a claim no independent analyst has yet audited, and which almost certainly uses the same narrow “final run” accounting. Moonshot claims Kimi K3’s Kimi Delta Attention architecture delivers up to 75% KV-cache reduction and up to 6x decoding throughput at 1M-token context (Morph; Tom’s Hardware, July 2026). Treat all lab-reported training costs, US or Chinese, as marketing until the weights and the accounting are both open. As of July 2026, only the weights ever are.

Why Asia gets cheaper subscriptions

Consumer pricing tells the same story from the other end. ChatGPT Plus holds at $20/month in the US (Pro at $200), as of July 2026 App Store data (OpenTheRank). But across Asia, OpenAI sells ChatGPT Go — a budget tier launched in India in August 2025 and rolled out globally by January 2026 — at ₱300 in the Philippines, ₫132,000 (~$5) in Vietnam, IDR 75,000 in Indonesia and S$13 in Singapore (Marketing-Interactive, January 2026). In India, Go (list price ₹399, ~$4.50) has been free for 12 months for all users since late October 2025 (TechCrunch) — with those free years due to start converting to paid billing from late 2026.

Why the generosity? Because in Asia the counterfactual is not “no AI” — it is free, open-weight Chinese AI, plus Google undercutting at $4.99/month with AI Plus (Fello AI, June 2026). Budget tiers are a land-grab for the next billion users in markets where DeepSeek, Qwen and GLM chatbots cost nothing, and where OpenAI is now preparing to monetise Free and Go tiers with ads (Marketing-Interactive). Which models are actually available in which Asian market — and at what local price — is tracked in our LLM access and pricing tracker for Asia, with the wider demand picture in our Asia AI market hub.

The capex question hanging over all of it

Here is the tension that makes this tracker worth keeping. US hyperscalers’ 2026 AI capital-expenditure plans have topped $700 billion, per Reuters (Yahoo Finance, May 2026). Amazon alone spent $44.2 billion in Q1 2026; Microsoft’s quarterly capex hit $30.9 billion, up 84% year on year; Meta guides to $125–145 billion for the full year. Meanwhile Chinese labs give away weights that capture 46% of tokens on the biggest neutral routing platform, and Amazon’s trailing free cash flow has collapsed 95% to $1.2 billion (Yahoo Finance).

The bull case: inference demand is exploding — OpenRouter’s throughput quadrupled from ~5 trillion to over 20 trillion tokens a week between April 2025 and April 2026 — so capacity, not models, is the scarce asset, and DeepSeek’s peak-hour surcharge is early evidence that even the cheapest providers hit compute walls. The bear case: if open weights keep compressing frontier margins to DeepSeek levels, $700 billion a year is a lot of depreciation chasing $0.87-per-million-token revenue. Our view, as of July 2026: token prices will keep falling faster than capex, and the squeeze lands hardest on mid-tier US models — the Terras and Flashes — which compete head-on with Chinese open weights without the frontier-capability premium of a Sol or an Opus. The strategic backdrop sits in our China AI market hub.

FAQ: AI model costs in 2026

What is the cheapest capable AI model API in 2026?

Among frontier-adjacent models, Qwen3.5 Flash at $0.10 input / $0.40 output per million tokens with a 1M-token context window is the cheapest we track, followed by DeepSeek V4 Flash at $0.14/$0.28 off-peak, as of July 2026 (BenchLM). Zhipu also offers GLM-4.7 Flash free via API (Fello AI).

Is DeepSeek really 10x cheaper than GPT-5.6?

Per token, yes: DeepSeek V4 Pro’s $0.435/$0.87 is roughly a tenth of GPT-5.6 Sol’s $5/$30, as of July 2026. Per completed task, often not: Artificial Analysis found GPT-5.6 uses about one-ninth the output tokens on coding-agent work, which can erase most of the gap (TNW).

Did DeepSeek really train its model for $6 million?

Only in the narrowest sense. The ~$5.6M figure covers GPU time for one final pre-training run. SemiAnalysis estimates DeepSeek’s total server capex at ~$1.6 billion with over $500 million in GPUs (SemiAnalysis). Baidu’s “6% of industry cost” claim for ERNIE 5.1 uses similarly narrow accounting.

Why is ChatGPT cheaper in Asia than in the US?

OpenAI’s ChatGPT Go tier prices to local purchasing power and to competition from free Chinese models: ~$4–5/month in Vietnam, the Philippines and Indonesia versus $20 for Plus in the US, and free for 12 months in India, as of July 2026 (Marketing-Interactive; TechCrunch).

Are Chinese open-source models actually free?

The weights are free to download and self-host (DeepSeek V4, GLM-5.2, Qwen3.5, and Kimi K3 from 27 July 2026). Running them is not: you pay for GPUs or a hosted API, and verbose outputs — GLM-5.2 averages ~43,000 tokens per benchmark task — inflate compute bills. “Free” means no licence fee and no vendor lock-in, not zero cost.

Update log

July 2026: Tracker launched. Baseline prices captured for GPT-5.6 family, Claude Opus 4.8, Gemini 3.1/3.6, DeepSeek V4 (including new peak/off-peak regime), Qwen3.5, GLM-5.2 and Kimi K3. Next scheduled checks: Kimi K3 weights release (27 July), India ChatGPT Go free-year conversions (late 2026), Q3 hyperscaler capex guidance.

Share this article

Discover more from Digital in Asia

Subscribe to get the latest posts sent to your email.

Tom Simpson

Tom Simpson is an investor, advisor, and writer working across AI, markets, media, and culture — tracking where value and attention are moving. He is the founder of AK3R, working selectively with founders, investors, and companies on strategy, while investing in and building businesses in digital markets. He writes the Hyperfuture Memo on Substack, on how AI is reshaping markets, media, and culture. He is also the founder and editor of Digital in Asia, an independent publication covering Asia's digital markets since 2013. He splits time between Vietnam, Singapore, and the UK.