Last updated: 25 July 2026. This is a living tracker, refreshed after each major release wave. For the full strategic picture — investment, chips, talent — see our China vs US AI race analysis. This page tracks one thing only: raw capability.
How far behind are Chinese AI models on the AI benchmarks? As of July 2026, roughly three to five months — and the gap is decreasing. Nathan Lambert of Interconnects, whose gap estimate has become the industry reference, wrote after the Kimi K3 launch that “the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months” (Interconnects, 20 July 2026). Moonshot AI’s Kimi K3, a 2.8-trillion-parameter mixture-of-experts model announced on 16 July 2026, ranks #2 of 39 models on the Vals AI Index and #3 on the Artificial Analysis Intelligence Index — behind only Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 (Vals AI; Artificial Analysis, July 2026). The market has noticed: Chinese-origin models peaked at 46% of weekly enterprise tokens routed through OpenRouter in mid-2026, up from 4.5% in H1 2025, and have held at least 30% every week since 8 February 2026 (CNBC, 7 July 2026). US closed models still hold the top two slots on every major leaderboard. But the daylight beneath them has nearly vanished.
The headline metric: a 3-5 month gap, as of July 2026
The “gap in months” is the tracker’s north-star metric: how long ago did the best US closed model perform at the level of today’s best Chinese open model? There is no single formula. Lambert’s 3-5 month figure is an observational judgement across leaderboard composites; analysts triangulate it from index scores, arena Elo and pretraining-quality estimates. Zvi Mowshowitz, surveying the Kimi K3 evidence, put the gap at “at least four” months with “my median guess is six” (Don’t Worry About the Vase, 20 July 2026). Redwood Research chief scientist Ryan Greenblatt landed higher, estimating K3’s pretraining quality sits “about halfway between Opus 4 and Opus 4.5” — roughly eight months behind the frontier (via Zvi, July 2026). The honest range is therefore 3-8 months, with the centre of gravity around four to five.
Whichever end you take, the direction is unambiguous. In mid-2025 the consensus gap was 6-9 months; Lambert now calls Kimi K3 “the strongest open model ever released” and “the closest open models have been to the frontier since DeepSeek R1” (Interconnects, 20 July 2026). Ethan Mollick’s caveat is worth keeping pinned to this tracker: K3 lands “roughly where I would expect on the curve rather than an unexpected leap” — but “a good model that is still months behind looks like the future to many” (via Zvi, July 2026).
The scoreboard: frontier models ranked, 25 July 2026
The table below combines the Artificial Analysis Intelligence Index (AA), Vals AI Index rank and headline strengths for the current frontier set. US models are closed; all listed Chinese models are open-weight or have open weights committed.
| Model (release) | Lab / origin | AA Index (July 2026) | Standout result |
|---|---|---|---|
| Claude Fable 5 (2026) | Anthropic, US | ~60 — #1 | #1 Vals AI Index; leads GDPval-AA v2 (1760 Elo) |
| GPT-5.6 (Jul 2026) | OpenAI, US | ~59 — #2 | Top-2 across AA composites (Sol variant) |
| Kimi K3 (16 Jul 2026) | Moonshot AI, China | 57 — #3 | #2/39 Vals AI Index (74.7%); #1 Frontend Code Arena; #1 AutomationBench-AA (53%) |
| Claude Opus 4.8 (2026) | Anthropic, US | ~57 — comparable to K3 | SWE-bench Verified 88.6%; Terminal-Bench 2.1 71.91 (BenchLM/Vals neutral harness) |
| GPT-5.5 (Apr 2026) | OpenAI, US | ~56-57 | Held AA #1 at launch, April 2026 |
| Gemini 3.1 Pro (current Pro) | Google, US | top-8 arena tier | Multimodal and long-context strength (3.x line) |
| GLM-5.2 (16 Jun 2026) | Z.ai (Zhipu), China | 51 | SWE-bench Pro 62.1 — beats GPT-5.5 (58.6) |
| DeepSeek V4 Pro (Apr 2026, official Jul 2026) | DeepSeek, China | 44 | SWE-bench Verified 80.6% at $0.87/M output tokens |
On LMArena’s crowd-judged text leaderboard, the pattern holds: US closed models occupy the top four slots as of July 2026, with Kimi K3 the highest Chinese entry — provisionally top-five, with votes still stabilising — and Alibaba’s Qwen 3.7 Max the second Chinese model in the top ten (LMArena aggregations, July 2026). For what these models cost to run — where the rankings invert dramatically — see our AI model cost tracker.
Capability by capability: where China leads and lags
Coding: Chinese models now win individual benchmarks
Coding is where the gap is narrowest — and occasionally inverted. Z.ai’s GLM-5.2 scored 62.1 on SWE-bench Pro against GPT-5.5’s 58.6, the first time an open Chinese model has beaten a current US flagship on a major software-engineering benchmark (Z.ai published results, June 2026; Apidog analysis). Kimi K3 sits #1 on the Frontend Code Arena, ahead of Claude Fable 5, and #3 of 74 models on Vals AI’s SWE-bench run (Artificial Analysis; Vals AI, July 2026). The US retort: US models still top the harder end. On SWE-bench Verified, Claude Fable 5 leads overall at 95.0% (Morph leaderboard, 2026), with Opus 4.8 at 88.6% versus DeepSeek V4 Pro’s 80.6%. On Terminal-Bench 2.1’s neutral harness, GPT-5.6 Sol leads at 85.77, ahead of Claude Fable 5 (80.52), Opus 4.8 (71.91) and GLM-5.2 (67.79) — GLM-5.2’s oft-quoted 81.0 comes from Z.ai’s self-run harness (BenchLM/Vals; Z.ai results, June 2026). Verdict: parity on mainstream coding, US lead on the hardest agentic-coding evals.
Agentic and long-horizon tasks: the US lead is real but shrinking
Long-horizon agentic work — multi-step tasks over hours, not minutes — remains the clearest US advantage. Claude Fable 5 leads Artificial Analysis’s agentic composites (GDPval-AA v2 Elo 1760 vs Kimi K3’s 1668), yet K3 took #1 on AutomationBench-AA with 53% and second place on AA’s long-horizon agentic Elo (Artificial Analysis, July 2026). Practitioner reports consistently find Chinese models’ real-world agentic reliability lags their benchmark scores — OpenAI’s Roon predicted exactly this pattern for K3 (via Zvi, July 2026). Expect this row of the tracker to move fastest.
Reasoning and knowledge: a one-tier difference
On reasoning composites, Kimi K3’s 57 on the AA Intelligence Index makes it “comparable to Opus 4.8 and GPT-5.5 but behind Fable 5 and GPT-5.6” (Artificial Analysis, July 2026) — i.e. level with the US flagships of two months ago. GLM-5.2 posted 54.7 on Humanity’s Last Exam with tools, edging GPT-5.5’s 52.2 — though both figures come from Z.ai’s own tool-augmented harness; on the independent no-tools leaderboard the order reverses, with GLM-5.2 at 40.1 against GPT-5.5’s 44.3 — and 91.2 on GPQA-Diamond (Z.ai, June 2026). Chinese models also burn more tokens to get there: K3 is notably reasoning-heavy, though 21% more token-efficient than its predecessor K2.6 (Simon Willison, 16 July 2026).
Multimodal and long context: mixed picture
Multimodal remains a US — chiefly Google — stronghold: the Gemini 3.x line leads vision and video understanding, and GLM-5.2 ships text-only with no vision variant (Apidog, June 2026). Kimi K3 narrows this with native image and video input and strong vision scores (Simon Willison, July 2026). On long context, China has arguably levelled: Kimi K3, GLM-5.2 and DeepSeek V4 all ship 1M-token windows as standard, while Google’s reported 2M-token Gemini 3.5 Pro was still unreleased as of early July (TechTimes, 13 July 2026). How these models handle Asian languages specifically is a different question — see our Asian-language benchmark tour.
Safety and refusals: the widest gap of all
The largest US-China difference isn’t capability but guardrails. Multiple researchers reported Kimi K3 showing effectively absent safeguards on biology-risk tasks, and the UK AI Security Institute found Chinese models’ cyber capabilities narrowing towards US levels, with GLM-5.2 reaching Claude Opus 4.5-level on longer cyber reasoning tasks (via Zvi, 20 July 2026). For enterprises, this cuts both ways: fewer refusals in benign use, materially weaker safety assurance — a growing factor in Western procurement decisions.
The release timeline that closed the gap
The gap compressed because Chinese labs out-shipped their US rivals between releases. The cadence since early 2025: DeepSeek R1 (January 2025) started the cycle; Kimi K2 and GLM-4.5 (July 2025) established trillion-parameter open MoE models; OpenAI’s GPT-5 (August 2025), Google’s Gemini 3 and Anthropic’s Claude Opus 4.5 (November 2025) re-extended the US lead; Kimi K2 Thinking (November 2025) closed within weeks of it. Then 2026 accelerated: DeepSeek V4 preview with a 1.6T-parameter Pro and 284B Flash (24 April 2026, per DeepSeek’s API docs), GPT-5.5 (April 2026), Claude’s Opus 4.7/4.8 line and Fable 5, GLM-5.2 (open weights 16 June 2026, hosted access from 13 June), GPT-5.6 (rolled out from 9 July 2026), Gemini 3.6 Flash (21 July 2026), Kimi K3 (16 July 2026) and DeepSeek V4’s move from preview to official release with peak/off-peak pricing in mid-July 2026 (TechNode and SCMP reports; legacy API aliases retired 24 July 2026 per TechTimes).
Policy now reinforces cadence: at WAIC 2026 in Shanghai — the same week as the K3 launch — Xi Jinping committed China’s AI ecosystem to open-source release and global diffusion as explicit state strategy (Interconnects; AI Weekly, July 2026). Every major Chinese frontier model since has shipped, or committed to ship, open weights, most under MIT licences. The wider industrial context sits in our China AI market hub.
What the benchmarks don’t capture
Four caveats keep this scoreboard honest. First, self-reported numbers inflate: Moonshot’s own K3 figures (88.3 on Terminal-Bench — on the older 2.0 suite, not 2.1 — above Fable 5’s 84.6) exceed what independent harnesses reproduce, and Moonshot published no SWE-bench Verified score; circulating third-party figures diverge widely by harness (OpenClaw analysis, 18 July 2026). Second, benchmark deltas overstate practical parity — the analyst consensus is that Chinese models’ day-to-day agentic reliability trails their scores. Third, distillation muddies attribution: researchers note Chinese models’ jagged cyber profile may reflect training on outputs of safety-constrained US models (via Zvi, July 2026). Fourth, price-performance — not raw capability — is driving adoption: DeepSeek V4 Pro delivers output tokens at roughly 28x below Claude Opus 4.8 ($0.87 vs $25 per million), and CNBC pegs Chinese models at 60-90% cheaper overall (Morph, 2026; CNBC, 7 July 2026). A model four months behind at a tenth of the price is a different competitive object than the gap metric implies.
What’s next on both sides
The next release wave is already scheduled. China: Z.ai has said GLM-5.5 — over 1 trillion parameters, 1M-token context, trained with an eye on domestically produced chips — lands in August 2026 (Geeky Gadgets, 22 July 2026; AIBase); Alibaba’s Qwen 3.8, a 2.4T-parameter open-weights model, is announced as imminent (Interconnects, July 2026); Kimi K3’s open weights are promised by 27 July 2026; and DeepSeek R2 remains the perennial watch item — unconfirmed, with Reuters reporting founder Liang Wenfeng has repeatedly held it back on quality grounds. US: Google’s Gemini 3.5 Pro, reportedly targeting mid-July with a 2M-token context, had not shipped as of this update (TechTimes, 13 July 2026); OpenAI’s post-5.6 flagship and Anthropic’s next Claude release have no public dates. Each of these triggers a tracker refresh.
Update log
July 2026: Tracker launched after the Kimi K3 release wave. Baseline gap set at 3-5 months (Lambert) with a 3-8 month analyst range; scoreboard seeded with Kimi K3, GLM-5.2, DeepSeek V4 vs Claude Fable 5, GPT-5.6, Opus 4.8 and the Gemini 3.x line. Next scheduled review: after GLM-5.5 and the Kimi K3 weights release, August 2026.
FAQ: US vs Chinese AI model capability
How far behind are Chinese AI models in 2026?
Roughly 3-5 months behind the best US closed models as of July 2026, per Interconnects’ Nathan Lambert, with analyst estimates ranging from three to eight months. In mid-2025 the consensus was 6-9 months.
What is the best Chinese AI model right now?
Moonshot AI’s Kimi K3 (released 16 July 2026): #2 of 39 on the Vals AI Index and #3 on the Artificial Analysis Intelligence Index, behind only Claude Fable 5 and GPT-5.6. GLM-5.2 leads on some coding benchmarks; DeepSeek V4 leads on price-performance.
Do Chinese models beat US models on any benchmarks?
Yes. As of July 2026: GLM-5.2 beats GPT-5.5 on SWE-bench Pro (62.1 vs 58.6); Kimi K3 ranks #1 on the Frontend Code Arena and AutomationBench-AA. US models still lead overall composites and the hardest long-horizon agentic evals.
Why are companies using Chinese models if US models are better?
Price and open weights. Chinese models run 60-90% cheaper (CNBC, July 2026), most ship MIT-licensed weights for private deployment, and the capability gap is now months, not generations. That combination pushed Chinese models to up to 46% of weekly OpenRouter enterprise tokens in mid-2026.
How is the “gap in months” calculated?
It asks: when did US closed models last perform at the level of today’s best Chinese open model? Analysts triangulate from index scores (Artificial Analysis, Vals AI), arena Elo and pretraining-quality estimates. It is a judgement call, not a formula — which is why estimates span 3-8 months.
Track the broader regional picture in our Asia AI market hub, and the money behind these models in the China vs US AI race deep dive.
Discover more from Digital in Asia
Subscribe to get the latest posts sent to your email.