Scope
This assessment examines the proposition that the strategic contest in artificial intelligence is moving away from a narrow competition over which company possesses the strongest individual foundation model and toward a competition over which technological-industrial architecture can deliver usable intelligence most cheaply, reliably and frictionlessly across enterprises, governments and industrial systems, with particular attention to China, the United States and allied technology ecosystems between 2024 and September 2026, and with a forward analytical horizon extending through 2030. The underlying research mandate explicitly requires the argument to be tested against industrial organisation, network economics, path dependence, standard-setting, semiconductor constraints and enterprise integration rather than reduced to a comparison of chatbots.
AI Will Be Chosen as Infrastructure, Not as a Chatbot
Artificial intelligence is entering the phase in which technology choices harden into industrial structure. By September 2026, China had 1,112 filed generative-AI services, OpenAI said its Stargate programme had secured more than 10 GW of U.S. AI infrastructure, Anthropic had committed more than $100 billion over ten years to AWS technologies, Alibaba had pledged at least RMB 380 billion to AI and cloud infrastructure, and Alphabet was guiding to $175–185 billion of 2026 capital expenditure. These figures describe different instruments and cannot be added into a synthetic “AI investment” total, but they expose the same shift: the contest is leaving the chatbot screen and entering data centres, grids, semiconductor supply chains, enterprise identity systems and industrial workflows. The decisive question is no longer which model wins a benchmark, but which architecture makes intelligence cheapest, safest and hardest to dislodge once deployed.
China has already moved beyond the DeepSeek story
The Chinese market can no longer be reduced to DeepSeek. By 31 August 2026, the Cyberspace Administration of China recorded 1,112 generative-AI services completing filing procedures, up from 302 at the end of 2024 and 748 at the end of 2025, while 731 downstream applications or functions had registered direct use of filed models through APIs or equivalent mechanisms. Behind those numbers sits an increasingly differentiated market: Alibaba’s Qwen, Tencent’s Hy, Baidu’s ERNIE, Zhipu’s GLM, Moonshot’s Kimi, MiniMax, Huawei’s Pangu and DeepSeek pursue different combinations of open weights, sparse architectures, long context, agents, industrial deployment and low-cost inference.
The important fact is not proliferation by itself but the architecture surrounding it. Alibaba’s September 2026 strategy explicitly ran from chips and cloud infrastructure through models to agents; Tencent distributed its own Hy models and rival systems through TokenHub; Baidu linked ERNIE to PaddlePaddle, Qianfan and a cloud business that reported RMB 8.8 billion of AI Cloud Infrastructure revenue in Q1 2026, up 79% year on year; Huawei tied Pangu to Ascend accelerators, CANN software and industrial deployments. China is therefore building competing routes from model intelligence to production, not merely producing alternative chatbots.
Cheap intelligence matters only after it becomes good enough
The Chinese price structure already points toward commoditisation. DeepSeek V4.1-Flash was listed in September 2026 at $0.30 per million cache-miss input tokens and $1.20 per million output tokens at peak pricing; MiniMax M2.7 used the same $0.30/$1.20 structure; Tencent listed GLM-5.3-Flash at RMB 0.8 per million input tokens and RMB 2.8 per million output tokens through TokenHub. Kimi K3, by contrast, charged $3 per million cache-miss input tokens and $15 per million output tokens, reflecting a different performance and context architecture.
Those figures do not identify a winner because token price is not the correct economic denominator. In manufacturing, finance, healthcare or public administration, the relevant cost is the price of a verified completed task after retries, human checking, latency, integration, security and failure risk. A cheaper model that requires more supervision can cost more than a premium system; a more expensive model becomes irrational when several alternatives already cross the required performance threshold. The commercial battlefield therefore shifts from “best intelligence” toward capability-adjusted deployment cost.
The real lock-in is moving above and below the model
Open-weight models weaken the assumption that the model provider necessarily controls the ecosystem. Qwen, GLM, Hy and other Chinese models can be hosted through Chinese clouds, foreign hyperscalers, sovereign infrastructure or local inference engines; the same weights can therefore spread Chinese technical influence without creating Chinese cloud or hardware dependence. Meta’s Llama offers the Western precedent: by March 2025 Meta reported more than one billion downloads, while AWS, Azure, Google Cloud, NVIDIA and other providers captured much of the surrounding infrastructure value.
The deeper lock-in sits elsewhere. By 2026 Microsoft Foundry connected models and agents to Entra identity, Agent 365, Defender and Purview; AWS AgentCore combined runtime, identity, policies, tools, state and observability while remaining model-agnostic; OpenAI’s Agents API separated the managed agent harness from execution environments; Alibaba and Tencent similarly built multi-model control layers. Once thousands of agents hold corporate permissions, access internal databases, retain operational state and execute workflows, changing the underlying model can be easier than changing the control plane. AI’s durable standard may therefore be an identity, agent or orchestration architecture rather than a model family.
The West is no longer competing with frontier models alone
The idea that China is building an ecosystem while the West competes through a few premium models no longer survives the 2026 evidence. OpenAI has moved into physical infrastructure through Stargate, enterprise execution through Agents API, coding through Codex and organisational deployment through ChatGPT Work and Presence. Microsoft combines Azure, Foundry, Microsoft 365, GitHub, Entra and enterprise security. AWS combines Bedrock, AgentCore, IAM and its own Trainium accelerator family. Google combines Gemini, TPUs, Google Cloud, Workspace and an enterprise agent platform. NVIDIA sits beneath competing models through GPUs, CUDA, NIM, NeMo and AI Enterprise.
Capital reinforces that architecture. Microsoft indicated approximately $175 billion of calendar-2026 capital expenditure after lease-accounting adjustments; Alphabet guided to $175–185 billion; OpenAI said in April 2026 that Stargate had already secured more than 10 GW of U.S. AI infrastructure capacity; Anthropic’s AWS agreement covered up to 5 GW of additional compute and more than $100 billion over ten years of AWS technology commitments. The Western counter-model is not a single vertically integrated champion but a set of interlocking companies that compete at one layer while reinforcing one another at another.
Semiconductor sovereignty remains China’s hardest ceiling
China’s deployment architecture is more complete than the chatbot narrative suggests, but production autonomy remains less certain. Zhipu reported in September 2026 that GLM-5.3-Flash production inference was running on a cluster of more than 100,000 Chinese-made AI accelerators, while Huawei continued expanding Ascend, CANN and SuperPoD systems. That is evidence of meaningful domestic compute deployment, not merely strategic aspiration.
But the constraint has shifted rather than disappeared. U.S. export controls introduced in December 2024 targeted 24 categories of semiconductor-manufacturing equipment, three categories of software tools and high-bandwidth memory, precisely the upstream layers required to scale advanced AI systems. Domestic accelerators do not by themselves resolve fabrication yields, HBM supply, advanced packaging, lithography or software maturity. The distinction matters: China already possesses an increasingly integrated deployment stack, but the dossier does not establish a fully autonomous production stack.
Electricity and capital are becoming geopolitical inputs
The International Energy Agency projects global data-centre electricity consumption at approximately 945 TWh by 2030, roughly double current levels, with China and the United States together accounting for nearly 80% of projected global demand growth. China alone is projected to add approximately 175 TWh of data-centre demand between 2024 and 2030. At that scale, AI stops being only a software industry and becomes a grid-planning problem involving generation, substations, transmission, cooling and long-duration power contracts.
The same transformation is visible in capital markets. The OECD reported that AI firms captured 61% of global venture-capital value in 2025, receiving $258.7 billion, while AI infrastructure and hosting companies alone attracted $109.3 billion. Alibaba committed at least RMB 380 billion over three years to AI and cloud infrastructure; the European Union launched InvestAI with a €200 billion mobilisation objective, including a €20 billion AI-gigafactory facility; and the United Kingdom committed to expanding sovereign public AI compute by at least 20 times by 2030, alongside a £750 million heterogeneous AI supercomputer. The relevant geopolitical question is increasingly how quickly financial capital can be converted into energised, utilised compute.
By 2030, the winning architecture will combine intelligence, integration and availability
The next 12–24 months will show whether today’s infrastructure commitments become productive capacity or stranded ambition. Three tests matter. First, whether Chinese accelerator, HBM and packaging capacity grows fast enough to support domestic model expansion without creating a persistent cost penalty. Second, whether Microsoft, AWS, Google, OpenAI, Anthropic and NVIDIA can convert extraordinary capital expenditure into lower unit inference costs before grid and power constraints erase part of that advantage. Third, whether open standards such as MCP and A2A make models and agents sufficiently portable to prevent any one cloud or national stack from owning the control layer.
The cost of inaction will not fall on model laboratories alone. Governments that fail to secure sovereign or diversified compute will pay through dependency; utilities that cannot deliver grid capacity will lose data-centre investment; enterprises that allow one agent platform to absorb identity, data and workflow without portability will pay through switching costs; and industrial economies that treat AI as a procurement of software licences rather than as a redesign of infrastructure will discover that the decisive capital has already been committed elsewhere.
The 2030 standard will therefore not be determined by intelligence alone, because capability is diffusing; not by price alone, because cheap models that fail the task threshold destroy value; and not by infrastructure alone, because idle compute produces no productivity. It will be determined by the architecture that most efficiently combines sufficient intelligence, low integration friction and reliable physical capacity. AI will not ultimately be bought as a chatbot. It will be financed, powered and governed as infrastructure.
Navigational Index
Pillar I — From National AI Strategy to Industrial Infrastructure
Chapter 1 — Institutional, Historical and Physical Baseline
How China moved from the 2017 New Generation Artificial Intelligence Development Plan to the 2025–2026 “AI+” industrialisation architecture; the underlying digital economy, telecommunications network, server base, industrial internet, computing infrastructure, power system and policy institutions that make large-scale AI deployment possible.
Chapter 2 — Models, Compute and the Architecture of Deployment
The actual structure of the Chinese AI market beyond DeepSeek: Qwen, Hy, ERNIE, GLM, Kimi, MiniMax, Pangu and other model families; open-weight versus proprietary strategies; API, cloud and on-premise deployment; agents; inference economics; domestic accelerators; HBM; advanced packaging; semiconductor manufacturing constraints; and the degree to which China possesses an integrated stack rather than merely an extensive application layer.
Chapter 3 — Industrial Embedding and the Economics of Integration
How AI moves from foundation models into manufacturing, software development, logistics, finance, healthcare, telecommunications, public administration and industrial control; where integration costs arise; how enterprise switching costs accumulate; and under what conditions “good enough and cheap to integrate” becomes economically superior to frontier capability.
Pillar II — Standards, Platforms and the Geoeconomics of AI
Chapter 4 — Network Effects, Lock-In and the Battle for the Control Layer
Where AI network effects actually reside: weights, APIs, developer frameworks, agent systems, enterprise data, retrieval architectures, cloud regions, identity, security, compiler stacks, accelerators and industrial software; whether open-weight Chinese models create Chinese ecosystem dependence or instead commoditise the model layer and strengthen platform neutrality.
Chapter 5 — The Western Counter-Architecture
A like-for-like comparison of the Chinese ecosystem with OpenAI, Microsoft, Anthropic, Amazon, Google, Meta and NVIDIA; why the Western system can no longer be described as a collection of isolated frontier laboratories; how hyperscalers, model developers, accelerator vendors and enterprise software increasingly form competing integrated stacks.
Chapter 6 — Stress-Testing the Oil, Telecom, Solar and Battery Analogies
A rigorous examination of which historical mechanisms actually transfer to AI: commodity fungibility, infrastructure dependence, industrial learning curves, economies of scale, overcapacity, standards formation, installed-base effects, vertical integration, supply-chain chokepoints and strategic control.
Pillar III — Capital Allocation, Strategic Dependency and the 2030 Contest
Chapter 7 — Geopolitics as Capital Allocation
AI capital expenditure as geopolitical infrastructure: semiconductor fabs, accelerators, data centres, cloud capacity, electricity generation, grid connections, networking, public procurement, sovereign compute, venture finance and long-duration corporate commitments; how sunk capital can turn temporary technological preferences into persistent strategic alignment.
Chapter 8 — Cognitive Infrastructure, Scenarios and the 2030 Standard
The conditions under which AI becomes infrastructure rather than software; alternative pathways for Chinese, U.S.-allied and multi-homed ecosystems; falsifiers of the central thesis; observable indicators; strategic dependencies; capital-allocation implications; and the final assessment of whether the dominant AI architecture will be determined principally by intelligence, integration cost, infrastructure availability or some combination of all three.
Executive reconstruction of the thesis
The strongest defensible version of the thesis is not that Chinese artificial-intelligence models are inevitably superior, nor that China has already established a self-sufficient AI technology stack, but that the economically decisive unit of competition is expanding from the foundation model toward the deployable AI system, incorporating inference cost, cloud capacity, accelerator availability, developer tooling, agent infrastructure, enterprise software, regulatory compatibility, local deployment, industrial data and the human cost of integration. Chinese policy and corporate behaviour increasingly conform to this interpretation: China’s State Council formally directed the expansion of an “AI+” architecture across science, industry, consumption, public services and governance, while explicitly calling for AI-chip development, software ecosystems, large-scale intelligent-computing clusters, a nationally integrated computing network and standardised scalable cloud services. State Council Opinions on Deepening Implementation of the “Artificial Intelligence Plus” Action — State Council of the PRC — Aug 2025
The empirical record also supports the argument that China’s generative-AI sector can no longer reasonably be represented as a contest among one or two national champions. The Cyberspace Administration of China reported that 1,112 generative-AI services had completed filing procedures by 31 August 2026, compared with 302 at the end of 2024 and 748 at the end of 2025; it additionally reported 731 registered applications or functions that directly call filed models through APIs or other mechanisms, which is evidence of substantial proliferation at both model-service and application layers even though regulatory filings cannot by themselves establish commercial success, quality or active utilisation. Announcement on Filed Generative AI Services, July–August 2026 — Cyberspace Administration of China — Sep 2026 Announcement on 2025 Filed Generative AI Services — Cyberspace Administration of China — Jan 2026
The thesis becomes substantially weaker, however, if “complete ecosystem” is interpreted as technological autarky from semiconductor design through leading-edge fabrication, HBM, lithography, manufacturing equipment and software tooling. United States export-control architecture continues explicitly to target China’s access to advanced computing chips, high-bandwidth memory, semiconductor manufacturing equipment and supporting software, while the January 2026 licensing adjustment allowing H200-, MI325X- and comparable-class accelerators to approved Chinese customers on a case-by-case basis reinforces rather than eliminates the fact that leading AI compute remains strategically exposed to external licensing decisions. Department of Commerce Revises License Review Policy for Semiconductors Exported to China — BIS — Jan 2026 Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors — BIS — Dec 2024
The proposition that the West competes principally through a small number of isolated frontier models is also no longer supported by the 2026 product architecture. Microsoft Foundry integrates models, agents, knowledge tools, monitoring, evaluation, networking and enterprise governance; AWS has moved Bedrock toward model and agent orchestration while adding OpenAI models and managed agents; NVIDIA AI Enterprise explicitly spans application and infrastructure layers; and Google’s Vertex/agent platform combines model access with enterprise deployment and orchestration, meaning that both Chinese and U.S.-allied systems are moving toward full-stack competition, albeit from different industrial structures and regulatory environments. What is Microsoft Foundry? — Microsoft — 2026 NVIDIA AI Enterprise — NVIDIA — 2026 Amazon Bedrock now offers OpenAI models, Codex, and Managed Agents — AWS — Apr 2026
The principal judgment is therefore narrower and stronger than the original slogan: AI competition is increasingly a contest over the total cost and institutional friction of converting compute and models into dependable workflow-level intelligence, but neither China nor the United States can yet be reduced to a single competitive logic, and model capability remains economically decisive wherever a capability threshold cannot be crossed by a cheaper alternative. The emerging standard will consequently be shaped not by benchmark leadership alone and not by price alone, but by the interaction among capability thresholds, inference economics, compatibility, developer ecosystems, enterprise switching costs, compute availability, sovereign requirements, power infrastructure and industrial deployment. AI alone won’t change your business. The system running it will — Microsoft — Jun 2026 Energy and AI — IEA — 2025
Claim-by-claim audit
| Proposition | Type | Assessment from verified record | Principal evidence |
|---|---|---|---|
| China has moved beyond a single national AI champion | Empirical | Strongly supported as a statement about product and provider proliferation; commercial concentration remains a separate question | CAC filing data — Sep 2026 |
| Chinese firms are competing through model + cloud + applications + industrial deployment | Empirical | Supported for Alibaba, Tencent, Baidu and Huawei, although architectures differ materially | Alibaba full-stack AI roadmap — Sep 2026 |
| The West competes mainly through frontier-model quality | Empirical/analytical | Overstated by 2026; Microsoft, AWS, Google, NVIDIA and OpenAI expose extensive enterprise/platform layers | Microsoft Foundry — 2026 |
| Cheapest-to-integrate AI will defeat the highest-capability AI | Causal | Conditional rather than universal; valid above task-specific capability and reliability thresholds | Microsoft enterprise-agent architecture — Jun 2026 |
| Open weights produce ecosystem lock-in | Causal | Ambiguous; they can diffuse a model family while simultaneously weakening model-layer rents and increasing framework portability | Tencent Hy4 preview — Aug 2026 |
| China possesses a complete AI production stack | Empirical | Not established; advanced compute, HBM and manufacturing equipment remain exposed to foreign controls | BIS semiconductor-control package — Dec 2024 |
| Geopolitics is increasingly capital allocation | Analytical | Substantially supported if “capital” includes compute, power, data-centre, semiconductor, cloud and procurement allocation | Energy and AI — IEA — 2025 |
| AI will be “chosen like oil” | Analogy | Useful for infrastructure dependence and strategic supply, misleading if interpreted as commodity fungibility | Energy Technology Perspectives 2026 — IEA |
From a few champions to a differentiated Chinese AI market
China now exhibits several overlapping competitive layers rather than a single vertically coordinated model hierarchy. DeepSeek competes through open models, aggressive inference economics and API compatibility; Alibaba operates Qwen as a model family embedded within Alibaba Cloud Model Studio and, as of September 2026, explicitly describes its strategy as extending from proprietary AI chips through cloud infrastructure and foundation models to agents; Tencent has rebuilt Hunyuan, now Hy, around agent use, open sourcing and deep product integration; Baidu combines ERNIE, its Qianfan platform, accelerator infrastructure and AI applications; Huawei’s Pangu architecture concentrates more heavily on industrial and sector-specific models; while MiniMax, Moonshot/Kimi and Zhipu’s GLM family increasingly compete around multimodality, coding, long-context reasoning and agents. Alibaba Unveils Roadmap on Full-Stack AI Strategy from Chips, Cloud Infrastructure, Models to Agents — Alibaba Cloud — Sep 2026 Tencent Releases and Open-Sources Tencent Hy4 preview — Tencent — Aug 2026 ERNIE 5.1 Officially Released — Baidu — May 2026
This diversity matters because the relevant segmentation increasingly occurs by deployment mode and workflow, not merely model intelligence. Alibaba Cloud exposes Qwen alongside third-party models through OpenAI-compatible APIs and regional endpoints; Tencent’s TokenHub routes between proprietary and third-party models and is integrated with enterprise agent infrastructure; Baidu Qianfan exposes more than 60 models and 200 APIs across multiple modalities; Huawei combines Pangu models with ModelArts Studio, data engineering, model development and application-development toolchains; and Zhipu makes GLM models available through APIs, open weights and local-serving frameworks including vLLM and SGLang. What is Alibaba Cloud Model Studio — Alibaba Cloud — Sep 2026 Tencent Cloud Productivity Agent Suite — Tencent — Jun 2026 Baidu Qianfan Developer Center — Baidu AI Cloud — 2026 What Is PanguLM? — Huawei Cloud — Jan 2026
The market is nevertheless not synonymous with spontaneous laissez-faire competition, because model proliferation operates inside a state-defined regulatory and industrial-policy architecture. The CAC filing regime establishes conditions for public generative-AI services, while the State Council’s August 2025 “AI+” programme explicitly links models to national computing infrastructure, chips, data, industry, public governance and standards; the appropriate analytical description is therefore competitive plurality inside a strategically directed national development framework, rather than either a single state champion or a conventionally fragmented private market. Artificial Intelligence Plus Action — State Council — Aug 2025 CAC Generative AI Filing Announcement — Sep 2026
Model quality versus completeness of stack
The economically meaningful choice between “best model” and “best-integrated stack” is not binary, because model quality behaves as a constraint before it behaves as a differentiator: when two models both exceed the accuracy, reliability, latency and domain-performance threshold required for a workflow, differences in integration cost, inference cost, data locality, security certification and tool compatibility can dominate procurement, whereas below that threshold no amount of cheap inference transforms an inadequate model into an acceptable production system. Microsoft’s own 2026 enterprise architecture articulates this transition explicitly by arguing that production AI requires contextualisation, governance, observability and workflow integration around the model rather than model access alone. AI alone won’t change your business. The system running it will — Microsoft — Jun 2026
Current API pricing demonstrates why this distinction matters, while simultaneously showing why simple token-price tables cannot establish an overall winner. DeepSeek’s current pricing page lists DeepSeek-V4.1-Flash peak rates of $0.30 per million cache-miss input tokens and $1.20 per million output tokens, with lower off-peak rates; Alibaba’s Qwen3.6-Plus Beijing pricing lists $0.276 input and $1.651 output per million tokens for contexts up to 256K; OpenAI lists GPT-5.6 Sol at $4 input and $20 output, Terra at $2 and $12, and Luna at $0.20 and $1.20, while Google’s pricing architecture includes materially cheaper Flash tiers than premium reasoning models. These prices concern products with materially different capabilities, caching rules, context limits, regional availability and service characteristics, so they establish price segmentation, not quality-adjusted superiority. Models & Pricing — DeepSeek API — Sep 2026 Qwen3.6-Plus Model Pricing — Alibaba Cloud — 2026 GPT-5.6 Sol — OpenAI API — 2026 Agent Platform Pricing — Google Cloud — 2026
The strongest challenge to the original thesis is consequently that Western firms have themselves moved aggressively toward horizontal and vertical stack integration. Microsoft Foundry provides more than model inference by combining agent hosting, model catalogues, knowledge integration, monitoring, evaluation, networking, identity and policy; NVIDIA AI Enterprise places NIM, NeMo and application frameworks over CUDA, drivers, Kubernetes operators, orchestration and infrastructure management; AWS increasingly combines models from competing vendors with AgentCore and enterprise controls; and Anthropic’s 2026 agreement with Amazon secures up to 5 GW of future compute while extending Claude deeper into AWS enterprise infrastructure. Microsoft Foundry — Microsoft — 2026 NVIDIA AI Enterprise Overview — NVIDIA — 2026 Anthropic and Amazon expand collaboration for up to 5 gigawatts — Anthropic — Apr 2026
The resulting competitive model is therefore better expressed as quality-adjusted total deployment cost rather than model quality versus ecosystem completeness: total cost includes inference, accelerator utilisation, data movement, engineering time, fine-tuning or retrieval infrastructure, compliance, observability, security, model-switching costs, latency, error remediation and the economic value of tasks successfully completed, because an apparently expensive model can be cheaper per successful workflow if it requires fewer retries or less engineering, while an inexpensive model can dominate high-volume workloads once capability differences cease to matter. NVIDIA AI Enterprise — NVIDIA — 2026 GPT-5.6 — OpenAI — 2026
Historical analogy stress-test
The oil analogy is useful principally at the infrastructure and geopolitical layers, because both oil and AI involve capital-intensive upstream capacity, bottlenecks, strategic dependence, state intervention and downstream economies that cannot function without reliable supply; it becomes misleading when extended to the product layer because crude grades remain substantially fungible commodities whereas foundation models differ in capability, behaviour, languages, tool use, safety policies, context, deployment requirements and compatibility, while copied or open model weights can have near-zero marginal reproduction cost in a way that physical hydrocarbons cannot. The “AI will be chosen like oil” formulation should therefore be retained as a geopolitical metaphor for infrastructure dependence, not treated as a literal model of market organisation. Energy and AI — International Energy Agency — 2025
The photovoltaic analogy is considerably stronger at the industrial-policy layer, because the IEA documented how China’s combination of manufacturing scale, supply-chain integration, industrial policy and domestic demand produced more than an 80% share across major solar manufacturing stages and contributed to very large global cost reductions, while Energy Technology Perspectives 2026 still records China at approximately 85% of solar and 80% of lithium-ion battery supply-chain production capacity. The transferable mechanism is not “China always wins through low prices”; it is that repeated investment, supplier clustering, production learning, infrastructure and overcapacity can drive costs downward sufficiently to restructure international adoption patterns even where other jurisdictions retain technological capabilities. Solar PV Global Supply Chains — IEA — 2022 Energy Technology Perspectives 2026 — IEA
The battery analogy is similarly instructive because integration can migrate from manufacturing scale toward upstream materials, components and downstream deployment, with the IEA reporting that China accounted for more than 80% of global battery-cell production in 2025 and still larger shares of some active-material production; the analogy nevertheless weakens at the foundation-model layer because software can move between hardware platforms, APIs and jurisdictions far faster than electrochemical manufacturing plants can be relocated, while interoperability standards can deliberately reduce switching costs that physical supply chains cannot eliminate. Electric Vehicle Batteries — Global EV Outlook 2026 — IEA
The telecommunications analogy is arguably the most analytically useful for the model-plus-platform layer, because AI adoption increasingly involves standards, APIs, interfaces, developer familiarity, tooling, certification and installed-base effects rather than the purchase of an isolated commodity; however, AI currently lacks the same degree of mandatory global technical standardisation found in telecommunications, and OpenAI-compatible APIs, open-weight models, containerisation, Kubernetes, vLLM and other portability mechanisms can reduce the degree to which adopting one model irrevocably commits an organisation to one national ecosystem. Alibaba’s explicit use of OpenAI-compatible interfaces and Microsoft’s support for models from multiple competing providers demonstrate that interoperability can be a competitive strategy in itself. Alibaba Cloud Model Studio — Alibaba Cloud — Sep 2026 Microsoft Foundry Models and Agents — Microsoft — 2026
Where network effects and lock-in actually reside
The evidence does not support locating the principal network effect exclusively in model weights. The stronger lock-in candidates are the developer and agent toolchain, proprietary enterprise data and retrieval architecture, identity and security environment, observability layer, cloud networking, accelerator/compiler stack and business-process redesign surrounding the model, because those assets accumulate organisation-specific integration effort and cannot be replaced merely by changing an API endpoint. NVIDIA’s CUDA-centred application and infrastructure stack illustrates hardware-software complementarity, while Microsoft’s Foundry architecture illustrates how identity, networks, policies, tools and monitoring can become part of the adoption decision independently of the underlying model. NVIDIA AI Enterprise Overview — NVIDIA — 2026 Microsoft Foundry — Microsoft — 2026
Open-weight Chinese models therefore have two opposing economic effects that must not be conflated: releasing weights can increase ecosystem reach because developers can fine-tune, self-host and incorporate the architecture into downstream products, but the same openness can commoditise the model layer, allowing Western clouds and infrastructure providers to serve Chinese-origin models without adopting a Chinese cloud, chip or enterprise stack. Microsoft’s own Foundry documentation provides DeepSeek access, while Alibaba Cloud hosts third-party models and OpenAI-compatible interfaces; this is evidence that model diffusion can increase cross-ecosystem substitution as readily as national lock-in. Microsoft Foundry Documentation — Microsoft — 2026 Alibaba Cloud Supported Models — Alibaba Cloud — Sep 2026
DeepSeek is particularly important because it illustrates this ambiguity. Its models have combined open releases with API compatibility and aggressive pricing, while NVIDIA itself has incorporated technology distilled from DeepSeek-R1 into its own Llama Nemotron products; rather than demonstrating a simple migration into a Chinese-controlled technology stack, this shows how open-model innovation can propagate into a multinational infrastructure ecosystem whose value may ultimately be captured by chip vendors, cloud providers, agent platforms or enterprise integrators rather than by the originating laboratory. Introducing DeepSeek-V3 — DeepSeek — Dec 2024 NVIDIA AI Enterprise — NVIDIA — 2026
Capital allocation as geopolitics
The phrase “geopolitics is capital allocation” becomes analytically meaningful when translated into observable investment decisions: accelerator procurement, semiconductor-fabrication capacity, data-centre construction, grid connections, power generation, cloud commitments, sovereign-compute programmes, public procurement, venture finance and the allocation of engineering talent. The IEA estimates that global data-centre electricity consumption rises to approximately 945 TWh by 2030 in its base case, more than double current consumption, while China and the United States together account for nearly 80% of projected global data-centre electricity-demand growth through 2030; this means that AI competition already extends directly into electricity systems, grid infrastructure, permitting and capital-intensive physical assets. Energy Demand from AI — IEA — 2025
The same logic is visible at company level. Anthropic and Amazon announced in April 2026 an agreement providing for up to 5 GW of new compute capacity, with Anthropic stating that it would commit more than $100 billion over ten years to AWS technologies; NVIDIA separately described new capital structures for multi-tenant “AI factories”; and Alibaba’s September 2026 roadmap explicitly linked proprietary chips, cloud infrastructure, foundation models and agents, demonstrating that what appears at the application layer as a contest among models is supported underneath by increasingly large commitments of financial, electrical and semiconductor capital. Anthropic and Amazon Expand Collaboration for up to 5 Gigawatts — Anthropic — Apr 2026 NVIDIA Unlocks AI Compute at Scale — NVIDIA — Jul 2026 Alibaba Full-Stack AI Strategy — Alibaba Cloud — Sep 2026
This produces a more precise interpretation of the thesis: geopolitical advantage in AI is partly the ability to determine where capital becomes irreversible. Once an enterprise has redesigned workflows around a model-access platform, a government has built sovereign data centres around a hardware/software architecture, or a cloud provider has contracted power and accelerators for years, future technological choices are constrained by sunk cost, compatibility and switching burdens, which makes infrastructure and procurement decisions today potentially more strategically consequential than temporary differences in benchmark rankings. Energy and AI — IEA — 2025 NVIDIA AI Cloud Ecosystem Expands Worldwide — NVIDIA — May 2026
Cognitive infrastructure: from chatbot to production system
“Cognitive infrastructure” is defensible as an analytical category when defined operationally as the layer through which models become repeatable organisational production inputs: agents embedded in ERP, CRM, MES and development environments; retrieval over corporate information; automated document and case handling; industrial visual inspection; software generation; procurement and logistics assistance; scientific computation; public-service processing; and other systems in which inference is consumed continuously rather than through occasional chatbot interaction. Huawei reports Pangu deployment across more than 500 scenarios in more than 30 industries, while its cement-industry collaboration describes an AI operating architecture spanning central training, edge inference, continuous learning and more than 40 production scenarios, demonstrating a form of industrial embedding materially different from consumer chatbot usage. Huawei Cloud Announces Pangu Models 5.5 — Huawei Cloud — Jun 2025 Conch Group and Huawei Cement Industry AI Model — Huawei — Apr 2025
Tencent provides another observable example of this shift: its June 2026 enterprise suite combines WorkBuddy, CodeBuddy, an Agent Development Platform, TokenHub, runtime infrastructure and sector solutions across more than 20 industries, while Tencent reports that Hy3 integration into WorkBuddy reduced first-response time by 54% and average task-completion time by 47%; these are first-party performance claims rather than independently audited measurements, but they illustrate the increasingly relevant metric—workflow completion economics rather than benchmark score in isolation. Tencent Cloud Debuts Productivity Agent Suite — Tencent — Jun 2026
Baidu similarly demonstrates that the Chinese market is moving into infrastructure and monetisation rather than remaining a research showcase: the company reported RMB 8.8 billion of AI Cloud Infrastructure revenue in Q1 2026, up 79% year on year, with GPU Cloud revenue up 184%, while simultaneously deploying ERNIE and agent products; because these are company-reported financial categories they should not be treated as an independent measure of national adoption, but they provide concrete evidence that AI infrastructure is becoming a revenue-producing business rather than merely a benchmark race. Baidu Announces First Quarter 2026 Results — Baidu Investor Relations — 2026
The production-layer constraint
The strongest structural weakness in the “complete Chinese ecosystem” thesis remains the distinction between application-stack completeness and production-stack completeness. China can increasingly offer models, APIs, cloud infrastructure, agents, industry applications and domestic accelerators, but the U.S. export-control framework specifically identifies advanced-node semiconductor manufacturing equipment, relevant software and HBM as strategic control points, with the December 2024 rules adding controls on 24 categories of semiconductor manufacturing equipment, three types of software tools and high-bandwidth memory, followed by further foundry due-diligence measures in January 2025. Commerce Strengthens Export Controls on Advanced Semiconductors — BIS — Dec 2024 Commerce Strengthens Restrictions on Advanced Computing Semiconductors — BIS — Jan 2025
The January 2026 U.S. revision permitting H200, MI325X and comparable semiconductor exports to approved Chinese customers under case-by-case review changes the degree of constraint but not its strategic character, because access remains conditioned by licensing, purchaser compliance procedures, capacity considerations and third-party testing. An ecosystem dependent on policy-contingent access to externally controlled accelerators cannot yet be described as fully autonomous, even when domestic substitutes are improving; conversely, continued restrictions can create powerful incentives for Chinese accelerator, compiler and inference-software development, so the same constraint may accelerate localisation over a longer horizon. Department of Commerce Revises License Review Policy for Semiconductors Exported to China — BIS — Jan 2026
The correct distinction is therefore between a commercially increasingly complete application and deployment stack and a production stack whose leading-edge self-sufficiency remains unproven. This distinction is strategically decisive because successful industrial embedding can proceed despite an upstream bottleneck for some time, particularly through efficiency improvements, model sparsity, domestic accelerators and allocation of imported compute, but rapid increases in training or inference demand can restore the bottleneck whenever compute supply becomes binding. State Council Artificial Intelligence Plus Action — Aug 2025 BIS Advanced Semiconductor Controls — Dec 2024
Why the chatbot-race framing is increasingly inadequate
The chatbot framing survives because consumer interfaces produce visible, comparable moments of competition, benchmark releases generate simple rankings, and model launches can be reported more easily than enterprise migration, data-centre investment, compiler ecosystems or integration costs; it remains useful for measuring consumer mindshare and selected model capabilities, but it is increasingly inadequate for evaluating where durable economic rents and strategic dependencies are forming. Alibaba’s September 2026 Model Studio catalogue, for example, simultaneously offers Qwen, DeepSeek, GLM, Kimi, MiniMax and other models, meaning that the cloud platform can capture developer and enterprise activity even when the underlying model provider changes. Alibaba Cloud Model Studio Supported Models — Sep 2026
Western platforms reveal precisely the same phenomenon: Microsoft Foundry markets access to models from Microsoft, OpenAI, Anthropic, Meta and others under one governance and deployment environment, while AWS increasingly treats Bedrock as a multi-model enterprise control plane. The strategic competition is consequently becoming partly a contest over who owns the layer above the model, because that layer can preserve customer relationships even as the identity of the technically preferred model changes. Microsoft Foundry — Microsoft — 2026 Amazon Bedrock and OpenAI — AWS — Apr 2026
Competitive-logic model: quality-max versus stack-min-cost
The two competitive logics can therefore be formalised without assuming that either universally dominates. A quality-max strategy maximises expected economic value from each difficult task even at a higher inference price and is rational where failure costs are high, reasoning complexity is extreme, labour substitution value is large or capability differences remain substantial; a stack-min-cost strategy minimises the total cost of reliably deploying adequate intelligence and becomes rational once several models clear the relevant capability threshold and the dominant costs move toward integration, infrastructure, latency, governance and scale. Current model catalogues on Alibaba Cloud, Microsoft Foundry and AWS demonstrate that major platforms are already institutionalising this logic by letting customers select among models rather than assuming one model should serve every workload. Alibaba Model Studio — Sep 2026 Microsoft Foundry — 2026
| Condition | Quality-max tends to dominate | Stack-min-cost tends to dominate |
|---|---|---|
| Capability gap | Large and economically consequential | Small above required threshold |
| Failure cost | High | Low/moderate and recoverable |
| Volume | Moderate | Very high |
| Latency sensitivity | Secondary | High |
| Data sovereignty | Manageable through selected provider | Local/on-prem deployment decisive |
| Workflow integration | Limited | Extensive |
| Switching cost | Low | High after embedding |
| Hardware scarcity | Secondary to capability | Strong incentive for efficient models |
| Regulation | Model can satisfy required controls | Local ecosystem enjoys compliance advantage |
| Task profile | Frontier research, hard coding, complex reasoning | High-volume enterprise automation, retrieval, routine agents |
These conditions explain why the original thesis is strongest in public administration, manufacturing, customer operations, logistics, coding assistance and other high-volume settings where “adequate intelligence × low cost × integration” can be economically more important than the final increments of benchmark performance, while it is weaker for scientific research, advanced cybersecurity, difficult software engineering or high-consequence analysis where incremental capability can have disproportionately high economic value. Huawei Pangu Models 5.5 — Huawei Cloud — Jun 2025 GPT-5.6 — OpenAI — 2026
Capital-allocation and standard-setting implications
For sovereign and institutional capital, the relevant exposure is therefore no longer confined to laboratories developing frontier models, because value can accumulate in semiconductor equipment, accelerators, memory, networking, data-centre construction, electricity generation, grid infrastructure, inference software, cloud orchestration, industrial integration and domain-specific applications. The IEA’s projection that data-centre electricity consumption reaches approximately 945 TWh by 2030, together with its warning that around 20% of planned data-centre projects could face delays unless grid constraints are addressed, demonstrates that access to electricity and physical infrastructure has become part of AI competitiveness rather than an external utility assumption. Energy and AI Executive Summary — IEA — 2025
For corporate buyers, the key metric should consequently migrate from “best model” toward cost per successfully completed governed workflow, accompanied by time-to-deployment, engineering hours, accelerator utilisation, latency, error rate, data-governance requirements, portability and vendor concentration. This does not favour a national ecosystem automatically; it instead creates a procurement environment in which a Chinese open-weight model running on a Western cloud, a Western model exposed through an Asian cloud, or multiple models dynamically routed inside one enterprise platform can all be economically rational outcomes. Tencent’s TokenHub and Microsoft’s multi-model Foundry architecture already make this model-routing logic explicit. Tencent Cloud Productivity Agent Suite — Jun 2026 Microsoft Foundry — 2026
This also weakens the strongest interpretation of a “Chinese standard” spreading through open models, because compatibility can generate model adoption without stack adoption. If DeepSeek, Qwen, GLM or Kimi weights are deployed through NVIDIA hardware, Western inference frameworks, Microsoft Azure or an independent sovereign cloud, their technical diffusion does not automatically shift control of the infrastructure layer to China; conversely, repeated exposure to Chinese model architectures, tool interfaces and open-source components can still influence developer practice and reduce dependence on proprietary Western APIs. DeepSeek-V3.2 Open Release — DeepSeek — Dec 2025 GLM-5.2 — Z.ai — Jun 2026
Falsifiers and alternative hypotheses
Alternative hypothesis: capability remains the dominant bottleneck. The ecosystem thesis would weaken materially if frontier systems continued to expand the range of economically valuable tasks faster than lower-cost competitors could close the gap, because enterprises would repeatedly pay large premiums for models capable of completing work that cheaper alternatives simply cannot perform. OpenAI’s differentiated Sol, Terra and Luna architecture itself implies that capability segmentation remains commercially material rather than disappearing into commodity inference. GPT-5.6 Sol — OpenAI API — 2026 GPT-5.6 Luna — OpenAI API — 2026
Alternative hypothesis: AI becomes multi-homed rather than locked in. If standard APIs, open weights, containerised inference, routing platforms and common agent protocols continue to reduce model-switching costs, enterprises may deliberately avoid choosing a single ecosystem and instead arbitrage among providers on capability, cost and jurisdiction; Alibaba’s OpenAI-compatible interfaces, Microsoft’s multi-model platform and Tencent’s dynamic model-routing architecture provide observable mechanisms through which such multi-homing can develop. Alibaba Cloud Model Studio — Sep 2026 Tencent Cloud TokenHub — Tencent — Jun 2026
Alternative hypothesis: compute restrictions create a persistent Chinese capability ceiling. This explanation gains support if domestic accelerators, HBM, semiconductor manufacturing equipment and software ecosystems fail to substitute sufficiently for restricted foreign technology while AI workloads continue to increase their compute intensity, and it weakens if Chinese firms demonstrate sustained frontier-class training and inference at scale on domestically controlled hardware and toolchains; the continuing BIS controls identify exactly the technologies whose evolution should therefore be monitored. BIS Semiconductor Export Controls — Dec 2024
Alternative hypothesis: Western stack acceleration neutralises China’s integration advantage. This pathway has already gained evidence during 2026 because Microsoft, AWS, NVIDIA, Google and OpenAI are increasingly productising agents, orchestration, enterprise governance and inference infrastructure rather than relying on model leadership alone, which means the thesis should not be tested against a static assumption that Western firms remain structurally committed to a “few champions, best model” strategy. Microsoft’s System for the Agentic Enterprise — Jun 2026 NVIDIA AI Enterprise — 2026
Research agenda and indicator dashboard
| Indicator to monitor through 2030 | Why it matters | Evidence required |
|---|---|---|
| Quality-adjusted inference cost by task | Tests whether cheaper systems actually deliver cheaper completed work | Reproducible enterprise task suites plus live API prices |
| Chinese accelerator share of domestic inference | Tests production-layer autonomy | Vendor shipments, cloud instance catalogues, procurement data |
| Domestic HBM and advanced packaging availability | Tests upstream constraint | Manufacturer filings and official production statistics |
| Enterprise production deployments by model family | Distinguishes downloads from real adoption | Procurement records, audited company disclosures |
| Share of Chinese models served on non-Chinese clouds | Tests “Trojan standard” versus commoditisation | Cloud catalogues and usage data |
| Model-switching time and engineering cost | Measures actual ecosystem lock-in | Enterprise migration studies |
| Electricity and grid connection for AI data centres | Measures physical capital constraint | Utility, permitting and investment records |
| Agent transaction/workflow volume | Measures transition from chatbot to infrastructure | Platform disclosures and audited application data |
| AI public procurement by jurisdiction | Measures state-driven standard formation | Tender and contract databases |
| API and framework compatibility | Measures multi-homing potential | Platform technical documentation |
The most important missing dataset is a genuinely comparable, longitudinal measure of total cost per successfully completed enterprise task, because token prices, benchmark scores and model downloads each observe only one component of the economic problem. A rigorous test would combine task success, latency, token consumption, retries, human verification, integration labour, hardware utilisation and governance cost across identical workflows and across Chinese, U.S. and open-weight architectures, thereby measuring the precise proposition on which the thesis ultimately depends rather than substituting rankings or headline API prices for production economics. DeepSeek Models & Pricing — 2026 OpenAI API Pricing — 2026 Google Agent Platform Pricing — 2026
What the thesis gets right, what it overstates, and what matters next
The thesis gets its central structural intuition right: the strategic object called “AI” is becoming larger than the model, because usable intelligence increasingly depends on compute, inference software, cloud services, agents, enterprise data, applications, electricity, developer ecosystems and institutional integration, while China’s 2025–2026 policy and corporate trajectory provides substantial evidence of deliberate movement toward precisely this kind of vertically and horizontally connected architecture. The rapid expansion to 1,112 filed generative-AI services, Alibaba’s chip-to-agent roadmap, Tencent’s model-product co-design, Baidu’s growing AI infrastructure business and Huawei’s industrial deployments collectively make the “chatbot race” framing analytically inadequate. CAC Generative AI Filing Announcement — Sep 2026 Alibaba Full-Stack AI Roadmap — Sep 2026 Huawei Pangu Models 5.5 — Jun 2025
What the thesis overstates is the asymmetry between China and the West. China does not yet possess publicly demonstrated independence across every critical production layer, while Western technology companies are no longer competing only through a handful of premium models and have themselves constructed powerful integrated ecosystems spanning chips, cloud, models, agents, enterprise software and governance. The contest is therefore not frontier intelligence versus ecosystem completeness; it is increasingly ecosystem versus ecosystem, with different strengths, bottlenecks, ownership structures and geopolitical dependencies. BIS Advanced Semiconductor Controls — Dec 2024 Microsoft Foundry — 2026 NVIDIA AI Enterprise — 2026
The decisive question through 2030 is consequently not whether DeepSeek defeats ChatGPT, whether Qwen defeats Gemini, or whether one benchmark leaderboard temporarily changes order; it is whether one architecture can convert increasingly capable models into cheaper, more reliable and more institutionally embedded units of productive work while retaining sufficient access to semiconductors, electricity, developers, data, capital and international markets. If that transition occurs, AI will indeed start to resemble infrastructure more than software purchased model by model, although the better historical analogy will probably be a hybrid of telecommunications standards, cloud computing, semiconductor platforms and industrial supply chains rather than oil alone. Energy and AI — IEA — 2025 NVIDIA AI Enterprise Overview — 2026
The most decision-relevant signpost is therefore where organisations make investments that become difficult to reverse: accelerator fleets, data centres, electricity contracts, enterprise agent platforms, data architectures, developer toolchains and redesigned business processes. Model leadership can change within months; an installed technological-industrial architecture can shape procurement and behaviour for years, which is why the strongest version of the original proposition is ultimately not that “AI will be chosen like oil,” but that the geopolitical struggle over AI will increasingly be decided by who makes intelligence easiest to industrialise, finance, deploy, govern and reproduce at scale. Energy Demand from AI — IEA — 2025 State Council Artificial Intelligence Plus Action — Aug 2025
AI Competition Is Becoming Ecosystem Versus Ecosystem
The strategic contest is shifting from isolated model performance toward the total economics of embedding intelligence across cloud infrastructure, semiconductors, enterprise software, agents, industrial processes, power systems and institutional procurement.
China has developed a broad and increasingly integrable AI application ecosystem, but the verified record does not establish full production-layer technological autonomy. At the same time, the United States and allied technology sector can no longer be characterised as competing only through frontier model quality. The emerging contest is therefore not “Chinese ecosystem versus Western chatbot”, but competing full-stack architectures whose relative advantage depends on capability thresholds, inference economics, deployment friction, semiconductor access, power, enterprise lock-in and institutional trust.
China’s AI market is broadening rapidly
Regulatory filing counts do not measure commercial success or technological quality, but they demonstrate a shift away from a single-champion interpretation of the Chinese generative-AI landscape.
Scale: bar length is indexed to the August 2026 value of 1,112 = 100%. Source: Cyberspace Administration of China, September 2026 .
The competition is no longer model against model
Both Chinese and U.S.-allied technology groups increasingly compete across multiple layers of the AI production and deployment chain, although the structure and degree of upstream independence differ materially.
Chinese ecosystem trajectory
U.S. and allied ecosystem trajectory
Where the competitive advantage actually moves
Capability remains a threshold variable
Cost cannot compensate for a system that fails the accuracy, reasoning, security or reliability threshold required by the task. Frontier capability therefore remains economically decisive in high-complexity and high-consequence applications, even as cost becomes more important in mature workloads.
Integration cost can dominate once capability is sufficient
When several models meet the operational threshold, procurement can shift toward latency, inference cost, data location, integration effort, governance, application compatibility and time-to-deployment.
Physical infrastructure increasingly defines AI scale
Accelerators, memory, data centres, networking and electricity are becoming direct inputs into AI competitiveness. The IEA projects global data-centre electricity use of approximately 945 TWh by 2030.
International Energy Agency — Energy and AIThe strongest lock-in may sit above and below the model
Enterprise data, agent frameworks, identity systems, workflow redesign, cloud architecture, accelerator software and observability can create larger switching costs than model weights themselves, particularly where models remain interchangeable through compatible APIs.
Quality-max versus stack-min-cost
The relevant economic question is not simply which model is cheaper, but which architecture produces the lowest cost per successfully completed, governed and reliable workflow.
Quality-max logic
- Large capability differences remain economically consequential.
- Failure or hallucination costs are high.
- Tasks involve difficult reasoning, science, coding or high-value analysis.
- The value of a successful answer materially exceeds inference cost.
- Model performance remains the binding constraint.
Stack-min-cost logic
- Several models already exceed the required quality threshold.
- Deployment volume is very high.
- Latency and accelerator efficiency matter strongly.
- Data sovereignty or on-premise deployment is important.
- Integration and switching costs exceed marginal model-quality differences.
The strongest constraint on the Chinese thesis
Official source: U.S. Bureau of Industry and Security — Advanced Semiconductor Controls .
Historical analogy stress-test
Oil
Useful for understanding strategic supply, physical infrastructure, bottlenecks and state intervention, but weak as a literal product-market analogy because AI models are heterogeneous and software can be reproduced at near-zero marginal copying cost.
Photovoltaics
Stronger for understanding manufacturing scale, supplier clustering, industrial policy and learning curves that can turn production economics into international market power.
IEA solar supply-chain analysisBatteries
Relevant to vertically integrated supply chains and the migration of competitive advantage from end products into components, materials, manufacturing capacity and deployment infrastructure.
IEA Global EV Outlook 2026Telecommunications
The strongest analogy for APIs, compatibility, developer ecosystems, installed bases and standards, although AI currently permits substantially greater model portability and multi-homing than traditional telecom networks.
Claim audit
| Proposition | Status | Assessment | Primary evidence |
|---|---|---|---|
| China is no longer a single-model AI market. | Supported | Strong evidence of provider and service proliferation, although filing counts do not establish commercial success. | CAC |
| China is building model + cloud + application + industry stacks. | Supported | Visible across Alibaba, Tencent, Baidu and Huawei, with different levels of vertical integration. | Alibaba Cloud |
| The West competes mainly through frontier model quality. | Overstated | Western firms now expose extensive model, cloud, agent, governance and infrastructure layers. | Microsoft Foundry |
| The cheapest model will become the standard. | Conditional | Price dominates only after capability, reliability and regulatory thresholds are satisfied. | Microsoft |
| China possesses a fully autonomous AI production stack. | Not established | Advanced semiconductor manufacturing and HBM remain material strategic constraints. | U.S. BIS |
| AI geopolitics increasingly operates through capital allocation. | Supported | Compute, electricity, data centres and long-duration cloud capacity increasingly require large irreversible capital commitments. | IEA |
Decision indicators through 2030
| Indicator | What it tests | What would strengthen the ecosystem thesis |
|---|---|---|
| Quality-adjusted inference cost | Whether inexpensive models truly reduce the cost of completed work | Persistent cost advantage after controlling for accuracy, retries and verification |
| Chinese accelerator share of domestic AI workloads | Compute autonomy | Large-scale domestic inference and training without controlled foreign accelerators |
| Enterprise production deployments | Actual industrial embedding | Growth in long-duration production use rather than model downloads or trials |
| Model-switching time and engineering cost | Location of lock-in | High switching costs concentrated in platform, data and workflow layers |
| Chinese models running on non-Chinese clouds | Standard diffusion versus national-stack capture | Model adoption accompanied by Chinese tools, frameworks or infrastructure rather than model weights alone |
| AI data-centre power and grid capacity | Physical scalability | Rapid expansion of affordable, reliable power supporting inference growth |
Net assessment
The strongest version of the thesis is not that China will win because its models are cheaper, nor that frontier intelligence is becoming irrelevant. The more defensible proposition is that durable AI power will increasingly belong to the architecture that converts models into productive intelligence with the lowest combined burden of compute, electricity, integration, governance, switching, deployment and organisational redesign. The decisive competition through 2030 is therefore becoming ecosystem versus ecosystem, while the strategically critical question is where enterprises and governments make investments that become difficult to reverse.
Principal verified sources
- Cyberspace Administration of China — Generative AI filing announcement
- State Council of the People’s Republic of China — Artificial Intelligence Plus Action
- Alibaba Cloud — Full-Stack AI Strategy
- Microsoft — Microsoft Foundry
- NVIDIA — AI Enterprise Documentation
- U.S. Bureau of Industry and Security — Advanced Semiconductor Export Controls
- International Energy Agency — Energy and AI
- Anthropic — Amazon compute collaboration
Pillar I — From National AI Strategy to Industrial Infrastructure
Navigational Index
Pillar I — From National AI Strategy to Industrial Infrastructure
Chapter 1 — Institutional, Historical and Physical Baseline
How China moved from the 2017 New Generation Artificial Intelligence Development Plan to the 2025–2026 “AI+” industrialisation architecture; the underlying digital economy, telecommunications network, server base, industrial internet, computing infrastructure, power system and policy institutions that make large-scale AI deployment possible.
Chapter 2 — Models, Compute and the Architecture of Deployment
The actual structure of the Chinese AI market beyond DeepSeek: Qwen, Hy, ERNIE, GLM, Kimi, MiniMax, Pangu and other model families; open-weight versus proprietary strategies; API, cloud and on-premise deployment; agents; inference economics; domestic accelerators; HBM; advanced packaging; semiconductor manufacturing constraints; and the degree to which China possesses an integrated stack rather than merely an extensive application layer.
Chapter 3 — Industrial Embedding and the Economics of Integration
How AI moves from foundation models into manufacturing, software development, logistics, finance, healthcare, telecommunications, public administration and industrial control; where integration costs arise; how enterprise switching costs accumulate; and under what conditions “good enough and cheap to integrate” becomes economically superior to frontier capability.
Pillar II — Standards, Platforms and the Geoeconomics of AI
Chapter 4 — Network Effects, Lock-In and the Battle for the Control Layer
Where AI network effects actually reside: weights, APIs, developer frameworks, agent systems, enterprise data, retrieval architectures, cloud regions, identity, security, compiler stacks, accelerators and industrial software; whether open-weight Chinese models create Chinese ecosystem dependence or instead commoditise the model layer and strengthen platform neutrality.
Chapter 5 — The Western Counter-Architecture
A like-for-like comparison of the Chinese ecosystem with OpenAI, Microsoft, Anthropic, Amazon, Google, Meta and NVIDIA; why the Western system can no longer be described as a collection of isolated frontier laboratories; how hyperscalers, model developers, accelerator vendors and enterprise software increasingly form competing integrated stacks.
Chapter 6 — Stress-Testing the Oil, Telecom, Solar and Battery Analogies
A rigorous examination of which historical mechanisms actually transfer to AI: commodity fungibility, infrastructure dependence, industrial learning curves, economies of scale, overcapacity, standards formation, installed-base effects, vertical integration, supply-chain chokepoints and strategic control.
Pillar III — Capital Allocation, Strategic Dependency and the 2030 Contest
Chapter 7 — Geopolitics as Capital Allocation
AI capital expenditure as geopolitical infrastructure: semiconductor fabs, accelerators, data centres, cloud capacity, electricity generation, grid connections, networking, public procurement, sovereign compute, venture finance and long-duration corporate commitments; how sunk capital can turn temporary technological preferences into persistent strategic alignment.
Chapter 8 — Cognitive Infrastructure, Scenarios and the 2030 Standard
The conditions under which AI becomes infrastructure rather than software; alternative pathways for Chinese, U.S.-allied and multi-homed ecosystems; falsifiers of the central thesis; observable indicators; strategic dependencies; capital-allocation implications; and the final assessment of whether the dominant AI architecture will be determined principally by intelligence, integration cost, infrastructure availability or some combination of all three.
Chapter 1 — Institutional, Historical and Physical Baseline
Principal judgment
China’s present AI position cannot be understood primarily through DeepSeek, Qwen or any other contemporary foundation model, because the more important historical fact is that Beijing has spent close to a decade constructing the institutional, digital and industrial substrate into which foundation models can now be inserted. The critical transition occurred between the 2017 New Generation Artificial Intelligence Development Plan, which framed AI as a strategic technology requiring coordinated advances in algorithms, chips, high-performance computing, data, robotics and industrial applications, and the 2025 “Artificial Intelligence Plus” action, which moved the centre of gravity from research capability toward economy-wide diffusion, infrastructure and application. The State Council’s 2017 plan already connected AI research with core electronic components, high-end general-purpose chips, basic software, integrated-circuit equipment, supercomputing and intelligent manufacturing; eight years later, the State Council explicitly required AI to penetrate science and technology, industry, consumption, public services, governance and international cooperation while supporting intelligent computing infrastructure, software ecosystems, AI chips, data resources and scalable cloud services. The historical sequence therefore indicates that China’s current strategy is not simply to produce competitive models but to transform AI into a general-purpose layer of the industrial economy. New Generation Artificial Intelligence Development Plan — State Council — Jul 2017 Opinions on Deepening Implementation of the “Artificial Intelligence Plus” Action — State Council — Aug 2025 Governo Cinese
This institutional continuity matters because the central competitive proposition of the wider dossier depends on whether China possesses not merely strong AI laboratories but an economy capable of absorbing intelligence at scale. By 2025, China’s official statistics recorded RMB 140.19 trillion of GDP, an industrial sector generating RMB 41.68 trillion of value added, information transmission, software and IT services producing RMB 7.06 trillion of value added, software-industry revenue of RMB 15.48 trillion, 5.97 million servers produced during the year, 4.84 million operational 5G base stations, 1.204 billion 5G mobile subscribers, 2.888 billion mobile Internet-of-Things terminals, and 690.82 million fixed-broadband users. These figures do not establish AI leadership, but they establish something analytically different and essential: an unusually large physical and digital surface across which AI applications can potentially be distributed. Statistical Communiqué of the People’s Republic of China on the 2025 National Economic and Social Development — National Bureau of Statistics — Feb 2026 Ufficio Nazionale di Statistica
The 2017 strategy was already an ecosystem strategy
The 2017 State Council plan is important less for its forecasts than for its architecture. It did not define artificial intelligence as an isolated software sector; instead, it placed fundamental theory, common technologies, hardware, intelligent manufacturing, robotics, high-performance computing, data, open-source technologies and application demonstrations inside a coordinated national programme. The document proposed a “1+N” research structure centred on a major national AI science-and-technology project but explicitly connected that programme to national initiatives covering core electronic components, high-end general-purpose chips, basic software, integrated-circuit equipment, high-performance computing, intelligent manufacturing, robotics, big data and other strategic technologies. This is direct documentary evidence that the conceptual unit of Chinese AI policy had already moved beyond algorithms long before the generative-AI boom beginning in 2022–2023. New Generation Artificial Intelligence Development Plan — State Council — Jul 2017 Governo Cinese
The distinction is fundamental for the present thesis. A state strategy organised around model research alone would be vulnerable to repeated technological discontinuities because leadership could change whenever a competing laboratory released a better architecture; a strategy organised around complementary assets attempts instead to preserve value through chips, computing infrastructure, industrial deployment, data, standards and applications even when individual model rankings change. The 2017 document cannot prove that China has successfully achieved this ambition, but it establishes that complementarity between software, hardware and industrial adoption was an explicit policy objective rather than a retrospective interpretation constructed after DeepSeek’s emergence. New Generation Artificial Intelligence Development Plan — State Council — Jul 2017 Governo Cinese
Institutional trajectory from research policy to economic infrastructure
| Period | Institutional step | Primary economic object | Strategic significance | Official source |
|---|---|---|---|---|
| 2017 | New Generation AI Development Plan | Research, chips, software, computing, applications | Establishes AI as a coordinated national technology programme rather than an isolated software sector | State Council, 2017 |
| 2017 onward | “1+N” AI science-and-technology programme | Basic theory plus related national technology programmes | Connects AI development to integrated circuits, supercomputing, robotics and intelligent manufacturing | State Council, 2017 |
| 2023–2025 | Expansion of generative-AI regulatory filing architecture | Public generative-AI services and applications | Creates a governed route from model creation toward commercial public deployment | CAC, 2026 |
| 2025 | “Artificial Intelligence Plus” action | Economy-wide adoption | Moves policy emphasis toward embedding AI across production, consumption, public services and governance | State Council, 2025 |
| 2027 target | AI integration across six priority areas; intelligent-terminal and agent penetration above 70% | Adoption | Introduces explicit diffusion objectives | State Council, 2025 |
| 2030 target | Intelligent-terminal and agent penetration above 90% | Economy-wide diffusion | Treats AI increasingly as general economic infrastructure | State Council, 2025 |
| 2035 target | Entry into a broadly intelligent economy and society | Systemic transformation | Extends AI policy beyond a technology-industry horizon | State Council, 2025 |
The 2025 State Council programme therefore represents a change of scale rather than an entirely new doctrine. Its explicit targets call for artificial intelligence to achieve extensive integration across six major areas by 2027, with penetration of new-generation intelligent terminals and agents exceeding 70%, rising above 90% by 2030, before the stated objective of entering a new stage of an intelligent economy and society by 2035. These are policy targets rather than verified outcomes and should be treated as such, but they reveal that the Chinese government is attempting to measure success increasingly through diffusion rather than only technological invention. Opinions on Deepening Implementation of the “Artificial Intelligence Plus” Action — State Council — Aug 2025 Ministero del Commercio
The decisive precondition: a large digital-industrial substrate
The scale of China’s existing digital economy gives this diffusion strategy a materially different starting point from one built on AI laboratories alone. According to the National Bureau of Statistics, the value added of the core industries of China’s digital economy reached RMB 14.0891 trillion in 2024, equivalent to 10.5% of GDP. Digital technology application industries contributed RMB 6.1928 trillion, or 44.0% of that digital-economy core; digital product manufacturing contributed RMB 4.8145 trillion, or 34.2%; digital-factor-driven industries generated RMB 2.6519 trillion, or 18.8%; and digital product services contributed RMB 429.8 billion. These categories are broader than artificial intelligence and must not be presented as an AI-market valuation, but they measure the industrial environment into which AI services can be sold and integrated. Value Added of China’s Core Industries of the Digital Economy Takes up 10.5% of GDP in 2024 — National Bureau of Statistics — Dec 2025 Ufficio Nazionale di Statistica
The enterprise population underneath those aggregates is similarly consequential. China’s Fifth National Economic Census recorded 2.916 million corporate enterprises in the core digital-economy industries at the end of 2023, employing 36.159 million people and generating RMB 48.4485 trillion of business revenue during the year. Digital technology application alone accounted for 1.430 million firms and 14.609 million employees, while digital-product manufacturing employed 13.372 million people. Again, none of those numbers demonstrates that firms are using foundation models, but they establish the scale of the potential producer, integrator and customer ecosystem through which generative AI can diffuse. Communiqué on the Fifth National Economic Census, No. 6 — National Bureau of Statistics — Dec 2024 Ufficio Nazionale di Statistica
China’s pre-existing digital industrial base
| Indicator | Verified value | Reference period | Why it matters for AI industrialisation | Source |
|---|---|---|---|---|
| Core digital-economy value added | RMB 14.0891tn | 2024 | Measures scale of digitally native economic substrate | National Bureau of Statistics |
| Share of GDP | 10.5% | 2024 | Indicates macroeconomic weight of core digital industries | National Bureau of Statistics |
| Digital-technology application value added | RMB 6.1928tn | 2024 | Large potential integration and application layer | National Bureau of Statistics |
| Digital-product manufacturing value added | RMB 4.8145tn | 2024 | Hardware/manufacturing complement to software diffusion | National Bureau of Statistics |
| Core digital-economy enterprises | 2.916m | End-2023 | Potential supply and adoption network | Fifth National Economic Census |
| Employment in those enterprises | 36.159m | End-2023 | Depth of digital labour and integration base | Fifth National Economic Census |
| Business revenue | RMB 48.4485tn | 2023 | Commercial scale of digital industrial ecosystem | Fifth National Economic Census |
The implication is not that Chinese AI firms automatically inherit the entire digital economy, but that AI deployment can occur through a pre-existing industrial structure containing millions of technology enterprises, manufacturing suppliers, telecommunications networks, software companies and digital users. Economies of scope are therefore potentially important: the same cloud providers, telecommunications operators, industrial-software vendors and device manufacturers that already maintain enterprise relationships can add model inference or agent functions without building distribution from zero. The empirical question for later chapters will be how frequently this potential is converted into durable production deployments rather than pilots, demonstrations or subsidised installations.
Telecommunications infrastructure converts models into distribution
China’s telecommunications infrastructure is another important component of the baseline because distributed AI systems require reliable connectivity between users, edge devices, enterprises, clouds and data centres. At the end of 2025 China had 12.87 million mobile base stations, including 4.84 million 5G base stations; 5G subscribers numbered 1.204 billion; fixed broadband users reached 690.82 million, of whom 238.39 million subscribed to connections of 1 Gbps or above; and mobile Internet-of-Things terminals reached 2.888 billion. Internet users totalled 1.125 billion, corresponding to 80.1% national internet penetration. Statistical Communiqué on 2025 National Economic and Social Development — National Bureau of Statistics — Feb 2026 Ufficio Nazionale di Statistica
The significance for AI is not that 5G itself creates intelligence, but that a dense communications layer reduces the marginal infrastructure required to connect AI-enabled devices, factories, vehicles, logistics systems, retail networks and public services. The National Development and Reform Commission stated in its interpretation of the “AI+” initiative that China had developed more than 17,000 “5G + Industrial Internet” projects covering all 41 major industrial categories, alongside what it described as the world’s largest 5G network. Because this is a government characterisation of national infrastructure rather than an independent audit of project effectiveness, the figures establish breadth of rollout but not productivity impact. “Artificial Intelligence Plus” Opens a New Chapter of Intelligent Development with Chinese Characteristics — NDRC — Aug 2025 ndrc.gov.cn
The National Bureau of Statistics subsequently stated that industrial-internet integration applications had achieved coverage across all 41 major industrial categories by 2025, while the output of industrial robots increased 28% and high-technology manufacturing value added increased 9.4%. These indicators should not be causally attributed to AI, but together they document the presence of a manufacturing system in which digital connectivity, automation and software were already expanding before generative AI became deeply embedded. National Bureau of Statistics Director on China’s 2025 Economic Performance — Jan 2026 Ufficio Nazionale di Statistica
Physical distribution layer available to AI systems
| Infrastructure indicator | 2025 level | Change / contextual measure | AI relevance | Source |
|---|---|---|---|---|
| 5G base stations | 4.84m | Part of 12.87m total mobile base stations | Edge/cloud connectivity | NBS |
| 5G subscribers | 1.204bn | — | Large connected user/device market | NBS |
| Fixed broadband users | 690.82m | +20.99m during 2025 | Enterprise/household access layer | NBS |
| ≥1 Gbps broadband users | 238.39m | +31.57m | High-bandwidth services and cloud access | NBS |
| Mobile IoT terminals | 2.888bn | +232m | Potential machine-to-AI interface layer | NBS |
| Internet users | 1.125bn | 80.1% penetration | Consumer and service distribution | NBS |
| “5G + Industrial Internet” projects | >17,000 | Covering 41 industrial categories | Potential industrial AI deployment channels | NDRC |
This infrastructure changes the economic meaning of a model release. Where large populations of connected factories, devices and enterprises already exist, a new model can potentially be distributed through existing cloud, telecommunications and industrial-software relationships; the competitive issue becomes less the creation of access and more the ability to convert access into sufficiently reliable automation. This is one of the principal reasons the Chinese ecosystem thesis needs to be evaluated as an industrial diffusion problem rather than only a model-development problem.
Server production and software capacity provide an additional layer
China’s 2025 production statistics show 5.97 million servers produced during the year, an increase of 12.6%, while software and information-technology service revenue increased 13.2% to RMB 15.4831 trillion. Information transmission, software and information-technology services generated RMB 7.0599 trillion in value added, up 11.1% from 2024. These figures are broader than AI-specific compute and do not identify what share of servers contained AI accelerators, but they demonstrate that the country possesses a large domestic information-technology production and service base capable of supplying installation, maintenance, systems integration and software layers around AI infrastructure. Statistical Communiqué on 2025 National Economic and Social Development — National Bureau of Statistics — Feb 2026 Ufficio Nazionale di Statistica
The distinction between servers and AI compute must nevertheless remain explicit. Server production does not establish availability of frontier-class GPUs, high-bandwidth memory or leading-edge accelerators; nor does a large software sector establish the ability to replicate CUDA-class tooling or the performance of advanced semiconductor systems. The physical baseline therefore contains both a major advantage and a major analytical trap: China has exceptional scale in digital equipment, connectivity and industrial software, but those aggregate indicators cannot be used to infer self-sufficiency at the most advanced compute layer.
The electrical system has become part of the AI stack
Artificial intelligence also changes the meaning of energy infrastructure because high-density training and inference clusters transform electricity from a generic operating cost into a potential capacity constraint. The International Energy Agency projects worldwide data-centre electricity consumption to exceed 945 TWh in 2030, more than twice the current level, with the United States and China accounting for nearly 80% of global data-centre electricity-demand growth through 2030. The IEA estimates that China’s consumption increases by approximately 175 TWh between 2024 and 2030, equivalent to growth of roughly 170% from the 2024 level. Energy Demand from AI — International Energy Agency — 2025 IEA
China’s power mix creates both scale advantages and structural complications. The IEA estimates that close to 70% of electricity physically consumed by Chinese data centres currently originates from coal, with renewables supplying nearly 20%, nuclear close to 10% and natural gas the remainder; between 2024 and 2030, coal-fired generation serving data centres increases by almost 90 TWh while renewables add a similar amount. The same analysis points to policies encouraging data-centre construction in renewables-rich western provinces and anticipates a larger contribution from nuclear and renewables after 2030. These are scenario-based IEA projections rather than realised outcomes, but they establish that power availability, generation geography and grid transmission are integral elements of AI scale. Energy Supply for AI — International Energy Agency — 2025 IEA
AI’s emerging physical-resource baseline
| Indicator | China / global measure | Period | Analytical significance | Source |
|---|---|---|---|---|
| Global data-centre electricity consumption | ~945 TWh | 2030 projection | Defines scale of emerging physical constraint | IEA |
| China’s additional data-centre demand | ~175 TWh | 2024–2030 | Indicates very rapid compute-related electricity growth | IEA |
| Increase relative to China’s 2024 level | ~170% | To 2030 | Shows infrastructure challenge is nonlinear | IEA |
| U.S. + China contribution to global demand growth | ~80% | To 2030 | Concentrates AI infrastructure competition geographically | IEA |
| Coal share of electricity serving Chinese data centres | ~70% | Current IEA estimate | Creates carbon, location and grid implications | IEA |
| Additional renewable generation for Chinese data centres | ~90 TWh | 2024–2030 projection | Illustrates simultaneous expansion of computing and power capacity | IEA |
The physical consequence is that an AI ecosystem capable of inexpensive model inference cannot be evaluated independently of electricity generation, grid connections, cooling, network infrastructure and geographic distribution of computing facilities. If models become cheaper but electricity or accelerator access becomes constrained, the theoretical cost advantage may fail to translate into deployable capacity; conversely, an architecture that combines efficient models with abundant compute and inexpensive power can convert apparently modest model-level advantages into much larger system-level advantages.
“AI+” formally shifts policy from invention toward penetration
The State Council’s August 2025 “AI+” action is particularly important because its language places application diffusion at the centre of policy. The programme identifies six areas—science and technology, industrial development, consumption, people’s livelihoods, governance capability and global cooperation—and calls for large-scale commercial application, new infrastructure, new technology systems and new industrial ecosystems. The National Development and Reform Commission interpreted this programme as a transition from the earlier “Internet+” paradigm toward an AI-driven transformation of production factors, industrial structures and value chains. Opinions on Deepening Implementation of the “Artificial Intelligence Plus” Action — State Council — Aug 2025 NDRC Interpretation of the “Artificial Intelligence Plus” Action — Aug 2025 Governo Cinese
This is materially different from a policy that simply subsidises frontier-model laboratories. The official strategy calls for strengthening AI chips, software ecosystems and intelligent-computing clusters; improving the nationally integrated computing-power network; expanding standardised and scalable cloud services; developing data resources; and accelerating application across sectors. In institutional terms, the intended system therefore extends from inputs—chips, computing capacity and data—through models and infrastructure to applications and social adoption. Opinions on Deepening Implementation of the “Artificial Intelligence Plus” Action — State Council — Aug 2025 Governo Cinese
The Chinese policy stack as of September 2026
| Layer | Institutional objective evidenced in official policy | Existing physical/economic substrate | Critical unresolved question |
|---|---|---|---|
| Fundamental research | AI theory and common technologies | National AI science-and-technology programmes | Can research remain near the frontier under compute constraints? |
| Chips | Strengthen AI-chip capability | Large electronics-manufacturing base | Can domestic supply close gaps in leading-edge logic, HBM and toolchains? |
| Computing | Intelligent-computing clusters and nationally integrated computing network | Large server production and data-centre buildout | What fraction is frontier-capable AI compute? |
| Connectivity | High-speed digital infrastructure | 4.84m 5G base stations; 690.82m fixed broadband users | How effectively can connectivity translate into production AI? |
| Data | Development and use of data resources | Massive connected economy and digital services | Availability, quality, governance and cross-sector interoperability |
| Cloud | Scalable cloud services | Alibaba, Tencent, Baidu, Huawei and other platforms | Market concentration and switching costs |
| Models | Broad model-development ecosystem | Hundreds of filed generative-AI services | Quality distribution and sustainable economics |
| Agents | Explicit 2027/2030 diffusion targets | Fast-growing enterprise agent layer | Production reliability and workflow economics |
| Industrial applications | “AI+” integration across the economy | 41 industrial categories already connected through industrial internet | Depth rather than nominal breadth of deployment |
| Governance | Safety, regulation and public-sector integration | CAC filing and content-governance regime | Compatibility with global deployment markets |
Regulatory architecture is part of deployment infrastructure
The institutional baseline also includes a regulatory gate through which public-facing generative-AI services must pass. By 31 August 2026, the Cyberspace Administration of China reported 1,112 generative-AI services completing the relevant filing procedures, while 731 applications or functions that directly call filed models through APIs or other methods had completed registration. These numbers should not be interpreted as 1,112 commercially successful foundation models; they include services and represent regulatory status rather than market share. Their importance lies elsewhere: they demonstrate the rapid construction of a formalised national deployment layer linking model providers, downstream applications and regulatory oversight. Announcement on Filed Generative AI Services, July–August 2026 — Cyberspace Administration of China — Sep 2026
The growth is substantial even before questions of quality and commercial survival are addressed. Official CAC notices recorded 302 filed generative-AI services by the end of 2024, 748 by the end of 2025, and 1,112 by 31 August 2026. The increase from 302 to 1,112 represents approximately 3.68 times the end-2024 total, calculated from the CAC’s published figures; it indicates rapid proliferation of registered services but cannot by itself determine whether the market is becoming competitive, concentrated, fragmented or economically sustainable. 2025 Generative AI Filing Announcement — Cyberspace Administration of China — Jan 2026 July–August 2026 Filing Announcement — Cyberspace Administration of China — Sep 2026
Expansion of China’s regulated generative-AI service layer
| Date | Filed generative-AI services | Change from previous cited point | Interpretation |
|---|---|---|---|
| End-2024 | 302 | — | Early regulatory-commercial population |
| End-2025 | 748 | +446 / +147.7% | Rapid provider/service expansion |
| 31 Aug 2026 | 1,112 | +364 / +48.7% versus end-2025 | Continued expansion at a slower proportional rate |
| Aug 2026 | 731 downstream registered applications/functions | Separate category | Evidence of an application layer calling approved models through APIs or equivalent mechanisms |
Calculated percentage changes use the official service totals published by the Cyberspace Administration of China and its 2025 filing announcement.
The distinction between service filing and downstream API registration is especially important for the wider thesis because it provides institutional evidence of a layered ecosystem: a growing number of applications can consume models as upstream infrastructure rather than each application developing a proprietary foundation model. That architecture is conceptually consistent with the emergence of AI as a utility-like input, although later chapters must determine whether commercial behaviour actually follows that model.
Industrial capacity does not remove semiconductor dependence
The greatest danger in interpreting the preceding baseline is to move from “large ecosystem” to “complete autonomous ecosystem.” The public record does not support that conclusion. The U.S. Bureau of Industry and Security’s December 2024 semiconductor-control package specifically added controls covering 24 categories of semiconductor manufacturing equipment, three categories of software tools, high-bandwidth memory, additional Entity List designations and other measures intended to constrain China’s ability to manufacture advanced-node semiconductors used in advanced computing and AI. Whatever judgment is made about the policy itself, the existence and design of these controls identifies upstream technologies that U.S. authorities consider material external dependencies or chokepoints. Commerce Strengthens Export Controls to Restrict China’s Capability to Produce Advanced Semiconductors for Military Applications — U.S. Bureau of Industry and Security — Dec 2024 Bis
The analytical baseline must therefore distinguish at least four different propositions that are frequently collapsed into one. China has a very large digital economy; it has a large and rapidly proliferating AI application ecosystem; it possesses substantial domestic computing, cloud and semiconductor capabilities; but the evidence currently examined does not establish complete independence in leading-edge semiconductor production, HBM or semiconductor-manufacturing equipment. Those propositions can all be true simultaneously. The distinction will be decisive in Chapter 2 because the ability to industrialise AI at enormous scale depends both on downstream demand and on whether compute supply grows rapidly enough to meet it.
Corporate architecture increasingly mirrors national policy architecture
The corporate layer provides evidence that at least some major Chinese firms are organising themselves in a manner consistent with the policy emphasis on vertical and horizontal integration. On 22 September 2026, Alibaba announced what it explicitly described as a full-stack AI strategy “from chips, cloud infrastructure, models to agents,” including proprietary AI chips, its Qwen model family, an agent-oriented cloud architecture and agent products. Because this is a first-party company announcement, its claims about technological leadership require independent validation; nevertheless, the structure of the announced roadmap is itself evidence that Alibaba is competing across several complementary layers rather than presenting Qwen as an isolated model product. Alibaba Unveils Roadmap on Full-Stack AI Strategy from Chips, Cloud Infrastructure, Models to Agents — Alibaba Cloud — Sep 2026 AlibabaCloud
Huawei demonstrates another version of this architecture, focused more heavily on industrial deployment. Huawei Cloud reported in September 2025 that its Pangu models had been applied in more than 500 scenarios across more than 30 industries, embedded in a broader architecture combining cloud infrastructure, data and AI engines. These figures are corporate claims and should not be read as independently audited evidence of productivity gains, but they illustrate the strategic relevance of China’s pre-existing industrial base: the potential advantage is not simply the ability to offer a model API but the ability to attach models to established relationships in manufacturing, energy, logistics and other sectors. Huawei Cloud: Fostering the Fertile Ground for Compute, Empowering AI Pioneers for Industries — Huawei — Sep 2025 Huawei
Baseline distinction: what is established and what is not
| Proposition | Evidentiary status at end of Chapter 1 | Basis |
|---|---|---|
| China has treated AI as a strategic national technology since at least 2017 | Established | State Council policy record |
| Chinese AI strategy explicitly connects models to chips, computing, data and industrial applications | Established | 2017 plan and 2025 “AI+” action |
| China possesses a very large digital-industrial adoption surface | Established | NBS digital-economy, telecom and enterprise statistics |
| AI diffusion is an explicit national objective rather than a by-product of private model competition | Established | 2025 State Council targets |
| The Chinese generative-AI service population has expanded rapidly | Established | CAC filing data |
| China’s digital infrastructure provides potential economies of distribution for AI | Evidence-supported analytical judgment | Telecommunications, industrial-internet and digital-economy scale |
| China already possesses complete upstream AI technological autonomy | Not established | External controls remain directed at advanced semiconductor equipment, HBM and advanced compute |
| High filing counts demonstrate successful commercial adoption | Not established | Filing is a regulatory status, not usage or profitability |
| Large server output equals frontier AI compute capacity | Not established | Server statistics do not identify accelerator class |
| Industrial AI deployment has already produced economy-wide productivity gains | Not established from the official record examined here | Requires firm-level and sector-level productivity evidence |
The baseline changes the central question
The relevant starting question for the remainder of the dossier is therefore no longer whether China can produce individual models competitive with Western systems; that question is too narrow to explain the economic architecture now visible in the official record. The stronger question is whether China’s combination of industrial scale, digital infrastructure, manufacturing depth, cloud capacity, regulatory coordination and enormous potential deployment market can convert sufficiently capable models into lower-cost, more deeply embedded and more difficult-to-displace productive systems, while overcoming upstream constraints in advanced compute.
That formulation also prevents the opposite analytical mistake: assuming that infrastructure scale guarantees technological leadership. Physical distribution networks, a huge industrial base and millions of digital enterprises create opportunities for diffusion, but they do not eliminate capability gaps, poor model economics, integration failures, semiconductor constraints, security problems or enterprise reluctance. The baseline establishes capacity for ecosystem formation, not the inevitable success of the ecosystem.
Key judgments
| Judgment | Evidence base | Confidence |
|---|---|---|
| China’s AI strategy predates the generative-AI boom and has long treated AI as a multi-layer industrial system | 2017 and 2025 State Council policy documents | High |
| China’s telecommunications, digital-economy and industrial base provides an unusually large potential surface for AI diffusion | NBS and NDRC statistics | High |
| Policy emphasis has shifted materially from technological development toward economy-wide AI penetration | “AI+” objectives and 2027/2030 targets | High |
| The regulatory environment is developing a layered model-provider/application architecture | CAC filing and API-application registration data | High |
| China’s downstream and application-layer ecosystem is substantially more complete than a chatbot-centred interpretation suggests | Policy record, infrastructure indicators and corporate architecture | High |
| Production-layer autonomy remains materially less certain than application-layer completeness | BIS controls and absence of sufficiently granular official Chinese supply data | Moderate–High |
| The decisive test is no longer model availability but the economics of embedding models into productive workflows | Synthesis of infrastructure and institutional evidence | Moderate–High |
What would change the assessment
The assessment would strengthen materially if official or independently auditable evidence showed that a rising share of China’s large-scale AI training and production inference workloads is running on domestically controlled accelerators, memory, networking and software stacks at costs and reliability levels comparable with foreign alternatives; if industrial deployments demonstrated sustained productivity improvement rather than pilot-stage adoption; and if Chinese model families increasingly became embedded in non-Chinese enterprise and sovereign infrastructures while retaining Chinese toolchains or standards rather than being absorbed as interchangeable weights into foreign platforms.
It would weaken if the increase in filed AI services proved to conceal high attrition or consolidation without substantial production usage; if advanced-compute scarcity materially limited enterprise inference capacity; if model portability prevented durable ecosystem lock-in; if industrial customers treated models as easily replaceable commodities; or if U.S.-allied providers reduced integration costs sufficiently that China’s scale of domestic deployment no longer translated into an international ecosystem advantage.
Open official record
The most important unresolved official records for the next stage are AI-specific accelerator deployment by cloud provider; domestic versus foreign accelerator shares in production inference; HBM availability and origin; advanced packaging capacity allocated specifically to AI accelerators; utilisation rates of national and regional intelligent-computing centres; verified enterprise inference volumes; sector-specific productivity effects; and the geographic distribution and power intensity of AI-focused data centres. The aggregate evidence reviewed in this chapter cannot legitimately substitute for those missing measurements.
Chapter 2 — Models, Compute and the Architecture of Deployment
Principal judgment
China’s artificial-intelligence market in September 2026 is no longer adequately described through DeepSeek, or even through a small group of nationally prominent foundation-model developers, because the competitive structure has become differentiated simultaneously by model architecture, openness, inference economics, cloud distribution, agent tooling, enterprise deployment mode and access to computing infrastructure. Alibaba’s Qwen family now occupies several capability and price tiers inside Model Studio; Tencent’s Hy architecture combines open weights with TokenHub, WorkBuddy and CodeBuddy distribution; Baidu links ERNIE to PaddlePaddle, Qianfan and a rapidly expanding AI Cloud Infrastructure business; Zhipu’s GLM family competes through open deployment and long-horizon agentic workloads; Moonshot’s Kimi pushes extremely large sparse models and long context; MiniMax combines inexpensive model APIs with multimodal and agent products; Huawei links Pangu to Ascend compute, ModelArts, industry applications and increasingly ambitious domestic AI infrastructure; and DeepSeek remains an important force in cost-efficient, portable and API-compatible inference. The resulting market is therefore not a collection of interchangeable chatbots but an emerging portfolio of different routes from model intelligence to deployed productive systems.
The strongest new evidence nevertheless requires a precise qualification. China is moving from application-layer breadth toward genuine domestic compute-stack integration: Huawei reports Ascend development across chips, CANN software, SuperPoDs, interconnect and cloud infrastructure, while Zhipu reported in September 2026 that production inference for GLM-5.3-Flash was operating on a cluster of more than 100,000 Chinese-made AI accelerators. These are significant first-party claims because they document production deployment rather than an experimental benchmark, but they do not independently establish national self-sufficiency in leading-edge logic fabrication, high-bandwidth memory, lithography, advanced manufacturing equipment or semiconductor-design software. The correct assessment is consequently that China’s deployable AI stack is becoming substantially more integrated, while its semiconductor production stack remains uneven and exposed to external chokepoints.
The Chinese model market has differentiated into competing economic strategies
The most important structural change is that Chinese developers no longer compete according to a single model-development logic. Some pursue scale; others emphasise sparse activation, inference efficiency, coding, extremely long context, multimodality, open weights or deep integration into a proprietary commercial platform. This differentiation matters because enterprise buyers do not consume “intelligence” abstractly: they purchase a particular combination of task performance, context length, latency, data location, tool compatibility, infrastructure requirement, price and operational support.
Chinese model architecture as of September 2026
| Ecosystem | Current representative family | Deployment orientation | Openness / portability | Distinctive architecture or commercial logic | Verified distribution layer |
|---|---|---|---|---|---|
| Alibaba | Qwen 3.6 / 3.7 / 3.8 families | Cloud API, dedicated deployment, agents, global cloud regions | Mixture of proprietary API models and broader Qwen ecosystem | Multiple capability/price tiers; multimodal and agentic workloads | Alibaba Cloud Model Studio and AgentCore |
| Tencent | Hy4 Preview | API, productivity agents, coding, consumer products | Hy4 open-sourced; multi-model platform | 770B total / 49B active parameters; >1M context | TokenHub, WorkBuddy, CodeBuddy, Yuanbao, ima |
| Baidu | ERNIE 5.1 | Cloud, enterprise, agents, consumer applications | Primarily proprietary commercial stack plus third-party models through Qianfan | Smaller architecture than predecessor with claimed training-efficiency gains | Baidu AI Cloud / Qianfan / PaddlePaddle |
| Zhipu / Z.ai | GLM-5.2 / 5.3 | Open deployment, coding, long-horizon agents | Weights available for GLM-5.2; local deployment supported | 1M context; inference-system optimisation; domestic-accelerator production experiments | Z.ai, coding plans, local vLLM/SGLang and cloud distribution |
| Moonshot AI | Kimi K3 | API, coding, enterprise, knowledge work | K3 announced as open-weight | 2.8T parameters, sparse MoE, 1M context | Kimi API, Kimi Code, Kimi Work |
| MiniMax | M2.7 / M3 | API, coding agents, multimodal products | Open-compatible integrations and developer plans | Aggressive inference pricing; agentic engineering and office workflows | MiniMax API, MiniMax Code, external coding tools |
| Huawei | Pangu + third-party models | Industrial AI, sovereign/private infrastructure, cloud | More vertically integrated hardware/software path | Domain models tied to Ascend, CANN, ModelArts and Industry AI Foundry | Huawei Cloud / ModelArts / Ascend |
| DeepSeek | V4.1-Flash / V4-Pro | API, coding, agents, third-party clouds | Strong portability and OpenAI/Anthropic-compatible APIs | Very low token prices, caching and flexible reasoning | DeepSeek API plus Chinese and international cloud platforms |
Sources: Alibaba’s Model Studio currently lists Qwen together with DeepSeek, GLM, Kimi and MiniMax models, demonstrating explicitly that its commercial layer is multi-model rather than Qwen-exclusive; Tencent describes Hy4 Preview as a 770-billion-parameter model with 49 billion active parameters and more than one million tokens of context; Kimi describes K3 as a 2.8-trillion-parameter sparse model; and Z.ai documents one-million-token context and publicly available weights for GLM-5.2.
The resulting diversity makes the term “Chinese model” analytically weak unless the deployment model is specified. Qwen consumed through Alibaba Model Studio is economically different from Qwen weights deployed locally; Hy4 inside WorkBuddy is different from an open Hy4 implementation; GLM running through Z.ai is different from GLM weights served through vLLM; Kimi K3 consumed through Moonshot’s API is different from the same architecture embedded by an inference partner. The locus of commercial control can therefore reside in the model developer, cloud platform, inference provider, agent layer or enterprise integrator depending on the deployment path.
Alibaba is building the clearest commercial full-stack architecture
Alibaba currently presents the most explicit Chinese corporate version of a cloud-to-model-to-agent architecture. Its Model Studio catalogue, last updated on 24 September 2026, includes multiple Qwen generations while also distributing DeepSeek V4, GLM 5.x, Kimi K3, MiniMax models and multimodal systems, which means the platform can retain enterprise customers even if the preferred model changes. This is strategically important because a multi-model cloud can convert commoditisation at the foundation-model layer into demand for orchestration, data, identity, inference and agent services at higher layers of the stack.
Qwen itself is increasingly segmented. Alibaba documents Qwen3.6 and later families across dense and sparse architectures, long contexts, multimodal inputs, tool calling and reasoning modes; for example, Qwen3.6-27B is a dense vision-language model with a 262,144-token context window, while other current Qwen 3.6 and 3.7 variants reach one million tokens and support text, images, video, function calling and built-in tools. The significance is not the benchmark positioning claimed by Alibaba but the availability of different cost-performance envelopes under one deployment platform.
Alibaba’s 2026 infrastructure announcements further extend the architecture downward. In May it announced upgrades covering Qwen3.7-Max, cloud model services, new T-Head chips and the Panjiu AL128 Supernode server; in September it publicly described its strategy as extending “from chips, cloud infrastructure, models to agents,” while announcing additional proprietary AI chips and a purpose-built agentic cloud. These are first-party announcements and therefore establish corporate strategy and announced product architecture rather than independently verified relative performance, but they provide direct evidence that Alibaba’s competitive unit is intentionally broader than the Qwen model family.
Alibaba is also offering dedicated model deployments rather than only shared API access. Its September 2026 documentation lists hourly and monthly prices for dedicated Qwen, GLM and DeepSeek deployments, including Qwen3.6-Plus configurations starting at US$88 per hour / US$41,832 per month in Singapore and Qwen3.6-35B-A3B configurations in Beijing beginning at approximately US$59.41 per hour for one documented unit specification. Dedicated deployment is economically important because large enterprise workloads frequently require predictable capacity, stronger isolation and controllable latency rather than best-effort public APIs.
Alibaba’s deployment ladder
| Layer | Current mechanism | What the customer buys | Lock-in mechanism |
|---|---|---|---|
| Model | Qwen plus third-party catalogue | Inference capability | Model-specific prompts and behaviour |
| API | Model Studio | Usage-based access | API configuration, data pipelines, monitoring |
| Dedicated inference | Dedicated deployments | Reserved/model-specific capacity | Infrastructure configuration and spend commitment |
| Agent layer | AgentCore | Agent runtime, tools, teams, governance | Identity, permissions, skills, MCP connections |
| Cloud | Alibaba Cloud | Compute, storage, databases, networking | Broader cloud architecture |
| Proprietary compute | T-Head / announced AI chips | Hardware-supported AI capacity | Hardware/software optimisation |
| Enterprise workflow | Applications and partner integrations | Completed business processes | Process redesign and organisational switching cost |
Alibaba’s new AgentCore is especially relevant to ecosystem economics because it defines agents as managed execution units connected to models, skills and external tools while introducing workspaces, teams, identity, observability and governance as platform services. Once an enterprise stores tool permissions, agent definitions, authentication, monitoring and business logic in such a platform, switching costs migrate away from the model alone and toward the orchestration layer.
Tencent illustrates a different model: open foundation model plus controlled productivity layer
Tencent’s Hy4 Preview demonstrates a highly sparse architecture: the company reports 770 billion total parameters but only 49 billion active parameters, with a context window exceeding one million tokens. Hy4 is open-sourced, but Tencent simultaneously distributes it through its own WorkBuddy, CodeBuddy, Yuanbao and ima applications and through TokenHub APIs. The combination illustrates how open weights need not imply surrendering the downstream commercial relationship; the model can circulate openly while the developer monetises convenience, orchestration, subscriptions and enterprise tooling.
TokenHub is particularly important because Tencent explicitly markets it as a unified large-model service entrance that combines its own Hy models with third-party systems spanning GLM, Kimi, MiniMax and DeepSeek. It supports usage-based calls, guaranteed-resource products and dedicated deployment, while WorkBuddy can switch among multiple models and CodeBuddy provides IDE, plug-in and CLI forms. The commercial architecture therefore separates the model from the platform: Tencent can capture enterprise workflow even when a customer selects a competitor’s model.
Current Tencent pricing also demonstrates how extreme model-level price differentiation has become. TokenHub lists Hy4 Preview in Guangzhou at RMB 6 per million input tokens, RMB 18 per million output tokens and RMB 0.3 per million cached tokens, while GLM-5.3-Flash is listed at RMB 0.8 input / RMB 2.8 output, Kimi K3 at RMB 20 input / RMB 100 output, and DeepSeek V4.1-Flash at time-dependent prices beginning around RMB 1 input and RMB 4 output during off-peak periods. These are not comparable quality-adjusted costs, because model capabilities, context, cache behaviour and task success differ; nevertheless, the spread proves that the Chinese market already contains several distinct inference price tiers.
Baidu’s architecture is increasingly monetised at the infrastructure layer
Baidu’s competitive structure is more vertically coordinated around AI infrastructure → PaddlePaddle → ERNIE → Qianfan → applications. The company describes its infrastructure as combining domestic and international high-performance computing resources, PaddlePaddle as its in-house framework, ERNIE as the foundation-model layer and Qianfan as the enterprise model-and-agent platform. Qianfan itself has evolved from a proprietary-model gateway toward a multi-model environment containing ERNIE, DeepSeek and other third-party systems.
ERNIE 5.1, released in May 2026, is particularly relevant to the cost-efficiency thesis because Baidu reports that it reduced total parameters to approximately one-third and active parameters to approximately one-half of ERNIE 5.0 while using roughly 6% of what Baidu describes as the pre-training cost of comparable models. These are company claims and cannot be treated as independent comparative evidence, but they demonstrate that architectural efficiency has itself become a competitive design objective rather than a secondary concern.
More important than benchmark claims is Baidu’s financial evidence. In the first quarter of 2026, Baidu reported RMB 8.8 billion of AI Cloud Infrastructure revenue, up 79% year over year, with GPU Cloud revenue increasing 184% year over year, while AI Applications generated RMB 2.5 billion. These figures are unaudited quarterly company disclosures, but they show that infrastructure monetisation is scaling more rapidly than the company’s reported application business, reinforcing the proposition that substantial economic value may accumulate below the visible chatbot layer.
Baidu’s Q1 2026 AI business indicators
| Indicator | Q1 2026 | YoY change | Analytical interpretation |
|---|---|---|---|
| Core AI-powered business revenue | >RMB 13.6bn | +49% | AI becoming economically central inside Baidu |
| AI Cloud Infrastructure revenue | RMB 8.8bn | +79% | Strong infrastructure demand |
| GPU Cloud revenue | Not separately disclosed | +184% | Rapid expansion in accelerator-backed cloud usage |
| AI Applications revenue | RMB 2.5bn | Approximately flat | Application monetisation growing less rapidly than infrastructure |
Source: Baidu’s Q1 2026 financial release.
GLM provides unusually important evidence on domestic compute
Zhipu’s GLM development matters because it offers one of the most concrete public descriptions of the relationship between Chinese models and domestic accelerators. GLM-5.2 introduced a one-million-token context window, optimisations such as IndexShare and revised speculative decoding, and public support for local deployment through Transformers, vLLM, SGLang, xLLM and ktransformers. The openness of these deployment routes means the model can circulate independently of Z.ai’s hosted API and therefore strengthens the commoditisation side of the ecosystem thesis.
More consequentially, Z.ai reported on 17 September 2026 that GLM-5.3-Flash production inference had been built on a cluster containing more than 100,000 Chinese-made AI accelerators, and that all production inference for that model was running on the system. The company explicitly acknowledged limited chip memory capacity, bandwidth constraints and immature ecosystem support, while reporting that a combination of tensor parallelism, quantisation, memory optimisation and disaggregated serving increased end-to-end serving performance by about three times. Because the evidence comes from the model developer itself, the claims require external validation, but the engineering detail makes this materially stronger evidence of domestic accelerator deployment than generic statements of strategic intent.
This development materially changes, but does not erase, the semiconductor constraint. It shows that Chinese accelerators can support high-volume production inference for at least one leading domestic model family, and it demonstrates an engineering pathway in which software optimisation compensates partly for memory and ecosystem disadvantages. It does not establish comparable economics for frontier training across the entire Chinese model sector, nor does it reveal the fabrication nodes, HBM origins, yields, packaging capacity or supply volumes underlying the accelerator cluster. Those unresolved variables remain essential.
Kimi makes scale and inference architecture part of the product
Moonshot’s Kimi K3 represents a different route: extreme model scale combined with sparse activation and long context. Kimi describes K3 as a 2.8-trillion-parameter model using 16 of 896 experts under its sparse architecture, with a one-million-token context window; unusually, the company explicitly acknowledges that K3’s overall performance still trails the most powerful proprietary models it compares against. This admission is analytically useful because it demonstrates that Chinese developers themselves do not assume scale or openness automatically eliminates frontier-quality differences.
Kimi’s inference economics are also structured around caching. The current API lists K3 at US$3 per million cache-miss input tokens, US$0.30 per million cache-hit input tokens and US$15 per million output tokens, while Moonshot states that its Mooncake disaggregated inference architecture achieves a cache-hit rate above 90% for coding workloads on its official API. The latter figure is a first-party operational claim, but it demonstrates why quoted input-token prices are increasingly insufficient: effective cost depends on workload structure, repeated context, cache hit rates and the inference architecture underneath the model.
MiniMax is explicitly competing on intelligence-per-dollar
MiniMax has placed cost efficiency near the centre of its strategy. Its current M2.7 API price is US$0.30 per million input tokens and US$1.20 per million output tokens, with a cache-read rate of US$0.06 per million; the high-speed version doubles the standard input and output price. The company also sells subscription token plans supporting multiple concurrent agents and integration with external coding tools, indicating a commercial transition from individual model calls toward persistent agent workload consumption.
The business evidence is notable. MiniMax reported first-half 2026 revenue of US$116.6 million, up 283.1% year over year, while Open Platform and other AI enterprise-service revenue increased 703.1% to US$73.9 million, representing 63.4% of total revenue. The company simultaneously reported that token consumption in July 2026 had reached 20 times its January level. These are company-reported financial numbers, but they provide evidence that inexpensive model access and enterprise APIs are becoming an actual revenue architecture rather than merely a promotional pricing strategy.
DeepSeek’s importance increasingly lies in portability and price architecture
DeepSeek remains structurally important because it has helped normalise the expectation that advanced reasoning and agentic models can be accessed at prices far below the earlier frontier-API norm and through familiar interfaces. DeepSeek’s September 2026 API supports both OpenAI-format and Anthropic-format access, one-million-token contexts, tool calls, structured output and tiered caching. Peak pricing for DeepSeek-V4.1-Flash is listed at US$0.30 per million cache-miss input tokens and US$1.20 per million output tokens, falling by half outside defined peak periods.
The strategic consequence is paradoxical. Compatibility increases DeepSeek’s distribution because developers can replace another provider with less engineering effort, but the same compatibility reduces DeepSeek-specific lock-in because another compatible model can subsequently replace DeepSeek. The value created by DeepSeek diffusion can therefore be captured by foreign or domestic clouds, routing platforms, accelerator suppliers and enterprise applications without requiring the customer to adopt a vertically integrated “DeepSeek ecosystem.”
Open weights are both a diffusion weapon and an anti-lock-in mechanism
This paradox applies across the Chinese open-model landscape. Hy4 is open-sourced; GLM-5.2 weights are available and supported by several local-serving frameworks; Kimi K3 was announced with public weights; and many Qwen releases circulate broadly. Open distribution reduces licensing friction, facilitates sovereign or on-premise deployment and allows system integrators to adapt models without sending sensitive information to the originating vendor. These characteristics are valuable in countries or industries where data residency and strategic autonomy matter.
At the same time, open weights weaken the proposition that adoption of a Chinese model automatically produces adoption of Chinese infrastructure. A GLM or Qwen model served through non-Chinese accelerators, an independent European cloud or an internal enterprise cluster can diffuse Chinese model architecture while transferring little economic rent or strategic control to the originating ecosystem. The distinction between model-standard diffusion and infrastructure-stack capture is therefore one of the principal variables that later chapters must track.
What open models can lock in — and what they can commoditise
| Layer | Effect of open weights | Strategic consequence |
|---|---|---|
| Model architecture | Faster international diffusion | Strengthens developer familiarity |
| API layer | Often weakened if compatible alternatives exist | Reduces proprietary model lock-in |
| Inference engine | Encourages vLLM/SGLang/etc. portability | Can strengthen neutral infrastructure |
| Accelerator | Model can potentially migrate across hardware | Limits national-stack capture |
| Fine-tunes / enterprise data | Local adaptation increases switching cost | Lock-in migrates downstream |
| Agent/workflow layer | Tool definitions and business logic persist independently of model | Platform can capture value |
| Cloud | Customer can change model without changing cloud | Cloud may become stronger than model provider |
| National standard influence | Chinese architectures can spread without Chinese hosting | Strategic influence and revenue may diverge |
Inference economics cannot be reduced to a token-price league table
The widening price dispersion in China strongly supports the claim that inference economics matter, but raw token pricing is an incomplete metric because models differ in reasoning depth, latency, success rate, cache behaviour, context, output verbosity, tool use and required human correction. A model costing one-fifth as much per token can be more expensive per successful workflow if it uses five times the tokens, fails more often or requires significantly greater verification.
Illustrative published API economics, September 2026
| Model / service | Published input price | Published output price | Important qualification | Source |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | US$0.30/M peak cache-miss; US$0.15 off-peak | US$1.20/M peak; US$0.60 off-peak | Cache-hit input far cheaper | DeepSeek |
| Qwen3.6-27B Beijing | ~US$0.413/M | ~US$2.475/M | Specific Alibaba region/model | Alibaba Cloud |
| Qwen3.6-27B Singapore | US$0.60/M | US$3.60/M | International region | Alibaba Cloud |
| Hy4 Preview Guangzhou | RMB 6/M | RMB 18/M | Tencent regional price | Tencent Cloud |
| GLM-5.3-Flash via TokenHub | RMB 0.8/M | RMB 2.8/M | Tencent distribution price | Tencent Cloud |
| Kimi K3 official API | US$3/M cache miss; US$0.30/M cache hit | US$15/M | 1M context; cache economics matter strongly | Kimi |
| MiniMax M2.7 | US$0.30/M | US$1.20/M | Cache read US$0.06/M | MiniMax |
Sources and current price documentation: DeepSeek, Alibaba Cloud, Tencent Cloud, Kimi and MiniMax.
The table demonstrates segmentation, not an overall ranking. In particular, GLM-5.3-Flash’s Tencent-hosted price cannot legitimately be compared with Kimi K3’s output price as if both represented equivalent intelligence, and a cached coding workload cannot be compared with a short uncached consumer query. Procurement decisions require a denominator based on completed work rather than tokens consumed.
Agents are changing the economic unit from invocation to task execution
The transition toward agents is visible across nearly every major Chinese platform. Alibaba AgentCore provides workspaces, identities, agent teams, skills, tool access, observability and evaluation; Tencent WorkBuddy Enterprise combines coding, office and managed agents under one enterprise layer; Baidu has launched DuMate and other agents while evolving Qianfan toward agent infrastructure; MiniMax emphasises agent teams and complex tool use; and Kimi offers Kimi Code and enterprise products while developing hosted-agent infrastructure.
This transition alters inference economics because a long-running agent may perform dozens or hundreds of model calls, web searches, database operations, code executions and verification steps for a single user objective. The commercially relevant metric becomes the cost and reliability of task completion, not the cost of a single prompt. Platform features such as caching, tool routing, model selection, persistent context and sandboxing can therefore have greater economic effects than modest differences in base model price.
Huawei is attempting to close the compute-stack gap vertically
Huawei provides the clearest Chinese attempt to integrate accelerators, memory, interconnect, software, supernodes, cloud and application deployment within one corporate architecture. Its Ascend ecosystem includes the CANN software layer and increasingly broad support for PyTorch, Triton, vLLM and related open-source projects; Huawei stated in September 2026 that more than 40 models had been natively pre-trained on Ascend/CANN and that Ascend supported more than 90 major third-party open-source projects. These are first-party claims, but they indicate a deliberate effort to solve one of China’s most consequential weaknesses: accelerator hardware without an equally mature developer software ecosystem.
Huawei’s infrastructure roadmap extends beyond individual accelerators. The company reports that the previously deployed Atlas 900/CloudMatrix384 architecture links 384 Ascend 910C chips into a logical supernode and that more than 300 such systems had been deployed across more than 20 customers by September 2025; the company subsequently announced Atlas 950 and a longer Ascend roadmap. Because these are Huawei disclosures, the comparative-performance claims require independent confirmation, but the deployment numbers are relevant evidence of commercialisation beyond laboratory prototypes.
The September 2026 roadmap is substantially more ambitious. Huawei stated that Ascend 960 development was ahead of its earlier schedule, that the Ascend ecosystem had moved toward annual generations, and that its new SuperCluster architecture could theoretically interconnect hundreds of thousands of NPUs. The strategic logic is explicit: where individual domestic chips cannot equal the strongest imported accelerators on every dimension, Huawei is attempting to compete through system-level scaling, interconnect, memory architecture and software optimisation.
HBM is not a peripheral constraint
Large-model inference and especially training depend not only on arithmetic throughput but on memory capacity and bandwidth. Long contexts intensify KV-cache requirements; mixture-of-experts systems move very large parameter sets; and high concurrency requires rapid movement of data between accelerator memory and compute units. Z.ai’s own GLM-5.2 engineering analysis states explicitly that increasing context to one million tokens shifts serving bottlenecks toward KV-cache capacity, long-context kernels and system overhead even where per-token FLOPs are reduced.
This is why U.S. restrictions on high-bandwidth memory remain strategically relevant. The December 2024 BIS package added controls over HBM alongside semiconductor-manufacturing equipment and software, explicitly linking those technologies to advanced computing and AI. Huawei’s response has included proprietary memory and system designs, such as the announced HiZQ 2.0 memory configuration for future Ascend generations, but future-product specifications are not equivalent to demonstrated mass production at required volumes.
The correct analytical unit is therefore usable memory bandwidth per deployed system, not merely chip FLOPS. A theoretically powerful accelerator that cannot obtain adequate HBM or cannot communicate efficiently across a large cluster may produce poor real-world token economics, while a weaker accelerator can become commercially viable if software, quantisation and system architecture compensate sufficiently.
Advanced packaging is becoming nearly as important as the processor die
Modern AI accelerators increasingly rely on complex packaging to combine logic dies, memory and high-bandwidth interconnect. As cluster scale rises, packaging yield, substrate availability, thermal management and optical/electrical interconnect become system-level constraints. Huawei’s focus on SuperPoDs, near-packaged optics and unified interconnect reflects this shift from “chip performance” toward system performance per rack, per watt and per unit of interconnect bandwidth.
Public evidence remains insufficient to establish China’s position across all advanced-packaging inputs and volumes at a granularity comparable with its model and cloud markets. This gap is important because successful domestic accelerator design does not automatically imply the ability to manufacture, package and deploy those accelerators in the quantities required by national AI demand.
Export controls remain a constraint, but no longer describe the whole system
U.S. semiconductor restrictions remain an important part of the Chinese compute environment because they target advanced computing chips, HBM, manufacturing equipment and supporting software. The restrictions identify precisely the upstream layers at which China’s AI ecosystem is least comparable with its far more mature downstream application and telecommunications base. The January 2026 shift to case-by-case treatment for certain H200-, MI325X- and comparable-class exports altered access conditions but preserved regulatory control rather than normalising unrestricted trade.
The analytical mistake would be to infer from those restrictions either that China lacks usable AI compute or that domestic substitution has already neutralised the restrictions. The first proposition is contradicted by domestic cloud infrastructure, Ascend deployment and Z.ai’s reported >100,000-accelerator inference system; the second is unsupported because neither aggregate domestic accelerator supply nor independent evidence of leading-edge memory, fabrication and advanced-equipment autonomy is sufficiently complete.
Compute-stack assessment, September 2026
| Layer | Chinese capability visible in public record | Degree of integration | Principal unresolved constraint |
|---|---|---|---|
| Foundation models | Broad and competitive portfolio | High | Capability distribution and economics |
| Model serving software | Rapidly improving; vLLM/SGLang integration, CANN, proprietary engines | Medium–High | Maturity and cross-hardware portability |
| Cloud distribution | Alibaba, Tencent, Baidu, Huawei and others | High | Capacity, concentration and regional reach |
| Domestic AI accelerators | Ascend and other domestic systems deployed | Medium–High | Supply volume and performance economics |
| Large accelerator clusters | Demonstrated in first-party disclosures | Medium | Independent validation and national scale |
| HBM / high-bandwidth memory | Domestic alternatives being developed | Medium–Low / uncertain | Volume, bandwidth, yields and technology generation |
| Advanced packaging | Growing domestic capability | Medium / insufficiently transparent | Scale, yield, substrates, thermal systems |
| Leading-edge fabrication | Significant domestic capability but constrained | Incomplete | Equipment, node economics and yield |
| Semiconductor manufacturing equipment | Material localisation effort | Incomplete | Critical foreign tool dependencies |
| Full production autonomy | Not established | Incomplete | Combined upstream dependencies |
China now has an integrated deployment stack, but not yet a demonstrably autonomous production stack
The evidence therefore permits a more precise answer to the chapter’s central question. China unquestionably possesses more than an “application layer”: it has multiple advanced foundation-model families, domestic and international APIs, open-weight ecosystems, agent platforms, cloud infrastructure, domestic accelerators, inference software, industrial applications and increasingly large domestic compute clusters. Huawei and Alibaba are explicitly integrating hardware and software vertically, while Tencent and Alibaba are simultaneously building multi-model horizontal control layers.
What China does not yet publicly demonstrate at comparable confidence is independent control of every critical upstream production input at the quantities and economics needed to sustain unconstrained frontier training and national-scale inference. The difference between those statements is central: a country does not need technological autarky to possess a powerful ecosystem, but an ecosystem exposed to HBM, lithography, equipment or fabrication constraints carries a different strategic risk profile from one whose critical production inputs can all be expanded domestically.
Key judgments
| Judgment | Evidence | Confidence |
|---|---|---|
| China’s AI market is structurally multi-model rather than DeepSeek-centric | Current platform catalogues and model releases | High |
| Chinese developers pursue substantially different strategies across openness, scale, sparsity, price and deployment | Product documentation | High |
| Multi-model cloud and agent platforms are becoming strategically more important than individual models | Alibaba, Tencent, Baidu architectures | High |
| Open-weight diffusion can expand Chinese model influence without producing Chinese infrastructure lock-in | Cross-platform deployment evidence | High |
| Domestic accelerator systems are now supporting meaningful production inference | Z.ai and Huawei disclosures | Moderate–High |
| Huawei is building a more vertically integrated domestic compute/software stack | Ascend/CANN/SuperPoD record | High as corporate strategy; moderate for comparative performance |
| HBM, advanced packaging and manufacturing equipment remain materially important constraints | Technical workload requirements and export-control architecture | High |
| China’s AI deployment stack is substantially integrated | Multi-layer deployment evidence | High |
| Complete upstream technological autonomy has been achieved | Public record does not establish this | Low / not established |
What would change the assessment
The assessment would strengthen significantly if independently auditable evidence established sustained frontier-model training across multiple Chinese laboratories on domestic accelerators, domestically sourced high-bandwidth memory and domestic manufacturing toolchains at competitive cost; if Chinese cloud providers disclosed rising shares of domestic accelerator utilisation without deterioration in availability or price; and if independent international customers adopted Ascend/CANN or another Chinese accelerator stack rather than merely consuming Chinese model weights on non-Chinese infrastructure.
It would weaken if announced domestic accelerators proved difficult to manufacture at scale, if HBM or advanced packaging constrained cluster growth, if software incompatibility substantially increased engineering costs, or if open-model portability caused infrastructure value to migrate toward hardware-neutral global cloud and inference platforms.
Open official record
The highest-value missing records are accelerator shipment volumes by vendor and generation; cloud workload shares by accelerator type; domestic HBM output, bandwidth and yields; advanced-packaging capacity dedicated to AI; semiconductor manufacturing equipment localisation by process step; power consumption and utilisation of large AI clusters; cost per delivered token on domestic versus imported accelerators; and independently verified training runs for frontier-scale models on entirely domestic stacks.
Chapter 3 — Industrial Embedding and the Economics of Integration
Principal judgment
The decisive economic transition in Chinese AI is beginning to occur after model inference, where intelligence is attached to enterprise data, software environments, industrial equipment, regulatory procedures and human workflows. Evidence from manufacturing, coding, telecommunications, government services, drug regulation and enterprise-agent platforms shows that deployment is already moving beyond standalone chat interfaces; however, adoption depth remains highly uneven, and many public examples are vendor case studies rather than independently measured productivity programmes. The most defensible conclusion is therefore that China possesses a rapidly expanding industrial experimentation and deployment surface, while the still-unresolved issue is whether these installations generate sufficiently durable productivity, reliability and switching costs to turn model adoption into structural ecosystem advantage.
The underlying economic mechanism is straightforward but frequently misunderstood. A foundation model has relatively little organisational value until it is connected to proprietary data, authorisation systems, software tools, production equipment, human approval processes and measurable business outcomes. Every such connection requires engineering and governance expenditure, but every successful connection can also create switching costs, because replacing the model later may require revalidation, retraining, prompt or workflow redesign, security review and operational testing. The principal competitive question is consequently not whether a Chinese model can answer the same benchmark questions as a Western model, but whether an ecosystem can repeatedly convert adequate model capability into lower total deployment cost and lower organisational friction.
Industrial embedding has several distinct technical layers
AI adoption inside an enterprise should not be treated as a single implementation event. At minimum, production deployment normally requires a model layer, access to enterprise data, retrieval or domain adaptation, tool invocation, identity and access control, monitoring, human escalation and integration with existing systems. Industrial environments add additional requirements such as deterministic control, edge inference, machine interfaces, operational technology security, uptime and physical-safety constraints.
The industrial AI integration stack
| Layer | Function | Typical integration cost | Principal source of switching cost |
|---|---|---|---|
| Foundation model | Reasoning / generation / perception | API or infrastructure cost | Model-specific behaviour |
| Enterprise data | Context and proprietary knowledge | Cleaning, permissions, connectors | Data schemas and retrieval architecture |
| RAG / knowledge layer | Controlled information retrieval | Indexing, evaluation, maintenance | Knowledge pipelines |
| Agent orchestration | Multi-step execution | Tools, state, planning, routing | Agent logic and tool definitions |
| Identity / security | Access control | IAM, auditing, policy design | Enterprise governance architecture |
| Application connectors | ERP, CRM, MES, PLM, office systems | Custom APIs and middleware | Workflow integration |
| Edge / OT interface | Machines, sensors, production control | Hardware, latency, certification | Physical infrastructure |
| Human validation | Review and exception handling | Labour and process redesign | Organisational routines |
| Monitoring / evaluation | Reliability and compliance | Logs, benchmarks, incident handling | Governance and operational tooling |
| Training / change management | User adoption | Education and process change | Human capital and organisational learning |
The table explains why model substitution can become inexpensive at one layer but expensive at the system level. If a firm already exposes internal tools through standard interfaces and maintains model-independent evaluation, changing the model can be relatively easy; if business logic, permissions, retrieval, prompts and quality controls are tightly coupled to a particular vendor environment, the apparently simple API substitution becomes an organisational migration project.
Manufacturing is the strongest test because physical processes impose hard constraints
Manufacturing is strategically important precisely because it is less tolerant of hallucination and latency than consumer chat. AI used for equipment monitoring, inspection, process optimisation or safety interacts with physical assets whose failure creates direct cost, so integration requires model accuracy, sensor reliability, edge/cloud coordination and operational fallback mechanisms.
Huawei and Conch Group’s cement-industry deployment provides a concrete example. Huawei reported that the jointly developed system uses Pangu computer-vision models and distributed optical-fibre sensing for 28 equipment-management scenarios, including roller anomalies and belt tearing, while safety-production functions monitor more than 20 categories of personnel violations and equipment malfunction, with Huawei claiming 95% identification accuracy for those events. The same deployment includes natural-language access to industry knowledge and expert experience. These remain first-party/vendor claims, but they demonstrate the architecture of industrial AI: model + sensors + domain data + continuous monitoring + production workflow.
Earlier Huawei evidence from Shandong Energy illustrates the same pattern in mining: Pangu was reported as operating across nine processes and 21 scenarios including digging, drivage, equipment control, transport, ventilation and coal washing, with the company attributing an additional 8,000 tonnes of cleaned coal per year to one deployment. The productivity figure requires independent validation, but the breadth of processes demonstrates that industrial embedding can create more persistent integration than generic model access because knowledge, sensors and operating procedures accumulate around the system.
Why manufacturing integration is difficult to reverse
| Component | Initial deployment burden | Why it can create persistence |
|---|---|---|
| Machine/sensor connectivity | High | Physical interfaces remain in place |
| Historical maintenance data | Medium–High | Model learns plant-specific failure patterns |
| Computer-vision calibration | High | Camera placement and thresholds become site-specific |
| MES/SCADA integration | High | AI becomes part of production workflow |
| Safety approval | High | Replacement may require recertification |
| Operator training | Medium | Human procedures adapt to system |
| Edge compute | Medium–High | Hardware installed near production assets |
| Continuous monitoring | Medium | Operational data accumulates over time |
The lock-in here is not necessarily to a Chinese model. It is initially to the integrated production architecture. If the architecture supports model substitution behind a stable interface, models may remain contestable while the system integrator captures the economic rent; if the software and hardware are tightly coupled, model and infrastructure lock-in can reinforce each other.
Industrial AI therefore rewards companies with pre-existing enterprise relationships
The importance of China’s cloud, telecommunications and industrial-equipment companies becomes clearer at this layer. A foundation-model startup may possess a technically superior model yet still lack access to factory data, ERP systems, regulated customer relationships or edge hardware. Huawei, Alibaba, Tencent and Baidu can approach AI deployment through pre-existing infrastructure, cloud contracts, enterprise communication systems and industrial partnerships, potentially reducing customer acquisition and integration friction.
Huawei’s September 2026 Industry AI Foundry illustrates this logic. Huawei states that the platform has accumulated more than 1,000 industry assets and supports more than 1,000 deployed projects, while its AgentArts enterprise agent platform serves more than 100 enterprises. Those figures are vendor disclosures and do not establish independent economic impact, but they show how a platform attempts to transform reusable industry assets into lower marginal deployment cost across subsequent customers.
Huawei’s newly published seven-step adoption method similarly reflects the fact that model selection represents only one stage of enterprise AI. The company’s framework centres on transforming enterprise data into knowledge, knowledge into models, models into agents and agents into scaled business operations; this is commercially self-interested guidance, but it accurately reveals where system integrators expect implementation revenue and switching costs to accumulate.
Software development is emerging as the fastest laboratory for agent economics
Software engineering differs from manufacturing because the environment is already digital, tool interfaces are relatively accessible and output can often be automatically tested. It therefore provides an unusually favourable environment for agents capable of reading repositories, editing code, executing commands, observing failures and retrying.
Tencent CodeBuddy illustrates vertical integration from model access into the development environment. Tencent offers CodeBuddy in IDE, plug-in and CLI forms, with support for its own models and third-party systems through TokenHub; enterprise deployments can use dedicated or private environments. This means model choice can be changed without abandoning the surrounding development interface, which strengthens Tencent’s control of the application layer even while weakening any single model’s lock-in.
Kimi follows a similar strategy through Kimi Code, which supports both OpenAI-compatible and Anthropic-compatible protocols and exposes K3 through terminal-based coding workflows. MiniMax likewise integrates its models into external coding environments and sells token plans designed for persistent coding-agent consumption. The common strategic direction is clear: compatibility is becoming a customer-acquisition mechanism, while durable value is sought in recurring agent usage rather than proprietary API syntax.
MiniMax’s own internal M2.7 case study makes the economics particularly visible. The company states that M2.7 was used to construct research-agent infrastructure, assist with model-development workflows and reduce recovery times for some live production incidents to under three minutes; these are first-party claims, but they demonstrate the direction of travel from code suggestion toward closed-loop engineering agents that operate within live development environments.
The enterprise-agent layer converts interoperability into a competitive weapon
Alibaba AgentCore, Tencent WorkBuddy Enterprise and Baidu Qianfan all increasingly abstract the model behind an enterprise control plane. AgentCore can connect multiple model providers and external tools while maintaining unified agent governance and monitoring; WorkBuddy Enterprise supports multi-model access across coding, office and managed-agent workflows; Qianfan has expanded from ERNIE access into an agent-oriented platform containing proprietary and third-party models.
This architecture can produce a counter-intuitive market outcome: the more interchangeable models become, the more valuable the orchestration layer may become. A company that can swap Qwen for DeepSeek, Hy for GLM or another provider without rebuilding authentication, tool access and business processes has weak model lock-in but strong platform lock-in.
Where switching costs accumulate in an agentic enterprise
| Asset | Model-specific? | Switching burden after deployment | Likely rent-capture layer |
|---|---|---|---|
| Prompt templates | Partly | Low–Medium | Model / integrator |
| Fine-tuning | Often | Medium–High | Model vendor |
| RAG knowledge base | Usually portable | Medium | Data/orchestration platform |
| Tool definitions | Often portable in principle | Medium | Agent platform |
| Identity and permissions | Platform-specific | High | Cloud / enterprise platform |
| Agent memory | Platform-dependent | Medium–High | Agent platform |
| Observability | Platform-specific | Medium | Cloud / platform |
| Evaluation datasets | Portable | Low–Medium | Enterprise |
| Workflow integrations | Often custom | High | Integrator / platform |
| Employee training | Product-specific | Medium | Application platform |
| Regulatory approval | Deployment-specific | High | Installed system |
The table shows why the strategic standard may emerge above the model. If enterprises deliberately make evaluation data, tools and model interfaces portable, they can preserve bargaining power; if they allow the entire agent stack to become proprietary, the model provider or cloud platform can obtain much stronger customer control.
Logistics exposes the importance of latency, routing and real-world exceptions
Logistics is another environment in which value depends less on conversational quality than on persistent access to orders, inventories, routes, documentation, customs data and exception management. Alibaba’s September 2026 global enterprise announcement identifies logistics among the sectors being targeted by its full-stack cloud and AI services and cites collaborations including Lion Parcel, while Alibaba’s broader agent architecture supports external tools and multi-agent workflows. These are corporate deployment announcements rather than independently measured productivity studies, but they illustrate how cloud providers seek to place models inside operational systems rather than sell them as standalone assistants.
The economic advantage of a sufficiently capable lower-cost model can be substantial in logistics because workloads are repetitive, high-volume and time-sensitive. Document extraction, shipment-status interpretation, customer support, exception classification, routing assistance and procurement can generate enormous token volumes, meaning that modest per-request differences compound materially. However, logistics also punishes reliability failures because incorrect customs classifications, addresses or routing instructions can impose real financial costs; therefore the “cheapest model wins” proposition is valid only after acceptable error thresholds are satisfied.
Telecommunications transforms AI from an application into network workload
Telecommunications operators occupy a unique position because they provide both connectivity and increasingly compute, storage and edge infrastructure. Huawei’s June 2026 live-network validation with China Mobile Hubei reported an up to 372% increase in token throughput for long-sequence inference using OceanStor A800 storage, Ascend A3 SuperPoD and Unified Cache Manager. This is a joint vendor/operator disclosure rather than an independent benchmark, but it directly illustrates how storage architecture and KV-cache management can influence inference economics independently of the foundation model itself.
This point is strategically important: if AI agents consume very long contexts, the bottleneck can migrate from arithmetic throughput toward memory, storage and communication. Telecom groups that already control network infrastructure and regional data centres can therefore become AI infrastructure providers even without owning the highest-performing foundation model.
Finance illustrates why governance can outweigh raw model price
Financial applications expose a different integration problem because access control, auditability, data residency, operational resilience and human accountability can dominate inference price. An inexpensive model that cannot operate within institutional identity controls, private networks and traceable approval workflows may be unusable regardless of benchmark quality.
Huawei’s 2026 AI data-centre work with China Construction Bank is one example of infrastructure providers targeting financial-sector workloads, while the broader Huawei enterprise stack emphasises private and controlled deployment; Baidu, Alibaba and Tencent similarly offer enterprise platforms capable of isolation and dedicated resources. The evidence supports the presence of sector-targeted infrastructure but remains insufficient for a sector-wide quantitative assessment of productivity gains, so claims that Chinese AI has already transformed bank economics would exceed the current official record.
The implication is that finance strengthens the case for total integration cost over API price. Procurement decisions can involve model evaluation, cybersecurity, model-risk governance, private data connectivity, audit trails, staff controls and regulatory review, meaning the integration project can cost orders of magnitude more than the marginal inference bill.
Healthcare demonstrates the limits of “good enough”
Healthcare is precisely the kind of sector in which the cheapest adequate model principle requires a particularly high threshold for “adequate.” The National Health Commission has documented continued expansion of AI-enabled smart healthcare networks, while China’s broader policy framework is promoting AI in medical-service and public-health infrastructure. Nevertheless, diagnostic, treatment and regulatory applications impose higher verification requirements than office automation, and deployment evidence should not be conflated with clinically demonstrated superiority.
The National Medical Products Administration’s August 2026 “Artificial Intelligence + Drug Regulation” implementation framework is more concrete from an institutional perspective because it formally embeds AI into pharmaceutical regulation and explicitly links the initiative to a nationally integrated smart-regulation system. This indicates that AI is entering regulatory workflows themselves, not merely commercial healthcare applications.
Healthcare therefore illustrates a critical boundary condition for the central thesis: where false outputs carry high clinical or legal costs, frontier capability, validation quality and traceability can remain more important than token price. “Good enough and cheap” becomes economically dominant only once the system passes a substantially higher reliability threshold.
Public administration is moving toward governed domain deployment
China’s October 2025 guidance for large-model deployment in government explicitly identifies government services, social governance, internal office work and decision support as application categories. The guidance describes intelligent question answering, form pre-filling, assisted review, case handling and knowledge retrieval, while emphasising data governance, controlled deployment and safety. This is particularly significant because public administration is a large market in which models can become deeply embedded through domain knowledge bases, procedural rules and internal document systems.
A subsequent 2026 policy on “AI + Human Resources and Social Security” provides quantifiable adoption objectives. The programme calls for approximately 20 application scenarios and associated high-quality datasets in 2026, around 50 high-value scenario pathways by 2027, and broad sectoral AI deployment by 2030, while explicitly linking compute, models, data and application scenarios into a coordinated ecosystem. These remain policy objectives, not achieved outcomes, but they demonstrate that public-sector implementation is increasingly being designed as a domain-stack problem rather than generic chatbot procurement.
Public-sector embedding architecture
| Government use case | Required data layer | Integration burden | Principal risk |
|---|---|---|---|
| Citizen Q&A | Regulations, procedures, FAQs | Medium | Incorrect guidance |
| Assisted application processing | Forms, eligibility rules, historical cases | High | Administrative error |
| Document drafting | Internal templates and records | Medium | Confidentiality / factual error |
| Case triage | Case files and classification rules | High | Bias / improper prioritisation |
| Decision support | Multi-source administrative data | Very high | Automation bias |
| Regulatory inspection | Laws, filings, operational data | Very high | Incorrect enforcement action |
| Benefits administration | Identity, employment, eligibility records | Very high | Rights and entitlement errors |
Government applications can create particularly strong switching costs because models become embedded in procedures, secure data environments, document templates and administrative accountability structures. Yet these same characteristics make model substitution possible only after extensive revalidation, meaning that the initial platform architecture can have long-lived procurement consequences.
Integration cost has five economically distinct components
The proposition that “cheapest to integrate” can defeat “best model” becomes rigorous only when integration cost is decomposed rather than treated rhetorically.
The five components of total AI deployment cost
Technical integration cost includes APIs, data connectors, retrieval systems, fine-tuning, tool calling, identity, observability, edge deployment and systems testing; it is usually highest during initial installation but reappears whenever models, infrastructure or workflows change.
Inference operating cost includes input and output tokens, reserved accelerator capacity, cache storage, networking, vector databases, agent execution and retries; this component grows directly with workload volume and can therefore make inexpensive models economically attractive in large repetitive processes.
Verification cost includes human review, automated evaluation, exception handling and error remediation; a cheap model with a higher failure rate can become more expensive than a premium model after these costs are included.
Governance cost includes cybersecurity, privacy, audit, legal assessment, model-risk management, regulatory compliance and operational-resilience controls; in finance, government, healthcare and critical infrastructure this component can dominate model cost.
Organisational transition cost includes employee training, redesigned responsibilities, revised incentives, process changes and resistance management; this can create the deepest long-term lock-in because workflows and organisational knowledge become adapted to the installed system.
Total-cost formulation
A useful decision framework is therefore:
Total cost of deployed intelligence = model/inference expenditure + infrastructure + integration engineering + verification + governance + human transition + expected failure cost.
The denominator should be successfully completed useful work, not tokens generated.
This framework explains why a model with a higher API price can still be economically superior if it requires fewer retries and less human supervision, while a cheaper model can dominate when task complexity is moderate, volume is enormous and integration is already standardised.
When “good enough + cheap to integrate” wins
The stack-min-cost logic becomes strongest when several conditions occur simultaneously: the task has a clear minimum capability threshold; multiple models exceed it; workload volume is large; error costs are limited or cheaply detectable; enterprise data provide much of the task-specific advantage; and the model can be integrated without extensive proprietary dependencies.
Typical examples include document classification, structured extraction, routine customer correspondence, code transformation, retrieval over controlled internal documents, first-line support and repetitive administrative assistance. In these settings, moving from 98 to 99 units of abstract model quality may have little commercial value if both systems reliably meet the operational requirement, while a large difference in inference or deployment cost persists.
Conditions favouring stack-min-cost
| Condition | Why it favours lower-cost integrated AI |
|---|---|
| Very high request volume | Small unit-cost differences compound |
| Stable repetitive task | Capability threshold easier to define |
| Structured output | Automated validation reduces failure cost |
| Strong enterprise data | Domain context narrows model-quality gap |
| Effective caching | Large reduction in recurring context cost |
| Low-latency requirement | Efficient local inference gains value |
| On-premise requirement | Open weights / domestic infrastructure gain value |
| Model portability | Enterprise can arbitrage providers |
| Mature integration platform | Switching models becomes inexpensive |
| Human review already embedded | Model errors can be intercepted cheaply |
DeepSeek, MiniMax, Qwen and GLM price structures make this strategy technically plausible, while multi-model platforms from Alibaba and Tencent create an environment in which enterprises can route different workloads toward different cost-performance tiers rather than standardise on one “best” model.
When frontier capability still dominates
The opposite logic applies when task failure is expensive, capability differences remain large or errors are difficult to detect. Scientific reasoning, advanced engineering, complex autonomous coding, high-stakes legal analysis, medical applications, sophisticated cyber defence and open-ended strategic analysis can all create high economic value from incremental capability because a single correctly completed difficult task may be worth substantially more than the inference bill.
The growing use of effort levels in GLM, Kimi and other systems itself demonstrates market recognition that users want to allocate more computation selectively to harder problems rather than consume maximal reasoning for every request. GLM-5.2 supports selectable reasoning effort, while Kimi K3 is designed around multiple thinking-effort modes. This architecture suggests that the market may converge not toward one cheap model but toward adaptive compute allocation, where cheap inference handles routine work and expensive reasoning is invoked only when marginal capability has economic value.
Conditions favouring quality-max
| Condition | Why premium capability remains valuable |
|---|---|
| Difficult reasoning beyond cheaper models | Capability threshold not met |
| High cost of undetected error | Reliability dominates price |
| Low task volume / high task value | Token savings economically trivial |
| Scientific or engineering discovery | Incremental reasoning has asymmetric value |
| Complex codebase transformation | Failure remediation expensive |
| Weak automated validation | Human review burden rises |
| High regulatory or safety exposure | Conservative model choice preferred |
| Multi-step autonomous execution | Error propagation compounds |
Enterprise lock-in may form even when model lock-in does not
One of the most important conclusions from current Chinese platform architecture is that model competition and ecosystem competition can move in opposite directions. Alibaba and Tencent actively distribute rival Chinese models; Kimi and MiniMax expose compatibility with external tools; GLM can run locally; DeepSeek supports multiple API conventions. This lowers barriers to model switching.
Yet enterprises simultaneously become more dependent on the platform that manages model access, identity, agent memory, tool permissions, evaluation, billing, data and deployment. The market could therefore become highly competitive at the model layer while becoming concentrated at the cloud or agent-control layer.
This pattern resembles database and cloud markets more than oil. Enterprises can technically move workloads, but migration becomes expensive after schemas, applications, permissions and operating knowledge accumulate around an installed platform.
Integration creates learning effects that benchmark comparisons miss
Every production deployment generates information about where models fail, which prompts work, what data are missing, how employees behave and which processes should be redesigned. These feedback effects create an organisational learning curve. A company with hundreds of industrial deployments may therefore improve implementation efficiency even if its underlying model is not always the strongest.
Huawei’s more than 500 reported Pangu scenarios across over 30 industries, its thousand-plus Industry AI Foundry assets and Alibaba’s global enterprise expansion are relevant primarily for this reason rather than as proof of technological superiority: repeated deployment potentially generates reusable templates, connectors, evaluation procedures and domain expertise that lower future implementation costs.
The critical empirical test is whether those assets actually shorten deployment time and reduce cost across customers. Public vendor disclosures currently provide only partial evidence, so this remains a major research requirement rather than an established conclusion.
Scale can produce a self-reinforcing deployment loop
If integration learning is real, a plausible economic mechanism emerges:
more deployments → more domain data and implementation experience → reusable tools and lower deployment cost → more competitive bids → more deployments.
This differs from the classic consumer-network effect because the advantage does not require every user to interact with every other user. It is closer to learning-by-doing and economies of scope across industrial implementation.
Such a loop could favour Chinese providers domestically because their enormous manufacturing and public-sector deployment surface provides many opportunities to refine domain-specific integrations; however, it could remain geographically bounded if data regulation, cybersecurity concerns, procurement restrictions or incompatible enterprise software prevent those implementation assets from transferring internationally.
International expansion is therefore the decisive test of ecosystem power
Domestic deployment alone cannot establish that China is setting a global AI standard. A Chinese company can achieve enormous domestic scale while international enterprises continue to use Chinese models through non-Chinese clouds or avoid Chinese infrastructure entirely. The stronger test is whether the integrated stack—rather than only the model weights—wins outside China.
Alibaba is directly attempting this transition: by June 2026 it reported 105 availability zones across 32 regions, and in September announced additional regions and data-centre expansion across Europe, the Middle East and Asia. These are first-party infrastructure disclosures, but they demonstrate that international cloud presence is being treated as a prerequisite for exporting AI services.
Kimi, MiniMax, DeepSeek and open Qwen/GLM models can diffuse internationally with much less physical infrastructure because developers can host them elsewhere. Their success would therefore demonstrate Chinese model influence, but not necessarily Chinese cloud or hardware influence.
Four different forms of Chinese AI internationalisation
| Export form | Example | Chinese control retained | Strategic significance |
|---|---|---|---|
| Hosted Chinese API | Kimi / DeepSeek direct API | High at model-service layer | Recurring revenue and telemetry |
| Chinese cloud stack | Alibaba Cloud / Huawei Cloud | High across multiple layers | Stronger infrastructure dependence |
| Open model on foreign cloud | Qwen / GLM / DeepSeek hosted elsewhere | Limited | Model standard spreads without full-stack capture |
| Locally hosted open weights | Sovereign/private deployment | Low operational control | Architectural influence but weak recurring rent |
The distinction should be maintained throughout the dossier because claims that a Chinese open model is “winning globally” can mean fundamentally different things depending on where inference, data, infrastructure and revenue actually reside.
Industrial adoption also creates new countervailing forces
The same integration that creates lock-in creates risk. A tightly embedded agent can amplify errors across many processes; a platform outage can halt multiple workflows; proprietary agent memory can complicate migration; centralised data access increases cybersecurity exposure; and autonomous execution can produce actions rather than merely incorrect text.
This means enterprises have incentives to preserve modularity. Multi-model routing, open APIs, portable evaluation sets, independent data layers and standard tool protocols can become deliberate anti-lock-in strategies. Alibaba, Tencent, Kimi and MiniMax themselves increasingly support compatibility because reducing initial adoption friction is commercially valuable, even if that compatibility later gives customers greater bargaining power.
The emerging competition is therefore over the control plane
The evidence from Chapters 2 and 3 together points toward a deeper conclusion than “China has many cheap models.” Foundation models are becoming one layer inside a broader control architecture comprising model routing, enterprise identity, data, agent memory, tools, governance, infrastructure and workflow automation.
The actor controlling that layer gains several advantages: it observes enterprise workload demand; can route workloads among models; can optimise cost against latency; manages permissions and compliance; can introduce proprietary services without replacing the underlying application; and becomes expensive to remove after workflows have been reconstructed around it.
This explains why Alibaba is building AgentCore, Tencent is integrating TokenHub with WorkBuddy and CodeBuddy, Baidu is evolving Qianfan into agent infrastructure and Huawei is combining Industry AI Foundry with Ascend and AgentArts. Their strategies differ, but all move toward owning the execution environment around intelligence, not simply selling intelligence itself.
Sector-by-sector assessment
| Sector | Deployment maturity visible in public record | Main economic mechanism | Principal constraint | Stack-min-cost potential |
|---|---|---|---|---|
| Manufacturing | Medium–High | Automation, inspection, predictive operations | Safety, OT integration | High after validation |
| Software development | High / rapidly expanding | Labour augmentation, autonomous task execution | Reliability on long tasks | Very high |
| Logistics | Medium | High-volume documents, routing, exception handling | Real-world error cost | High |
| Finance | Medium | Document, analysis, service and operations automation | Governance and audit | Moderate |
| Healthcare | Medium but highly controlled | Knowledge support and workflow assistance | Clinical safety and validation | Low–Moderate for high-stakes tasks |
| Telecommunications | High at infrastructure layer | Compute/network/inference optimisation | Capex and power | High |
| Public administration | Policy-driven expansion | Case handling, Q&A, document and benefits workflows | Accountability and data protection | High for routine processes |
| Industrial control | Selective but deep | Sensor interpretation and operational optimisation | Determinism, safety, edge reliability | High only after stringent validation |
A more useful economic metric: cost per verified completed task
The critical metric for enterprise AI procurement should therefore be neither benchmark score nor dollars per million tokens, but cost per verified successfully completed task at required service quality.
A simplified measurement framework would include:
| Variable | Measurement |
|---|---|
| Model inference | Tokens × live model price |
| Cache efficiency | Effective repeated-context discount |
| Tool calls | API, database and external-service expense |
| Compute reservation | Dedicated GPU/NPU capacity |
| Integration labour | Engineering hours × labour cost |
| Human review | Reviewer time per completed task |
| Retry rate | Failed/repeated agent runs |
| Error cost | Expected financial consequence of undetected failure |
| Governance | Security, legal, compliance and audit cost |
| Latency | Economic value of response time |
| Availability | Cost of outage / degraded performance |
| Switching cost | Expected future migration expense |
The denominator must then be the number or economic value of verified useful outcomes, producing a metric such as cost per successfully resolved case, completed code task, verified industrial inspection, processed administrative application or resolved customer request.
Only such measurements can establish whether Chinese model economics actually translate into superior industrial economics.
Falsifiers of the integration thesis
The thesis that ecosystem integration will become more decisive than isolated frontier-model quality would weaken materially if enterprises systematically demonstrated low switching costs after deep deployment; if standard agent protocols made platforms as interchangeable as models; if frontier capability gaps remained so large that firms repeatedly paid premium prices regardless of infrastructure; if Chinese domestic accelerators materially raised rather than lowered full-system cost; or if highly integrated industrial deployments failed to produce measurable productivity improvements.
Conversely, it would strengthen if enterprises increasingly routed routine work toward inexpensive models while reserving frontier systems for exceptional tasks; if agent platforms became persistent control planes independent of the selected model; if model prices continued to decline faster than integration costs; and if large Chinese industrial deployments generated reusable assets that demonstrably reduced subsequent implementation time and cost.
Key judgments
| Judgment | Evidence base | Confidence |
|---|---|---|
| Chinese AI deployment has moved materially beyond consumer chat | Industrial, agent and public-sector deployments | High |
| Enterprise integration costs increasingly matter more than nominal token price | Architecture and deployment evidence | High |
| Manufacturing and industrial control create particularly deep integration | Huawei industrial cases | Moderate–High |
| Coding is currently one of the clearest environments for agentic economic scaling | Tencent, Kimi, MiniMax, GLM products | High |
| Public administration is becoming an explicit domain-model market | Chinese government deployment guidance | High |
| Multi-model platforms reduce model lock-in while potentially increasing platform lock-in | Alibaba, Tencent and Baidu architecture | High |
| “Good enough + cheap” can dominate after capability thresholds are satisfied | Cost structure and workflow mechanics | High as conditional mechanism |
| It will dominate all AI workloads | Evidence does not support this | Low / rejected as universal proposition |
| Repeated industrial deployment can create learning-by-doing advantages | Mechanism supported; economy-wide magnitude unproven | Moderate |
| Chinese model diffusion automatically produces Chinese-stack dependence | Evidence contradicts automatic equivalence | Low |
What would change the assessment
The assessment would become substantially stronger if firms published independently verifiable before-and-after measures of labour hours, cycle time, defect rates, downtime, throughput, energy use or administrative processing cost following AI deployment; if those gains persisted after subsidy and pilot periods; and if providers demonstrated that previously built industry assets systematically lowered implementation cost for later customers.
It would weaken if most cited deployments remained demonstrations rather than sustained production systems, if high human-verification requirements offset inference savings, if enterprises standardised on hardware- and model-neutral agent interfaces that sharply reduced switching costs, or if organisations discovered that process redesign rather than model access constituted the dominant bottleneck and therefore limited rapid ecosystem scaling.
Open official record
The most decision-useful missing evidence is now operational rather than technological: verified AI penetration by industrial sector; production versus pilot deployment counts; task-level labour and productivity changes; inference volume by sector; human-review ratios; error rates before and after deployment; implementation time; integration expenditure; model-switching cost; share of workloads dynamically routed among models; public-sector procurement values; failure and incident data from production agents; and longitudinal measurements showing whether AI implementation becomes cheaper as integrators accumulate experience.
The answer to the broader thesis will depend increasingly on these indicators, because the strategic contest ceases to be a model race precisely when the majority of economic cost and organisational value migrate outside the model itself.
Pillar II — Standards, Platforms and the Geoeconomics of AI
Chapter 4 — Network Effects, Lock-In and the Battle for the Control Layer
Principal judgment
The most consequential network effects in artificial intelligence are migrating away from foundation-model weights themselves and toward the control layer surrounding the model: identity, permissions, enterprise data, retrieval architecture, agent memory, tool connections, observability, evaluation, security policy, cloud networking, inference optimisation, accelerator software and workflow state. Model weights can remain strategically important, particularly when capability differences are large, but the architecture emerging during 2025–2026 increasingly allows the model to be substituted while preserving the surrounding enterprise system. Microsoft now describes Agent 365 explicitly as an enterprise control plane in which agents receive identities, permissions, lifecycle controls and security policies; Amazon Bedrock AgentCore describes its managed harness as model-agnostic and capable of changing models during a session; OpenAI’s Agents API separates model choice from the agent harness, persistent session and execution environment; and Google presents Gemini Enterprise Agent Platform as an end-to-end environment for building, scaling, governing and optimising enterprise agents. These are not peripheral features: they indicate where durable platform power is beginning to accumulate.
This changes the central China-versus-West question. Open-weight Chinese models can spread rapidly through global developer ecosystems without forcing users onto Chinese clouds, Chinese accelerators or Chinese enterprise platforms, just as Meta’s Llama models have historically diffused through AWS, Azure, Google Cloud, NVIDIA and numerous independent inference providers without transferring all downstream economic value to Meta. China’s open-model strategy can therefore increase architectural influence while simultaneously commoditising the very model layer from which national ecosystem dependence would otherwise arise. The resulting struggle is not simply over which model developers attract the largest number of users; it is increasingly over which platform becomes the persistent intermediary between models and organisational activity.
Network effects in AI are not located in one place
The phrase “network effect” can obscure more than it reveals unless the underlying mechanism is specified. Classic direct network effects arise when a product becomes more useful as more users join the same network; many AI systems instead exhibit learning effects, complementor effects, ecosystem economies, installed-base effects and switching costs, which are economically related but not identical. A foundation model does not necessarily become more useful to one customer merely because another enterprise uses the same weights, whereas an agent platform can become more valuable when a growing ecosystem produces compatible tools, connectors, evaluation frameworks and integrations.
The relevant analytical question is therefore not “does AI have network effects?” but where increasing returns accumulate and who can capture them.
The principal locations of increasing returns in the AI stack
| Layer | Principal increasing-return mechanism | Strength of direct lock-in | Portability | Likely rent-capturing actor |
|---|---|---|---|---|
| Model weights | Model familiarity, fine-tuning ecosystem, derivative models | Low–Medium for open weights; higher for closed APIs | High for open models | Model developer |
| API interface | Developer integration and application code | Low if standardised | High where OpenAI-compatible or equivalent | API/model provider |
| Developer framework | Libraries, examples, tooling, skills | Medium | Medium–High | Framework/platform owner |
| Agent harness | Persistent state, orchestration, recovery, tool execution | High | Medium | Agent platform |
| Tool protocol | Connector ecosystem | Low if open standard | High | Ecosystem broadly |
| Enterprise data | Proprietary context and historical information | Very high organisational value | Potentially high technically | Enterprise / data platform |
| RAG architecture | Indexes, embeddings, retrieval policies, metadata | Medium–High | Medium | Data/cloud platform |
| Identity and security | Permissions, credentials, policy, audit | Very high | Low–Medium | Cloud / identity provider |
| Agent memory | Persistent workflow and user state | High | Medium–Low | Agent platform |
| Observability and evaluation | Traces, metrics, test suites, incident history | Medium–High | Medium | Platform / enterprise |
| Cloud region | Data residency, network proximity, commitments | High | Medium–Low | Hyperscaler |
| Inference engine | Hardware-specific optimisation | Medium–High | Medium | Infrastructure provider |
| Accelerator/compiler stack | Drivers, kernels, libraries, operator expertise | Very high | Low | Hardware/software platform |
| Industrial software | ERP/MES/PLM/process integration | Very high | Low | Enterprise-software vendor / integrator |
The central implication is that open models weaken one potential source of lock-in without eliminating the others. If a Qwen, DeepSeek or GLM model can be served through Alibaba Cloud, Azure, AWS, an independent European provider, NVIDIA NIM or a private cluster, then the model can achieve global diffusion while the user remains economically attached to a non-Chinese infrastructure provider. Alibaba itself demonstrates this multi-model architecture by offering Qwen alongside DeepSeek, GLM, Kimi and MiniMax in Model Studio, meaning that Alibaba can preserve the cloud relationship even when the preferred model changes.
Model weights are becoming less durable as a lock-in mechanism
A closed model accessed only through one provider can create switching costs because organisations accumulate prompts, evaluation datasets, model-specific behaviour assumptions and fine-tuned applications around the service. Yet the rapid spread of compatible APIs and open-weight models is pushing in the opposite direction. Alibaba exposes several competing Chinese model families through a common platform; Tencent’s WorkBuddy supports multi-model switching; NVIDIA NIM packages many different open models behind standard APIs; and Meta’s Llama strategy has deliberately encouraged deployment across competing clouds and hardware environments.
This produces an important distinction between model installed base and platform installed base. A company can become a major user of a model family without becoming dependent on the originating company’s infrastructure. Conversely, an enterprise can become highly dependent on a cloud or orchestration platform while replacing the underlying model repeatedly. In such an environment, model providers face the same strategic problem historically encountered by component suppliers: technological importance does not automatically imply control of the customer relationship.
Meta’s Llama experience is particularly instructive. By March 2025 Meta reported more than one billion Llama downloads, while the models were simultaneously available across AWS, Azure, Google Cloud, NVIDIA, Databricks and other platforms. That scale gave Meta considerable influence over model architecture and developer familiarity, but the distribution architecture deliberately permitted other companies to capture cloud, inference, integration and enterprise revenues. The Chinese open-weight strategy can produce a similar outcome.
APIs are becoming interfaces rather than ecosystems
The API itself can generate substantial switching cost when it is proprietary and applications depend heavily on provider-specific semantics, but the diffusion of broadly compatible interfaces reduces this mechanism. OpenAI-compatible interfaces have become common across third-party model hosts, while OpenAI’s own current agent architecture supports standardised MCP connections to external tools. The economic consequence is that the API may increasingly resemble a socket through which different intelligence providers can be substituted rather than a durable proprietary platform.
This does not eliminate model-specific engineering. Different systems continue to vary in tool-call behaviour, structured output, context management, reasoning modes, safety restrictions and tokenisation, so switching is not frictionless. It does, however, lower the cost sufficiently that cloud providers can increasingly offer model marketplaces where the customer relationship survives model substitution.
Model substitution versus platform substitution
| Change | Technical burden | Governance burden | Typical switching cost |
|---|---|---|---|
| Model A → Model B through same cloud/model gateway | Low–Medium | Low–Medium | Low–Medium |
| Open model A → open model B on same inference infrastructure | Medium | Low | Medium |
| Cloud A → Cloud B while retaining same model | High | High | High |
| Agent platform A → Agent platform B | High–Very high | High | Very high |
| Identity system A → Identity system B | Very high | Very high | Very high |
| Accelerator/compiler stack A → stack B | Very high | Medium–High | Very high |
| Enterprise workflow platform replacement | Very high | Very high | Very high |
The table illustrates why the battle for the model endpoint may ultimately be less strategically important than the battle for identity, execution and organisational workflow.
Agent harnesses are emerging as the new middleware
AWS now describes an agent as something substantially larger than a model: AgentCore’s harness manages the orchestration loop, tool execution, context, persistent state, recovery and isolated execution environment. Critically, AWS states that the harness is model agnostic and can switch models during a session. This transforms Bedrock from a model marketplace into a potential operating layer that can capture workloads independently of the underlying model supplier.
OpenAI’s September 2026 Agents API follows the same architectural logic from a different direction. OpenAI manages sessions, orchestration, context compaction and recovery, while developers can choose hosted or external execution environments and connect MCP servers, custom functions and built-in tools. The durable resource is therefore no longer only the model invocation but the persistent agent session and its environment.
Microsoft goes further by integrating agent infrastructure with the enterprise identity system itself. Foundry agents are represented through Microsoft Entra identities; administrators can govern their permissions, access reviews and lifecycle, while Agent 365 maintains an organisational registry of agents across environments and links them to Microsoft Defender and Purview controls. Once an agent is assigned identity, authorised against corporate applications and incorporated into compliance procedures, model substitution becomes relatively easier than platform substitution.
Google’s Gemini Enterprise Agent Platform similarly combines frontier models, development tools, enterprise data grounding, security, deployment and orchestration into what Google describes as an end-to-end agentic system. The common direction across the major Western providers is unmistakable: the competitive unit is becoming the managed agent system rather than the model endpoint.
Enterprise identity may become the strongest lock-in mechanism of all
Identity is particularly powerful because an enterprise agent that can merely generate text has limited operational value; an agent that can open files, issue refunds, update CRM records, execute code, query financial systems or modify infrastructure must possess authenticated rights. Those permissions need owners, policies, revocation procedures, audit trails, least-privilege rules and incident response.
Microsoft’s architecture provides a particularly explicit example: Foundry provisions agent identities through Entra ID and uses those identities for both governance and downstream tool authentication, while Agent 365 permits administrators to inventory agents, enforce access controls and apply security policies. This moves AI into the same institutional machinery that already governs employees, applications and service accounts.
AWS AgentCore similarly combines runtime, identity and access control, policy management, session persistence, tool connectivity, evaluation and observability at infrastructure level. These characteristics produce a form of path dependence that benchmark comparisons cannot measure: once thousands of corporate agents have identities and permissions inside one governance system, migration requires much more than changing a model identifier.
Why identity generates deeper lock-in than model quality
| Enterprise asset | Model replacement required? | Platform migration required? |
|---|---|---|
| Prompts | Usually | No |
| Evaluation datasets | Sometimes | No |
| User identities | No | Yes |
| Agent identities | No | Yes |
| Role-based access control | No | Yes |
| Audit policies | No | Yes |
| Security monitoring | No | Yes |
| Data-loss prevention | No | Yes |
| Compliance workflows | No | Yes |
| Tool credentials | No | Yes / revalidation required |
The implication is that identity architecture can outlive many model generations. A model may be replaced every few months while the same identity, security and governance system remains embedded for years.
Enterprise data create an asymmetric form of network value
Proprietary enterprise data do not create conventional public network effects, but they create private increasing returns. As an organisation connects more repositories, databases, applications, transaction histories and operational records to its AI environment, the agent becomes more useful inside that organisation. The value is cumulative and organisation-specific.
This is why Anthropic’s original MCP proposition was strategically important. MCP was designed as an open standard connecting AI systems to external data and tools, replacing separate bespoke connectors with a common protocol. Anthropic later added remote MCP integrations and donated the protocol into Linux Foundation governance through the Agentic AI Foundation. OpenAI now supports MCP directly in the Agents API, demonstrating that a protocol originating with one model vendor can become infrastructure used by its competitors.
This is one of the strongest pieces of evidence against a simple national-stack lock-in theory. A universal connector standard allows Chinese and Western models to access the same enterprise systems without forcing the enterprise to rebuild each integration. The more MCP-like standards succeed, the more value migrates from proprietary integration toward portable data connectivity.
Open standards can deliberately destroy proprietary network effects
Google’s Agent2Agent protocol provides a parallel example at the agent-to-agent layer. Google launched A2A specifically to permit agents built by different vendors and frameworks to communicate; in June 2025 the protocol was transferred into Linux Foundation governance alongside AWS, Cisco, Microsoft, Salesforce, SAP and ServiceNow. The stated aim was to create an open interoperable ecosystem rather than a proprietary Google-controlled communication standard.
The International Telecommunication Union’s broader description of standards captures the economic mechanism: widely recognised standards can lower trade barriers, promote compatibility, support economies of scale and generate network effects because systems can interact across organisations and markets. In AI, an open standard can therefore create network effects around the protocol itself while reducing lock-in to any individual supplier.
Proprietary lock-in versus open-standard network effects
| Architecture | Value increases as ecosystem grows | Switching cost | Strategic beneficiary |
|---|---|---|---|
| Proprietary model API | Yes | Medium | Model vendor |
| Proprietary agent runtime | Yes | High | Platform vendor |
| Open model weights | Yes | Low–Medium | Broad ecosystem |
| Open MCP connector ecosystem | Yes | Low | Protocol ecosystem / users |
| Open A2A agent ecosystem | Yes | Low | Multi-vendor ecosystem |
| Proprietary identity platform | Yes | Very high | Cloud / enterprise platform |
| Standardised inference API | Yes | Low | Model-neutral infrastructure |
| Proprietary compiler/runtime | Yes | High | Accelerator vendor |
The strategic point is subtle but important: network effects do not always produce monopoly-like proprietary lock-in. Open protocols can create powerful positive network effects precisely by making individual vendors replaceable.
Retrieval architectures are sticky because they contain institutional memory
Retrieval-augmented generation systems create another form of persistence. Once enterprise documents have been classified, chunked, embedded, permissioned, indexed and tied to citations or access rules, the resulting retrieval system contains substantial implementation capital. Changing the foundation model does not necessarily require rebuilding that corpus; changing the retrieval platform may.
This encourages cloud providers to treat knowledge retrieval as part of the platform control layer. Microsoft’s Foundry Agent Service connects agents to SharePoint, Fabric, Azure Blob Storage and other sources while preserving data-location and governance requirements; OpenAI’s agent infrastructure supports persistent files, vaults and MCP sources; Anthropic’s enterprise integrations connect Claude to internal repositories and external information services.
The economic rent may therefore migrate toward whoever owns the knowledge plane rather than the language model.
Compiler stacks create the deepest technical path dependence
At the infrastructure level, NVIDIA remains an unusually strong example of platform economics because CUDA, libraries, drivers, orchestration, inference engines, NIM microservices and enterprise support form a layered system around the GPU. NVIDIA AI Enterprise explicitly spans an application layer containing NIM, frameworks and SDKs and an infrastructure layer containing drivers, Kubernetes operators and cluster management; NIM can package models from different developers behind standard APIs while optimising them for NVIDIA hardware.
This architecture allows NVIDIA to benefit even when the model vendor changes. NVIDIA currently distributes or optimises open models including systems originating from Meta, OpenAI and other laboratories; its economic position therefore derives partly from being underneath the model contest.
The same logic explains why Huawei’s CANN ecosystem is strategically important for China. Domestic accelerator substitution is not complete when a chip can execute matrix multiplication; a competitive stack requires compilers, kernels, frameworks, profilers, debugging tools, operators, libraries and developer expertise. The deeper those assets accumulate around one architecture, the more expensive migration becomes.
Open Chinese models can strengthen Western infrastructure
The paradox becomes particularly visible when Chinese weights are deployed through NVIDIA-optimised inference, Western clouds or independent providers. Alibaba itself distributes rival Chinese models; NVIDIA’s infrastructure can optimise open models from many origins; and OpenAI-compatible APIs reduce switching friction. A Chinese model can therefore gain millions of users while strengthening the commercial position of a U.S. accelerator or cloud company.
This produces four distinct outcomes that must never be conflated:
| Outcome | Chinese model used? | Chinese cloud used? | Chinese hardware used? | Chinese ecosystem dependence |
|---|---|---|---|---|
| Qwen on Alibaba Cloud | Yes | Yes | Possibly | Medium–High |
| DeepSeek on Tencent Cloud | Yes | Yes | Variable | Medium |
| Qwen on foreign cloud / NVIDIA | Yes | No | No | Low at infrastructure layer |
| GLM self-hosted on sovereign infrastructure | Yes | No | Variable | Low operational dependence |
The first outcome produces genuine stack penetration; the latter two produce model diffusion without equivalent infrastructure control.
The control layer is therefore the decisive strategic asset
The emerging AI control layer can be defined as the combination of identity + model routing + data access + tools + agent runtime + state + security + observability + evaluation + billing + infrastructure allocation. Whoever controls that layer can substitute models, monitor workload demand, enforce policies, optimise cost, capture telemetry, cross-sell infrastructure and make customer migration increasingly expensive.
The contest is therefore unlikely to be decided simply by which country produces the most widely downloaded weights. The more model interchangeability increases, the more valuable this intermediary layer becomes.
Control-layer power matrix
| Control asset | Chinese examples | Western examples | Strategic durability |
|---|---|---|---|
| Multi-model gateway | Alibaba Model Studio, Tencent TokenHub | AWS Bedrock, Microsoft Foundry | High |
| Agent runtime | Alibaba Agent platforms, Tencent WorkBuddy | AgentCore, OpenAI Agents API, Google Agent Platform | High |
| Enterprise identity | Cloud IAM / enterprise suites | Entra, AWS IAM, Google IAM | Very high |
| Open tool standard | Growing MCP compatibility | MCP ecosystem | High ecosystem value / low vendor lock-in |
| Agent interoperability | Emerging | A2A | Potentially high |
| Accelerator runtime | CANN and domestic stacks | CUDA / NIM | Very high |
| Enterprise productivity surface | DingTalk / WeCom / cloud ecosystems | Microsoft 365, Google Workspace, ChatGPT Work | Very high |
| Industrial software | Huawei and sector partners | Siemens, Dassault, SAP, Palantir and others | Very high |
Key judgments
| Judgment | Confidence |
|---|---|
| Model weights are not the principal durable source of AI lock-in | High |
| Enterprise identity, agent orchestration and workflow integration create deeper switching costs | High |
| Open-weight Chinese models can spread globally without creating equivalent Chinese infrastructure dependence | High |
| Model commoditisation can strengthen hyperscalers and accelerator platforms | High |
| Open protocols such as MCP and A2A materially reduce proprietary integration lock-in | High |
| CUDA-style software ecosystems remain among the deepest technical network effects in AI | High |
| Chinese open-model diffusion is therefore strategically irrelevant | Rejected |
| Chinese model diffusion automatically creates Chinese-stack dependence | Rejected |
| The most consequential future competition is likely to occur at the control layer | Moderate–High |
What would change the assessment
The assessment would change materially if model behaviour remained sufficiently differentiated that enterprises could not substitute models without extensive application redesign, because this would restore the foundation-model provider as the dominant control point. It would also change if MCP, A2A or comparable standards failed to achieve broad implementation, leaving tool and agent integration fragmented into proprietary ecosystems.
Conversely, the control-layer thesis would strengthen if enterprises increasingly switched foundation models while retaining the same cloud, identity and agent infrastructure, if model routing became commonplace, and if procurement contracts increasingly separated “intelligence provider” from “agent/control platform.”
Open official record
The most valuable missing evidence consists of enterprise-level model-switching frequency, cost of model substitution versus platform migration, distribution of workloads across multi-model gateways, share of agent-tool connections using open protocols, migration costs between identity platforms, and actual enterprise dependence on accelerator-specific software libraries.
Chapter 5 — The Western Counter-Architecture
Principal judgment
The proposition that the United States and its allies continue to compete mainly through a handful of frontier models is no longer empirically defensible in 2026. OpenAI, Microsoft, Anthropic, Amazon, Google, Meta and NVIDIA collectively span nearly every layer relevant to an AI industrial system: hyperscale compute, accelerators, custom silicon, model training, open and closed model distribution, cloud regions, model marketplaces, inference runtimes, agent harnesses, enterprise identity, productivity software, data integration, consulting partnerships, governance, developer ecosystems and end-user distribution. The Western system is not vertically integrated into one national champion, but it is increasingly interconnected through overlapping corporate stacks that compete at some layers and complement one another at others.
This distinction matters because the Western architecture may be more modular than the Chinese policy model but is not less systemic. Anthropic models run on AWS, Google Cloud and Microsoft infrastructure; OpenAI models are now accessible through AWS Bedrock; Meta’s open models run across rival clouds and NVIDIA infrastructure; NVIDIA provides hardware and inference software to almost all of them; and Microsoft combines OpenAI and third-party models with Azure identity, governance and Microsoft 365 distribution. Competitive advantage can therefore emerge from federated complementarities rather than ownership by a single integrated corporate group.
The Western system is best understood as overlapping stacks
A comparison with China becomes misleading if “China” is represented as an integrated ecosystem while Western firms are represented only as isolated laboratories. The appropriate unit of comparison is a collection of partially overlapping platforms.
Western AI architecture, September 2026
| Actor | Models | Compute / infrastructure | Agent layer | Enterprise control | Distribution advantage |
|---|---|---|---|---|---|
| OpenAI | GPT family | Stargate + Oracle/NVIDIA + partner infrastructure | Agents API, Codex, ChatGPT Work, Frontier | Enterprise policies, tools, connectors, Presence | Large consumer + developer + enterprise base |
| Microsoft | Own + OpenAI + third-party catalogue | Azure | Foundry Agents | Entra, Agent 365, Defender, Purview | Microsoft 365, Azure, GitHub |
| Anthropic | Claude | AWS Trainium + Google/Microsoft access | Claude Code, Managed Agents, SDKs | MCP, enterprise controls, partner network | Multi-cloud + consulting/integration network |
| AWS | Multi-model Bedrock incl. Anthropic/OpenAI | AWS + Trainium/Inferentia | AgentCore | IAM, PrivateLink, CloudTrail, policies | Largest cloud installed base |
| Gemini + open/partner models | Google Cloud + TPUs | Gemini Enterprise Agent Platform | Cloud IAM, enterprise grounding and governance | Search, Workspace, Android, Google Cloud | |
| Meta | Llama | Partner infrastructure | Llama Stack / ecosystem | Mostly decentralised | Open-weight distribution + social platforms |
| NVIDIA | Nemotron + third-party models | GPUs, networking, systems | NIM, NeMo, blueprints, agent tooling | AI Enterprise | Hardware/software developer ecosystem |
This architecture means that the Western side can simultaneously compete on frontier intelligence, openness, cloud neutrality and vertical integration, because different firms specialise in different combinations.
OpenAI has moved from model provider toward an integrated intelligence platform
OpenAI’s 2026 strategy provides perhaps the clearest evidence against the “frontier laboratory only” description. In April 2026 the company stated that enterprise customers accounted for more than 40% of revenue, its APIs were processing more than 15 billion tokens per minute, and Codex had reached millions of weekly users. The same announcement described the company’s intended architecture as spanning infrastructure, models, agents and unified employee interfaces. These figures are OpenAI’s own disclosures rather than independently audited operational statistics, but they establish the company’s commercial direction.
OpenAI’s infrastructure layer has also become increasingly explicit. Stargate began with a commitment to secure 10 GW of U.S. AI infrastructure by 2029; OpenAI stated in April 2026 that it had already surpassed that milestone. Its flagship Abilene site uses Oracle Cloud Infrastructure and NVIDIA GB200 systems, while a separate partnership with SB Energy involves a 1.2 GW data-centre site and substantial associated energy investment. Whatever the future utilisation of those sites, the magnitude establishes that OpenAI now treats physical compute and electricity as strategic assets rather than external cloud inputs.
The Agents API adds another layer. OpenAI now provides persistent agent sessions, subagents, tool connections, MCP, file and code environments, context management and recovery, while permitting execution through OpenAI-hosted or partner infrastructure. The company is therefore moving from “model inference as API” toward managed cognitive execution infrastructure.
OpenAI Presence extends further into enterprise implementation by combining models with company knowledge, policies, guardrails, escalation rules and forward-deployed engineering support. This is especially significant because it enters a category traditionally occupied by consulting firms and enterprise integrators: workflow redesign and long-term operating deployment.
OpenAI’s expanding stack
| Layer | Evidence visible by Sep 2026 |
|---|---|
| Physical infrastructure | Stargate, multi-gigawatt compute projects |
| Frontier models | GPT series |
| Developer API | Model APIs, Responses, Agents API |
| Agent runtime | Durable agent sessions and hosted harness |
| Coding | Codex |
| Enterprise work surface | ChatGPT Work |
| Workflow deployment | Frontier / Presence |
| Data/tool connectivity | MCP and plugin/connectors |
| Integration services | FDEs + global systems-integrator alliances |
| Consumer distribution | ChatGPT |
This architecture is structurally much closer to a full-stack ecosystem than to an isolated frontier-model laboratory.
Microsoft possesses perhaps the strongest enterprise control-plane advantage
Microsoft’s position differs from OpenAI’s because its deepest advantage resides in the installed enterprise stack surrounding intelligence: Azure, Entra, Microsoft 365, GitHub, security, compliance and existing corporate identity.
Foundry now provides models, native and hosted agents, evaluation, observability, red teaming, governance and control-plane functions. Microsoft’s 2026 documentation explicitly describes the Foundry control plane as the mechanism for governance across models, tools and agents rather than a model development interface alone.
Agent 365 then pushes this architecture into enterprise administration. Agents can be registered, assigned identities, governed through access policies and monitored through Microsoft’s security systems; Foundry agents integrate automatically with that registry. This can create a formidable installed-base advantage because enterprises already using Microsoft identity and productivity systems do not need to construct a parallel governance structure merely to introduce AI agents.
Microsoft’s position therefore illustrates a form of adjacent-market network effect: dominance or strength in identity, office productivity, enterprise security and cloud computing can lower the cost of adopting Microsoft-mediated AI even if the underlying model was developed by another company.
AWS is attempting to become the neutral operating environment for agents
AWS’s competitive strategy is increasingly model-neutral. Bedrock already aggregates models from multiple developers, and in April 2026 AWS announced that OpenAI models, Codex and OpenAI-powered Managed Agents would become available through Bedrock alongside existing providers. OpenAI workloads inherit AWS IAM, PrivateLink, guardrails, encryption and CloudTrail controls, demonstrating that model intelligence can be inserted into AWS governance rather than requiring customers to leave it.
AgentCore reinforces this neutrality. AWS states that its managed harness can switch models within a session and manages reasoning loops, tools, memory, state and recovery; the service includes runtime, identity, policy, tool connectivity, evaluations and observability. This is precisely the control-layer strategy identified in Chapter 4: make the model substitutable while making the execution environment persistent.
AWS also possesses a distinctive vertical-integration route through its Trainium accelerators. Anthropic reports that it currently uses more than one million Trainium2 chips, that over 100,000 customers run Claude on Amazon Bedrock, and that the April 2026 collaboration with Amazon provides up to 5 GW of additional compute while Anthropic commits more than $100 billion over ten years to AWS technologies. These are first-party disclosures, but they show the depth at which model, accelerator and cloud economics are becoming entangled.
Anthropic is building an ecosystem through standards and integrators rather than cloud ownership
Anthropic represents a different strategic architecture because it does not own a hyperscale cloud. Instead, it is attempting to become a model and agent layer that travels across infrastructure providers while shaping integration standards.
Claude is available through AWS, Google Cloud and Microsoft, while Anthropic states that more than 100,000 customers run Claude through Bedrock alone. The company has simultaneously invested in MCP, Claude Code, partner networks, consulting alliances and enterprise agent deployments.
MCP is the most strategically unusual element. Anthropic created the protocol in 2024, made it open source and later donated it to the Linux Foundation’s Agentic AI Foundation. By January 2026 Anthropic reported 100 million monthly MCP downloads, while OpenAI and other competitors had adopted the protocol. If MCP becomes a durable industry standard, Anthropic will have influenced the architecture of enterprise AI without retaining proprietary control over it.
Anthropic is also building a labour-intensive deployment ecosystem. It committed $100 million to the Claude Partner Network in March 2026; by June it reported more than 40,000 firms applying and over 10,000 consultants earning Claude certification. Separate partnerships with DXC, Cognizant, PwC and others place Claude within existing systems-integration channels. These are corporate figures, but they show that integration expertise itself is being treated as a strategic complement to model capability.
This is important for the comparison with China: China may possess advantages from dense domestic industrial relationships, but Western model developers can partially reproduce integration scale by mobilising global consulting and cloud ecosystems rather than owning every implementation channel directly.
Google combines model, TPU, cloud, agent protocol and workplace distribution
Google’s architecture has perhaps the widest theoretically integrated surface: it controls Gemini models, TPUs, Google Cloud, Workspace, Android, search infrastructure, data services and a rapidly evolving enterprise agent platform.
In April 2026 Google repositioned its enterprise architecture around Gemini Enterprise Agent Platform, describing it as a comprehensive system for building, scaling, governing and optimising enterprise agents and as an evolution of Vertex AI. The platform combines frontier models, enterprise data, development tooling, deployment and governance rather than presenting Gemini solely as a model API.
Google simultaneously pursued an open interoperability strategy through A2A. By transferring A2A into a Linux Foundation project alongside AWS, Microsoft, Salesforce, SAP and other firms, Google reduced the likelihood that agent interoperability itself would be controlled by a single hyperscaler.
This apparently contradictory combination—vertically integrated proprietary platform plus support for open interoperability—is likely to become characteristic of the AI market. Providers want customers deeply embedded in their control planes while supporting enough interoperability to reduce the fear of initial adoption.
Meta demonstrates that open models can be an ecosystem strategy even without cloud capture
Meta represents the strongest Western analogy to China’s open-weight model diffusion strategy. Llama’s broad availability permits developers to deploy models through AWS, Azure, Google Cloud, NVIDIA, Groq, Databricks, local infrastructure and numerous other environments. Meta reported more than one billion downloads by March 2025, demonstrating extraordinary model-layer diffusion.
Meta has consistently described Llama as part of a system rather than a standalone model, including Llama Stack interfaces, safety tools and an ecosystem of deployment partners. Llama 4 continued that approach, with partners spanning cloud platforms, accelerator manufacturers, consulting firms, inference providers and enterprise software companies.
The Meta example provides a critical empirical warning for the China thesis: extreme open-model adoption does not necessarily create infrastructure capture by the model originator. Meta can influence standards, model design and developer practice while AWS, Microsoft, NVIDIA and other companies capture significant economic value from the surrounding deployment.
NVIDIA may control the layer with the deepest technological path dependence
NVIDIA’s strategic position differs from every model developer because its platform captures value across competing models. NVIDIA AI Enterprise combines NIM microservices, NeMo, frameworks and SDKs at the application layer with drivers, orchestration, Kubernetes operators and cluster-management software at the infrastructure layer.
NIM is particularly important because it packages open and proprietary models into optimised enterprise containers exposing standard APIs while allowing deployment across cloud, data-centre and edge environments. The model can change while the NVIDIA inference and hardware architecture remains constant.
NVIDIA’s June 2026 Agent Toolkit expands further upward into agent runtimes, blueprints, security and domain-specific CUDA-X skills, with partnerships involving engineering and enterprise-software firms. NVIDIA therefore increasingly occupies not just the chip layer but the translation layer between models and industrial computation.
This gives NVIDIA a structural advantage analogous to a platform supplier whose products are complements to nearly every competing application vendor.
Western vertical integration is real, but fragmented across corporate boundaries
The Western system’s distinguishing characteristic is not lack of integration but integration through contracts, standards and overlapping ecosystems rather than single-firm ownership.
Integrated-stack comparison
| Functional layer | China | U.S./allied ecosystem |
|---|---|---|
| Frontier / advanced models | Qwen, DeepSeek, Hy, ERNIE, GLM, Kimi, MiniMax, Pangu | OpenAI, Anthropic, Gemini, Llama, Nemotron |
| Model marketplaces | Alibaba, Tencent, Baidu | AWS Bedrock, Microsoft Foundry, Google |
| Agent runtimes | Alibaba/Tencent/Baidu/Huawei ecosystems | OpenAI Agents, AgentCore, Foundry, Gemini Agent Platform |
| Identity/governance | Domestic cloud/enterprise systems | Entra, AWS IAM, Google IAM |
| Open interoperability | Increasing compatibility | MCP, A2A, open API ecosystems |
| Accelerators | Ascend and other domestic vendors | NVIDIA, Google TPU, AWS Trainium, AMD |
| Accelerator software | CANN and domestic runtimes | CUDA, NIM, XLA, AWS stacks |
| Hyperscale cloud | Alibaba, Tencent, Huawei, Baidu | AWS, Azure, Google Cloud, Oracle |
| Productivity distribution | WeCom, DingTalk and domestic applications | Microsoft 365, Google Workspace, ChatGPT, Meta |
| Systems integration | Domestic cloud/telecom/industrial firms | Accenture, PwC, Cognizant, DXC, Capgemini, BCG, McKinsey and others |
| Consumer distribution | Large domestic super-app ecosystems | ChatGPT, Meta platforms, Google, Microsoft |
| Global cloud regions | Expanding but less extensive | Very extensive |
| Leading-edge semiconductor production | Material constraints | Stronger allied supply chain, still geographically concentrated |
The conclusion is not that the two systems are equivalent. China’s state policy can coordinate infrastructure, industrial policy and domestic adoption in ways that differ from the more corporate and contractual Western architecture, while the allied semiconductor supply chain has different upstream strengths and dependencies. But the evidence no longer permits an analytical contrast between a Chinese “ecosystem” and a Western “model race.”
Western openness is itself strategically heterogeneous
Meta’s open weights, Anthropic’s MCP, Google’s A2A, NVIDIA’s support for open models and hyperscaler model marketplaces create significant openness, but this does not mean the Western architecture lacks lock-in. In many cases openness at one layer supports proprietary capture at another.
AWS benefits from model choice because it makes Bedrock more attractive. NVIDIA benefits from open models because more models can run on CUDA/NIM. Microsoft can offer competing model families while retaining identity, governance and productivity-suite relationships. Google can support A2A while operating the agent infrastructure beneath it.
The architecture therefore resembles open complements surrounding proprietary control points.
Openness as a platform strategy
| Company | Open / interoperable element | Proprietary control point |
|---|---|---|
| Meta | Llama weights / ecosystem | Consumer distribution and Meta platforms |
| Anthropic | MCP | Claude models / products |
| A2A | Gemini + Google Cloud / Workspace | |
| NVIDIA | Open-model support and standard APIs | CUDA / GPU infrastructure |
| AWS | Multi-model Bedrock | AWS infrastructure / IAM |
| Microsoft | Multi-model Foundry | Azure / Entra / Microsoft 365 |
| OpenAI | MCP support and multiple execution environments | OpenAI models / agent harness / ChatGPT |
This strategic pattern is central to the standards war: vendors can promote interoperability where interoperability expands the market while maintaining proprietary advantage at layers with greater switching costs.
The Western system has a powerful global-region advantage
Cloud geography matters because enterprise and government workloads are constrained by data residency, latency, sovereignty and regulatory requirements. OpenAI reported in late 2025 that eligible enterprise customers could select data residency across Europe, the United Kingdom, United States, Canada, Japan, South Korea, Singapore, India, Australia and the UAE. Microsoft, AWS and Google operate much broader cloud-region footprints.
This geographic infrastructure is strategically relevant because globally distributed enterprises often prefer a provider capable of supplying one governance framework across many jurisdictions. Chinese cloud providers are expanding internationally, but the Western hyperscalers retain a substantial installed-base and geographic advantage outside China.
The Western architecture also possesses a powerful distribution flywheel
OpenAI stated in September 2026 that its products reached more than one billion weekly active users and 2.5 million businesses, while describing consumer familiarity as a distribution channel into enterprise use. These figures are company disclosures, but the mechanism is strategically significant: consumer adoption can reduce training and behavioural friction when the same interface enters the workplace.
Meta possesses an even broader consumer communications footprint through WhatsApp, Instagram, Facebook and Messenger; Google controls search, Android and Workspace; Microsoft controls large portions of office productivity and enterprise identity. China’s equivalent advantage exists through domestic super-apps and large consumer platforms, but it transfers internationally less automatically.
The contest is therefore partly a battle over existing user interfaces, not merely new AI applications.
Capital and compute scale remain core Western advantages
The infrastructure scale underpinning the Western ecosystem is extraordinary. OpenAI reported more than 10 GW of Stargate capacity secured by April 2026; Anthropic announced up to 5 GW of additional AWS capacity and more than $100 billion of ten-year AWS commitments; NVIDIA continues to expand an international network of AI-cloud providers; Google and AWS operate custom accelerator programmes in addition to NVIDIA deployments.
This scale means Western providers can compete simultaneously through model capability and declining unit inference cost. The proposition that China alone is pursuing an “industrial” AI strategy therefore understates the degree to which Western capital markets and hyperscalers are building physical intelligence infrastructure.
Selected disclosed infrastructure signals
| Indicator | Disclosed value | Source character |
|---|---|---|
| OpenAI Stargate capacity secured | >10 GW by Apr 2026 | Company disclosure |
| OpenAI Michigan campus | 1 GW | Company/project disclosure |
| OpenAI–SB Energy site | 1.2 GW | Company/project disclosure |
| Anthropic–AWS new capacity agreement | Up to 5 GW | Company disclosure |
| Anthropic AWS technology commitment | >$100bn / 10 years | Company disclosure |
| Trainium2 chips reportedly used by Anthropic | >1 million | Company disclosure |
| Claude customers reported on Bedrock | >100,000 | Company disclosure |
| OpenAI business customers reported Sep 2026 | 2.5 million | Company disclosure |
Sources: OpenAI and Anthropic corporate announcements.
These numbers are not directly comparable measures of operational compute and should not be aggregated, but they establish the capital intensity of the Western response.
The stronger comparison is therefore architecture against architecture
The Chinese and Western systems differ principally in how integration is organised, not in whether integration exists.
China’s architecture exhibits relatively strong coordination between industrial policy, telecommunications, domestic cloud providers, public deployment and localisation objectives. The Western architecture is more decentralised, but contractual complementarities allow specialised firms to assemble powerful end-to-end systems.
Structural comparison
| Variable | Chinese architecture | Western architecture |
|---|---|---|
| Coordination mechanism | State strategy + corporate competition | Corporate markets + contracts + standards |
| Model market | Highly competitive | Highly competitive |
| Open weights | Very important | Important through Meta and others |
| Hyperscaler strength | Strong domestically | Exceptionally strong globally |
| Enterprise identity | Fragmented among domestic ecosystems | Strong incumbent platforms |
| Accelerator independence | Increasing but constrained | Stronger, especially through NVIDIA/TPU/Trainium |
| Global systems integrators | Developing international reach | Extensive established network |
| Public-sector mobilisation | Strong central policy direction | Jurisdictionally fragmented |
| Developer standards | Growing compatibility | Strong open-protocol activity |
| Global cloud footprint | Expanding | Extensive |
| Consumer distribution | Exceptional domestic scale | Exceptional global scale |
| Industrial embedding | Strong domestic opportunity | Strong via enterprise software/integrators |
Key judgments
| Judgment | Confidence |
|---|---|
| The Western AI system cannot accurately be described as frontier laboratories competing only on model quality | High |
| Western firms increasingly compete through full-stack or federated-stack architectures | High |
| Microsoft and AWS possess especially strong enterprise control-plane advantages | High |
| NVIDIA captures value across rival model ecosystems | High |
| Anthropic’s open-protocol strategy can create influence without cloud ownership | High |
| Meta demonstrates that open-model diffusion need not create infrastructure dependence | High |
| Western integration is weaker because it crosses corporate boundaries | Not established |
| China uniquely understands AI as infrastructure | Rejected by current evidence |
| The global contest is increasingly ecosystem versus ecosystem | High |
What would change the assessment
The Western counter-architecture would appear weaker if multi-vendor integration produced persistent coordination failures, if enterprise customers experienced severe interoperability costs across clouds and models, if electricity or accelerator shortages prevented announced infrastructure from being utilised, or if regulatory fragmentation made international deployments substantially more expensive.
It would appear stronger if model interchangeability continued increasing while Microsoft, AWS, Google and NVIDIA retained control of the surrounding infrastructure, because this would allow Western platforms to absorb Chinese open models without surrendering customer relationships.
Open official record
The highest-value missing evidence concerns actual cross-cloud migration costs, enterprise shares by agent control plane, realised rather than announced compute utilisation, cost per token across proprietary accelerator systems, usage distribution between model providers inside hyperscaler marketplaces, and international adoption of Chinese versus Western control-layer technologies.
Chapter 6 — Stress-Testing the Oil, Telecom, Solar and Battery Analogies
Principal judgment
None of the four historical analogies is sufficient by itself to describe artificial intelligence. Oil is useful for infrastructure dependence, chokepoints and geopolitical supply exposure but weak on product economics; telecommunications is strongest for interoperability, standards and installed-base effects; solar photovoltaics is strongest for manufacturing learning curves, scale, cost compression and overcapacity; batteries are strongest for multi-stage supply chains, vertical integration and midstream chokepoints. The most defensible historical model for AI is therefore a composite rather than a single analogy.
The slogan that “AI will be chosen the way oil is chosen” is analytically provocative but literally inaccurate. Oil itself is not homogeneous: crude streams differ in density and sulphur content and require different refinery configurations, yet they remain physical feedstocks whose economic value depends heavily on extraction, transport, refinery compatibility and geographic chokepoints. AI models, by contrast, are reproducible software artefacts whose weights can sometimes be copied almost costlessly, whose marginal inference cost is driven by compute rather than physical depletion, and whose performance can change rapidly without replacing the distribution network.
The analogy becomes stronger only when “oil” refers not to the model but to compute-enabled intelligence as a strategic input to economic activity.
Oil: useful for chokepoints, weak for fungibility
Oil markets demonstrate how an economy can be globally supplied by a commodity while remaining exposed to physical chokepoints. The U.S. Energy Information Administration estimated that approximately 20 million barrels per day, equivalent to around 20% of global petroleum-liquids consumption, moved through the Strait of Hormuz in 2024. Subsequent disruptions in 2026 contributed to substantial volatility in Brent prices.
The AI equivalent is not a shipping strait but a collection of bottlenecks: advanced lithography, HBM, advanced packaging, accelerator production, data-centre power, networking and perhaps particular software ecosystems. If a critical input is difficult to substitute, control over that layer can matter more than the downstream brand.
Oil mechanisms that transfer to AI
| Oil mechanism | AI analogue | Transfer quality |
|---|---|---|
| Strategic upstream supply | Advanced semiconductors / accelerators | Strong |
| Transport chokepoints | HBM, lithography, packaging, data-centre power | Strong conceptually |
| Refinery compatibility | Model/hardware/software compatibility | Moderate |
| Spare capacity | Available compute capacity | Strong |
| Long-lived infrastructure | Data centres / grids | Strong |
| Commodity fungibility | Interchangeable models | Weak–Moderate |
| Physical depletion | Token inference | Very weak |
| Marginal extraction cost | Marginal inference cost | Partial |
| Storage inventories | Compute capacity | Weak |
The first major problem with the analogy is fungibility. The EIA emphasises that even crude oil differs by API gravity and sulphur content and that refineries require different processing configurations, yet these variations remain far narrower than the functional differences among AI models.
The second problem is reproduction. One barrel cannot be copied; open model weights can. Once weights have been released, the originator may no longer control where they run. This makes the model itself much less analogous to a strategic raw material.
The third problem is rapid substitutability. A refinery, pipeline or oil field can remain economically relevant for decades. An AI model generation may be superseded within months. The long-lived asset is therefore more likely to be the data centre, accelerator software, enterprise integration or agent control plane than the model.
The oil analogy becomes stronger at the compute layer
OpenAI’s infrastructure strategy reinforces this distinction. The company explicitly argues that more compute supports stronger models, lower delivery cost and broader intelligence access, while Stargate represents a long-term physical infrastructure project involving multiple gigawatts of electricity and data-centre capacity. At this layer, AI increasingly resembles an energy-intensive industrial input rather than traditional software.
Anthropic’s up-to-5-GW AWS agreement and more than $100 billion technology commitment display the same logic: access to future compute is being contracted like a strategic capacity asset rather than purchased opportunistically from a spot software market.
The economically useful formulation is therefore not “models are oil,” but “compute capacity increasingly exhibits some infrastructure and geopolitical characteristics historically associated with energy supply.”
Telecom is the strongest analogy for standards and interoperability
Telecommunications offers a more powerful analogy for the layer where AI systems communicate with tools, agents and external applications. The ITU notes that technical standards enable interoperability across devices, networks and services and can support economies of scale, network effects, competition and international trade.
AI is beginning to develop precisely these dynamics. MCP establishes a common method for connecting AI systems with tools and data; A2A establishes an architecture for agents from different vendors to communicate. Both have moved toward independent open governance rather than remaining proprietary extensions of their originating companies.
Telecom mechanisms that transfer to AI
| Telecom mechanism | AI equivalent | Transfer quality |
|---|---|---|
| Interoperability standards | MCP, A2A, standard APIs | Very strong |
| Installed base | Existing enterprise cloud / identity platform | Strong |
| Device/network complementarity | Model/tool/cloud complementarity | Strong |
| Standards organisations | Linux Foundation, ITU, industry consortia | Strong |
| Network effects | Connector / developer ecosystems | Strong |
| Spectrum scarcity | Compute scarcity | Weak analogy |
| Physical network ownership | Cloud/data-centre ownership | Moderate–Strong |
| Roaming/interconnection | Cross-platform agent communication | Strong conceptually |
The telecom analogy also clarifies why standards wars do not necessarily produce one monopolistic winner. A common standard can create a large interoperable market in which multiple equipment, network and service companies continue competing.
This is particularly important for the China thesis because global adoption of a Qwen or DeepSeek model does not necessarily imply adoption of a Chinese “AI standard” if those models communicate through protocols, APIs and agent frameworks controlled by open communities or foreign infrastructure providers.
A standard can win while its originator captures little economic rent
MCP illustrates this possibility unusually clearly. Anthropic invented the protocol, but its adoption by OpenAI and donation into Linux Foundation governance means Anthropic deliberately gave up exclusive control in exchange for a larger interoperability ecosystem.
Google did something similar with A2A. The protocol originated at Google but was transferred to the Linux Foundation with participation from rival technology companies.
This has a direct implication for geopolitical analysis: technical influence, economic rent and sovereign control can diverge. A country or company may invent an important standard while competitors make more money implementing it.
The same distinction must be applied to Chinese open models.
Solar photovoltaics provide the strongest cost-curve analogy
Solar PV offers perhaps the strongest historical analogy for the original thesis’s claim that a “good enough and cheap enough” technology can restructure global markets. According to the International Energy Agency, China invested more than $50 billion in new PV supply capacity after 2011 and built a manufacturing position exceeding 80% across every major solar-panel manufacturing stage by the early 2020s. The IEA concluded that Chinese industrial policies, large-scale production and continuous innovation contributed to an over-80% decline in PV manufacturing costs over the relevant decade.
The mechanism transfers meaningfully to AI: enormous domestic demand, capital investment, supplier density, engineering repetition and competitive pressure can drive the cost of a technology downward even when other countries retain strong research capability.
The 2026 IEA assessment shows that this structural advantage remains substantial: China accounts for around 85% of solar supply-chain production capacity and 95% of wafer capacity.
Solar mechanisms that transfer to AI
| Solar PV mechanism | AI analogue | Transfer quality |
|---|---|---|
| Scale economies | Data-centre / inference scale | Strong |
| Learning-by-doing | Repeated model/deployment engineering | Strong |
| Supplier clustering | AI compute / cloud / software ecosystem | Strong |
| Vertical integration | Chip-to-cloud-to-agent integration | Strong |
| Manufacturing overcapacity | Excess inference / model supply | Moderate |
| Falling unit costs | Falling inference cost | Strong |
| Export-led expansion | Model/cloud exports | Strong |
| Physical manufacturing | Software reproduction | Important limitation |
The greatest transferable lesson is not that China inevitably dominates AI because it dominated PV. It is that cost curves can become strategic. If thousands of deployments, large compute fleets and intense competition continuously reduce the cost of delivering adequate intelligence, market structure can shift even without undisputed technological leadership.
The PV analogy also exposes the danger of overcapacity
Solar demonstrates that scale advantages can generate destructive competition. The IEA reported that global PV manufacturing capacity expanded so rapidly that manufacturing supply substantially exceeded demand and module prices fell sharply; in 2023, spot module prices dropped by almost 50% year over year amid expanding capacity.
A comparable dynamic is plausible in AI. Multiple Chinese providers already price model inference aggressively, while large cloud and accelerator investments create pressure to maintain high utilisation. If model capability converges while capacity continues expanding, inference may become structurally commoditised.
That outcome would benefit downstream application developers but could compress margins for model laboratories.
AI overcapacity transmission mechanism
Large capital investment → abundant compute/model supply → aggressive token pricing → declining model-layer margins → increased application consumption → greater importance of platforms and downstream integration.
This is one reason the thesis that “the winner will be the cheapest model” is incomplete. In an overcapacity environment, model economics may deteriorate for everyone, moving profit elsewhere in the stack.
Batteries provide the strongest supply-chain analogy
Battery production is particularly useful because competitive advantage is distributed across many technically interdependent stages: raw materials, refining, cathodes, anodes, cells, packs, electronics and final vehicles. AI similarly relies on semiconductor equipment, fabrication, HBM, advanced packaging, accelerators, networking, data centres, inference runtimes, models and applications.
The IEA reports that China accounted for more than 80% of global battery-cell production in 2025, approximately 85% of cathode active-material production and more than 90% of anode active-material production for electric-car batteries.
This is precisely why focusing only on the final battery cell would miss the strategic structure. Similarly, focusing only on a foundation model misses the components upstream and downstream that determine scale and cost.
Battery mechanisms that transfer to AI
| Battery mechanism | AI analogue | Transfer quality |
|---|---|---|
| Multi-stage value chain | Semiconductor → compute → model → application | Very strong |
| Midstream bottleneck | HBM / packaging / fabrication equipment | Very strong |
| Scale manufacturing | Accelerator and data-centre scale | Strong |
| Vertical integration | Chip/cloud/model/application integration | Strong |
| Process learning | Inference optimisation / deployment learning | Strong |
| Materials dependence | Critical semiconductor inputs | Moderate–Strong |
| Logistics dependence | Data/network infrastructure | Weak–Moderate |
| Physical inventory | Model/software availability | Weak |
The battery analogy therefore strengthens the distinction developed in Chapter 2 between application-stack completeness and production-stack completeness. A country can dominate downstream integration while remaining dependent on a few upstream inputs, just as an electric-vehicle industry can be constrained by cathode materials or battery cells despite strength in final assembly.
Batteries also show how middle layers can capture disproportionate power
Energy Technology Perspectives 2026 demonstrates that clean-technology supply chains often contain one or more steps where non-Chinese capacity is insufficient to replace Chinese supply, even if final assembly exists elsewhere. The IEA’s “N-1” analysis finds multiple clean-energy stages where less than one-quarter of demand could be met outside the largest supplier.
The AI analogue is direct: a country can possess model developers and data centres yet remain strategically dependent if one critical layer—HBM, lithography, advanced packaging, accelerator software or power equipment—cannot scale independently.
This is why AI sovereignty cannot be measured by counting domestic models.
Solar and batteries also demonstrate the difference between company nationality and production geography
The IEA notes that Chinese solar firms increasingly own capacity outside China and that Chinese battery producers are expanding internationally. The relevant strategic unit is therefore not always where the factory is geographically located but who controls technology, capital, equipment, intellectual property and supply relationships.
The same issue will emerge in AI. A Chinese model hosted in Germany on U.S. accelerators is not purely “Chinese” infrastructure; an American model deployed in a sovereign Middle Eastern data centre on locally controlled infrastructure is not operationally identical to a U.S.-hosted API. National labels therefore become progressively less informative as the stack globalises.
Oil and batteries explain chokepoints; telecom explains standards; solar explains cost
The four analogies can therefore be ranked by mechanism, not by overall resemblance.
Analogy stress-test
| Analytical question | Oil | Telecom | Solar PV | Batteries |
|---|---|---|---|---|
| Commodity fungibility | Medium | Low | Medium | Low |
| Strategic infrastructure | Very high | High | High | High |
| Network effects | Low | Very high | Low | Low |
| Standards formation | Low | Very high | Medium | Medium |
| Economies of scale | High | High | Very high | Very high |
| Learning curves | Medium | Medium | Very high | Very high |
| Vertical integration | High | High | High | Very high |
| Supply chokepoints | Very high | Medium | High | Very high |
| Overcapacity dynamics | Medium | Low | Very high | High |
| Installed-base effects | High | Very high | Medium | High |
| Software portability | Very weak analogy | Medium | Weak analogy | Weak analogy |
| Best fit for AI | Compute/infrastructure | Control protocols | Cost curve | Full value chain |
No single analogy dominates because AI simultaneously exhibits software, network, manufacturing and infrastructure characteristics.
What does not transfer from solar or batteries
The greatest limitation of the clean-technology analogy is that physical manufacturing determines marginal supply in solar and batteries, whereas model weights can be copied and deployed widely. China could produce one competitive open model and have it replicated globally without constructing factories in every country.
This makes open AI fundamentally different from PV modules. A Chinese solar manufacturer usually captures revenue when its module is sold; a Chinese model developer may capture little or no revenue when an open model is downloaded and deployed on another company’s hardware.
The economic object that behaves more like a manufactured product is therefore inference capacity, not the weights.
What does not transfer from telecommunications
Telecom standards operate under extensive formal standardisation and hard interoperability requirements. AI standards remain much more fluid. MCP and A2A are young protocols; API conventions are not equivalent to legally or technically mandatory telecom standards; agent architectures can change quickly.
The telecom analogy therefore becomes stronger only if a small number of protocols stabilise and persist over multiple model generations.
What does not transfer from oil
Oil demand consumes physical material; AI inference consumes electricity and computing time but not model weights. Higher use therefore does not deplete the intellectual asset itself.
Oil markets also possess long-established global pricing benchmarks, while intelligence quality is intrinsically multidimensional and task-specific. There is no meaningful single global “price of intelligence” analogous to Brent.
A model with a nominally lower token price can be more expensive per completed task if it requires more tokens, retries or human supervision.
The most useful synthesis is a layered analogy
The AI system can be understood more accurately by assigning a different historical analogy to each layer.
Layered historical model of AI competition
| AI layer | Best historical analogy | Why |
|---|---|---|
| Electricity / data-centre capacity | Oil / energy infrastructure | Strategic supply and chokepoints |
| Semiconductor manufacturing | Battery supply chain | Multi-stage physical dependency |
| Accelerators + software runtime | Telecom equipment/platform | Installed base and compatibility |
| Foundation-model weights | Software / partially commodity-like | High replicability |
| Inference service | Cloud + industrial utility | Capacity and unit economics |
| APIs / protocols | Telecommunications standards | Interoperability |
| Agent control layer | Operating-system / cloud platform | Ecosystem and switching costs |
| Enterprise integration | Industrial automation | Process-specific sunk cost |
| Large-scale model commoditisation | Solar PV | Learning curves and price compression |
This layered approach resolves much of the confusion created by the oil metaphor.
The decisive historical lesson: complements determine technology power
Across all four analogies, one recurrent principle survives. Technologies rarely dominate because of one component alone. Oil requires transport, refining and distribution. Telecommunications requires equipment, spectrum, standards and devices. Solar requires polysilicon, wafers, cells, modules, inverters and installation. Batteries require minerals, materials, cells, packs, power electronics and vehicle integration.
Artificial intelligence likewise requires compute, memory, networking, models, data, tools, applications, identity and organisational integration.
The decisive strategic asset is therefore the ability to coordinate complements.
This reframes the original thesis
The thesis “AI will not be bought; it will be chosen the way oil is chosen” should therefore be reformulated.
The empirical record does not support treating foundation models as equivalent to crude oil. It does support a stronger proposition:
AI will increasingly be adopted as an infrastructure system whose economics depend on the interaction of intelligence, compute, standards, platform compatibility, capital intensity and organisational integration.
The historical mechanisms are consequently distributed:
Oil explains strategic capacity and bottlenecks.
Telecommunications explains standards, interoperability and installed bases.
Solar photovoltaics explain industrial learning, scale, overcapacity and cost compression.
Batteries explain vertical integration, supply-chain depth and intermediate chokepoints.
Together, these analogies provide a far stronger framework than any one of them alone.
Cross-analogy evidence table
| Historical evidence | Verified measure | AI implication |
|---|---|---|
| Chinese solar manufacturing share | ~85% of solar supply-chain capacity in 2026 assessment | Scale + supplier clustering can create durable cost advantage |
| Chinese PV wafer capacity | ~95% | Strategic power can concentrate in intermediate stages |
| Chinese battery-cell production | >80% in 2025 | Downstream adoption can depend on upstream industrial concentration |
| Chinese cathode production | ~85% | Midstream matters |
| Chinese anode production | >90% | Bottlenecks can exist outside final product |
| Solar cost decline associated with scale/industrial development | >80% over prior decade in IEA analysis | Learning curves can reshape global markets |
| Solar module price decline in 2023 | ~50% YoY | Overcapacity can rapidly compress margins |
| Hormuz petroleum flow in 2024 | ~20m b/d, ~20% of global petroleum-liquid consumption | Concentrated infrastructure can become geopolitical leverage |
| MCP monthly downloads reported Jan 2026 | 100m | Open standards can create network effects without proprietary ownership |
| OpenAI Stargate capacity reported secured | >10 GW | AI competition already requires utility-scale physical infrastructure |
| Anthropic–AWS capacity agreement | Up to 5 GW | Long-term compute is becoming a strategic contracted input |
Sources: International Energy Agency, U.S. Energy Information Administration, OpenAI and Anthropic.
Implications for the China thesis
The analogies collectively support five conclusions.
First, price matters most after capability becomes sufficiently standardised. Solar became globally transformative because basic product functionality was mature enough for manufacturing cost to dominate buyer decisions. AI reaches the same condition only task by task, not universally.
Second, industrial scale can become a technology advantage in itself. Repeated deployment generates learning, supplier expertise, utilisation efficiency and capital confidence.
Third, the most strategic bottleneck may not be the visible product. Batteries demonstrate the power of cathodes and anodes; AI may ultimately be constrained more by HBM, packaging, accelerator software or electricity than by foundation-model availability.
Fourth, open standards can prevent ecosystem capture. Telecom demonstrates that interoperability can expand an entire market while supporting competition among suppliers.
Fifth, overcapacity can commoditise the layer where investment is greatest. If training and inference supply expand faster than differentiated demand, model providers may experience the same margin pressure that affected solar manufacturing, shifting value toward applications and platform control.
Final stress-test of the original proposition
| Proposition | Assessment after stress-test |
|---|---|
| AI will become a commodity like oil | Too strong |
| AI infrastructure will develop strategic bottlenecks analogous to energy | Strongly supported |
| Cost can overwhelm modest quality differences after functionality matures | Supported conditionally |
| China can repeat solar-style cost compression in AI | Plausible mechanism, not established outcome |
| Chinese open models automatically create Chinese standards | Not supported |
| Standards and interoperability can determine ecosystem structure | Strongly supported |
| Midstream AI layers may matter more than model brands | Strongly supported |
| Full-stack integration can produce durable competitive advantage | Strongly supported |
| One national ecosystem must eventually dominate globally | Not established |
Key judgments
The strongest historical analogy depends on the analytical question being asked. Telecommunications is the strongest analogy for the control and standards layer; solar is strongest for cost competition and industrial learning; batteries are strongest for value-chain dependency; oil is strongest for physical capacity and geopolitical chokepoints.
The original oil metaphor therefore identifies something real but locates it at the wrong layer. The model is not the barrel of oil. Compute capacity, electricity and the infrastructure required to manufacture intelligence at scale are closer to the energy analogy, while models behave more like rapidly evolving software components operating inside that physical system.
The most important lesson from solar and batteries is that China’s strategic advantage does not need to arise from possessing the single best product. It can emerge from scale, supplier density, learning-by-doing, integration and sustained capital deployment. The most important lesson from telecommunications is the opposite: interoperability can prevent industrial scale from turning automatically into proprietary standards control.
The resulting contest is therefore not deterministic. China possesses credible mechanisms through which a domestic model and deployment ecosystem can generate lower integration costs and enormous learning effects. The U.S.-allied system possesses equally credible mechanisms through which open standards, hyperscaler scale, global identity platforms, accelerator software and multi-model control layers can absorb Chinese model innovation without surrendering infrastructure control.
What would change the assessment
The historical assessment would change decisively if foundation models became highly interchangeable across most economically significant workloads, because the solar and commodity dimensions would become much stronger. It would move in the opposite direction if frontier systems continued opening qualitatively new classes of high-value work that cheaper models could not perform, because capability leadership would remain structurally scarce.
The telecom analogy would strengthen if MCP, A2A and a small number of related interfaces stabilised into durable cross-vendor standards. It would weaken if major providers progressively closed their agent ecosystems.
The battery analogy would strengthen if HBM, advanced packaging and accelerator supply remained persistent binding constraints. It would weaken if hardware abstraction and abundant compute substantially reduced those dependencies.
The oil analogy would strengthen at the infrastructure level if AI companies increasingly secured long-duration power, compute and data-centre capacity through strategic bilateral contracts rather than ordinary cloud consumption—a trend already visible in the multi-gigawatt infrastructure commitments documented in 2026.
Open official record
The decisive missing evidence is no longer another benchmark leaderboard. The records capable of changing the Pillar II assessment are enterprise model-switching data, control-plane market shares, MCP and A2A production usage, long-term agent-platform retention, cloud migration costs, accelerator-software portability, international distribution of Chinese open-model inference, model-routing volumes, realised data-centre utilisation, inference-capacity pricing and the share of total enterprise AI expenditure captured by models versus infrastructure, software and integration.
Without those data, no defensible analysis can conclude that either China or the U.S.-allied ecosystem has already established the global standard. What can be established as of 25 September 2026 is narrower but strategically more important: the contest has moved decisively beyond the chatbot, and the assets most capable of determining durable AI power increasingly sit in standards, control planes, physical compute, enterprise identity and the complementary infrastructure surrounding intelligence.
The Battle Is Moving From the Model to the Control Layer
Open models are reducing lock-in at the foundation-model layer while identity, agent orchestration, cloud infrastructure, enterprise data, accelerator software and workflow integration increasingly determine where durable economic and geopolitical power accumulates.
The emerging AI market is not converging toward a simple contest between Chinese integrated ecosystems and Western frontier laboratories. Both sides increasingly operate through multi-layer architectures, while open weights and interoperability standards make the model itself progressively more replaceable. The strategically durable assets are increasingly the control plane, compute infrastructure, enterprise identity, developer runtime, industrial integration and standards that determine how intelligence is connected to real organisations.
Network Effects, Lock-In and the Control Layer
Where increasing returns and switching costs actually reside across weights, APIs, data, agents, identity, cloud, compilers, accelerators and industrial software.
The Western Counter-Architecture
How OpenAI, Microsoft, Anthropic, AWS, Google, Meta and NVIDIA increasingly form competing but interconnected full-stack architectures.
Oil, Telecom, Solar and Battery Stress-Test
Which historical mechanisms genuinely transfer to AI and which analogies fail when exposed to the economics of software, compute and standards.
Where AI network effects actually reside
The deepest switching costs increasingly sit outside the foundation model itself. The bars below are qualitative architecture indicators, not numerical scores, and represent the relative persistence described in the accompanying report.
Important for capability and developer familiarity, but highly portable when weights are open and inference APIs are compatible.
Persistent state, execution logic, tool access, recovery and orchestration create significantly deeper operational dependence.
Permissions, access control, audit, DLP and compliance architecture can survive many generations of underlying models.
Compilers, kernels, libraries, tooling and operator expertise can produce extremely persistent technical path dependence.
Internal repositories and historical operational knowledge generate organisation-specific value that increases as more systems are connected.
Indexes, metadata, access rules and knowledge pipelines represent implementation capital that is more durable than many model generations.
Residency, networking, committed spend and operational architecture make infrastructure migration expensive even when models remain portable.
ERP, MES, PLM, safety certification and physical-process integration can create the deepest organisational switching costs.
Qualitative architecture assessment only; widths are visual encodings of relative persistence described in Chapter 4 and are not empirical risk scores.
From model competition to control-plane power
Model layer
Qwen, DeepSeek, GLM, Hy, Claude, GPT, Gemini, Llama and other systems compete on capability, cost and availability.
Control layer
Model routing, identity, memory, tools, governance, evaluation, observability and infrastructure allocation become persistent.
Enterprise lock-in
Organisational dependence increasingly accumulates around the execution environment rather than the replaceable model endpoint.
Open standards are a counterforce to proprietary lock-in
Model Context Protocol — MCP
Anthropic introduced MCP as an open method for connecting models with external data and tools; the protocol was subsequently transferred into Linux Foundation governance and adopted beyond Claude, including by OpenAI.
Sources: Anthropic — Model Context Protocol, OpenAI — MCP supportAgent2Agent — A2A
Google’s A2A protocol is designed to let agents built by different vendors and frameworks communicate and was transferred into Linux Foundation governance with participation from competing enterprise vendors.
Source: Google Developers Blog — A2A donationOpen Chinese models can diffuse without Chinese infrastructure capture
| Deployment | Chinese model | Chinese cloud | Chinese hardware | Likely infrastructure dependence |
|---|---|---|---|---|
| Qwen on Alibaba Cloud | Yes | Yes | Variable | Higher Chinese stack capture |
| DeepSeek through Tencent | Yes | Yes | Variable | Intermediate |
| Qwen on foreign cloud / NVIDIA | Yes | No | No | Low Chinese infrastructure capture |
| GLM self-hosted on sovereign infrastructure | Yes | No | Variable | Low operational dependence |
This distinction separates model-standard diffusion from cloud, hardware and operational dependence.
The Western system is also an ecosystem
Western integration increasingly occurs through overlapping corporate stacks, contracts and open standards rather than through a single national champion.
Chinese architecture
U.S. / allied architecture
Selected Western control points
| Actor | Primary model layer | Infrastructure / platform advantage | Control-layer mechanism |
|---|---|---|---|
| OpenAI | GPT family | Stargate + partner infrastructure | Agents API, Codex, ChatGPT Work, enterprise deployment |
| Microsoft | OpenAI + third-party catalogue | Azure + Microsoft 365 + GitHub | Foundry, Entra, Agent 365, Defender, Purview |
| Anthropic | Claude | Multi-cloud distribution | MCP, Claude Code, partner/integrator ecosystem |
| AWS | Multi-model Bedrock | AWS + Trainium / Inferentia | AgentCore, IAM, PrivateLink, CloudTrail, policies |
| Gemini + partner models | Google Cloud + TPUs | Gemini Enterprise Agent Platform + Workspace | |
| Meta | Llama | Partner infrastructure | Open-model distribution and developer ecosystem |
| NVIDIA | Nemotron + third-party models | GPU, networking, AI systems | CUDA, NIM, NeMo, AI Enterprise |
Physical scale is part of the Western counter-architecture
Open complements around proprietary control points
| Company | Open / interoperable component | Principal proprietary control point |
|---|---|---|
| Meta | Llama open-weight ecosystem | Consumer platforms and distribution |
| Anthropic | MCP | Claude models and products |
| A2A | Gemini + Google Cloud / Workspace | |
| NVIDIA | Open-model support / standard APIs | CUDA + GPU infrastructure |
| AWS | Multi-model Bedrock | AWS infrastructure + IAM |
| Microsoft | Multi-model Foundry | Azure + Entra + Microsoft 365 |
| OpenAI | MCP and multiple execution environments | Models + agent harness + ChatGPT distribution |
No single historical analogy is sufficient
Oil
Strong for strategic supply, chokepoints, spare capacity and long-lived infrastructure. Weak for model fungibility and software reproduction.
Telecommunications
Strongest for interoperability, installed-base effects, protocol adoption, network compatibility and standards competition.
Solar PV
Strongest for scale, supplier clustering, learning-by-doing, overcapacity and dramatic unit-cost compression.
Batteries
Strongest for multi-stage dependencies, vertical integration, intermediate bottlenecks and strategic control of midstream inputs.
Which mechanisms actually transfer?
| Mechanism | Oil | Telecom | Solar | Batteries | AI interpretation |
|---|---|---|---|---|---|
| Strategic infrastructure | Very strong | High | High | High | Data centres, grids and accelerator supply |
| Network effects | Low | Very strong | Low | Low | Protocols, developer tools, agent ecosystems |
| Economies of scale | High | High | Very strong | Very strong | Compute and deployment scale |
| Learning curves | Medium | Medium | Very strong | Very strong | Model/inference optimisation and implementation learning |
| Supply chokepoints | Very strong | Medium | High | Very strong | HBM, lithography, packaging, power |
| Standards formation | Low | Very strong | Medium | Medium | MCP, A2A and API conventions |
| Overcapacity | Medium | Low | Very strong | High | Potential inference commoditisation |
| Software portability | Very weak analogy | Medium | Weak analogy | Weak analogy | Open weights can move across infrastructure |
Selected empirical anchors
The layered historical model
| AI layer | Best analogy | Mechanism captured |
|---|---|---|
| Electricity / data-centre capacity | Oil / energy infrastructure | Capacity, long-lived assets, geopolitical supply |
| Semiconductor manufacturing | Batteries | Multi-stage supply chains and chokepoints |
| Accelerator + runtime | Telecommunications equipment | Installed base and compatibility |
| Foundation weights | Software / partial commodity | Replicability and portability |
| Inference service | Cloud / industrial utility | Capacity and unit economics |
| APIs / protocols | Telecommunications standards | Interoperability |
| Agent control layer | Operating system / cloud platform | Ecosystem and switching costs |
| Enterprise integration | Industrial automation | Process-specific sunk cost |
| Model commoditisation | Solar PV | Learning curves and price compression |
Net assessment
The original proposition survives only after being reformulated. Foundation models are not equivalent to barrels of oil. The strategically durable object is the wider system that manufactures, distributes, governs and embeds intelligence. Oil explains capacity and chokepoints; telecommunications explains interoperability and standards; solar explains learning curves, scale and overcapacity; batteries explain vertical integration and intermediate dependencies. The competitive struggle therefore increasingly concerns who controls the interfaces between models, compute, data, identity, tools and organisational workflows. Chinese open models can become globally influential without automatically producing Chinese infrastructure dependence, while Western hyperscalers can absorb Chinese-origin models and still retain the more durable control plane. The emerging global standard is consequently more likely to be decided by the architecture around intelligence than by any single benchmark-leading model.
Principal verified sources
- Microsoft — Agent 365 integration with Foundry
- AWS — Amazon Bedrock AgentCore managed harness
- OpenAI — Agents API
- Anthropic — Model Context Protocol
- Google — Agent2Agent and Linux Foundation
- NVIDIA — AI Enterprise
- OpenAI — Intelligence infrastructure and Stargate
- Anthropic — AWS compute collaboration
- International Energy Agency — Solar PV Global Supply Chains
- International Energy Agency — Electric Vehicle Batteries
- International Energy Agency — Energy Technology Perspectives 2026
- U.S. Energy Information Administration — Crude oil characteristics and refining
- International Telecommunication Union — Standardization and interoperability
Pillar III — Capital Allocation, Strategic Dependency and the 2030 Contest
Chapter 7 — Geopolitics as Capital Allocation
Principal judgment
Artificial intelligence has entered a phase in which geopolitical capability is increasingly determined by the ability to convert financial capital into usable compute, and usable compute in turn depends on a chain of physical and institutional assets that extends far beyond the semiconductor itself: fabrication equipment, advanced packaging, high-bandwidth memory, accelerators, networking, land, data-centre shells, cooling, transformers, substations, transmission capacity, electricity generation, cloud orchestration, sovereign-compute allocation and long-duration customer commitments. The essential geopolitical fact is therefore not simply that unprecedented amounts of money are being invested in AI, but that those investments are becoming long-lived, geographically anchored and difficult to reverse. Once a company signs a multi-year compute contract, a government finances a sovereign supercomputer, a utility constructs transmission infrastructure for an AI cluster or an enterprise reorganises around one cloud platform, what began as a technological choice becomes a capital-allocation decision with strategic persistence.
This transition is visible across radically different political-economic systems. Alibaba committed at least RMB 380 billion, approximately US$53 billion, over three years to AI and cloud infrastructure; OpenAI reported in April 2026 that its Stargate programme had already secured more than 10 GW of U.S. AI infrastructure capacity against its original 2029 target; Anthropic committed more than US$100 billion over ten years to AWS technologies under an arrangement securing up to 5 GW of additional compute; Microsoft indicated during fiscal 2026 that calendar-year capital expenditure would reach roughly US$175 billion after lease-accounting adjustments; Alphabet guided to US$175–185 billion of 2026 capital expenditure, with the large majority directed toward technical infrastructure; the European Union is attempting to mobilise €200 billion through InvestAI, including a €20 billion AI-gigafactory facility; and the United Kingdom has committed to expanding sovereign public AI compute by at least 20 times by 2030. These figures are not directly additive because they describe different accounting categories, time horizons and instruments, but together they establish that AI competition has become a contest over fixed capital on a scale previously associated with telecommunications, energy and heavy industry. Alibaba to Invest RMB380 Billion in AI and Cloud Infrastructure — Alibaba Group — Feb 2025 Building the Compute Infrastructure for the Intelligence Age — OpenAI — Apr 2026 Anthropic and Amazon Expand Collaboration for up to 5 GW — Anthropic — Apr 2026 Microsoft FY2026 Q4 Earnings Call — Microsoft — Jul 2026 Alphabet 2025 Q4 Earnings Call — Alphabet — Feb 2026 InvestAI — European Commission — Feb 2025 UK AI Opportunities Action Plan — GOV.UK
The consequence is that AI geopolitics can increasingly be read through a capital-allocation map. Capital placed in model research can be redeployed relatively quickly; capital placed in a multi-gigawatt data-centre campus, accelerator fleet, transmission line or semiconductor fab cannot. The more of the AI stack that becomes physically embodied, the more today’s investment decisions determine tomorrow’s strategic options.
Capital is moving from software expenditure toward industrial infrastructure
Traditional software companies could scale globally with comparatively low physical marginal investment once a product had been written and distributed. Frontier AI reverses part of that economic model because greater usage and more sophisticated models require additional compute, and compute requires semiconductor production, data centres and electricity.
The International Energy Agency estimates that electricity consumed by data centres rises from approximately 460 TWh in 2024 to more than 1,000 TWh of associated generation by 2030, while end-use data-centre electricity consumption reaches approximately 945–950 TWh by 2030, roughly double current levels. AI-accelerated servers account for almost half of the net increase in the IEA’s base case, while electricity consumption by accelerated servers grows around 30% annually. The physical consequence is that AI scaling requires energy-sector investment on timelines far longer than software development cycles. Energy Demand from AI — International Energy Agency Energy Supply for AI — International Energy Agency
This difference in lead times is strategically decisive. The IEA notes that data centres can often be deployed within roughly two to three years, whereas generation, transmission and other energy infrastructure may require much longer planning and construction cycles. AI demand can therefore accelerate faster than the system supplying it, creating bottlenecks even where sufficient investment capital exists. Energy Demand from AI — International Energy Agency
The AI capital stack
| Capital layer | Typical asset | Economic life / persistence | Primary bottleneck | Strategic consequence |
|---|---|---|---|---|
| Model R&D | Training runs, datasets, research teams | Short–Medium | Talent, compute | Rapid capability change |
| Accelerators | GPU/NPU fleets | Medium | Chip supply, memory | Direct compute availability |
| Networking | High-speed fabric, switches, optics | Medium | Supply and deployment | Cluster efficiency |
| Data centres | Buildings, cooling, power distribution | Long | Permitting, land, equipment | Geographic anchoring |
| Grid connection | Substations and transmission | Very long | Queue and construction lead time | Regional capacity constraint |
| Generation | Renewables, gas, nuclear, storage | Very long | Regulation, equipment, financing | Energy-security linkage |
| Semiconductor fabs | Advanced manufacturing plants | Very long | Tools, yields, process IP | Sovereign technology capacity |
| Cloud regions | Integrated compute/data/network estate | Long | Capital + demand density | Platform lock-in |
| Enterprise integration | Workflow, security, data architecture | Medium–Long | Organisational change | Customer persistence |
| Sovereign compute | Public supercomputers / national capacity | Long | Procurement, funding | Strategic autonomy |
The critical shift is that increasingly large shares of the AI system cannot be redeployed without significant loss. A GPU can sometimes move between workloads; a hyperscale data-centre campus cannot easily move to another country; a transmission line cannot follow a model developer; and an enterprise identity and data architecture cannot be rebuilt overnight.
Venture capital is already reflecting the infrastructure turn
The concentration of private capital into AI is unusually large. The OECD reported that AI companies accounted for 61% of worldwide venture-capital investment in 2025, receiving US$258.7 billion out of US$427.1 billion of total global VC investment. More importantly for the thesis of this report, AI companies focused specifically on IT infrastructure and hosting attracted US$109.3 billion in 2025, the largest volume among AI sectors, while cumulative VC investment in that category reached US$256.1 billion between 2012 and 2025. AI Firms Capture 61% of Global Venture Capital in 2025 — OECD — Feb 2026
The geographic distribution is highly asymmetric. The OECD reports that U.S.-based investors accounted for approximately 56% of outgoing global AI VC value in 2025, compared with 9% for UK investors, 8% for China and 7% for the EU27. Firms based in the United States attracted approximately 75% of worldwide AI venture deal value, compared with 6% for the EU27, 5% for China and 5% for the United Kingdom. These figures measure venture investment rather than total AI capital formation and therefore exclude much hyperscaler capex, public investment and Chinese state-directed finance, but they indicate that private risk capital remains heavily concentrated in the United States. AI Firms Capture 61% of Global Venture Capital in 2025 — OECD — Feb 2026
Global AI venture-capital concentration, 2025
| Indicator | Value | Interpretation |
|---|---|---|
| Global VC investment | US$427.1bn | Total OECD-measured VC universe |
| AI-company VC | US$258.7bn | 61% of worldwide VC |
| AI infrastructure / hosting VC | US$109.3bn | Largest AI investment category |
| U.S. investor share of outgoing AI VC | 56% / US$124bn | Deepest private financing pool |
| UK investor share | 9% / US$20.7bn | Significant relative to economy |
| China investor share | 8% / US$17.2bn | VC only; excludes other financing channels |
| EU27 investor share | 7% / US$14.5bn | Fragmented compared with U.S. |
| U.S.-based firms’ share of AI VC received | 75% / US$194bn | Extreme concentration of private AI financing |
Source: OECD, February 2026.
The geopolitical implication is that national AI strength depends partly on the cost and availability of capital before any model is trained. Deep pools of risk capital allow experimentation, overbuilding and the absorption of technological failure; systems with thinner capital markets require public finance, development banks, industrial policy or concentrated corporate balance sheets to compensate.
Hyperscaler capital expenditure is becoming macro-industrial
Microsoft’s fiscal 2026 disclosures demonstrate the extent to which AI infrastructure has altered the economics of a large software company. Microsoft reported US$41 billion of capital expenditure in FY2026 Q4, roughly two-thirds directed toward shorter-lived assets primarily consisting of CPUs and GPUs, while the remaining expenditure supported long-lived infrastructure including large data-centre sites. The company stated that its calendar-year 2026 capex expectation, adjusted for lease classification, was approximately US$175 billion, while it remained capacity constrained despite rapid expansion. Microsoft FY2026 Q4 Earnings Conference Call — Jul 2026
Earlier in the fiscal year Microsoft had described plans to increase total AI capacity by more than 80% during the year and roughly double its total data-centre footprint within two years, while its Wisconsin Fairwater facility alone was expected to scale to approximately 2 GW. In Q3 the company guided to approximately US$190 billion of calendar-2026 capex before the later lease-accounting adjustment, illustrating both the scale and the accounting sensitivity of these infrastructure numbers. Microsoft FY2026 Q1 Earnings Conference Call Microsoft FY2026 Q3 Earnings Conference Call
Alphabet disclosed a comparable industrial transition. Its 2025 capital expenditure totalled US$91.4 billion, with approximately 60% of technical-infrastructure spending going to servers and 40% to data centres and networking, while management guided to US$175–185 billion of capex in 2026. Alphabet explicitly linked the spending to AI compute for Google DeepMind, consumer services and Google Cloud demand. Alphabet 2025 Q4 Earnings Call — Feb 2026
Selected disclosed AI-related infrastructure commitments
| Actor / jurisdiction | Commitment or guidance | Period | What the figure actually represents |
|---|---|---|---|
| Microsoft | ~US$175bn | Calendar 2026 | Expected capex after lease-accounting adjustment |
| Alphabet | US$175–185bn | 2026 | Total capex, overwhelmingly technical infrastructure |
| OpenAI Stargate | >10 GW secured | Reported Apr 2026 | AI-infrastructure capacity, not a dollar capex figure |
| Anthropic–AWS | >US$100bn commitment | 10 years | Anthropic commitment to AWS technologies |
| Anthropic–AWS | Up to 5 GW | Forward capacity | Training and inference capacity |
| Alibaba | ≥RMB380bn / US$53bn | Three years from 2025 | Cloud and AI infrastructure investment |
| EU InvestAI | €200bn mobilisation target | Multi-year | Public/private investment mobilisation, not realised spending |
| EU AI Gigafactory facility | €20bn | Multi-year | Facility to support up to five gigafactories |
| UK compute commitment | £1bn | Five-year programme | Expansion of national public compute |
| UK new heterogeneous AI supercomputer | £750m | Announced 2026 | Sovereign public compute investment |
Sources: Microsoft, Alphabet, OpenAI, Anthropic, Alibaba, European Commission, GOV.UK and DSIT.
These amounts should not be summed into a fictitious global AI-capex total. Microsoft’s number is company-wide capital expenditure with heavy AI/cloud exposure; Alphabet’s is company capex with technical-infrastructure dominance; OpenAI’s number measures power capacity rather than dollars; InvestAI is a mobilisation objective; and Alibaba’s three-year programme explicitly covers both AI and cloud infrastructure. Their analytical value lies in demonstrating capital intensity and irreversibility, not in generating a headline aggregate.
Alibaba demonstrates the Chinese corporate-capital model
Alibaba’s February 2025 commitment to invest at least RMB 380 billion in cloud and AI infrastructure over three years was larger, according to the company, than its cumulative spending in those areas over the preceding decade. Alibaba subsequently stated at Apsara 2025 that investment would rise beyond the original programme. Alibaba to Invest RMB380 Billion — Feb 2025 Alibaba Cloud Apsara Conference 2025 — Sep 2025
By August 2026 Alibaba reported US$7.1 billion of quarterly AI Cloud and Compute Services revenue, up 45% year over year, while the cloud segment’s adjusted EBITA rose 133%. These are company financial disclosures, but they matter because they show an investment loop in which infrastructure expenditure is increasingly matched by commercial utilisation rather than functioning solely as strategic capacity. Alibaba Full-Stack AI Accelerates Monetization — Aug 2026
The Chinese capital model therefore should not be understood purely as state-directed infrastructure. Large platform companies are themselves undertaking multi-year private commitments, while public strategy aligns data centres, power systems, telecommunications and industrial adoption around the same direction.
Power infrastructure is becoming part of national AI strategy
China formalised this relationship in September 2025 when the National Development and Reform Commission and National Energy Administration issued guidance on “AI+ Energy”, explicitly calling for coordinated development of intelligent computing capacity and electricity and for AI deployment to support energy-system reliability and efficiency. Implementation Opinions on Promoting High-Quality Development of “AI+ Energy” — NDRC / National Energy Administration — Sep 2025
The IEA projects that renewables meet almost half of the additional electricity generation required by data centres globally through 2030, while natural gas, coal and nuclear remain significant contributors. Data-centre electricity supply therefore becomes a national generation-planning problem rather than merely a corporate procurement decision. Energy Supply for AI — International Energy Agency
The distinction between contractual “green power” claims and physically available grid supply is important. The IEA explicitly evaluates the fuel mix of electricity physically consumed rather than the contractual procurement mix reported by data-centre operators. For geopolitical assessment, physical supply is the more important variable because grid congestion cannot be resolved through certificates.
Grid connection is becoming a hidden strategic bottleneck
The industry can manufacture more GPUs faster than many power systems can build substations and transmission infrastructure. The IEA’s updated analysis states that bottlenecks across energy equipment and chip manufacturing are reducing the probability of more aggressive near-term data-centre-growth scenarios even as investment pipelines expand. Key Questions on Energy and AI — International Energy Agency
This creates a strategic distinction between nominal compute investment and energised compute capacity. A planned data centre does not constitute available AI infrastructure until accelerators have been delivered, cooling installed, network links established and sufficient electricity can actually be supplied.
From announced investment to productive compute
| Stage | Capital committed? | Physical capacity available? | Strategic value |
|---|---|---|---|
| Corporate announcement | Yes | No | Signals intent |
| Land/site secured | Partly | No | Option value |
| Grid connection reserved | Increasingly | No | Critical future capacity |
| Data-centre shell constructed | Yes | No | Fixed infrastructure |
| Accelerators installed | Yes | Partly | Compute inventory |
| Power energised | Yes | Yes | Operational capacity |
| Cluster validated | Yes | Yes | Usable AI compute |
| Workload contracted | Yes | Yes | Monetised capacity |
This distinction should govern all infrastructure comparisons through 2030. Announced gigawatts, permitted gigawatts, connected gigawatts and fully utilised gigawatts are economically different objects.
Semiconductor fabs are the most capital-intensive sovereignty asset
Semiconductor fabrication creates even deeper irreversibility. Leading-edge fabs require long development cycles, specialised tooling, complex supplier ecosystems and high utilisation to achieve economic yields. Once constructed, the facility is geographically fixed and dependent on a surrounding network of utilities, chemicals, equipment, talent and logistics.
For China, this means semiconductor localisation is both a technological and capital-allocation question. Investment in domestic accelerators can reduce external dependence only if the surrounding manufacturing and memory ecosystem can provide sufficient volume and economic performance. For the United States and allies, subsidies and industrial-policy programmes likewise represent attempts to reshape the geography of semiconductor capital rather than merely support individual chip designs.
The strategically relevant variable is therefore not nominal fabrication capacity but economically usable advanced-node output multiplied by yield, packaging, memory availability and software efficiency.
Sovereign compute is emerging as a new category of public infrastructure
Governments increasingly treat AI compute as a strategic resource analogous to supercomputing, telecommunications or energy infrastructure. The European Union’s AI Factory programme now encompasses 19 AI Factories and 13 antennas, while EU and participating states have committed more than €2.6 billion to the AI Factories and antenna initiative and approximately €10 billion to EuroHPC supercomputing infrastructure and AI Factories over 2021–2027. AI Factories — European Commission EU Expands Network of AI Factories — Oct 2025
Europe has simultaneously launched InvestAI with a target of €200 billion in mobilised AI investment, including a €20 billion facility supporting up to five AI Gigafactories. By July 2026 the Commission had launched a gigafactory call intended to unlock more than €30 billion of investment. These amounts represent different policy mechanisms rather than cumulative realised spending, but they illustrate Europe’s attempt to correct a perceived compute-capital deficit through coordinated public-private finance. InvestAI — European Commission AI Gigafactories Call — European Commission — Jul 2026
The United Kingdom provides a smaller but conceptually similar example. The AI Opportunities Action Plan calls for at least a 20-fold expansion of public AI compute by 2030, while the government’s compute programme includes a £1 billion commitment to that expansion and a separately announced £750 million heterogeneous AI supercomputer. The existing AIRR includes Isambard-AI with 5,448 NVIDIA GH200 Grace Hopper superchips and Dawn with 1,024 Intel Data Centre GPU Max 1550 units. AI Opportunities Action Plan — GOV.UK AIRR Advanced Supercomputers — GOV.UK AIRR Heterogeneous Supercomputer — GOV.UK — Jul 2026
The UK also created a Sovereign AI Unit, with its next phase backed by up to £500 million, intended to support domestic companies in strategically important portions of the AI value chain. AI Opportunities Action Plan: One Year On — GOV.UK — 2026
Sovereign-compute strategies compared
| Jurisdiction | Instrument | Scale / target | Strategic objective |
|---|---|---|---|
| China | National computing + AI+ infrastructure programmes | Distributed national system | Domestic scale, industrial embedding, localisation |
| EU | EuroHPC AI Factories | 19 factories + 13 antennas | Shared public compute and SME access |
| EU | InvestAI | €200bn mobilisation objective | Close infrastructure/capital gap |
| EU | Gigafactory facility | €20bn | Up to five large AI gigafactories |
| UK | AIRR expansion | ≥20× by 2030 | Sovereign public AI compute |
| UK | Compute investment | £1bn | Public capacity expansion |
| UK | New heterogeneous AI supercomputer | £750m | Frontier research and large-scale inference |
| U.S. | Primarily private hyperscaler/model-company buildout | Multi-hundred-billion-dollar corporate capex | Frontier capacity and cloud leadership |
The institutional structures differ profoundly. Europe and the UK are explicitly building shared public access; U.S. capacity expansion is dominated by private hyperscalers and frontier-model partnerships; China combines large private platforms with state infrastructure coordination.
Public procurement can create standards before markets fully mature
Compute sovereignty is not only about owning machines. Public authorities can shape standards by determining what architectures receive procurement contracts, which clouds may host sensitive workloads, what data-residency rules apply, which models qualify for public deployment and which security frameworks become mandatory.
Public procurement can therefore create demand-side industrial policy. An accelerator, cloud or agent architecture that enters government at scale can accumulate integration assets, certification experience, security credentials and reference customers that improve its position in adjacent regulated sectors.
This is particularly important because the AI market remains technically fluid. Procurement decisions taken before standards stabilise can produce institutional path dependence extending well beyond the useful life of the original model.
Long-duration contracts transform temporary technological leadership into alignment
Anthropic’s ten-year AWS commitment provides a clear example. The agreement does not merely purchase current cloud capacity; it aligns Anthropic’s future training and inference requirements with AWS infrastructure and Trainium development across multiple hardware generations. Anthropic and Amazon Expand Collaboration — Apr 2026
Microsoft disclosed a similarly consequential relationship with OpenAI: by FY2026 Q1, OpenAI had contracted an additional US$250 billion of Azure services, while Microsoft retained specific API, intellectual-property and commercial rights under their updated relationship. Microsoft FY2026 Q1 Earnings Conference Call
Such commitments create mutual path dependence. The model developer optimises software around the infrastructure provider; the infrastructure provider finances capacity on the basis of future demand; suppliers design hardware for the expected workload; and the resulting technical integration makes later separation increasingly expensive.
Capital lock-in mechanisms
| Commitment | Immediate effect | Long-term lock-in mechanism |
|---|---|---|
| Multi-year cloud contract | Secures capacity | Workload optimisation around one platform |
| Dedicated accelerator architecture | Improves efficiency | Compiler/kernel dependence |
| Data-centre lease | Adds compute footprint | Geographic fixed cost |
| Power-purchase agreement | Secures energy | Long-duration energy exposure |
| Grid buildout | Enables capacity | Location becomes strategically privileged |
| Sovereign-compute procurement | Creates public capacity | Government tooling and standards |
| Enterprise migration | Enables AI workloads | Workflow/data/identity switching cost |
| Fab investment | Expands chip supply | Multi-decade industrial ecosystem |
Capital therefore acts as a mechanism for turning preferences into infrastructure.
The geographical value of electricity is changing
AI creates a new form of geographic arbitrage. Historically, cloud services benefited heavily from access to users and telecommunications networks. Large AI-training clusters increasingly value power availability, land, cooling conditions, transmission capacity and regulatory speed.
A region capable of providing gigawatts of reliable power can become strategically valuable even if it is not a traditional software centre. Conversely, established technology hubs with constrained grids may find that financial capital cannot be converted quickly into additional compute.
This is one reason governments increasingly identify AI Growth Zones, dedicated computing regions and accelerated permitting mechanisms. The UK’s Compute Roadmap explicitly links AI Growth Zones with large-scale infrastructure deployment and public-compute expansion. UK Compute Roadmap — GOV.UK
The key variable is therefore capital conversion efficiency
The decisive metric is not simply how much capital is announced, but how efficiently financial capital becomes productive intelligence.
A jurisdiction with US$100 billion of investment but severe grid delays, equipment constraints and low utilisation may create less usable intelligence than one investing considerably less into infrastructure that can be energised and monetised quickly.
Capital conversion chain
Financial capital → semiconductor supply → data-centre construction → grid connection → electricity → operational compute → model capability → integrated applications → productive output
Failure at any intermediate stage lowers the return on all preceding investment.
This chain explains why geopolitical AI analysis increasingly requires expertise in electricity, industrial construction, semiconductor economics and infrastructure finance, not only computer science.
Key judgments
| Judgment | Evidence | Confidence |
|---|---|---|
| AI investment has moved into infrastructure-scale capital formation | Hyperscaler, OpenAI, Alibaba, EU and UK commitments | High |
| Private AI capital remains disproportionately concentrated in the United States | OECD VC evidence | High |
| Compute capacity is increasingly constrained by energy and physical infrastructure | IEA analysis | High |
| Announced gigawatts should not be treated as operational compute | Infrastructure-development sequence | High |
| Long-duration cloud contracts create strategic alignment | OpenAI–Microsoft and Anthropic–AWS | High |
| Sovereign compute has become an explicit government-policy instrument | EU and UK programmes | High |
| The country spending the most automatically wins | Not supported | Rejected |
| Conversion of capital into energised, utilised compute is more important than headline capex | Evidence-supported analytical judgment | High |
What would change the assessment
The capital-allocation thesis would weaken if model efficiency improved so rapidly that frontier capability required substantially less physical compute, if inference shifted strongly toward inexpensive distributed devices rather than large data centres, or if open markets made compute capacity sufficiently fungible that geography ceased to matter.
It would strengthen if accelerator demand continued exceeding supply, grid connection became a persistent binding constraint, multi-gigawatt bilateral compute contracts became standard, or governments increasingly restricted strategically sensitive workloads to sovereign infrastructure.
Open official record
The records most capable of altering this assessment are operational rather than promotional: energised versus announced gigawatts; grid-connection queues; accelerator utilisation rates; data-centre construction completions; power prices paid by AI facilities; realised rather than pledged corporate capex; semiconductor fab yields; HBM supply; public AI-procurement values; sovereign-compute utilisation; and the proportion of multi-year cloud commitments actually consumed.
Chapter 8 — Cognitive Infrastructure, Scenarios and the 2030 Standard
Principal judgment
Artificial intelligence becomes infrastructure rather than software when organisations cease treating models as discrete products selected individually and begin depending continuously on a persistent system that mediates knowledge, decisions, software execution, administrative processes and physical operations. That transition is already underway, but it is uneven across sectors and jurisdictions. The likely 2030 architecture will therefore not be determined by model intelligence alone, nor by inference cost alone, nor by compute capacity alone. The stronger assessment is that the dominant systems will be those capable of optimising capability × integration × infrastructure availability simultaneously while keeping regulatory, security and switching burdens within acceptable limits.
This means the central thesis survives, but in modified form. AI will not simply be “chosen like oil,” because foundation models are too heterogeneous and portable. Nor will the global system necessarily select one national AI ecosystem. The most credible 2030 structure is a layered market in which frontier intelligence remains scarce at the upper end, routine intelligence becomes increasingly commoditised, enterprises multi-home across model providers, and durable strategic power accumulates in compute capacity, control planes, identity, data, agent infrastructure and industrial integration.
When AI becomes infrastructure rather than software
A technology becomes infrastructure when several conditions coincide.
First, consumption becomes continuous rather than episodic. Electricity is infrastructure because firms do not purchase a separate generation technology for each task; they consume a service continuously through an underlying system. AI begins to resemble infrastructure when agents, search, coding, customer service, document processing, analytics and industrial control invoke models continuously within existing workflows.
Second, the technology becomes embedded in complementary systems. An AI assistant is software; an AI control layer connected to identity, ERP, databases, communications, industrial sensors and public-service systems becomes infrastructure.
Third, replacement requires system migration rather than product substitution. When changing models requires revalidating workflows, permissions, data pipelines and regulatory controls, the organisation has acquired infrastructure dependence.
Fourth, capacity planning becomes strategic. Once an organisation must reserve accelerator capacity, electricity, cloud commitments or sovereign compute, AI enters the domain of infrastructure planning.
Fifth, the technology becomes a general input into other production processes rather than a standalone product.
Infrastructure transition test
| Criterion | Software-like AI | Infrastructure-like AI |
|---|---|---|
| Usage | Occasional | Continuous |
| Purchasing | Per product / seat | Capacity + service architecture |
| Integration | Shallow | Deep |
| Model switching | Easy | Model may be easy, system difficult |
| Data connection | Limited | Enterprise-wide |
| Identity | User login | Operational agent identities |
| Capital requirement | Mostly operating expense | Significant fixed capital |
| Failure impact | Localised | Systemic workflow disruption |
| Governance | Application-specific | Enterprise-wide |
| Capacity planning | Secondary | Strategic |
| Replacement | Product migration | Organisational/infrastructure migration |
The most mature enterprise-agent architectures examined in Chapters 3–5 already satisfy several of these criteria.
The 2030 market is unlikely to converge on one universal model
The technological market contains powerful forces toward commoditisation and equally powerful forces toward differentiation. Open models, standard APIs, model routing and agent platforms reduce switching costs; at the same time, very difficult tasks can continue to reward frontier capability.
The resulting equilibrium is likely to resemble a tiered intelligence market.
Potential 2030 intelligence layers
| Layer | Likely economic characteristic | Principal competition |
|---|---|---|
| Frontier reasoning | Scarce, expensive, capability-sensitive | Model intelligence |
| Professional agents | Integrated, persistent, governed | Model + control plane |
| Enterprise routine inference | High-volume, cost-sensitive | Integration + inference economics |
| Edge / industrial inference | Latency and sovereignty-sensitive | Hardware + efficiency |
| Commodity language tasks | Highly substitutable | Price |
| Sovereign / regulated workloads | Jurisdiction-sensitive | Trust + control + infrastructure |
Such segmentation allows Chinese and Western systems to coexist while competing at different layers.
Scenario architecture for 2030
The current evidence supports four genuinely distinct pathways. Numerical probabilities are not assigned because no defensible base rate or validated forecasting model exists for a technology market changing at this speed.
Pathway A — Chinese integrated-stack expansion
In this pathway, Chinese providers continue narrowing hardware constraints while maintaining aggressive inference economics and translating domestic deployment scale into reusable industrial AI systems. Qwen, DeepSeek, GLM, Hy, Kimi or successor families remain broadly adopted internationally, while Alibaba, Huawei or other Chinese platforms secure greater infrastructure participation outside China.
The pathway becomes stronger if domestic accelerator supply expands rapidly, HBM and advanced-packaging constraints diminish, Chinese inference systems remain significantly cheaper on quality-adjusted workloads and Chinese cloud providers win enterprise deployments outside their home market.
Its greatest obstacle is that open-model success can occur without stack success. Chinese weights hosted on NVIDIA systems or Western hyperscalers may expand Chinese technical influence while leaving control-plane economics elsewhere.
Pathway B — U.S.-allied platform consolidation
In this pathway, frontier capability remains economically valuable while Microsoft, AWS, Google, OpenAI, Anthropic and NVIDIA capture increasing shares of enterprise AI through global cloud infrastructure, identity, accelerator software and agent control planes.
Chinese open models may remain important but become inputs inside Western-managed platforms rather than drivers of Chinese infrastructure adoption.
This pathway strengthens if enterprise agents become deeply linked to Microsoft Entra, AWS IAM, Google Workspace or comparable installed platforms; if NVIDIA or allied accelerators retain strong software advantages; and if hyperscalers maintain sufficient capital and electricity access to keep inference prices falling.
Pathway C — Multi-homed modular AI
In this pathway, no single national or corporate ecosystem dominates because open weights, standard APIs, MCP, A2A and model-routing layers make intelligence increasingly interchangeable.
Enterprises maintain one data and control architecture while dynamically selecting models by cost, capability, geography and regulatory requirement.
This would substantially weaken both Chinese and American proprietary-stack theories while increasing the power of cloud-neutral orchestration, sovereign infrastructure, open standards and enterprise-owned data.
Pathway D — Sovereign fragmentation
In this pathway, governments increasingly require local compute, local model validation, jurisdictional data storage and sovereign control of sensitive AI workloads. The global AI economy fragments into national or regional stacks.
Europe’s AI Factories and gigafactories, the UK’s sovereign-compute programme and China’s domestic ecosystem already demonstrate elements of this logic. AI Factories — European Commission UK AIRR — GOV.UK
Fragmentation would increase duplication and potentially raise costs, but it could also create markets for sovereign clouds, local accelerator systems and open models capable of running independently of foreign APIs.
Scenario comparison
| Variable | Chinese integrated expansion | U.S.-allied consolidation | Multi-homed modularity | Sovereign fragmentation |
|---|---|---|---|---|
| Model openness | High | Mixed | Very high importance | High |
| Cloud concentration | Chinese platforms rise | Western hyperscalers dominate | Moderate | Regional |
| Model switching | Moderate | Moderate | High | Moderate |
| Sovereign compute | Important | Secondary in private sector | Important | Central |
| Hardware constraint | Must improve in China | Allied advantage persists | Abstracted through routing | Local availability decisive |
| Global standards | Mixed | Western/open protocols | Open standards dominate | Fragmented standards |
| Enterprise lock-in | Platform-specific | Platform-specific | Lower | Jurisdiction-specific |
| Cost competition | Very strong | Strong | Very strong | Weaker due duplication |
| Geopolitical alignment | Higher | Higher | Lower | Very high regionalisation |
These pathways are not necessarily mutually exclusive across every sector. Highly regulated government and defence workloads can fragment while consumer and software-development markets remain globally multi-homed.
The decisive variable is capability-adjusted integration cost
The central thesis becomes most defensible when reframed around a simple economic relationship:
Effective AI value = useful task capability ÷ total deployed cost
where total deployed cost includes model inference, compute, integration, governance, verification, infrastructure, failure and switching burdens.
The relevant denominator is not token cost.
A system offering extraordinarily cheap inference but poor reliability can have low economic value. A frontier model with expensive tokens can have high economic value when it solves a task other systems cannot. A slightly weaker model can dominate a repetitive enterprise process if it meets the required reliability threshold at one-fifth of the total operating cost.
The 2030 winner therefore cannot be identified from model benchmarks alone.
Capability thresholds produce different markets
Model quality matters non-linearly. Below a task’s minimum reliability threshold, cheaper inference has little value. Above that threshold, additional capability can exhibit diminishing commercial returns.
This produces a threshold structure:
Below threshold → capability dominates.
Near threshold → capability and verification cost dominate.
Well above threshold → integration and infrastructure economics increasingly dominate.
This framework explains why frontier systems can remain indispensable for some activities while routine enterprise inference becomes heavily commoditised.
Infrastructure availability can override both price and intelligence
Even the best quality-adjusted model cannot be deployed at scale without available compute.
The IEA’s updated assessment explicitly states that bottlenecks in chips and energy equipment are constraining near-term data-centre expansion. Its central projection remains around 950 TWh of data-centre electricity consumption by 2030, but the organisation notes meaningful upside after 2030 if investment relieves those constraints. Key Questions on Energy and AI — International Energy Agency
This introduces a third competitive variable beyond quality and integration: availability.
A model priced at US$0.20 per million tokens is irrelevant if capacity is unavailable when needed; an expensive but guaranteed dedicated deployment can be economically superior for high-value production systems.
The 2030 contest is therefore triangular
The central competition can be represented through three dimensions.
Intelligence
Can the system perform economically valuable tasks that alternatives cannot?
Integration
Can enterprises deploy and govern the system cheaply, quickly and reliably?
Infrastructure
Can sufficient compute, power and network capacity be supplied where and when required?
No durable ecosystem can ignore one of the three.
Strategic dominance conditions
| Competitive condition | Dominant variable |
|---|---|
| Large capability gaps | Intelligence |
| Small capability gaps + high volume | Integration / cost |
| Compute scarcity | Infrastructure availability |
| High regulatory burden | Governance / integration |
| Sovereign workload | Infrastructure control |
| Routine commodity task | Price / efficiency |
| High-consequence reasoning | Capability / reliability |
| Deep industrial process | Integration / installed base |
| Agentic enterprise | Control plane + identity |
The likely global structure will therefore be heterogeneous rather than winner-take-all.
China possesses a structural advantage in deployment density
China’s principal potential advantage is not simply model pricing. It is the combination of a very large industrial economy, dense telecommunications infrastructure, large cloud platforms, strong government coordination and a domestic market capable of generating enormous deployment volume.
If repeated deployments create learning-by-doing, integrators can accumulate reusable industry templates, data pipelines and operating procedures. Such learning effects can lower the marginal cost of future implementations.
Alibaba’s commitment of at least RMB 380 billion to cloud and AI infrastructure and its rapidly growing AI Cloud and Compute Services revenue demonstrate that the company is attempting to transform this scale into a self-reinforcing commercial system. Alibaba RMB380 Billion AI Infrastructure Programme Alibaba Q2 2026 AI Cloud and Compute Results
The constraint remains upstream production autonomy and international transferability.
The U.S.-allied architecture possesses a structural advantage in capital depth
The U.S.-allied system’s principal advantage is the combination of deep capital markets, leading accelerator/software ecosystems, hyperscale global cloud infrastructure and massive installed enterprise platforms.
OECD venture-capital evidence shows that U.S.-based AI companies attracted approximately US$194 billion of VC in 2025, three-quarters of the measured worldwide total, while Microsoft and Alphabet alone are planning approximately US$350 billion or more of calendar-2026 capex between them, though these figures cover broader technical infrastructure and cannot be classified entirely as AI investment. OECD — AI Venture Capital Microsoft FY2026 Q4 Alphabet 2025 Q4
The resulting advantage is not guaranteed technological supremacy; it is the ability to finance multiple competing technological bets simultaneously.
Europe is attempting to buy optionality
Europe’s strategic position differs from both China and the United States. The EU has strong research capability, industrial demand and public supercomputing infrastructure but historically less private frontier-AI capital and weaker hyperscaler scale.
InvestAI, AI Factories and AI Gigafactories represent an effort to purchase strategic optionality: Europe is creating compute capacity that can support domestic models and applications even if private markets alone do not finance the required scale. The EU currently reports 19 AI Factories, and the gigafactory programme is explicitly intended to support advanced frontier-model training on European infrastructure. AI Factories — European Commission AI Gigafactories Call — European Commission — 2026
This strategy does not require Europe to produce the single strongest global model. Its strategic value lies partly in ensuring that European companies and governments retain alternative compute, data and deployment routes.
The United Kingdom is pursuing a smaller sovereign optionality strategy
The UK’s explicit goal of increasing sovereign public AI compute by at least 20 times by 2030 reflects the same logic at national scale. The AI Opportunities Action Plan specifically recommends hosting multiple hardware providers to avoid vendor lock-in, demonstrating that compute sovereignty is being framed not only as capacity expansion but as architectural optionality. AI Opportunities Action Plan — GOV.UK
The £750 million heterogeneous-supercomputer programme announced in July 2026 reinforces this principle because its explicit objective is to combine established and novel AI hardware rather than hard-code national public compute around one accelerator architecture. AIRR Heterogeneous Supercomputer Host Site Selection — GOV.UK
This is a strategically important alternative to pure scale: diversification can itself be a sovereignty asset.
Multi-homing may become the rational corporate equilibrium
A sophisticated enterprise may ultimately refuse to choose one ecosystem.
Instead it can maintain:
a primary data architecture;
one enterprise identity system;
a model-routing layer;
multiple model suppliers;
multiple cloud or sovereign-compute options for critical workloads;
portable evaluation datasets;
standardised tools and MCP-like connectors;
and workload-specific procurement rules.
This architecture converts AI procurement from vendor selection into portfolio management.
The advantage is bargaining power and resilience; the disadvantage is greater integration complexity and potentially lower efficiency than a tightly vertically integrated system.
The most important 2030 standard may not be a model
Standards wars historically focus attention on visible formats: VHS versus Betamax, GSM versus CDMA, Blu-ray versus HD DVD. AI’s dominant standard may instead be an invisible architectural interface.
Candidates include:
the standard method by which models call tools;
the standard method by which agents communicate;
the standard API semantics for inference;
the identity model assigned to autonomous agents;
the portability format for agent memory;
the runtime interface between model and accelerator;
the security model used for autonomous execution;
or the dominant cloud control plane.
If those interfaces stabilise, model brands can change while the underlying standard remains.
This would make the 2030 AI market more analogous to the Internet, cloud computing and telecommunications than to an ordinary product market.
Observable indicators for the 2030 contest
The following indicators are more decision-useful than generic benchmark leadership.
Model-layer indicators
| Indicator | What it reveals |
|---|---|
| Quality gap across economically important tasks | Whether frontier intelligence remains scarce |
| Open-weight share of production workloads | Degree of model commoditisation |
| Model-switching frequency | Actual interchangeability |
| Price per verified completed task | Quality-adjusted economics |
| Long-context and agent reliability | Ability to automate complex work |
Platform indicators
| Indicator | What it reveals |
|---|---|
| Enterprise agents under each control plane | Platform installed base |
| Identity-linked autonomous agents | Depth of organisational integration |
| MCP / A2A production usage | Strength of interoperability |
| Model-routing share | Multi-homing maturity |
| Migration time between platforms | Real switching cost |
Infrastructure indicators
| Indicator | What it reveals |
|---|---|
| Energised AI data-centre GW | Real capacity |
| Accelerator utilisation | Capital efficiency |
| HBM supply and pricing | Memory bottleneck |
| Grid-connection lead time | Power constraint |
| Inference per watt | Energy efficiency |
| Compute cost per successful task | Infrastructure competitiveness |
Geopolitical indicators
| Indicator | What it reveals |
|---|---|
| Sovereign-compute capacity | National autonomy |
| Government AI procurement | State-created installed base |
| Cross-border data restrictions | Fragmentation pressure |
| Cloud-region expansion | International infrastructure reach |
| Export-control scope | Technology-denial intensity |
| Domestic accelerator share | Hardware sovereignty |
| Model usage outside home jurisdiction | International technological influence |
The thesis has identifiable falsifiers
A serious thesis requires conditions under which it would be rejected.
Falsifier: frontier quality remains overwhelmingly decisive
If the strongest models repeatedly unlock economically valuable capabilities unavailable to lower-cost systems, organisations will continue paying significant premiums for frontier intelligence and the “cheap integration wins” thesis will remain secondary.
Falsifier: switching costs collapse
If open standards make models, agents, memory, retrieval, identity and cloud execution genuinely portable, ecosystem lock-in could become much weaker than this report expects.
Falsifier: compute intensity falls dramatically
If algorithmic efficiency reduces training and inference requirements faster than demand expands, capital-intensive data-centre advantage may diminish.
Falsifier: industrial AI fails to generate measurable productivity
If enterprise AI remains largely an expensive interface layer without sustained effects on labour productivity, throughput, quality or innovation, the infrastructure thesis would be substantially overstated.
Falsifier: political fragmentation destroys scale economies
If jurisdictions require separate models, data centres and regulatory stacks, global network effects may weaken substantially and regional systems may replace worldwide ecosystem competition.
Critical dependencies by ecosystem
China
China’s principal dependencies remain domestic advanced-compute economics, HBM, leading-edge manufacturing, software maturity and the ability to export more than model weights.
United States and allied technology ecosystem
The principal dependencies are electricity, grid expansion, the economics of enormous capital expenditure, semiconductor supply concentration and continued enterprise willingness to remain within hyperscaler-centred architectures.
European Union
The principal dependencies are whether public compute creates private industrial scale, whether gigafactories are financed and utilised effectively, and whether European enterprises develop sufficiently large downstream AI businesses to justify the infrastructure.
United Kingdom
The principal dependencies are the successful scaling of AIRR, connection between public compute and high-growth domestic firms, access to international capital and the ability to preserve architectural flexibility while achieving sufficient scale.
Multi-homed global ecosystem
Its principal dependency is interoperability: open protocols must work reliably enough in production to prevent model and platform ecosystems from becoming closed silos.
Capital-allocation implications
The previous chapters imply that long-duration capital should not be analysed through a simple “which model company wins?” framework.
The highest-persistence assets are generally those with long physical lives or deep organisational embedding:
semiconductor manufacturing;
advanced packaging;
power generation and transmission;
high-speed networking;
data-centre campuses;
accelerator/software platforms;
enterprise identity;
agent control infrastructure;
industrial workflow software;
and sovereign-compute capacity.
Foundation-model leadership can create enormous economic value, but it also changes rapidly. Infrastructure investments therefore represent a different duration profile from model-company exposure.
Capital duration matrix
| Asset | Technological obsolescence risk | Physical / organisational persistence | Strategic relevance |
|---|---|---|---|
| Individual frontier model | Very high | Low | High but transient |
| Model laboratory | High | Medium | High |
| Agent platform | Medium | High | Very high |
| Cloud region | Medium | Very high | Very high |
| GPU/NPU fleet | High | Medium | High |
| Accelerator software ecosystem | Medium | Very high | Very high |
| Data-centre site | Low–Medium | Very high | Very high |
| Grid connection | Low | Very high | Very high |
| Power generation | Low | Very high | High |
| Semiconductor fab | Medium | Very high | Very high |
| Enterprise workflow integration | Medium | Very high organisationally | Very high |
| Open standard | Medium | Potentially very high | High if broadly adopted |
The strategic lesson is that durable AI advantage may be captured by assets that survive repeated model replacement.
The principal decision thresholds
Several observable thresholds would materially change the 2030 assessment.
Model commoditisation threshold: multiple independent model families consistently satisfy the same enterprise reliability requirements at similar quality.
Interoperability threshold: organisations can move agents, tool connections and memory between platforms without major re-engineering.
Domestic-compute threshold: China can train and deploy leading systems at large scale predominantly on domestically controlled accelerators and memory.
Power threshold: grid and generation investment grows rapidly enough that accelerator supply rather than electricity remains the primary physical bottleneck.
Productivity threshold: independently measured AI deployment produces material and persistent reductions in labour hours, cycle times or failure rates across major industries.
Sovereignty threshold: governments begin reserving strategically important AI workloads systematically for domestic or sovereign infrastructure.
Crossing different combinations of these thresholds would push the market toward different scenarios.
2030 signpost matrix
| Observable development | Chinese integrated pathway | U.S.-allied pathway | Multi-home pathway | Sovereign-fragmentation pathway |
|---|---|---|---|---|
| Chinese domestic accelerator economics improve sharply | Strengthens | Weakens | Neutral | Strengthens regionalisation |
| MCP/A2A become dominant production standards | Weakens exclusive stack lock-in | Weakens exclusive stack lock-in | Strongly strengthens | Weakens |
| Enterprise agents become tightly tied to cloud identity | Mixed | Strengthens | Weakens | Mixed |
| Governments mandate sovereign hosting | Strengthens domestically | Mixed | Weakens | Strongly strengthens |
| Frontier-model capability gaps persist | Mixed | Strengthens current leaders | Weakens | Mixed |
| Model-quality convergence accelerates | Strengthens cost competition | Strengthens hyperscaler platform layer | Strongly strengthens | Mixed |
| Grid shortages persist | Depends on local power | Weakens capacity expansion | Mixed | Strengthens local planning |
| Open Chinese models dominate international downloads | Strengthens model influence | May strengthen Western hosting | Strengthens | Mixed |
| Chinese clouds gain major non-China enterprise share | Strongly strengthens | Weakens | Weakens | Weakens |
| Western clouds host majority of Chinese-model inference outside China | Weakens Chinese stack capture | Strengthens | Strengthens | Neutral |
No numerical scores are assigned because the table represents directional effects rather than calibrated forecast probabilities.
The most likely structural outcome is layered competition rather than a single winner
The evidence available as of September 2026 does not support the assertion that one model, one company or one national ecosystem is on a defensible path toward universal dominance.
Several economic forces instead point toward layered competition.
Models are becoming more numerous and more portable.
Agent platforms are becoming more persistent.
Infrastructure is becoming more capital intensive.
Power availability is becoming geographically differentiated.
Governments are building sovereign alternatives.
Enterprises increasingly demand model choice.
Open standards are reducing some proprietary barriers while cloud identity and infrastructure deepen others.
These forces operate simultaneously rather than sequentially.
The resulting market can therefore support Chinese, U.S.-allied, European sovereign and globally interoperable architectures at the same time.
Final net assessment
The original thesis correctly identifies a historic shift: the strategic competition in artificial intelligence is no longer adequately represented by asking whether ChatGPT, DeepSeek, Gemini, Claude, Qwen or another individual model is “best.”
The economically decisive object is increasingly the system that makes intelligence available as a dependable productive input.
That system requires models, but also accelerators.
It requires accelerators, but also memory and packaging.
It requires data centres, but also electricity and grid connections.
It requires inference, but also enterprise data, identity and security.
It requires agents, but also workflows capable of using them productively.
And it requires capital capable of financing every layer simultaneously.
China’s potential advantage lies in enormous industrial deployment density, coordinated infrastructure development, aggressive model economics and a rapidly maturing domestic technology stack.
The U.S.-allied ecosystem’s potential advantage lies in frontier capability, extraordinary private-capital depth, globally dominant cloud platforms, accelerator software, enterprise identity and international systems-integration networks.
Europe’s emerging advantage is strategic optionality built through public compute, standards and regulatory-market scale.
The United Kingdom is attempting a narrower version of the same strategy through sovereign compute, heterogeneous hardware and concentrated support for domestic AI firms.
Open-model and open-standard ecosystems provide a fourth route in which national origin becomes less important because models, tools and agents travel across infrastructure boundaries.
The decisive variable through 2030 is therefore unlikely to be intelligence alone, because capability increasingly diffuses; it is unlikely to be cost alone, because low-cost systems that fail critical thresholds have little economic value; and it cannot be infrastructure alone, because unused compute does not create productivity.
The stronger conclusion is that strategic advantage will emerge from the interaction of all three:
sufficient intelligence to cross the task threshold;
sufficient integration efficiency to make deployment economically rational;
and sufficient physical infrastructure to provide that intelligence reliably at scale.
The architecture that minimises the total friction between these layers will possess the strongest claim to becoming the dominant cognitive infrastructure of the 2030s.
In that sense, the central thesis can be stated more precisely:
AI will not ultimately be chosen as a chatbot. It will be chosen as infrastructure.
And once intelligence becomes infrastructure, geopolitics increasingly becomes the question of who finances it, who powers it, who manufactures it, who governs its interfaces, who controls its dependencies and where the resulting capital becomes too deeply embedded to move.
Final key judgments
| Final judgment | Evidentiary status |
|---|---|
| AI competition has moved beyond standalone model performance | Strongly established |
| Capital allocation is becoming a central determinant of AI power | Strongly established |
| Compute and electricity are becoming geopolitical infrastructure | Strongly established |
| The U.S. currently possesses exceptional private-capital and hyperscaler depth | Strongly established |
| China possesses exceptional domestic deployment scale and growing stack integration | Strongly established |
| Europe and the UK are deliberately building sovereign-compute optionality | Established |
| Open Chinese models necessarily produce Chinese infrastructure dependence | Rejected |
| Western AI competition is merely a frontier-model race | Rejected |
| The cheapest model will automatically become the global standard | Rejected |
| Frontier intelligence alone will determine the global AI architecture | Not supported |
| One national ecosystem will necessarily dominate globally by 2030 | Not established |
| Capability-adjusted integration cost is likely to become a central procurement metric | Strong analytical judgment |
| Control-plane and infrastructure assets are more persistent than individual model leadership | Strong analytical judgment |
| The likely 2030 system is layered, multi-model and partly multi-homed | Evidence-supported analytical judgment; not a deterministic forecast |
What would change the final assessment
The assessment should be revised materially if any of five developments occurs before 2030: a breakthrough produces a durable frontier-capability monopoly; open standards collapse switching costs across the entire enterprise stack; Chinese semiconductor constraints are decisively removed or intensified; grid and power limitations become substantially less important because AI efficiency improves much faster than demand; or independently measured productivity evidence demonstrates that AI integration creates far smaller economic gains than current capital allocation assumes.
Open official record
The final unresolved record is concentrated in ten variables: energised AI compute capacity by jurisdiction; accelerator utilisation; domestic Chinese accelerator and HBM market shares; independently measured cost per completed enterprise task; enterprise model-switching frequency; agent-control-plane installed base; MCP and A2A production adoption; public-sector AI procurement values; grid-connection delays; and realised productivity gains from industrial AI deployment.
Those indicators, rather than chatbot download rankings or temporary benchmark leadership, should determine whether the central thesis is confirmed or falsified between 2026 and 2030.


















