A one-off shock, or a repeatable engine outside the United States—and where does profit accrue across the AI stack?
Alvin Lau
Equity Research Analyst
Across roughly eight weeks in mid-2026, five Chinese developers shipped six frontier-adjacent releases. That significance is not any single model, but the pace of delivery. Where January 2025 delivered a single, cost-led jolt with DeepSeek-R1, the 2026 sequence spans multiple firms and reaches beyond price into capability and scale. These include:
On the evidence, this looks structural, not a repeat surprise: China now has an industrialized pipeline that produces near-frontier models on a recurring basis under real compute limits.
Four developments may help explain what’s changing:
The clearest beneficiaries remain the companies that provide the infrastructure behind AI, where competitive advantages appear more durable than they do for model developers. Hyperscalers gain twice, from rising aggregate demand and from hosting open models and billing by usage. Model developers face the biggest challenge, as fast-rotating leadership implies a thinner moat and makes it harder to justify premium valuations. Enterprises and software deployers, by contrast, should keep a rising share of the surplus.
Four risks could challenge this view:
This is more than a rerun of the DeepSeek shock, but not yet proof that Chinese and US frontier models are on equal footing. The lead has narrowed, not disappeared. The clearest beneficiaries remain the companies that provide the compute and infrastructure, along with the enterprises best positioned to turn lower-cost intelligence into measurable productivity.
DeepSeek’s January 2025 shock was, in essence, about cost. It came from one company, made model training much cheaper by relying on reinforcement learning rather than supervised fine-tuning, and showed capable models could originate outside the US even if it still sat behind the frontier. Markets read it through Jevon’s lens: Microsoft’s Satya Nadella argued cheaper AI would send usage soaring, and Chinese equities rallied.
The phenomenon where increased efficiency in resource use leads to higher overall consumption of that resource.
The 2026 wave differs. In roughly two months, five firms delivered six near-frontier releases (DeepSeek shipped twice) across distinct architectures and modalities: Kimi K3, the open-weight capability leader; Alibaba’s enterprise-focused Qwen3.8-Max; DeepSeek’s ultra-cheap V4-Flash and, weeks later, its frontier-efficient V4 Pro; GLM-5.2, briefly the top open model; and ByteDance’s video leader Seedance (Figure 1).
Figure 1: Major Chinese AI model launches, June-August 2026
| Date | Model (developer) | Significance |
|---|---|---|
| June 16, 2026 | GLM-5.2 (Z.ai) | Opened the wave; briefly the top open-source model |
| July 16, 2026 | Kimi K3 (Moonshot) | Open-weight capability leader; sign-ups paused when GPUs ran out |
| July 31, 2026 | V4-Flash (DeepSeek) | Pricing breakthrough—a full benchmark for ~US$0.03; open weights on release |
| August 3, 2026 | Qwen3.8-Max (Alibaba) | 2.4T-param MoE, 1M context; enterprise / long-horizon work |
| Early August, 2026 | Seedance 2.5 (ByteDance) | Took the lead in AI video at a steep discount |
| August 13, 2026 | V4 Pro (DeepSeek) | Flagship on the efficiency frontier—GLM-5.2-level, far cheaper |
Source: State Street Investment Strategy and Research; company releases; Artificial Analysis. As of August 13, 2026. China’s 2026 model wave—six releases from five firms (the V4 family previewed in April 2026).
What really matters is not whether China surprised everyone once with a breakthrough, but whether it has built a system that can keep producing strong models repeatedly. And even people who were initially doubtful now increasingly agree that's what we're seeing.
Figure 2: From DeepSeek to the China AI wave
| Dimension | DeepSeek, January 2025 | China wave, June – August 2026 |
|---|---|---|
| Actor | One company, one event | Five firms, six releases (~8 weeks) |
| Nature | A cost breakthrough | Capability and scale, not just price |
| Interpretation | A statistical outlier | A repeatable production system |
| Frontier gap | Clearly behind the US | Competitive in some areas; US retains the lead overall |
Source: State Street Investment Strategy and Research, Artificial Analysis. As of August 13, 2026.
The evidence points to a maturing system, defined by breadth of firms and architectures, compressed development cycles, and adoption at production scale. The caveat is that “near-frontier” is not “frontier.” The most capable systems still lead. The gap has narrowed rather than closed, but the system now looks built to last.
Who actually keeps the economics of AI? Using NVIDIA President and CEO Jensen Huang’s “AI cake” analogy, the ecosystem reads as a series of layers (Figure 3). The key question is where profit accrues and where competitive pressures are most intense.
Figure 3: The AI value stack—Jensen Huang’s “AI cake”
Compute is the scarce, hard-to-replace base; the model layer is where competition is fiercest
At the base, hardware and infrastructure is where scarcity pricing and healthy margins live. Asia anchors it, supplying most leading-edge AI silicon so Western build-outs rest on Asian hardware, and its leaders (NVIDIA, TSMC, SK Hynix, Samsung, Broadcom) are hard to unseat.
Hyperscalers turn capital into capacity and can also host open models, metering by consumption. The model layer is the contested tier, split between closed models (OpenAI, Anthropic, xAI) and open-weight models (DeepSeek, Qwen, Moonshot, Z.ai, as well as Meta, and NVIDIA). The difference is how buyers obtain and pay for each: a closed model is rented through a paid, metered API, so the provider controls access and pricing, whereas an open-weight model publishes its parameters to download, fine-tune, and self-host, far cheaper at scale.
At the top is the application and enterprise layer, which captures the productivity dividend. In short, profit concentrates at the base while the surplus drifts to those who deploy.
Open models reshape how the market perceives AI economics, and adoption has followed. On estimates cited by a16z and the US–China review commission, roughly 80% of US startups built on open-source now run on Chinese models such as Qwen rather than OpenAI or Anthropic, with Airbnb among those running Qwen in production.4 The cost of a given capability has fallen just as sharply: Epoch AI tracks inference prices dropping up to several hundred-fold a year, and a GPT-3.5 equivalent is now two orders of magnitude cheaper to run than in late 2022 (Figure 4). Throughout this article we refer to that all-in cost of getting a task done as the price of intelligence.5
The all-in cost of completing a defined unit of work with an AI system, not the headline price per million tokens. It is the compute a task consumes, set by model architecture and deployment, multiplied by the price of that compute, set by accelerators, memory, and power. Because open competition compresses the first term while scarcity holds up the second, the price of intelligence falls quickly but has a floor beneath it.
Figure 4: Repricing intelligence
Capability vs. cost, mid-2026
The logic is sound: to win developers, a model must undercut the cheapest credible option or clearly outperform it.
Two caveats keep this from tipping into an outright open-model story:
Falling prices of intelligence have a floor. The binding limit has moved from model capability to compute availability. GPU rental prices remain elevated. DeepSeek has raised its API prices,6 citing capacity limits, and is building gigawatt-scale capacity. High bandwidth memory (HBM), the AI accelerators’ specialized memory, is the choke point, concentrated among a few suppliers. Data-center power use is set to more than double by 2030 as hyperscaler budgets climb toward half a trillion dollars.7
This is self-reinforcing. By Jevon’s logic, lower inference costs expand demand faster than they reduce cost per task, driving higher use of tokens, compute, and power. The counterpoint is also real, because Jevon’s explains demand, not unit economics: if returns lag, or if much demand is driven more by ecosystem funding rather than end-user adoption, the demand elasticity could deflate as fast as it inflated.
The combination of intensifying model-layer competition, rising AI usage, and persistent compute scarcity implies a lasting shift in pricing power toward the compute layer, potentially benefiting infrastructure providers and certain hyperscalers while squeezing standalone model economics.
Two caveats could alter this outcome:
With that floor price of intelligence in place, the price of intelligence trades within a band, bounded on each side.
The first—OpenAI’s Luna price cut—set a ceiling. On July 30th, days after Kimi K3’s weights became a free download, OpenAI cut GPT-5.6 Luna by roughly 80% yet left the flagship Sol untouched.8 Once a capable model can be deployed for free, no one can charge a premium for the same capability, and everything beneath the frontier is dragged toward the open-weight line.
The second move—DeepSeek’s V4 price increase—confirmed the floor.
Two weeks later, DeepSeek, which had started the price war, raised its V4 rates by 50 to 1,100%, citing the same compute strain.9 Prices cannot fall without limit, because cheap intelligence still rides on scarce, expensive hardware.
Figure 5: A ceiling and a floor—the intelligence price band
Open-weight competition caps the price of a given capability from above, while scarce compute holds it up from below. The ceiling pushes the model layer toward utility-like economics and compresses lab margins (See Risks and opportunities across the AI stack). The floor accrues to whoever owns compute, memory, and power.
In the three possible scenarios we consider, the decisive variable is how models are accessed and paid for. The compute-layer beneficiaries look similar across all three.
Whoever builds the computing backbone—chips, networking, and power—stays well-positioned no matter which paradigm prevails. The more durable approach is to invest in the layer that wins across outcomes.
The scenarios share one constant: the compute base, making it the natural starting point to begin a layer-by-layer assessment of risk and opportunity across the stack.
Hardware and infrastructure companies appear to be the clearest structural winners, combining rising aggregate demand with positions defended by capital intensity and know-how. Model leadership can turn over in weeks, leadership in advanced logic, HBM, and lithography is built over years. This layer owns the “floor” of open models pricing, and its pricing power is the stickiest in the stack, though still cyclical. China’s shift toward home-grown chips (Huawei Ascend, Cambricon) supports China-listed silicon even as it complicates the Western supply chain.
Open models help hyperscaler economics: providers not only sell compute capacity but also host open models and monetize usage, generating higher-value revenue streams. The near-term risk runs the other way: much of today’s compute demand is anchored by a few frontier labs training and serving their models, so if open competition compresses their monetization, that spend and hyperscaler backlog could cool. The decisive variable is the demand mix, because broad-based inference and enterprise workloads are more resilient than lumpy frontier-lab contracts, so watch it as closely as you do headline capital expenditures.
The labs face the hardest questions, because they sit directly beneath the ceiling of open models. With a freely deployable model now on the frontier, no lab can charge a premium for a capability available cheaply elsewhere. Pricing power survives only at the very frontier, and only for as long as that lead lasts. The flywheel that sustained the labs is now under pressure at every turn (Figure 6).
Figure 6: How open-weight competition can disrupt the traditional AI flywheel
| The traditional AI-lab flywheel | Where open-weight competition breaks it |
|---|---|
| Lead at the frontier | Leadership now turns over in weeks, not years |
| Charge a premium, high margins | Open rivals drag price toward marginal cost |
| Fund more compute and talent | The capital edge fades as “good enough” spreads |
| Command an even higher premium | The loop weakens; monetization and valuation wobble |
Source: State Street Investment Strategy and Research.
For OpenAI and Anthropic—both of which have filed to go public—those valuations lean on high margins that open competition now tests.12 This outcome may not be inevitable because frontier labs retain real advantages in capability, safety, brand, and integration. The question is whether current valuations adequately reflect the durability of those advantages and moat assumptions.
The flip side is opportunity for deployers. Cheaper inference and lower barriers push value toward the enterprises and software firms that operationalize AI, and teams already use a premium model to plan and structure a task, then pass the execution steps to a cheaper open model.
China’s progress highlights a US dilemma: spread the American AI stack so the world standardizes on it, or slow China through export controls. The two aims conflict because tightening controls nudges third-market customers toward cheaper Chinese options like Huawei, whereas loosening them weakens the grip on China’s compute access. The recent progress that came under tight compute constraints complicates the picture because it challenges the assumption that withholding compute reliably limits capability.
The commercial stakes are clear. When one side charges pennies for what the other prices in dollars, the contest for undecided markets shifts, and friction is already spilling into intellectual property disputes and sanction risk. Each tightening also accelerates China’s domestic-silicon push and could split the ecosystem into rival standards. Policy is now a first-order driver of value capture and a key swing factor for the open vs. closed model scenarios. Policy outcomes are hard to predict, but China’s AI ecosystem looks set to become an increasingly important part of the global landscape, which argues for a broad perspective on AI opportunities.
Open models are not unique to China either. US players such as Meta’s Llama and NVIDIA’s Nemotron are active too. A coalition of over thirty US firms formed the Open Secure AI Alliance,13 an aspiration that open models curb the risk of a few gatekeepers controlling access to intelligence.
Is this another DeepSeek moment? Our answer is a qualified no: this may be more than another DeepSeek moment, yet not proof that Chinese and US frontier models are on equal footing. A single, cost-led surprise has become a multi-company run that pairs capability with scale and is backed by production-scale adoption, so China now has a repeatable engine for near-frontier models despite compute constraints.
The evidence supports a durable system, value reallocated within the stack rather than destroyed, and compute as the binding constraint whose winners hold across outcomes. Three questions remain open: whether China closes the last capability gap, whether cheaper intelligence converts into validated returns, and how the US–China policy contest resolves.
Several indicators will help determine whether China’s recent AI progress marks a durable shift in industry economics and competitive positioning, or a temporary narrowing of the gap.
We claim no certainty on the model-layer winner. The higher-conviction view may be the compute-and-infrastructure base—the layer that benefits in every scenario—plus the hyperscalers and enterprise deployers best able to monetize cheap, abundant intelligence. The model layer offers the most optionality and the most risk, so size it accordingly.