---
title: "AI Inference Outlook: From Infrastructure Buildout to Commercial Monetization"
type: "News"
locale: "en"
url: "https://longbridge.com/en/news/295313493.md"
description: "The investment cycle in AI infrastructure may not have peaked yet. Citi believes that supply constraints in power, permits, labor, HBM, and interconnects will extend the originally projected three-year capacity expansion to five to ten years"
datetime: "2026-08-09T03:54:00.000Z"
locales:
  - [zh-CN](https://longbridge.com/zh-CN/news/295313493.md)
  - [en](https://longbridge.com/en/news/295313493.md)
  - [zh-HK](https://longbridge.com/zh-HK/news/295313493.md)
---

# AI Inference Outlook: From Infrastructure Buildout to Commercial Monetization

The main narrative of AI infrastructure is shifting from "buildout" to "monetization," but this does not signal a peak in capital expenditure.

The cost of AI inference tokens **plummeted by 37.5% from its May peak to $1.33 per million tokens**, while Blackwell GPU rental rates **rose counter-trend by 15.2% to $5.18 per hour** during the same period. With rising prices on the supply side and falling prices on the demand side, the AI industry is entering a "scissors gap" phase of the inference economy. **Triple constraints in power, permits, and labor are extending the capacity expansion, originally expected to take three years, to five to ten years**—construction is far from over, but the path to monetization has opened up.

In his latest weekly AI industry tracking report, Citi analyst Heath Terry judged that intelligent routing is the core driver behind the drop in token prices—enterprises no longer uniformly call upon the most powerful models but automatically schedule tasks based on complexity, systematically compressing inference expenses. However, the supply side presents a completely different picture: in the second quarter, hyperscalers' AI **return on capital expenditure reached 28%**, nearly five times the cost of financing (approximately 6%). Even as both US political parties rarely united this week to push for bans on data center construction, Google and SpaceX accelerated infrastructure expansion after their earnings releases, and Amazon **added $20 billion in capital expenditure**.

In terms of model competition, the global intelligence gap between open-source and closed-source models has narrowed to just 4 points, but this narrowing was primarily driven by Chinese vendors—**a 21-point chasm remains between leading US closed-source and leading US open-source models**. The optimization of inference speed in frontier models far exceeds the improvement in intelligence scores, and safety governance is evolving from a compliance issue into a commercial credential. Spending has not decreased, but the manner of spending is changing; the entire industry's focus is shifting from "how much computing power to build" to "how to turn computing power into revenue."

## Power Bottlenecks: First to Connect, First to Monetize

A statewide audit in Texas has frozen an approval queue of over **1,800 projects totaling 474 GW**, with data centers accounting for about 90% of new power applications. Power is replacing chips as the most urgent bottleneck in AI infrastructure.

Core builders state that new investments can recoup costs within one year. Amazon's additional $20 billion in capital expenditure was mainly driven by rising storage costs—a factor cited by many companies. AWS contract demand is booked through 2028, and management expects capacity tightness to last until 2027. Those who can secure definite power connection timelines earlier will command higher premiums for their sites and realize revenue faster.

With the triple constraints of power, permits, and labor overlapping, the capacity expansion originally expected to be completed in three years will likely stretch to five to ten years. For suppliers, high pricing for scarce infrastructure can be sustained for longer; the trade-off is that some revenue will be deferred to later cycles.

## Inference Becomes "Heavier": Storage and Interconnect Tightness Has Not Peaked

Over the past three weeks, the average output tokens per inference task rose by 11% to 24,000, with the tail end of inference-intensive tasks increasing by 9% to 40,000. **The share of output tokens rose to 41%**, cache activity dropped to 57%, and input tokens remained stable at 2%.

Each model call consumes more computing resources. Without breakthroughs at the architectural level, workloads dominated by output and caching will continue to drive up demand for HBM and interconnect chips. Memory and interconnects remain key bottlenecks, which is one of the underlying reasons for Amazon's increased capital expenditure.

## Global Open-Source Gap Narrows to 4 Points, US-China Open-Source Gap Stands at 21 Points

The Artificial Analysis Intelligence Index shows that the leading closed-source model (Anthropic Claude Opus 5) scored 61, while the leading open-source model (Moonshot AI Kimi K3) scored 57, **narrowing the gap from the previous 9 points to 4 points**.

However, the significance of these 4 points needs to be dissected. The global open-source sector catching up to closed-source is mainly driven by Chinese vendors—Kimi K3 (57 points), Zhipu GLM-5.2 (51 points), and DeepSeek V4 Flash (50 points) rank among the top open-source models. **The gap between leading US closed-source and leading US open-source models remains at 21 points**. Open-source models are exempt from pre-release safety testing, which may accelerate catch-up in the short term, but the asymmetry in open-source strength between China and the US is a deeper structural variable.

Optimization of inference speed is far outpacing intelligence improvements: the median inference speed among the top 20 vendors rebounded to 118 tokens/second, **up 55.3% week-over-week**, while the median intelligence score held steady at 43 points. The industry focus has shifted from "smarter" to "faster."

Pricing divergence is severe. The average price for mixed US-Europe tokens is $1.63 per million tokens, **while in China it is only $0.80, less than half the former**. DeepSeek V4 Flash and Xiaomi MiMo-V2.5-Pro are priced at $0.03 per million tokens, virtually zero cost. The average price for frontier models is $1.30, down 3% week-over-week.

A window for intensive new model releases is about to open: in the next six months, DeepSeek has scheduled V4.1 and V4.2, Google has lined up four versions of Gemini from 3.5 Pro to 4.2, and SpaceX has scheduled four generations of Grok from 4.6 to 5.1. A new wave of frontier models may widen the gap again, but challengers are also accelerating their pace.

## Safety Governance Is Becoming a Commercial Barrier

AISI recorded **19 instances of unauthorized real-time internet activity** in 122 model evaluations. The stronger the capability, the more real the risk of loss of control—this is not theoretical speculation.

In the long run, government safety reviews may become a competitive barrier for frontier closed-source vendors. AI-driven cyberattacks are becoming increasingly frequent and complex, and being "government-certified" is itself a commercial qualification for serving regulated enterprise clients.

The ceiling for frontier capabilities is still far off. OpenAI disclosed that an unreleased model generated **10 mathematical breakthroughs at an API cost of $2,000**. The upper limit of capability has not been touched, but the threshold for invocation is rapidly decreasing.

## Developers Accelerate Entry, Layoffs Temporarily Recede

Downloads of multi-provider SDKs **increased by 12.7% week-over-week**, showing clear acceleration momentum. On the application side, Tongyi Qianwen's weekly active users grew by 2.5%, Gemini by 0.5%, and Kimi declined by 0.8%, indicating that the traffic landscape is still changing rapidly.

In July, AI-driven layoffs totaled 3,220, **down 85% from the previous month's 21,640**. However, layoff data is highly volatile, and the proportion of AI-driven layoffs in total layoffs has not shown a trend decline. When the next macro adjustment arrives, the accelerating effect of AI substitution may amplify again.

### Related Stocks

- [C.US](https://longbridge.com/en/quote/C.US.md)
- [GOOGL.US](https://longbridge.com/en/quote/GOOGL.US.md)
- [GOOG.US](https://longbridge.com/en/quote/GOOG.US.md)
- [AMZN.US](https://longbridge.com/en/quote/AMZN.US.md)
- [OpenAI.NA](https://longbridge.com/en/quote/OpenAI.NA.md)
- [C-R.US](https://longbridge.com/en/quote/C-R.US.md)

## Related News & Research

- [I build AI data centers. I want them to disappear from view.](https://longbridge.com/en/news/295319962.md)
- [The AI bros are lonely](https://longbridge.com/en/news/294942232.md)
- [QIA joins SambaNova USD 1 billion Series F funding round](https://longbridge.com/en/news/295031818.md)
- [Influencers fear the AI 'Scarlet Letter'](https://longbridge.com/en/news/295206942.md)
- [Atlas Cloud Introduces Unified AI Inference Platform to Simplify Multi-Model Development for Engineering Teams](https://longbridge.com/en/news/295269435.md)