---
title: "To determine where AI will bottom out, the token patterns actually provide the answer."
type: "Topics"
locale: "en"
url: "https://longbridge.com/en/topics/43025853.md"
description: "Why aren't all these positive catalysts working anymore? The gains this year have been driven by the self-reinforcing cycle of AI coding, which, like all AI businesses, is governed by the same token dynamics. To determine whether there's a trend in AI going forward, we still need to look at the underlying token patterns."
datetime: "2026-07-29T06:20:43.000Z"
locales:
  - [en](https://longbridge.com/en/topics/43025853.md)
  - [zh-CN](https://longbridge.com/zh-CN/topics/43025853.md)
  - [zh-HK](https://longbridge.com/zh-HK/topics/43025853.md)
author: "[sofa硬科技](https://longbridge.com/en/profiles/2081951097110011904.md)"
---

# To determine where AI will bottom out, the token patterns actually provide the answer.

No good news works anymore, all positive news is treated as short-seller dumping.

**Google's earnings report caused a market dump, Corning's earnings met expectations but the market dumped anyway, and the market is trembling at the major tech companies' earnings reports this week.** Any positive news is seen as negative news. The market sentiment in the fourth quarter of last year was very similar to now. At that time, OpenAI had just completed a large round of financing, and the industry conditions were very strong, but the stock price continued to trade sideways. Market participants argued about capital expenditures, ROI, valuations, and financing every day—**exactly the same as now.**

The conclusion from this quarter's earnings reports remains overwhelmingly positive. Google Cloud grew by 82%, validating the ROI of cloud computing investments; Intel significantly exceeded expectations; ASML, TSMC, and Intel all raised their order or capital expenditure outlooks. It can be inferred that these companies also saw extremely strong downstream forecasts—**so strong that even the most conservative participants in the supply chain have the confidence to make aggressive bets.**

**If this information had been released in May or June, semiconductor stocks would almost certainly have surged.**

However, right now, every earnings release has become an opportunity for shorts to re-evaluate valuations and long-term investment logic. What we may be seeing is that the AI market is no longer in that "AI Summer" of May and June. **The same positive progress provides less support for stock prices.**

Take Google Cloud's 82% growth as an example. The market's first reaction used to be: "AI demand exceeded expectations—the catalyst is here." Now, the first reaction is: "So what? What about 2028? Can OpenAI be profitable? How many more years can GPU prices rise? Depreciation is rising, financing costs are increasing, and the entire capital expenditure argument needs to be repriced."

**In essence, the market has shifted from trading the growth of AI capital expenditures to trading its sustainability and ROI.**

In our view, market sentiment has become extremely bearish. Even long-term bulls are starting to question whether AI semiconductor stocks can continue to rise, and it is almost impossible to hear any calls for new highs. **When the market shifts from looking for further upside potential to searching for additional downside risks, it usually means that pessimistic expectations have already been largely priced in.**

Nevertheless, we have no doubts about the fundamentals and believe that ultimately facts determine stock prices. So the question becomes: **What will make these stocks rise again?**

Under the above framework, **additional capital expenditure plans alone are no longer enough to convince the bears. What is needed is verification of a new demand curve. And the most powerful and direct catalyst will be the emergence of a hit product.**

**If "Coding 1.0" proved that AI can improve developer productivity, then "Coding 2.0" must demonstrate that AI agents can truly replace part of the software development process.** Once new productivity use cases are verified, the market's concerns about AI ROI may be redefined—**AI infrastructure spending will once again be viewed as "productivity investment" rather than "cost."**

But "hit products" are only half the answer. To see the other half clearly, we must first calibrate the indicator that has been most misread in this debate.

* * *

## I. First, Calibrate That Misread Indicator

This debate on "whether AI has peaked" was largely ignited by one chart—Silicon Data's **LLM Token Expenditure Index**.

The market treats it as a downstream validation indicator for whether AI hardware trading can continue: if token pricing reverses, then from memory trading to the entire hardware and data center trading, this round may be coming to an end. The worry is understandable—AI hardware investment ultimately relies on downstream model revenue for support.

But there is a key misreading here that determines the direction of all subsequent inferences:

**This index is not a token usage index, nor is it an AI total expenditure index, but a "usage-weighted average token price index."**

In plain terms, it tracks not "how much token everyone used," but "how much everyone is willing to pay per million tokens on average."

Therefore, a decline in the index cannot be directly read as a decline in AI demand or a peak in hardware demand. It more accurately reflects—**whether the market is still willing to pay a premium for the strongest models.** As long as the total volume of tokens continues to grow, even if the average unit price falls cyclically, actual compute consumption does not necessarily decrease.

To be fair: although a decline in the index does not prove that AI demand has peaked, **it does affect the revenue structure, gross margin, and capital expenditure payback expectations at the model layer\*\*—which is precisely why it deserves serious attention.**

## **II. The Pattern Emerged from the Index: A "Strong Model Premium Cycle"**

**Looking at the trajectory over the past few months, it actually corresponds to a very regular cycle:**

****End of 2025**, when new-generation models like Gemini 3 Pro, Claude Opus 4.5, and GPT-5.2 were just released, the index was still low—people hadn't started using them yet.**

****January 2026**, as enterprises and developers gradually adopted them, usage flowed toward more expensive closed-source frontier models, **the index rose rapidly**.**

****February to early March**, a batch of cheaper open-source or open-weight models such as Kimi, DeepSeek, Qwen, and GLM were released intensively, and some routine tasks migrated to lower-priced models, **causing the index to fall**.**

****March to May**, GPT-5.4, GPT-5.5, Opus 4.7, Opus 4.8, Gemini 3.1, and Grok 4.3 were launched successively, and the market paid again for stronger capabilities, **pushing the index back up, breaking through 2.0**.**

****Early June**, it fell again.**

**The pattern is therefore clear:**

> ****New model release → Capability ceiling raised → Market willing to pay higher unit prices → Models understood, capability boundaries defined → Routine tasks sink to cheaper models → Premium naturally fades → Next stronger model appears → Willingness to pay moves up again.****

****So every decline is not entirely caused by open-source substitution, nor does it necessarily mean a decline in AI demand.** It is a regular step in this cycle. The most reasonable interpretation of the drop in early June is a **cyclical premium pullback** after the market fully understood the GPT-5.5 / Opus 4.8 generation. Reading it directly as the end of the AI hardware cycle is inaccurate.**

## **III. But This Round Has Something New: ROI Discipline**

**If there were only the cycle mentioned above, things would be simple. The complexity of this round lies in a structural change that occurred simultaneously on the demand side—**token maxxing has entered the stage of ROI discipline.****

**The approach in the early first half of the year was straightforward: encourage employees to use AI as much as possible, treating token consumption as a signal of AI penetration rate and organizational transformation speed. A monthly token budget of $10,000–$30,000 for senior engineers in Silicon Valley is quite common; Jensen Huang even said that if an engineer with an annual salary of $500,000 doesn't burn half of it on tokens in a year, they should worry that they aren't using AI seriously; OpenAI's Frontier team is organized as "one researcher + multiple agents operating on the same codebase in parallel," with a single team consuming over 1 billion tokens per day; Meta created a "Claudeonomics" leaderboard covering 85,000 employees, where token consumption rankings directly went into layoff decisions.**

****After April, enterprises started to struggle.****

**Enterprises found that token usage increased quickly, and bills increased even faster, but how much real output it translated into was hard to answer. **Tokens can reflect usage intensity, but do not directly equal productivity\*\*—teams consuming more tokens do not necessarily deliver more code, fewer bugs, or higher revenue. Amazon's internal staff even created meaningless AI tasks to inflate usage: once tokens are turned into leaderboards or KPIs, they change from "observation indicators" to "numbers that can be manipulated."****

****Billing pressure quickly pushed the debate to the forefront. Concentrated cases emerged from late May to early June: **Uber burned through its entire 2026 AI coding budget in four months**, subsequently setting a $1,500 per person per month cap on tools like Claude Code and Cursor, requiring internal approval for overages; **Microsoft canceled some internal Claude Code licenses**, requiring the Experiences + Devices department to migrate back to GitHub Copilot CLI; **Salesforce's Anthropic annual bill was approximately $300 million**, and Benioff began emphasizing the use of intelligent routing to assign different tasks to models of different price points; Priceline's Cursor renewal quote increased 4–5 times; and another large company without a cap reached a monthly Claude bill of $500 million.****

******It's not that AI spending disappeared, but that AI has moved from unconstrained trials to needing more disciplined productivity budget management.******

****Management frameworks are also upgrading: Salesforce launched "Agentic Work Units," shifting the focus of measurement from "how many tokens were consumed" to "how much work the agent completed"; the Tokenomics Foundation, planned to be established under the Linux Foundation, aims to promote token metering, cost auditing, and cross-vendor standardization. The direction is clear—**tokes will not exit enterprise management, but will transform from a vanity usage metric into an auditable, budgetable, and optimizable cost metric.******

******The implication for model-layer revenue needs to be stated precisely: the key variable is not just "can the strongest model do it," but "how many tasks are worth doing with the strongest model."** High-value engineering, research, security, and complex investment research tasks are willing to bear higher token costs; ordinary workflows rely more on cheaper models, mature workflows, or model routing. **The strongest capability does not equal the optimal cost efficiency.******

****The supply side also acknowledges this point. At an enterprise event hosted by OpenAI in early June, Sam Altman stated bluntly that cost optimization has become a very important demand for enterprise customers. **The questions customers ask have changed from "Is the model any good?" to "Can these capabilities be used more cheaply, controllably, and audibly?"******

## ****IV. Fable 5 Is the Ceiling, But It Crashed Into the Eastern Giant's Open-Source Models****

****According to the cycle mentioned earlier, June should have ushered in a new round of upward movement.****

****On June 9, Anthropic released Claude Fable 5 and Mythos 5, which share the same underlying model. Their parameter count and training compute are significantly larger than the flagship Opus series, with comprehensive benchmark leadership (SWE-Bench Pro 80.3% vs. Opus 4.8's 69.2%), **and the longer and more complex the task, the greater the relative advantage.******

****Its push on the index has three aspects, all mechanical: **First, it is inherently more expensive**, with input at $10 per million tokens and output at $50 per million tokens, approximately twice that of Opus 4.8; **Second, the area where it widens the gap is not ordinary Q&A, but long-cycle coding, complex knowledge work, full repository migration, and multi-turn agent tasks**—these tasks have higher value, consume more tokens, and have a higher proportion of output tokens, while the pricing for output is significantly higher than for input; **Third, after the trial period ends on June 22, Fable 5 will exit the monthly package and switch to consuming usage quotas**, meaning the strongest capabilities will no longer be bundled into a subscription pool, but will be reflected directly into the price index on a pay-per-use basis.****

****(The regulatory disturbance in between should also be clarified: Fable 5 / Mythos 5 were suspended from access on June 12 due to U.S. Department of Commerce export controls, the controls were lifted on June 30, and service resumed on July 1. It disrupted the ramp-up rhythm of high-priced models, but appears to be a one-time event that does not change the long-term narrative.)****

******The problem is that this move to raise the ceiling crashed into a reverse force that had never been so dense before.******

****On July 16, Moonshot AI released Kimi K3—a 2.8 trillion parameter MoE model activating 16 experts out of 896 per token, scoring 57.1 in independent evaluations, **ranking third globally, behind only Fable 5 (59.9) and GPT-5.6 Sol (58.9)**, topping the coding leaderboard at one point, with **full weights scheduled to be open-sourced on July 27**. Over the same weekend, Alibaba launched Qwen3.8-Max-Preview with 2.4 trillion parameters, also planning to open weights. Looking further back, DeepSeek, GLM, and MiniMax are all on the same track.****

******This compresses the originally "sequential" cycle into a "simultaneous" one.******

****The past rhythm had a sequence: new model launch → premium rises → gets understood → routine tasks sink to cheaper models → premium falls → next stronger model. The middle period of "premium rise" could last for several months.****

******Now, raising the ceiling and dismantling the tollbooth are happening in parallel.** Fable 5 raised the upper limit in June, and K3 turned a quasi-frontier alternative into a free file in July. To summarize this mechanism in one sentence: **open source turns models from "services" into "files"—and services have pricing power, while files only have distribution.** Once weights are public, pricing power shifts from "the author of the model" to "whoever can run it cheapest." The estimated gross margin of over 80% on closed-source APIs shrinks to only 40–45% on the path of open source plus hosting.****

******The half-life of premiums is therefore shrinking 急剧 ly.** The previous open-source iteration cycle caught up 13 index points within three months, leaving only a 2.8-point gap; some estimates suggest the best open-source model lags behind the frontier by only about four months. **Frontier premiums used to depreciate annually, but now they depreciate by release cycle.******

****Usage data is also moving in the same direction: On OpenRouter, the proportion of calls with a unit price below $1/million tokens rose from 18% to 41% over five months; the token share of Chinese open-source models rose from less than 2% to about 61% over 18 months, with all open-source weight models accounting for about 69%; the token share of U.S. labs dropped from about 70% to about 30% in one year.****

******This is the true difference between this round and previous cycles, and also the deep reason why the market no longer buys good earnings reports:** If every move to raise the ceiling is caught up by a free file within a few months, then the credibility of the chain "stronger models → higher unit prices → more model-layer revenue → supporting upstream capital expenditures" will be whittled down with each round. Shorts have never questioned the demand of this quarter, but **how many more times this chain can repeat.******

### ****But: Open Source Caps the Ceiling of Capability Premium, Not Everything****

******First, money still sits at the frontier.** Anthropic processes about 11% of tokens but takes home an estimated 42% of platform spending. A 25x price difference is still being paid. **Because enterprises don't pay for benchmark scores—they pay for "an agent that doesn't crash halfway through a fifty-step task," for reliability where failure costs are high, and for not having to re-verify everything on a new model.******

******Second, the strongest open-source models are not refined enough.** K3 burned 132 million output tokens running the full evaluation, more than double the median, because it comes with always-on maximum reasoning and lacks reasoning intensity control; the result is a single-task cost of $0.94—cheaper than Opus, but 71% more expensive than GPT-5.6 Terra ($0.55) and three times that of Grok 4.5 ($0.31); the hallucination rate also worsened from 39% in the previous generation to 51%. **So it is a price inflection point, not a true critical moment—the true critical point is a frontier model that is both open-source and token-efficient, which has not yet appeared.******

******Third, the relationship between China and the U.S. is a zigzag line, not a substitution curve.** The U.S. releases new models, widening the gap; Chinese labs catch up in the following months, narrowing the gap; the next U.S. model arrives. **The investable landing point is therefore simple: China and open-source models can cap the price ceiling for "near-frontier intelligence," even if the U.S. maintains absolute frontier status.** Because most commercial workflows need "good enough and appropriately priced," rarely needing the single strongest on Earth.****

******So what open source truly compresses is not the gross margin of frontier labs, but their mid-tier transaction volume.** Tasks with clear demand—document classification, extraction, summarization, simple coding—go first; what remains are the hardest tasks where capability differences are valuable. **The frontier price holds, but the mid-tier is being drained.******

******Putting these two forces together, the shape of the index emerges: it will still rise with new model releases, but the upward segments will become shorter and steeper, and the declines will come faster and faster.** Volatility will not stop, but the duration of the top segment is being compressed—and capital expenditures require exactly "duration."****

## ****V. This Year's Bull Market is Essentially the Self-Cycle of AI Coding****

****To judge when this adjustment will bottom out, we must first clarify what drove the rally this year.****

****The real change in the first half of 2026 was not that a certain model's raw capability took another step up, but that **agent systems began entering a self-iteration phase.******

****From the perspective of "what exactly is being scaled," AI progress has undergone four migrations: the first scaled pre-training data, parameters, and training compute; the second was reinforcement learning and verifiable rewards; the third was inference-time compute; **and the fourth, currently happening, scales the task environment, toolchain, and feedback loop in which agents operate.******

****This is the **Agent Scaling Law**. It is not a denial of old paradigms, but a natural extension—**the object of scaling has shifted from the model itself to the system in which the model operates.******

****The engineering premise supporting it is called **harness engineering**: building a real working environment for the models, doing three things—constructing the task environment, designing feedback mechanisms, and forming data loops, allowing agents to iterate and accumulate experience through trial and error. **Without an environment, a model can at best chat; with an environment, it starts working.******

****Four factors determine how fast a field's capabilities improve: whether tasks can be automatically verified, whether labs have made dedicated training investments, whether training data is sufficient, and whether users are willing to pay for such capabilities. **Fields that meet all four criteria see the largest and fastest capability jumps—and coding is a textbook case:** rewards are direct, goals are clear, auto-testable, results verifiable, environments simulatable, and data generatable.****

******This is why coding is the first scenario to run end-to-end successfully, and why this year's market rally is essentially driven by the "AI coding self-cycle." Then came the hardware facilities at the beginning of the year, and various hardware named after "bottlenecks" in March and April, including optics, 800V power supplies, PCBs, packaging and testing, glass bridges, etc. In fact, without the surge in token usage and confidence brought by coding, this wave in the first half of this year would not have happened; doubts about AI capital expenditures had already begun at the end of last year.******

## ****VI. So, When Will It Bottom Out?****

******This round is not a decline in demand, but a repricing of "is this money worth it, and when will it pay back?"** And it is particularly stubborn because of the new change mentioned in Section IV: **raising the ceiling and dismantling the floor are now happening simultaneously, and the half-life of frontier premiums has shrunk from "annual" to "by release cycle."** Against this backdrop, every piece of positive earnings data has become an opportunity for shorts to re-evaluate valuations—not because the data is bad, but because the market has shifted from trading growth to trading sustainability.****

****In this context, **any amount of capital expenditure plans does not constitute a catalyst.** Spending an extra two hundred billion only makes the question "how long until payback" louder.****

******What can truly end this adjustment is the verification of a new demand curve, and the most direct form is the emergence of a hit product.******

******If "Coding 1.0" proved that AI can improve developer productivity, then "Coding 2.0" must demonstrate that AI agents can truly replace part of the software development process.******

****The difference between these two things is fundamental. The ceiling of the former is the number of people—better tools and improved efficiency correspond to software budgets; the ceiling of the latter is the total volume of work itself—part of the work no longer requires humans, corresponding to labor budgets. **These are two pools of completely different magnitudes.******

******And it happens to bypass the ceiling capped by open source.** Open source suppresses "capability premiums"—for the same task, open-source models can do it pretty well, so you shouldn't pay extra. But Coding 2.0 sells not capability, but **reliability and replacement**: an agent that won't crash halfway through a fifty-step task, can truly be handed off, and bears the cost of failure. **This is precisely the reason why frontier models today take home 42% of spending with only 11% of tokens, and also the place where open source has yet to catch up\*\*—K3's 51% hallucination rate and double token consumption illustrate exactly this point.******

********So open source caps the ceiling of capability, but not the ceiling of reliability; and Coding 2.0 is precisely the step that changes the payment rationale from "capability" to "replacement."********

****Following this line of thought, the paths to end this adjustment are actually more than one.****

******The first is the successful run of the next self-cycle.** Coding succeeded first because it naturally satisfies three conditions: verifiable rewards, synthetic data, and self-iterating systems. Where is the next candidate? The feedback flywheel of coding agents is already spilling over to broader white-collar work; Codex evolving from a programmer tool to a knowledge work entry point is a sign; one step further out is physical AI—robots, autonomous driving, embodied intelligence. Once any scenario puts together the "environment + feedback + data loop" trio, it will be the second coding. The market is waiting for this signal.****

******The second path is closer, and already happening: the profit models of cloud providers themselves may see hope first.******

****This account is worth calculating in detail. Selling open-source model tokens is fundamentally different from the two types of business cloud providers have done in the past.****

****Compared to reselling frontier model tokens, **it saves a huge chunk of royalty expenses\*\*—on the old path, the bulk of what customers paid went to the frontier model vendors, with cloud providers only eating the underlying compute rental. With open-source weights being free, this royalty channel disappears directly.******

******Compared to simply renting out GPUs, **it adds a layer of premium.** Bare-metal GPU rental is the least differentiated business; but around reselling tokens for open-source models, cloud providers can stack a whole suite of things to charge for: cost advantages of self-developed ASICs, managed inference services, databases and vector retrieval, object storage, networking and data transmission, Agent runtime, security and identity, evaluation log monitoring, and enterprise support services.******

******In essence, this is commoditizing the model layer and then shifting value to the infrastructure-plus-platform-service layer—hosting, optimization, routing, governance, security. Gross margins will naturally increase significantly.** Open-source models are not a threat to cloud providers, but one of the most effective paths to improve the profit margins of their AI businesses.****

****A key fact in terms of timing: **super-large cloud providers only began large-scale deployment of open-source models in mid-to-late June, so the data is not yet sufficient.** So this logic is currently an inference, not a verification. But the verification point is clear—**if earnings reports in the following quarters confirm that deploying open-source models yields better margins and faster payback periods, then the sharpest skepticism from shorts—"is capital expenditure actually worth it?"—will get answered on the cloud providers' own income statements.******

******This is likely the main driving logic for the next wave of the bull market.******

****Until then, the index will continue to swing up and down within the "strong model premium" cycle, and the upward segments will get shorter and shorter each time; good earnings reports will continue to become opportunities for shorts. **The end of this decline is not in the capital expenditure numbers of the next earnings report—it is either in the next scenario that successfully runs a self-cycle, or in the gross margin of cloud providers selling open-source tokens. Current market sentiment is so poor that it must find new hope from these two sources.******

### Related Stocks

- [GOOG.US](https://longbridge.com/en/quote/GOOG.US.md)
- [GOOGL.US](https://longbridge.com/en/quote/GOOGL.US.md)
- [INTC.US](https://longbridge.com/en/quote/INTC.US.md)
- [ASML.US](https://longbridge.com/en/quote/ASML.US.md)
- [TSM.US](https://longbridge.com/en/quote/TSM.US.md)
- [GOOGN.US](https://longbridge.com/en/quote/GOOGN.US.md)
- [04335.HK](https://longbridge.com/en/quote/04335.HK.md)

## Comments (7)

- **摸鱼炒股达人 · 2026-07-30T03:32:16.000Z**: Does anyone know how to place an order? I've tried many times but it doesn't work.
  - **躺平坐等回本** (2026-07-30T03:32:33.000Z): Overseas SIM card or proxy, choose one of the two
- **不会辜负你的只有算法 · 2026-07-29T15:17:57.000Z**: If this article was written with the help of AI, then...
- **大魔王一刻 · 2026-07-29T15:06:22.000Z · 👍 1**: Would the vast majority of companies be willing to keep using tokens indefinitely, or would larger companies deploy their own local models for employees? That might be more cost-effective.
  - **Icloud** (2026-07-29T15:26:44.000Z): However, the server room still uses that of a major tech company; local models are likely mainly for security and specialization.
- **价值滑头 · 2026-07-29T07:00:59.000Z**: So much information, feels like an industry insider
- **桑不起666 · 2026-07-29T06:42:22.000Z · 👍 1**: The logic is clear, and the interpretation is reasonable.
