---
title: "SemiAnalysis: Kimi K3's KDA Mechanism Boosts Attention Efficiency but Will Require More GPUs, HBM, DRAM, and Networking, Not Less!"
type: "News"
locale: "en"
url: "https://longbridge.com/en/news/293114555.md"
description: "SemiAnalysis points out that while Moonshot AI's Kimi K3 employs a linear attention mechanism to alleviate bandwidth pressure, its scale of over 2.8 trillion parameters and large-scale expansion architecture still necessitate substantial high-end GPUs, HBM, and DRAM. Efficient inference will drive application adoption, thereby reinforcing long-term demand for AI hardware and disproving market concerns"
datetime: "2026-07-19T06:35:30.000Z"
locales:
  - [zh-CN](https://longbridge.com/zh-CN/news/293114555.md)
  - [en](https://longbridge.com/en/news/293114555.md)
  - [zh-HK](https://longbridge.com/zh-HK/news/293114555.md)
---

# SemiAnalysis: Kimi K3's KDA Mechanism Boosts Attention Efficiency but Will Require More GPUs, HBM, DRAM, and Networking, Not Less!

Moonshot AI's large language model, Kimi K3, adopts a linear attention mechanism, sparking market concerns that demand for NVIDIA chips, HBM, and networking equipment might weaken.

However, semiconductor research firm SemiAnalysis recently offered a starkly different assessment: **K3's massive parameter scale and inference architecture requirements will not only fail to dampen demand for high-end AI hardware but may instead further strengthen the need for NVIDIA's premium GPUs, HBM, and high-speed interconnect devices.**

SemiAnalysis noted that K3 boasts over 2.8 trillion parameters, with model weights requiring more than 1.5TB of HBM capacity. Even in scenarios with relatively limited concurrent users, KV caches still need to be heavily offloaded to CPU DDR5 memory and NVMe storage, leaving no significant surplus in HBM space.

More importantly, Moonshot AI previously revealed that efficient inference deployment of K3 requires a large-scale expansion domain architecture comprising at least 64 chips. This hardware requirement aligns closely with the design direction of rack-level AI systems such as NVIDIA's GB200/GB300 NVL72.

SemiAnalysis believes that the market's previous logic—interpreting linear attention as a factor that "weakens GPU demand"—was flawed. The actual impact may be quite the opposite: **more efficient model architectures lower AI inference costs, which will drive broader application adoption and, in turn, stimulate long-term demand for GPUs, HBM, DRAM, and network infrastructure.**

## NVIDIA Chip Demand Remains Robust Despite Linear Attention Iterations

Market concerns primarily stem from the Kimi Delta Attention (KDA) mechanism adopted by Kimi K3.

Compared to traditional Transformer attention mechanisms, KDA can significantly reduce data transmission requirements for KV caches, cutting network bandwidth pressure by up to 10 times. This development led some investors to recall the concerns about AI hardware demand following the release of DeepSeek R1, fearing that improved model efficiency might reduce reliance on high-end computing hardware.

**However, SemiAnalysis argues that this judgment overlooks another core requirement in large model inference: the computational and interconnect pressures driven by parameter scale.**

K3 features over 2.8 trillion parameters, meaning its model weights alone rely on large-scale distributed computing systems for deployment. Meanwhile, K3 employs a WideEP (Wide Expert Parallelism) optimization strategy, distributing 896 expert modules across multiple GPUs. This allows each single GPU to carry only a portion of the expert weights, thereby improving computational utilization.

Nevertheless, WideEP introduces new challenges: **frequent data exchange between experts requires stronger network interconnect capabilities.** SemiAnalysis pointed out that the copper backplane interconnect architecture used in GB200/GB300 NVL72 provides intra-rack bandwidth 18 times that of traditional DGX B200 systems, making it highly suitable for such large-scale expert parallel inference tasks.

**In other words, the KV cache communication savings from KDA may be partially offset by the weight exchange demands brought by WideEP, meaning the overall pressure on AI infrastructure has not decreased significantly.**

## The 64-Chip Expansion Domain Does Not Benefit NVIDIA Exclusively

However, there are differing views in the market. An insider known as GDP (@bookwormengr) pointed out that **the "64-chip expansion domain" mentioned by Moonshot AI does not necessarily imply NVIDIA's NVL72 solution.** Huawei's Ascend 950 SuperPod also utilizes a 64-chip configuration and possesses similar unified bus (UB) memory expansion capabilities akin to NVLink.

From an architectural capability perspective, the Ascend 950 SuperPod can support scaling across 16 racks to 1,024 NPUs, remaining competitive in meeting large-scale model inference needs. **Therefore, while the hardware demand growth driven by K3 does not mean NVIDIA will be the sole beneficiary, the trend toward strengthened demand for high-end AI interconnect systems remains clear.**

## Jevons' Paradox: AI Efficiency Gains May Drive Hardware Demand Growth

SemiAnalysis further cites Jevons' Paradox to explain trends in AI infrastructure.

This theory posits that when a technology improves resource utilization efficiency and lowers unit costs, demand often does not decline; instead, it may grow due to an expanded scope of application. **Applied to the AI sector, as linear attention mechanisms lower inference costs, they may encourage more enterprises to deploy AI applications, further expanding the global scale of AI inference and ultimately driving growth in demand for GPUs, HBM, DRAM, and high-speed networking equipment.**

However, GDP remains cautious about this view. He acknowledges the long-term logic of Jevons' Paradox but notes that KDA has practical significance for optimizing state storage in long-context tasks. Even with massive model weight scales, actual memory pressure may be lower than market intuition suggests, due to the use of 4-bit quantization and highly sparse designs.

He believes that **the variable truly worth watching is whether the AI companies with the largest global inference demands—OpenAI, Anthropic, and Google DeepMind—have adopted or will adopt linear attention schemes similar to KDA or DeepSeek's CSA/HCA.** If leading AI companies widely adopt such architectures, the demand for memory and interconnect resources in long-context inference could drop significantly, becoming a major factor influencing the future structure of AI hardware demand.

## Market Focus Shifts: Can AI Demand Growth Offset Architectural Efficiency Gains?

Overall, the emergence of Kimi K3 does not simply point to a decline in AI hardware demand. For NVIDIA, the true focus of competition may shift from "single-card performance" to "system-level capabilities"—including high-speed interconnects, rack-level scaling, and large-scale inference optimization.

As model scales continue to expand, AI infrastructure still faces pressures from parameter size, expert parallelism, data exchange, and increased inference throughput, even as attention mechanisms are continuously optimized.

**The core question for the market going forward is not whether linear attention reduces consumption of individual resources, but whether the new demand generated by the expansion of AI application scale can consistently exceed the resource savings brought by improvements in architectural efficiency.**

Risk Warning and Disclaimer

The market involves risks; investment should be approached with caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article align with their specific circumstances. Investment decisions made based on this content are the sole responsibility of the investor.

### Related Stocks

- [DRAM.US](https://longbridge.com/en/quote/DRAM.US.md)
- [SMH.US](https://longbridge.com/en/quote/SMH.US.md)
- [CLOU.US](https://longbridge.com/en/quote/CLOU.US.md)
- [588780.CN](https://longbridge.com/en/quote/588780.CN.md)
- [XDAT.US](https://longbridge.com/en/quote/XDAT.US.md)
- [159516.CN](https://longbridge.com/en/quote/159516.CN.md)
- [561980.CN](https://longbridge.com/en/quote/561980.CN.md)
- [512760.CN](https://longbridge.com/en/quote/512760.CN.md)
- [512720.CN](https://longbridge.com/en/quote/512720.CN.md)
- [DTCR.US](https://longbridge.com/en/quote/DTCR.US.md)
- [SMH.UK](https://longbridge.com/en/quote/SMH.UK.md)
- [PSI.US](https://longbridge.com/en/quote/PSI.US.md)
- [IDGT.US](https://longbridge.com/en/quote/IDGT.US.md)
- [AIQ.US](https://longbridge.com/en/quote/AIQ.US.md)
- [XSD.US](https://longbridge.com/en/quote/XSD.US.md)
- [159558.CN](https://longbridge.com/en/quote/159558.CN.md)
- [588200.CN](https://longbridge.com/en/quote/588200.CN.md)
- [159546.CN](https://longbridge.com/en/quote/159546.CN.md)
- [588170.CN](https://longbridge.com/en/quote/588170.CN.md)
- [512480.CN](https://longbridge.com/en/quote/512480.CN.md)
- [159325.CN](https://longbridge.com/en/quote/159325.CN.md)
- [159995.CN](https://longbridge.com/en/quote/159995.CN.md)
- [562820.CN](https://longbridge.com/en/quote/562820.CN.md)
- [SOXX.US](https://longbridge.com/en/quote/SOXX.US.md)
- [SOXL.US](https://longbridge.com/en/quote/SOXL.US.md)
- [USD.US](https://longbridge.com/en/quote/USD.US.md)
- [SSG.US](https://longbridge.com/en/quote/SSG.US.md)
- [SOXS.US](https://longbridge.com/en/quote/SOXS.US.md)
- [SOXQ.US](https://longbridge.com/en/quote/SOXQ.US.md)
- [FTXL.US](https://longbridge.com/en/quote/FTXL.US.md)
- [NVDA.US](https://longbridge.com/en/quote/NVDA.US.md)
- [OpenAI.NA](https://longbridge.com/en/quote/OpenAI.NA.md)
- [GOOGL.US](https://longbridge.com/en/quote/GOOGL.US.md)
- [GOOG.US](https://longbridge.com/en/quote/GOOG.US.md)
- [NVD.DE](https://longbridge.com/en/quote/NVD.DE.md)

## Related News & Research

- [AI Trade’s Favorite Chip ETF Sees Record Inflows Despite Worst Month Since 2008](https://longbridge.com/en/news/293222756.md)
- [JPMorgan Sees Recent AI Selloff as a Setup for Semiconductor Sector Upswing](https://longbridge.com/en/news/293197181.md)
- [SOXX Enters Bear Market: Why Ed Yardeni Says Semiconductor Stocks Could Fall Another 12%](https://longbridge.com/en/news/293182752.md)
- [After the AI Sector Pullback, Capital Begins Seeking New Certainty: Could Semiconductor Equipment Stocks Be the Next Major Theme?](https://longbridge.com/en/news/292839918.md)
- [SA asks: Are memory chipmakers building too much capacity?](https://longbridge.com/en/news/293128612.md)