---
title: "NVIDIA is cutting memory while betting on optics; CPO officially enters mass production."
type: "Topics"
locale: "en"
url: "https://longbridge.com/en/topics/43176339.md"
description: "NVIDIA is doing two opposite things within the same new architecture: cutting memory on one side while betting on optical interconnects on the other. On August 3, CPO officially announced mass production; that day, all four optical interconnect companies surged across the board, while storage stocks were split."
datetime: "2026-08-04T08:25:38.000Z"
locales:
  - [en](https://longbridge.com/en/topics/43176339.md)
  - [zh-CN](https://longbridge.com/zh-CN/topics/43176339.md)
  - [zh-HK](https://longbridge.com/zh-HK/topics/43176339.md)
author: "[热点君](https://longbridge.com/en/profiles/1450684.md)"
generator: "portal-rs"
---

# NVIDIA is cutting memory while betting on optics; CPO officially enters mass production.

Nvidia is doing two opposite things within the same new architecture: **cutting memory on one side while betting big on optics on the other.**

Yesterday (August 3 ET), the optics bet officially landed: Nvidia Senior Vice President Gilad Shainer announced that Co-Packaged Optics (CPO) has entered mass production, switches developed in collaboration with the supply chain have been delivered to closely cooperating customers, and deployment has begun in Nvidia's own facilities.

The market responded immediately. On August 3 at close, the four optical interconnect stocks surged across the board: $Applied Optoelectronics(AAOI.US) rose 16.85% $Coherent Corp.(COHR.US) rose 9.60%, $Lumentum(LITE.US) rose 9.24% $Fabrinet(FN.US) rose 4.72%; $NVIDIA(NVDA.US) itself rose 2.93%.

On the same day, storage stocks moved in all directions: $Sandisk(SNDK.US) rose 6.03% $Micron Tech(MU.US) rose slightly by 0.79%, $Seagate Tech(STX.US) fell 2.93% $Western Digital(WDC.US) fell 3.23%.

**Subtraction and addition are not two separate things; they are the same problem.** 

* * *

### I. Subtraction: Where do the three cuts come from?

To understand this subtraction, you first need to know that AI inference is divided into two steps.

For example, a large model working is like a student taking an exam: **Step 1: Reading the question**, reading tens of thousands of words of materials in one go and taking notes in your mind; **Step 2: Answering the question**, writing out word by word. Reading consumes brainpower (compute power) but doesn't consume much mouthwork (memory bandwidth); answering is exactly the reverse.

In the past, both steps ran on the same GPU, equipped with the most expensive HBM. The result was that during the reading phase, most of that expensive bandwidth sat idle.

Nvidia's solution is to split the two steps and equip each with its own hardware. This gives rise to the three cuts.

#### **Cut #1: No more HBM for the "reading" step**

Nvidia specifically created a Rubin CPX chip, equipped with 128GB GDDR7, delivering 30 PFLOPS compute power under NVFP4 precision. GDDR7 is the kind of video memory found on graphics cards; its bandwidth is inferior to HBM, but it is much cheaper. The official Rubin GPU still retains 288GB HBM4.

**HBM hasn't been cut at all; what's being cut is "assigning HBM to tasks that don't need HBM."**

#### **Cut #2: Halving memory on the CPU side**

According to reports from SemiAnalysis, Nvidia plans to reduce the SOCAMM module of the Vera CPU from 192GB to 96GB, bringing the total CPU memory per rack down from approximately 55TB to about 28TB. Reports from TrendForce indicate that with the original configuration, memory would account for about 29% of the bill of materials cost for the entire system.

Bernstein's calculations illustrate the motivation: An NVL72 Rubin rack might cost $9.1 million, primarily due to rising memory prices, with HBM4 unit prices potentially reaching $53 per GB by 2027.

> Key point to highlight: **Nvidia has never officially confirmed the downsizing; it is currently based on reports and institutional calculations.**

#### **Cut #3: Moving the notes to the cabinet**

The "notes" taken during answering are called KV Cache. The longer the context, the larger the notes, and HBM simply cannot hold them. Nvidia's solution presented at CES in January this year is the CMX Context Memory Storage platform: using BlueField-4 DPU to offload KV Cache to NVMe SSDs, creating a layer of storage between local drives and network storage.

BlueField-4 can complete encryption and data integrity checks at full line rate of 800Gb/s without consuming CPU compute power.

Together, the three cuts amount to one thing: **categorizing materials by use case instead of blindly applying the most expensive option everywhere.**

![NVIDIA BlueField-4 DPU Powers Gigascale AI Factories - NADDOD Blog](https://pub.pbkrs.com/uploads/2026/909462ab49809d9cfe36e7722f3cdabe?x-oss-process=style/lg)

* * *

### II. Is it really "storage demand" that's being cut?

The drop in storage stocks this round trades on the **loosening of expectations for storage shortages**—if even Nvidia is starting to allocate less, how much shortage is left?

This inference has two sides.

**Side 1: What's being cut is LPDDR5X, not HBM**

The configuration of 20.7TB of HBM4 per rack on the GPU side remains unchanged; SOCAMM2 is a JEDEC-standardized hot-swappable module, with high capacity still existing as a custom option.

Morgan Stanley semiconductor analyst Joseph Moore directly refuted the interpretation of weak demand in his client research report on July 20, stating verbatim: "This is not a normal cycle—memory is the bottleneck, and it is becoming increasingly the bottleneck for AI construction." According to his report, discussions with data center procurement managers show no signs of relief regarding the severity of the memory shortage, and he expects memory prices to rise at least 25% quarter-over-quarter in Q3.

Then there's the third cut layer—the sinking of KV Cache to SSDs itself represents a massive new demand for NAND. Management estimates from Western Digital suggest that just this Nvidia architecture could bring an incremental NAND demand of 75 to 100 EB by 2027, possibly doubling in 2028.

**Side 2: This sector has risen too much this year**

Based on the August 3 close, Western Digital has risen approximately 3.7 times year-to-date, and Micron has risen approximately 1.6 times.

But looking down from the year-to-date highs in mid-June, Western Digital has already retraced 41%, and Micron has retraced 27%. When gains have already priced in (price in) optimistic expectations for the next two years, any report of "downsizing" is enough to smash a hole in the price.

**So a more accurate statement is: The DRAM ledger got cut, the NAND ledger got added to, and they are not the same ledger.**

* * *

### III. Addition focused on optics: What does CPO add?

The money saved by subtraction is spent on the other end towards optical interconnects.

CPO (Co-Packaged Optics) put simply is: **packaging the optical engine directly next to the switching chip, eliminating the need for pluggable optical modules**, making electrical signal traces extremely short. According to the official stance given when Nvidia released Spectrum-X / Quantum-X Photonics in March 2025: laser usage is reduced to 1/4, energy efficiency is improved 3.5 times compared to traditional solutions and 5 times compared to pluggable optical modules, signal integrity is improved 63 times, large-scale network resilience is improved 10 times, and deployment speed is 1.3 times faster.

**Why must it be optics?** It's still Old Huang's "wall of copper": every time bandwidth doubles, the distance copper cables can travel is halved. Shainer's quantitative explanation this time is that the bandwidth demand of the Scale-up (vertical expansion) architecture is more than ten times higher than traditional horizontal expansion—**and Scale-up is precisely the territory inside racks, originally belonging to copper cables.**

The timeline is also moving forward. According to HPCwire, the industry originally expected CPO to be commercialized only by 2033, but Nvidia has brought it forward by a full 5 years; Jensen Huang stated in the analyst Q&A that the next-generation Feynman NVL1152 will be "fully CPO." TrendForce estimates the CPO market size will reach $39 billion by 2030, entering a volume ramp-up period in 2028-2029.

The money was bet on early. On March 2 this year, Nvidia announced investments of $2 billion each in Lumentum and Coherent, totaling $4 billion, accompanied by long-term purchase commitments, priority supply rights, and capacity reservations.

**Of course, the flip side must also be said.** This mass production is "beginning delivery to closely cooperating customers," not rolling out across the entire industry; pluggable optical modules remain the mainstay in the medium term, and ramp-up speed, yield rates, and maintainability still require time for verification.

Moreover, this news currently comes mainly from media paraphrasing of Shainer's public remarks; Nvidia's official website has not yet published a corresponding independent press release—**it is "confirmed mass production," not "official launch event," and the weight of these two differs.** Of the surge on August 3, how much was due to the news itself versus how much was a rebound after the severe oversold conditions of the previous week is still indistinguishable.

* * *

### From "Stacking Materials" to "Categorization"

Over the past two years, the narrative for AI hardware has been stacking materials: whose HBM is more abundant, whose copper cables are thicker, whose racks are denser. Now, Nvidia is doing categorization: assigning cheap materials to reading tasks and expensive materials to answering tasks; using copper for short distances and optics for long distances; placing hot data in HBM and warm data in SSDs.

For companies on the supply chain, this means not all segments will benefit synchronously—**position matters more than the windfall.**

Instead of agonizing over "whether storage can still be touched," it's better to first figure out which segment you believe in more.

-   **If you believe "AI data volume will only get bigger"**, look towards the NAND and large-capacity storage segment;

$Micron Tech(MU.US)$Sandisk(SNDK.US)$Western Digital(WDC.US)$Seagate Tech(STX.US)

-   **If you believe "connectivity is the next bottleneck"**, look upstream in optical devices;

$Lumentum(LITE.US)$Coherent Corp.(COHR.US)$Applied Optoelectronics(AAOI.US)$Fabrinet(FN.US)$Marvell Tech(MRVL.US) 

-   **If you believe "downsizing is just part of price negotiation"**, then this retracement in DRAM itself is a variable.$NVIDIA(NVDA.US)

* * *

Do you think Nvidia's current round of "cutting memory, betting on optical interconnects" is a signal that the storage cycle has peaked, or is it just moving money from one pocket to another? Let's chat in the comments 👇

Dollar-Cost Averaging Feature: Start your DCA plan now, steadily layout amidst volatility, [Click to jump](https://longbridge.cn/zh-CN/topics/37799438?invite-code=DXZBD9&channel=n37799438&app_id=longbridge&utm_source=longbridge_app_share).

Odd Lot Feature: Start from $1, build flexible positions, [Click to jump](https://longbridge.cn/zh-CN/topics/28739368?invite-code=DXZBD9&channel=n28739368&app_id=longbridge&utm_source=longbridge_app_share).

More content: Click \[My\] - \[Opportunity Insight\] to view selected opportunities

### Related Stocks

- [LITE.US](https://longbridge.com/en/quote/LITE.US.md)
- [STX.US](https://longbridge.com/en/quote/STX.US.md)
- [SNDK.US](https://longbridge.com/en/quote/SNDK.US.md)
- [WDC.US](https://longbridge.com/en/quote/WDC.US.md)
- [NVDA.US](https://longbridge.com/en/quote/NVDA.US.md)
- [AAOI.US](https://longbridge.com/en/quote/AAOI.US.md)
- [MU.US](https://longbridge.com/en/quote/MU.US.md)
- [COHR.US](https://longbridge.com/en/quote/COHR.US.md)
- [FN.US](https://longbridge.com/en/quote/FN.US.md)
- [MRVL.US](https://longbridge.com/en/quote/MRVL.US.md)

## Comments (11)

- **卖飞专业 · 2026-08-05T02:24:57.000Z**: As a professional who frequently sells too early, the third point resonates with me the most—the discussion around KV Cache sinking to SSDs is surprisingly scarce. While the first two points focused on cost-saving, the third actually creates a new demand out of thin air, going in the opposite direct
- **正在缓冲 · 2026-08-04T11:54:23.000Z · 👍 5**: All day long it's just rumors about memory, anyone with a brain knows it's because the memory is sold out and there's none left to install, so they have no choice but to reduce the memory usage per rack. Saying it in one sentence versus another yields vastly different effects.
- **一买就被套 · 2026-08-04T08:29:54.000Z**: How do I get in? I can't trade right now.
  - **巨米有财** (2026-08-04T08:31:16.000Z): You can try the HK ☎️ card or 🪜, you need to enable global mode. That's how I do it.
  - **啧啧啧** (2026-08-04T09:15:19.000Z): Is there any way for Android, bro?
  - **巨米有财** (2026-08-04T09:56:40.000Z): It's the same. First, set up a 🪜 and try enabling global mode.


---
> **Disclaimer: This article is for reference only and does not constitute any investment advice.**