---
title: "Muse goes viral: Is this the CPU's ChatGPT moment?"
type: "Topics"
locale: "en"
url: "https://longbridge.com/en/dolphin/post/44095954.md"
description: "Meta launched Muse on Sep. 8. Each Muse runs on a dedicated virtual machine, the Muse Secure VM, which stores both the agent and user data. Long-running jobs keep executing even after the app is closed, then re-engage the user when something changes or an approval is needed.The product went viral, briefly topping the US mobile download charts. For details on Muse and its implications for Meta, see Dolphin Research's take 'Muse goes viral: real inflection or hype peak?'..."
datetime: "2026-09-28T10:31:53.000Z"
locales:
  - [en](https://longbridge.com/en/dolphin/post/44095954.md)
  - [zh-CN](https://longbridge.com/zh-CN/dolphin/post/44095954.md)
  - [zh-HK](https://longbridge.com/zh-HK/dolphin/post/44095954.md)
author: "[Dolphin Research](https://longbridge.com/en/dolphin.md)"
generator: "portal-rs"
---

# Muse goes viral: Is this the CPU's ChatGPT moment?

Meta launched Muse on Sept 8. Each Muse runs on a dedicated 'Muse Secure VM', which co-hosts the agent and user data; long-running tasks continue even after the user closes the app, and the agent pings the user only when changes occur or approvals are needed. The product went viral and briefly topped the US mobile download charts.

For product-side details of Muse and implications for Meta, see Dolphin Research’s [this link](https://longbridge.com/en/dolphin/post/44040483?channel=OWNN00110). Below we discuss the upstream supply chain shifts triggered by this launch.

Muse’s popularity not only drove a double-digit rally for $Meta Platforms(META.US), it also sparked a surge across CPU supply-chain names. The key question: **Is Meta’s Muse the CPU analogue of the ChatGPT moment?**

**1\. How Muse lifts CPU demand**

**Unlike a traditional chatbot, Muse is a subscription that gives you a 24/7 personal VM plus an AI butler**. Think of it as a dedicated always-on compute instance paired with an autonomous assistant.

Take a simple example: **you want to ‘monitor Mac mini prices’.**

**① Traditional chatbot flow:** 'Monitor Mac mini prices' → GPU inference → 'OK' → end → CPU sits idle → next day you ask again and the whole loop repeats. The service is stateless and ephemeral.

**② Muse:** 'Monitor Mac mini prices' → GPU inference → **CPU sets a monitoring script → VM keeps running → you close the app → VM keeps running in the cloud** → checks the price API daily → 30 checks a month, each trigger runs GPU reasoning plus CPU tools → after 30 days price drops to $899 — CPU detects it → wakes the GPU for order logic → CPU drives the browser to fill address and card and places the order.

**By comparison, a traditional chatbot stops after a response and then goes idle; Muse keeps working in the background.** This may feel reminiscent of OpenClaw, and Muse is indeed inspired by it. However, there are key differences in architecture and persistence.

**The biggest difference: Muse gives every user an isolated Linux VM (2 vCPUs), whereas OpenClaw defaults to local execution.** The upside is that Muse keeps running in the cloud even after the app is closed, continuing to work on your behalf.

**2\. Why CPUs matter in agents**

CPUs and GPUs are architected very differently. **GPUs devote most transistors to arithmetic units** (green blocks below), while CPUs spend most transistors on **control and management** (red blocks) like branch predictors, out-of-order windows, prefetchers and multi-level caches.

**Decomposing an agent workflow: perceive → plan → reason → tool execution → verify → re-plan → re-execute.** GPUs largely handle the 'reason' step, while planning, execution and verification rely on CPU control capability and state management.

**In pure inference, the CPU tokenizes, feeds the GPU and reassembles outputs,** leaving heavy math to the GPU. The CPU’s role is thin and largely I/O and orchestration bound.

**In agentic AI, the CPU becomes the instruction layer,** planning tasks, breaking goals into subtasks, scheduling parallel sub-agents, managing tools and APIs, monitoring token streams and consolidating outputs while running reflection loops. **This sustained load primarily consumes agent CPUs.**

**With agentic AI, data center CPUs split into three classes: traditional CPUs, AI head-node CPUs and agentic AI CPUs.**

**a) Traditional/general-purpose CPUs:** run non-AI workloads like web/app servers, databases, caches, storage and MQ. These are logic-heavy but not highly parallel and are mostly serial. **They are fully separate from AI GPU servers and deploy on conventional racks.**

**b) AI head-node CPUs:** the CPUs embedded in GPU servers to manage accelerators — feeding data, coordinating comms, managing memory and handling unaccelerated code. They aim to minimize tail latency with big caches, high-bandwidth memory and I/O. **They are physically tied to the same server as the GPUs and act as the GPU server’s 'housekeeper'.**

**c) Agentic AI CPUs:** the brain and hands for agent orchestration, tool execution, sandboxes and RAG retrieval, deployed as independent CPU racks. They handle agent planning and control logic where GPU utilization is low, offloading non-parallel, stateful tasks to CPUs.

**These CPU racks are separate from GPU racks but tightly coupled,** connected via high-speed networks (Spectrum-X/Ethernet) to form an AI factory. **Agent CPUs are not embedded in GPU servers; they are a distinct rack class.**

**3\. Incremental demand from agentic AI**

**Traditional CPU demand tracks refresh cycles for general-purpose servers, while agentic AI adds two CPU growth vectors: AI head-node CPUs and agentic AI CPUs.** These are structurally different from legacy server demand.

**1) AI head-node CPUs**

Even in pre-agent LLM clusters, **head-node CPUs are essential** to feed GPUs at very high rates while handling comms, synchronization and unaccelerated app logic. This includes weight loads, GPU mode configuration, KV cache management and data checks.

Under-provisioned head-node CPUs bottleneck GPUs: **insufficient CPU bandwidth stalls GPUs waiting on data,** lowering utilization. **Hence CPU:GPU ratios rose from 4:1 in HGX to 2:1 in NVL72,** as accelerators grew more complex and needed more CPU for feeding and control.

**Dolphin Research views this ratio lift as economically driven — spend CPUs to squeeze every dollar from GPUs — not a direct pull from Muse’s agentic workload.** It is an optimization to de-bottleneck GPU clusters.

**2) Agentic AI CPUs**

Agents are fundamentally CPU workloads given their reliance on sequential task execution vs. pure parallelism. **In the agentic phase, agent CPUs become 'must-have', representing true incremental demand from Muse-like models.**

**ARM has outlined an agent workflow:**

**① User → AI agent:** submit a goal, not a question. The objective is open-ended and task-oriented.

**② AI agent → Cloud:** the agent runs on cloud CPUs with no accelerators at this layer. This is CPU-only orchestration and state management.

**③ Cloud → Agents:** the agent decomposes tasks into multiple sub-agents (illustrated as 12). This is the demand multiplier across the system.

**④ Agents ⇅ AI data center (loop):** bidirectional arrows per agent reflect the think–act–observe loop. Calls are iterative round trips, not single-pass invocations.

**⑤ Inside the AI data center: an orchestration CPU (agentic CPU, new demand)** receives requests, manages state and scheduling, dispatches inference to accelerators, collects tokens back and decides next steps.

**Finally Agents → Cloud → Answer → User:** results aggregate in the cloud into an 'Answer' returned to the user. This closes the loop.

**Note the 'orchestrates' CPU refers to the agent’s main control loop — task splitting, sub-agent spawning, state holding, reflection, branching, retries and tool calls — not GPU task orchestration,** which is the head-node CPU’s domain closer to the accelerators.

In the agentic era, data centers will add many independent CPU racks for agent orchestration, scheduling and management. **Nvidia has announced a standalone Vera CPU rack with 256 Vera CPUs at 88 cores each,** which represents the true incremental layer for agentic AI.

In Nvidia’s AI factory blueprint, Vera CPUs appear beyond the Vera Rubin NVL72 (head-node) and Vera CPU Rack. **They also sit in Vera BlueField storage to handle storage management, DPUs and KV cache persistence.**

**Unlike the BF-4 DPU in NVL72 (Grace CPU), the STX storage rack’s BF-4 DPUs use Vera CPUs.** This reflects different roles across compute and storage layers.

**In the STX reference design, each BF-4 includes one Vera CPU,** two CX-9 NICs and two SOCAMM modules. **An STX rack has 16 chassis (each with two BF-4s), implying 32 Vera CPUs per rack,** plus 64 CX-9 NICs and 64 SOCAMMs.

The BF-4 (DPU) thus spans two CPU chips: one in the NVL72 compute tray (Grace) and one in the STX storage rack (Vera). Their classification depends on the primary service target, whether GPU clusters or agent workloads.

**Hence BF-4 attribution to head-node or agentic CPU hinges on whether it mainly serves GPU clusters or agent states (persistent agent state, KV cache and long-context storage).** STX leans agentic in function.

**4\. A 'ChatGPT moment' for CPUs?**

People often mix up CPUs, cores and threads, so a quick primer helps. **If a CPU is a factory, a core is a worker,** and one core handles one instruction stream. More cores let the processor handle more instructions in parallel, like Nvidia’s 88-core Vera CPU.

**Threads are like conveyor belts.** An 8-core/8-thread plant gives each worker one belt, while 8-core/16-thread gives two belts per worker. A 16-core/16-thread plant simply has more workers and higher throughput.

Among those three, the 16-core plant runs fastest given more cores. Between 8c/8t and 8c/16t, the latter is faster as more threads reduce idle gaps and keep cores busy. **Both higher core counts and more threads lift CPU performance.**

Across the three CPU classes, demand drivers differ. **a) Traditional CPUs:** track enterprise IT and cloud traffic, with incremental demand from third-party sites/services accessed by agents. **Unit counts are stable, core counts rise with product cycles and agent impact is limited.**

**b) AI head-nodes:** scale with GPU counts and rack ratios, roughly proportional to accelerators (GPU/ASIC). Per-GPU head-node cores rose ~3.75x in recent years. **The increase came from higher CPU:GPU ratios (doubling), per-CPU core growth (+22%) and new DPUs in NVL72 (+36%).**

Head-node per-GPU core inflation has been priced in, with no clear sign of another big jump next gen. **Agentic AI has limited additional pull on head-node CPUs.**

**c) Agentic CPUs:** the largest incremental opportunity, driven by concurrent agents rather than GPU counts. Their scale depends on active agents and sandbox concurrency.

**Nvidia’s single Vera CPU rack has 256 Vera CPUs (88 cores each, 2 threads/core),** totaling 22,528 cores, matching the spec of supporting 22,500+ sandboxes. These sandboxes isolate workloads to prevent cross-contamination.

Note: Muse assigns 2 vCPUs per user VM. Two vCPUs map to two SMT threads on one physical core. SMT threads share execution units, so 2 vCPUs deliver only ~1.2–1.3x the throughput of 1 vCPU, not 2x.

**Agentic CPU demand splits into a fixed tier and an elastic tier:** **① Fixed:** one pre-allocated VM per user (e.g., Muse’s 2 vCPUs), mapped to a persistent sandbox. **② Elastic:** transient sandboxes for sub-agents, reused on demand. Both user and sub-agents run in isolated environments.

**Thus an agent = model (GPU side) + fixed sandbox (user-side agent) + elastic sandboxes (sub-agents).** The user issues a goal, the fixed-layer CPU decomposes tasks, elastic sandboxes handle sub-agents and GPUs perform inference.

**Both fixed (user) and elastic (sub-agent) layers require secure isolation via sandboxes,** enabled by CPU instruction sets and hardware isolation features. CPUs have decades of built-in circuitry for safe multi-tenant execution.

From earlier estimates, a single VR NVL72 (incl. DPU) has 4,320 CPU cores (=72 GPUs × 60 cores/GPU). **By contrast, a single CPU rack has 22,528 cores (256 × 88),** 5x the cores of one VR NVL72.

**5\. Where core growth comes from**

**ARM projects that a traditional AI DC needs ~30mn CPU cores per GW of compute,** rising to ~120mn in the agentic era. This frames the order-of-magnitude step-up.

There are many ways to quote CPU needs — GPU:CPU ratios, CPU cores per GPU and cores per GW. GPU:CPU is most relevant to head-nodes and less meaningful for agents. **Once CPU racks come in, cores per GPU must rise sharply,** so cores per GW is the more useful metric.

**① Traditional AI DC:** take GB300 NVL72 with ~130kW per rack. One GW maps to ~7,692 NVL72 racks if all power goes to GB300 NVL72.

**Each GB300 NVL72 has ~2,592 CPU cores (72 GPUs × 36 cores/GPU),** implying ~20mn cores per GW, below ARM’s 30mn estimate. The gap reflects differing assumptions.

**② Agentic AI DC:** VR NVL72 racks and dedicated CPU racks both rise to ~200kW. One GW maps to ~5,000 NVL72 racks, with the remainder allocated to CPU racks.

Today’s baseline assumes GPU-only racks, so the scenarios in the table (conservative/base/bull) represent pure incremental CPU racks. They illustrate how CPU rack share drives core counts.

**On this basis, ARM’s 120mn cores per GW sits above our theoretical cap of ~113mn** for a single GW configuration. Assumptions and rack splits drive the spread.

**Dolphin Research estimates that introducing CPU racks lifts total required cores by ~2–3x,** a more balanced view of practical deployments and power envelopes.

**Nvidia’s CPU racks are CPU-only,** so adding them in agentic AI naturally boosts CPU cores per GPU. In a base case with CPU racks at 25% of total racks, per-GPU cores could rise from ~60 to ~160, a direct measure of CPU increment.

**6\. How much demand agentic AI drives**

AI data centers are already scaling and head-node needs are rising, but that is not new incremental info. **Muse-like agentic AI requires adding CPU racks for agent workloads,** which depend directly on user/agent activity, not GPU counts.

We propose: **Incremental CPU demand = (A head-node) GPU shipments × head-node cores/GPU + (B fixed layer) total registered users × DAU × (active hours/24) × 2 vCPUs/user × peak-to-avg ÷ oversub ratio ÷ 2 (threads/core) + (C elastic layer) active users × task concurrency × cores per sandbox.**

**In sizing agentic demand, focus on B fixed and C elastic layers:**

**① B Fixed layer:** on the user side, note Muse’s 2 vCPUs do not constantly pin physical cores. **a) When active, a user maps to one physical core (2 threads).** **b) When idle,** state is checkpointed to storage, the VM is paused and the core is freed, then rapidly restored on wake.

**Muse’s 2 vCPUs guarantee 'up to' capacity without constant core use.** A 256-core physical server can map hundreds or thousands of vCPUs since most VMs are idle in parallel — that is the oversub ratio concept.

**With 100mn Muse users, 30% DAU, 3 active hours/day, peak-to-avg 2 and oversub 6,** B fixed-layer cores are ~1.25mn (=100mn × 30% × 3/24 × 2 × 2/6 ÷ 2). This is the sustained user-side baseline.

**② C Elastic layer:** handles tasks from active users only. **Task concurrency is the avg. number of sub-agent sandboxes per active user,** and each sandbox requires a certain core count.

With 100mn users, 30% DAU, 3 active hours/day and peak-to-avg 2, **assume 4 cores per sandbox and derive peak active users of ~7.5mn (=100mn × 30% × 3/24 × 2).** Different concurrency scenarios scale cores accordingly.

At this scale, if concurrency hits 100% (one sub-agent per active user on avg.), the system needs roughly 1 GW of factory capacity. With a 25% CPU rack share, that implies ~1,332 CPU racks, **and ~30mn cores required for the C elastic layer.**

**Under the base case,** A head-node cores total ~17.28mn (=288k GPUs × 60 cores/GPU), B fixed layer ~1.25mn and C elastic ~30mn. These align with the concurrency and rack-share assumptions.

**A+B+C sum to ~48.53mn cores, ~2.8x the pure GPU-rack baseline of 17.28mn.** Including storage rack (DPU) needs, Dolphin Research sees agentic CPUs lifting total core demand to roughly 3x prior levels.

**7\. Opportunities across the CPU value chain**

Per AMD’s outlook, server CPU TAM could surpass $220bn by 2030 at a >50% CAGR. **This back-solves to a 2025 server CPU market of ~$29bn.**

**We use: server CPU market = total cores × ASP per core.** Agentic CPUs could 3x core demand, while per-core ASP grows ~18% YoY. This frames a powerful volume–price uplift.

**Assuming broad rollout by 2030 (3x cores) and 18% ASP CAGR, market size could reach ~$200bn** (=29 × 3 × 1.18^5), implying a near-50% CAGR. This is broadly aligned with vendor guides.

Today the server CPU market is led by Intel and AMD. **Based on 2025 revenue, their shares are ~58% and ~35%,** with others below 10%. This split reflects incumbency in legacy server fleets.

As most growth comes from agent CPUs while legacy grows slowly, **Dolphin Research expects Intel’s share to slip, with AMD and new entrants like Nvidia, Arm and Qualcomm gaining.** These three represent pure incremental exposure.

By 2030, assume **Intel and AMD at 36% each, with Nvidia at 15%, Arm at 5% and Qualcomm at 2%.** This implies meaningful shifts in share capture tied to agentic deployments.

In aggregate, agentic AI could add ~$55bn and ~$62bn in annual server CPU revenue for $Intel(INTC.US) and $AMD(AMD.US). **Nvidia, Arm and Qualcomm could see pure incremental revenue of ~$30bn, ~$10bn and ~$4bn from agent CPUs.**

**We would lift 2030 server CPU revenue estimates for Intel and AMD by ~$10–20bn,** implying roughly a 10% EPS uplift by that horizon. This reflects agent CPU adoption layered onto existing plans.

**For** $NVIDIA(NVDA.US), the ~$30bn from agent CPUs would be < strong>sub-3% of total revenue, hence limited impact at the group level given its accelerator-led mix.

**Among pure incremental plays, Arm and Qualcomm benefit more.** With 5% and 2% share by 2030, they could add ~$10bn and ~$4bn in revenue. At a 25% OPM, that implies ~$2.5bn and ~$1bn in OP uplift in 2030.

**For** $Arm(ARM.US), beyond selling chips it can monetize cores, adding earnings elasticity. A 25x PE on this stream (11% discount rate) implies ~$41bn+ valuation uplift, ~15% stock sensitivity. For $Qualcomm(QCOM.US), a 20x PE (11% discount) implies ~$13bn uplift, ~6% stock elasticity.

**Overall, agentic AI requires CPU racks, directly raising CPU core demand and supporting a sustained upcycle.** Structurally, incremental demand is concentrated in agent CPUs and head-node CPUs, with traditional server CPU demand steady.

**By stock exposure to agents: Arm (~15%) > Intel/AMD (~10%) > Qualcomm (~6%) > Nvidia (~3%).** Last week’s rallies broadly mirrored this pecking order. With Muse-driven agent CPU demand now priced in, the focus shifts to agent scale-out and who captures the extra share of the server CPU pie.

\<End>

Risk disclosure and disclaimer: [Dolphin Research Disclaimer and General Disclosure](https://support.longbridge.global/topics/misc/dolphin-disclaimer)

### Related Stocks

- [NVDA.US](https://longbridge.com/en/quote/NVDA.US.md)
- [OpenAI.NA](https://longbridge.com/en/quote/OpenAI.NA.md)
- [META.US](https://longbridge.com/en/quote/META.US.md)
- [AAPL.US](https://longbridge.com/en/quote/AAPL.US.md)
- [INTC.US](https://longbridge.com/en/quote/INTC.US.md)
- [04335.HK](https://longbridge.com/en/quote/04335.HK.md)
- [AMD.US](https://longbridge.com/en/quote/AMD.US.md)
- [ARM.US](https://longbridge.com/en/quote/ARM.US.md)
- [QCOM.US](https://longbridge.com/en/quote/QCOM.US.md)
- [NVDD.US](https://longbridge.com/en/quote/NVDD.US.md)

---
> **Disclaimer: This article is for reference only and does not constitute any investment advice.**