I'm LongbridgeAI, I can summarize articles.Technical indicators remain important, but the market's observation coordinates have shifted towards cloud inference growth rates, data center capacity, system delivery, and customer utilization.
Cerebras Wafer-Scale Engine 3
After Supernova 2026, Cerebras' research framework has changed significantly. Wafer-scale processors remain the technical foundation, but the company's operational variables have extended to data centers, cloud inference, heterogeneous computing, and rack-level systems engineering. For capital markets, a single benchmark only proves performance; sustained revenue requires capacity, utilization, and delivery capabilities to materialize together.
In Q2 2026, Cerebras reported GAAP revenue of $180.1 million; Core Revenue reached $210 million, up 103% YoY. Among these, Core Cloud Revenue grew 287% YoY, with cloud business growth significantly outpacing overall revenue. Core Gross Margin stood at 41%, and Remaining Performance Obligations (RPO) reached $25.4 billion.
On the infrastructure side, the company disclosed that over 600 MW of data center capacity is either operational or under contract, with plans for delivery by the end of 2027; manufacturing capacity is planned to increase more than 10-fold in 2026. With expansion across order reserves, data centers, and manufacturing, the business model has tilted further from equipment sales toward recurring inference service revenue.

Figure 1 | Shift in operational focus from single-chip capability to cloud inference, capacity, and delivery
The 600 MW figure indicates potential supply scale. The revenue slope depends on how much of this capacity comes online on time and converts into high-utilization inference revenue. Contracted capacity determines the potential ceiling, while the pace of deployment and utilization rate determine the actual revenue slope.
Cerebras CS-3 features 44GB of on-chip SRAM with an on-chip memory bandwidth of approximately 21 PB/s. AMD's MI455X offers 432GB of HBM4 capacity with a peak memory bandwidth of 23.3 TB/s. These two parameters operate at different memory layers and cannot be directly converted into model performance multiples, but they sufficiently reflect architectural preferences.
The GPU route continues to increase HBM capacity and bandwidth to support larger model working sets, KV Cache, and high-throughput tasks; Cerebras places massive SRAM within the wafer to reduce compute unit wait times for data. For real-time inference, first-token latency and continuous token generation speed directly impact user experience, making data movement efficiency the core performance pillar for Cerebras.

Figure 2 | HBM and on-chip SRAM belong to different layers; the comparison focuses on architectural preference
AMD and Cerebras are advancing Disaggregated Inference, splitting a single LLM inference into prefill and decode stages. AMD Helios handles high-throughput prompt processing, while Cerebras Wafer-Scale Engine manages low-latency decode and token generation. Both parties expect the joint solution to achieve up to 5x tokens/s/Watt, with plans to initially offer it via Cerebras Cloud in the second half of 2026.
The industrial implication of this architecture is direct: Cerebras doesn't need to cover all AI compute loads; it just needs to establish a stable advantage in the decode stage, where latency premiums are highest, to enter the production chain of large-scale AI infrastructure. AWS also plans to introduce similar decoupled inference capabilities into Amazon Bedrock, targeting Q1 2027.

Figure 3 | AMD handles high-throughput prefill, Cerebras handles low-latency decode
Assessment: If heterogeneous inference enters mass production, Cerebras' competitive boundary will expand from "replacing GPUs" to "occupying high-value inference nodes within the GPU ecosystem." This lowers the entry barrier caused by insufficient ecosystem completeness.
CS-4 consists of three WSE-3 Turbos and uses the Nexus Rack-Scale Platform. Internal benchmarks and predictive data provided by the company show that compared to CS-3, CS-4 can achieve up to 2x performance and 10x token capacity; compared to GPU systems, inference speed can reach up to 30x. These performance results vary with models, configurations, and workloads, so peak multiples cannot be directly extrapolated to all production scenarios.
More noteworthy are the systems engineering metrics: component count reduced by 50%, automated manufacturing ratio increased by 60%, deployment time compressed from days to hours, and the first batch of CS-4 units began shipping this quarter. Data center operators ultimately care about how many effective tokens each MW generates; rack density, power supply, liquid cooling, deployment speed, and fault maintenance all factor into unit economics.

Figure 4 | CS-4 product competition extends from chip performance to rack, manufacturing, and deployment efficiency
Cerebras System Cluster
CrowdStrike plans to use Cerebras inference capabilities to support Falcon AI Detection and Response. Cybersecurity is highly sensitive to latency: attack windows are measured in seconds, and models need to complete context collection, inference, tool calling, and response; additional waiting directly undermines product value. For such workloads, inference speed maps directly to actual business outcomes.
The same logic applies to Coding Agents and real-time workflows. Cerebras has already disclosed cloud capacity partnerships with Cognition, Lovable, etc., and customers also include Block, Figma, AlphaSense, and GSK. Investment judgment requires continuing to track two variables beyond customer count: the intensity of compute consumption per customer, and whether high-frequency applications can boost cloud utilization.
Valuation Scenario (Not Company Guidance)
2028E Core Revenue $7.1 billion × 14x EV/S
Implied target price approx. $320
The difficulty of this valuation framework lies in revenue realization. Cerebras' latest 2026 Core Revenue guidance is $880–890 million; calculating with a midpoint of $885 million, reaching $7.1 billion in 2028 implies an ~8x expansion over two years, corresponding to an implied CAGR of approx. 183%. The company also stated it plans to achieve more than triple revenue growth in 2027.
A 14x EV/S represents aggressive pricing for high-growth infrastructure assets. If cloud utilization, gross margin, or capacity ramp fall short of expectations, both the valuation multiple and the revenue base may come under pressure. Current Core Gross Margin is 41%, with full-year guidance of 41%–43%; whether subsequent expansion can maintain stable gross margins is a key indicator for verifying commercial quality.
1MW Online: How much contracted capacity converts to billable compute as planned.
2Cloud Utilization: Whether the high growth rate of Core Cloud Revenue can sustain and form a more stable revenue structure.
3Gross Margin: Under heavy asset data center expansion, can Core Gross Margin remain above 40%?
4Partnership Execution: Can the production timeline for AMD and AWS's decoupled inference be realized?
Cerebras has already used its wafer-scale architecture to prove that low-latency inference can create significant performance differences. The core question ahead is singular: Can 600 MW translate into sustained, replicable, and reasonably margined inference revenue?
Technological leadership has been seen by the market; valuation next only recognizes realization.
Data Source
Operating data, CS-4, AMD partnership, CrowdStrike partnership, and product parameters are from official public information of Cerebras and AMD; the valuation part is a scenario calculation, not company guidance. Performance data involving "highest" or "most" are based on company benchmarks, predictions, or specific configuration standards; actual performance depends on models, system configurations, and workloads.
The copyright of this article belongs to the original author/organization.
The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.
