Equity research
2026.07.26 01:04

UBS: Open-Source Models Momentum

AI Demand & Enterprise Adoption

> Broadening Demand: AI demand remains exceptionally strong and is expanding rather than slowing down, despite heightened scrutiny around return on investment (ROI).

> Transition to Production: Enterprises are shifting from early testing and proof-of-concepts into full production and large-scale deployment.

> Incremental Categories of Spend: Companies like Perplexity, AlphaSense, ElevenLabs, Sierra, and Cursor reported exponential growth (e.g., Perplexity experiencing a 3x increase in ARR year-to-date) driven by new, non-cannibalistic categories of AI spending such as voice, digital coworkers, agents, and coding.

Pricing Shifts & Cost Optimization

> Consumption-Based Pricing: Vendors like Microsoft (with Copilot) are transitioning certain clients from traditional seat-based licensing to consumption-based pricing.

> Focus on Token Costs: This shift has intensified enterprise focus on maximizing ROI per token, making model-harness optimization, smart routing, and workload-specific model selection crucial components.

> Reallocation, Not Reduction: Demand is not shrinking; rather, it is being strategically reallocated toward the most cost-effective models and workloads.

Open-Source Momentum & Implications for NVIDIA

> Closing Performance Gap: Open-source models have significantly narrowed the performance gap against frontier solutions over the past year.

> Explosive Token Utilization: Open-source adoption has surged from sub-1% last year to 1% at the start of 2026, multiplied several times over by mid-2026, and is projected to potentially represent the majority of tokens used over time.

> Favorable for NVIDIA: This trend is viewed as highly positive for NVIDIA due to its software leadership (such as the Nemotron family) and the fact that most open-source models are trained and fine-tuned on NVIDIA hardware, resulting in superior inference performance.

> Expanded Infrastructure Demand: While open-source adoption pressures frontier-only economics, it ultimately expands total inference demand by turning lower-cost models into viable solutions for a broader range of mature workflows.

Fast-Inference Infrastructure

> Optimal Use Cases: Cerebras (CBRS) enterprise customer feedback indicates that wafer-scale engine (WSE) technology is ideally suited for latency-sensitive search and one-shot retrieval workloads.

> Long-Running Workloads: Long-running agentic workflows continue to prioritize model quality, tooling, and orchestration over raw inference speed.

> Market Share: NVIDIA's view that fast inference accounts for roughly 10-20% of the overall inference market remains intact, though use cases are expected to expand as token costs decline and architectural advancements (like disaggregated solutions and stacked-memory SRAM architectures) improve.

$NVIDIA(NVDA.US) $Cerebras(CBRS.US)

The copyright of this article belongs to the original author/organization.

The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.