NVIDIA plans to release a new "inference chip" that abandons expensive high-bandwidth memory to compete against the dual pressure from Google and Meta
I'm LongbridgeAI, I can summarize articles.NVIDIA plans to release a new AI inference chip at the upcoming "GTC 2026" conference, aimed at addressing challenges from competitors like Google and Meta. This chip will focus on the AI inference stage and will be based on the Groq technology it acquired. NVIDIA's CEO Jensen Huang has previously advocated for a single processor to handle multiple tasks, but this concept may change as AI tools become more complex. Additionally, the new chip will abandon expensive high-bandwidth memory in favor of more economical SRAM to reduce costs
The British Financial Times reported that AI giant NVIDIA (US: NVDA, Nvidia) is preparing to launch a new chip designed specifically to accelerate AI responses at the upcoming "GTC 2026" annual developer conference. This move indicates that Nvidia will break its long-standing strategy of using a single processor to handle multiple tasks. The report states that this new chip will focus on the AI "inference" stage (i.e., running rather than training models) and will be the first new product to debut after Nvidia's acquisition of the core team and technology of the startup Groq for $20 billion (approximately HKD 156.6 billion) last December.
First Launch of Groq Technology Chip Targeting AI Inference Market
The report indicates that Nvidia plans to introduce this language processing unit (LPU) based on Groq technology, which will work in conjunction with the upcoming flagship Vera Rubin GPU, aimed at countering competitors and addressing new AI applications.
Nvidia is currently facing challenges from startups and major clients like Google that are developing their own AI chips; competitors like Meta have also recently announced the launch of four new processors specifically designed for inference tasks. A Silicon Valley venture capital investor remarked, "We are entering an interesting phase where Nvidia no longer dominates."
The report notes that over the past three years, Nvidia's massive market value has primarily benefited from its GPUs becoming the backbone of the generative AI industry, used to train models like OpenAI's ChatGPT.
Nvidia CEO Jensen Huang has previously advocated that a single system can be used to train new AI models as well as run chatbots and coding tools built on these models. Although major tech giants have invested hundreds of billions of dollars in deploying these systems, they are also investing in the research and development of their own dedicated AI chips. Additionally, as AI tools become increasingly complex, such as intelligent agents, it may force Huang to abandon the idea of "a single processor handling multiple tasks."
Abandoning Expensive HBM for Cost-Effective SRAM
It is reported that Nvidia's existing Blackwell and upcoming Rubin systems heavily rely on high-bandwidth memory (HBM), which is expensive and in short supply. However, this new chip that integrates Groq technology will break from tradition by using static random-access memory (SRAM) instead of the dynamic random-access memory (DRAM) used in HBM. Since SRAM is relatively abundant in the market and its technical characteristics are more suitable for accelerating AI "inference" tasks, it is expected to significantly enhance computational efficiency and control costs.
Bank of America Estimates Inference Spending Will Reach 75% by 2030
As AI technology becomes increasingly widespread, the market demand for inference computing is rapidly growing. Bank of America analysts estimate that by 2030, when the global AI data center market reaches approximately $12 trillion (about HKD 93.6 trillion), spending related to inference applications will account for 75% of total expenditures, far exceeding last year's approximately 50%, reflecting that this field will become a battleground for major tech companies
