NVIDIA announced that the Groq rack will be launched this year, following the completion of its acquisition for $20 billion
I'm LongbridgeAI, I can summarize articles.NVIDIA announced that the Groq 3 LPX rack has been fully put into production, marking the commercialization of technology following its $20 billion acquisition of Groq assets. This rack will be deployed on the Neocloud Nebius platform, working in conjunction with the Vera CPU and Rubin GPU, and is expected to go live later this year. This move aims to meet the demand for low-latency AI inference and improve service quality for latency-sensitive customers
Nvidia Groq 3 LPU chip
Nvidia announced on Monday that its Groq 3 LPX rack has entered full production, marking the official commercialization of the technology involved in the company's largest acquisition to date.
Nvidia Senior Director Dion Harris stated that the Groq rack will be deployed on the Neocloud Nebius platform, working in conjunction with the Vera central processing unit and Rubin graphics processor, and will go live later this year.
Nvidia is ramping up the manufacturing of Groq chips and pushing them to customers, highlighting the growing importance of low-latency inference—key to enabling AI entities to respond quickly and avoid long wait times for users, especially in programming scenarios. Nvidia noted that cloud service providers can charge higher fees for such tokens.
Harris mentioned in a conference call, "For vendors providing token services, this offers them the possibility to provide a high-quality service tier to those customers who are most sensitive to latency, in order to meet the corresponding service level agreements."
In December last year, Nvidia acquired the assets of chip startup Groq for $20 billion, marking the company's largest acquisition to date.
The Groq architecture integrates 500 megabytes of high-speed SRAM on the chip die to reduce memory bottlenecks. Groq chips are manufactured by Samsung, while Nvidia's GPUs are produced by TSMC.
Nvidia has packaged 256 independent Groq 3 chips in the LPX rack. According to benchmarks from Artificial Analysis, Nvidia claims its Groq 3 LPX rack can deliver a throughput of 3,400 tokens per second.
This is a highly competitive field. Smaller GPU manufacturer AMD announced earlier this year that it will integrate its rack-level systems with the recently launched Cerebras chips, also focusing on low-latency inference. OpenAI's recently announced "ultra-fast mode" currently promises 750 tokens per second, and is "powered by Cerebras."
Low-latency chips will not replace GPUs—GPUs are the mainstay of AI chips, capable of both training and inference, and flexible enough to adapt to new technologies and models. Low-latency chips like Groq primarily focus on one aspect of model serving, namely the "decoding" phase.
Harris stated, "This is not about replacing GPUs, but about using the right processor at the right price for the appropriate workload." NVIDIA is currently accelerating the shipment of the Vera Rubin system, which went into production earlier this year. At the launch event for Vera Rubin and Groq 3 LPX in March, NVIDIA CEO Jensen Huang projected that cumulative sales from the current generation Blackwell chips to the new Vera Rubin system will reach $1 trillion by 2027.
Jensen Huang stated at the time that he would allocate a quarter of the data center space dedicated to programming applications for Groq chips.
Jensen Huang said, "The rest of my data center will be entirely using Vera Rubin."
NVIDIA will announce its financial results on Wednesday
