- NVIDIA announced that the Groq 3 LPX interactive AI inference accelerator is now in full production, extending the Vera Rubin platform to deliver ultrafast token generation for agentic AI.
- The accelerator achieved a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B, providing 4x faster responsiveness for latency-sensitive workloads.
- Nebius is the first AI cloud to adopt the technology for its Nebius Token Factory production inference platform, with Groq also planning early adoption.