NVIDIA Groq 3 LPX Enters Full Production as Vera Rubin Extension
NVIDIA said its Groq 3 LPX interactive AI inference accelerator entered full production as an extension of the NVIDIA Vera Rubin platform. The company said the update targets faster token generation for highly responsive agentic systems.
NVIDIA said agentic systems generate large volumes of tokens across many inference steps, so token generation speed affects how quickly an agent can complete work. The release said Vera Rubin NVL72 systems provide training and inference platform configurations that Groq 3 LPX extends through higher token generation rates.
NVIDIA Groq 3 LPX was described as purpose-built to extend Vera Rubin’s interactivity, defined in the release as the rate at which tokens are generated for an individual user. The accelerator was also described as addressing two computing challenges for agentic AI: processing large context amounts and generating tokens with low latency.
Nebius plans to bring NVIDIA Groq 3 LPX to production through its Nebius Token Factory. Jensen Huang, founder and CEO of NVIDIA, said, “Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” and Danila Shtan, chief technology officer of Nebius, said, “Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what NVIDIA Groq 3 LPX is built to accelerate.”
NVIDIA said the accelerator delivered record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B and that it plans additional early adoption by Groq after Nebius.
Provided by Globe Newswire on behalf of Nvidia. Read the original.