Nvidia announced on Monday that its Groq 3 LPX racks have entered full-scale production, marking the commercial rollout of the technology tied to the largest acquisition in the company's history.
Dion Harris, a senior director at Nvidia, said the Groq racks will be deployed on the Neocloud Nebius platform, working alongside Vera central processing units and Rubin graphics processing units, with availability slated for later this year.
The accelerated push to manufacture Groq chips and get them into customers' hands underscores the rising importance of low-latency inference—a critical factor in making AI agents respond quickly and sparing users long wait times, particularly in coding scenarios. Nvidia noted that cloud providers can charge premium rates for such tokens.
Harris said during a conference call: "For token providers, this gives them the ability to offer a premium service tier to those customers who are most sensitive to latency, allowing them to meet the corresponding service-level agreements."
In December, Nvidia closed a $20 billion deal to acquire the assets of chip startup Groq, the company's largest acquisition to date.
The Groq architecture integrates 500 megabytes of high-speed SRAM directly onto the chip die to alleviate memory bottlenecks. Groq chips are manufactured by Samsung, while Nvidia's GPUs are produced by TSMC.
Nvidia packages 256 individual Groq 3 chips into each LPX rack. Based on benchmarks from Artificial Analysis, Nvidia claims its Groq 3 LPX rack delivers throughput of 3,400 tokens per second.
This is a fiercely competitive space. Smaller GPU maker AMD announced earlier this year that it would integrate its rack-scale systems with the recently launched Cerebras chips, also targeting low-latency inference. OpenAI's newly unveiled "ultra-fast mode" currently promises 750 tokens per second and is "powered by Cerebras."
Low-latency chips are not meant to replace GPUs—GPUs remain the workhorses of AI computing, handling both training and inference while offering the flexibility to adapt to new technologies and models. Chips like Groq focus on a single segment of model serving: the "decode" phase.
Harris stated: "This is not about replacing GPUs, but rather using the right processor at the right price for the right portion of the workload."
Nvidia is also accelerating shipments of its Vera Rubin systems, which entered production earlier this year. At the March unveiling of Vera Rubin and Groq 3 LPX, Nvidia CEO Jensen Huang projected cumulative sales of $1 trillion by 2027, spanning from the current-generation Blackwell chips to the new Vera Rubin systems.
Huang said at the time that he would allocate a quarter of his data center space dedicated to programming applications to Groq chips.
Huang added: "The rest of my data center will be all Vera Rubin."
Nvidia is set to report its quarterly earnings on Wednesday.