Darius Baruo Aug 24, 2026 16:34

NVIDIA’s Groq 3 LPX enters full production, delivering record-breaking AI inference speeds for latency-critical workloads. Key to agentic AI development.

NVIDIA Groq 3 LPX Hits Full Production, Boosts AI Inference Speed

NVIDIA announced today that its Groq 3 LPX AI inference accelerator has entered full production, marking a significant milestone in the evolution of low-latency AI systems. Integrated into the NVIDIA Vera Rubin platform, the Groq 3 LPX is designed to address the growing demand for ultrafast token generation in agentic AI applications. Benchmarking results highlight its record-breaking speed of 3,400 tokens per second—four times faster than the nearest competing platform.

Agentic AI systems, which rely on real-time inference to execute complex tasks, have become a critical focus for NVIDIA. The Groq 3 LPX is tailored to handle latency-sensitive workloads by using deterministic execution and massive on-chip SRAM to eliminate bottlenecks. This positions it as a key player in NVIDIA’s AI factory architecture, which also includes its Vera Rubin NVL72 systems for broader training and inference workloads.

NVIDIA’s CEO Jensen Huang emphasized the importance of inference in AI development: “Inference is the growth engine of AI. The Groq 3 LPX advances the performance frontier, delivering ultrafast token generation just as demand for real-time AI computation accelerates globally.”

Nebius, an AI cloud platform, is the first adopter of the Groq 3 LPX. The company intends to deploy the accelerator in its Nebius Token Factory, enabling developers to leverage extreme token generation capabilities without changing APIs. Following Nebius, Groq’s own AI inference cloud is expected to adopt the platform, signaling broad industry interest in the technology.

From a technical perspective, the Groq 3 LPX brings significant advancements. Its rack-scale system includes 256 Groq 3 LP30 chips, offering 315 PFLOPS of FP8 compute and a staggering 40 PB/s on-chip SRAM bandwidth. These specifications allow the platform to handle large-context AI models efficiently, a necessity for agentic systems that generate hundreds of thousands of tokens in real time.

On the market side, NVIDIA’s stock is trading at $210.39, down 2.02% following today’s announcement. The company’s market capitalization stands at $5.13 trillion. While the stock’s dip aligns with general market volatility, the introduction of the Groq 3 LPX could serve as a long-term catalyst, especially as enterprise demand for high-performance AI inference solutions expands.

This production milestone also reflects a strategic pivot for NVIDIA. Earlier this year, the company removed its Rubin CPX accelerators from the roadmap, focusing instead on Groq 3 LPX as the cornerstone of its AI inference strategy. The move underscores NVIDIA’s commitment to dominating the latency-sensitive segment of the AI market, where token generation speed is critical for applications like coding, real-time decision-making, and multi-agent systems.

With the AI market projected to grow exponentially, NVIDIA’s Groq 3 LPX aims to capture enterprise demand for specialized inference solutions. The platform’s integration into both existing and emerging AI infrastructures could further solidify NVIDIA’s leadership in the AI acceleration space.

Image source: Shutterstock Source

LEAVE A REPLY

Please enter your comment!
Please enter your name here