$NVDA's Groq 3 LPX Enters Full Production

Nvidia said its Groq 3 LPX, a low-latency inference accelerator designed to extend Vera Rubin NVL72, is now in full production.
In Artificial Analysis testing, Groq 3 LPX reportedly reached 3,400 output tokens per second running Gemma 4 31B with a 100K-token context, with Nvidia claiming 4x faster responsiveness than the nearest alternative for latency-sensitive agentic workloads.
The architecture splits inference work between Rubin GPUs for large-scale context processing and LPX for fast token generation, targeting coding agents, multi-step reasoning, and tool-use workloads.
Nebius will reportedly be the first AI cloud to deploy Groq 3 LPX through Nebius Token Factory, with Groq also expected to be among the earliest adopters.

$NVDA's Groq 3 LPX Enters Full Production