NVIDIA Says Its New Inference Chip for AI Agents Is in Full Production

NVIDIA said on Aug. 24 that Groq 3 LPX, an accelerator it built for AI inference, has entered full production. The company announced it at the Hot Chips conference and describes the part as an extension of its Vera Rubin platform.
NVIDIA says the chip is aimed at how fast a system produces tokens for an individual user, the rate that governs how quickly an AI agent can finish each step of its work, rather than at training or at bulk throughput.
The performance figures in the announcement are the company's own. NVIDIA says the part reached 3,400 output tokens per second in benchmarking by the firm Artificial Analysis, running the open-source model Gemma 4 31B with a 100,000-token context, and calls that the fastest result recorded for that model. It also claims four times the responsiveness of what it calls "the nearest alternative platform," which the release does not name. Both figures come from NVIDIA's release and could not be independently verified.
NVIDIA names Nebius as the first AI cloud to adopt the accelerator. According to the release, Nebius plans to offer it through Nebius Token Factory, its production inference platform. Danila Shtan, the cloud provider's chief technology officer, is quoted saying developers will reach it "through the same API developers are already using, with no migration to a new stack." NVIDIA says the inference cloud provider Groq plans to be among the next adopters.
The release ends with NVIDIA's standing disclaimer that many of the products and features it describes "remain in various stages and will be offered on a when-and-if-available basis."
