Tools
NVIDIA puts Groq 3 LPX inference accelerator in production
NVIDIA said Groq 3 LPX, its interactive inference accelerator, is now in full production as an extension of Vera Rubin. The company cites 3,400 output tokens per second on Gemma 4 31B at 100,000-token context, four times the nearest alternative in an Artificial Analysis benchmark. This is NVIDIA hardware, not xAI's Grok model, and there is no affiliate block.