This page was machine-translated and may differ from the original. View original
NVIDIA Begins Mass Production of Groq 3 LPX…Achieves 3,400 Tokens Per Second Inference Performance
Vera Rubin NVL72 Platform Expansion, AI Cloud Companies Including Nebius and Groq Plan Initial Adoption
NVIDIA Groq 3 LPX has entered mass production. Recording 3,400 output tokens per second in the Artificial Analysis benchmark, it is evaluated as raising the performance standard in the agentic AI inference accelerator market.
NVIDIA announced on the 25th that it will begin full-scale mass production of Groq 3 LPX, an interactive AI inference accelerator.
This product is configured to expand the Vera Rubin NVL72 platform, focusing on increasing the token generation speed of agentic systems.
The previously mentioned benchmark figures were measured under conditions where 100,000 token context was applied to the open-source agentic model Gemma 4 31B, and NVIDIA stated that based on that model, it is currently the fastest performance achieved.
Groq 3 LPX Performance and Architecture
Agentic systems generate large-scale tokens during hundreds to thousands of inference steps.
Groq 3 LPX provides response speeds 4 times faster than the most similar alternative platform in this generation step, according to NVIDIA.
This means agentic tasks such as coding can be processed within minutes rather than hours.
The Vera Rubin platform features a rack-scale system composed of △NVIDIA BlueField®-4 DPU △Vera CPU rack △Vera BlueField-4 STX storage △SpectrumTM-6 SPX Ethernet.
NVIDIA Founder and CEO Jensen Huang stated, "Vera Rubin expands its vision with an AI factory configuration optimized for workloads in the agentic AI era, and LPX for ultra-fast token generation pushes the limits of performance even further."
Initial Adoption by AI Cloud Companies
AI cloud company Nebius plans to be the first to adopt Groq 3 LPX in its inference platform, Nebius Token Factory.
Nebius CTO Danila Shtan stated, "We provide developers with an experience where all steps of the agent loop happen instantaneously while using existing APIs as they are."
AI inference-specialized cloud company Groq is also planned to be among the early adopting companies, NVIDIA said.
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.















