Aiming at the low-latency AI inference market, Nvidia (NVDA.US) Groq 3 LPX enters full-scale production, with the first batch of systems to be deployed at NEBIUS (NBIS.US)
Nvidia announced on Monday that the Groq 3 LPX rack-level system has entered full-scale mass production, marking the official commercialization of this low-latency AI inference technology following the company's roughly $20 billion acquisition of Groq-related assets last year.
According to Zhitong Finance APP, Nvidia (NVDA.US) announced on Monday that the Groq 3 LPX rack-level system has entered full-scale mass production, marking the official commercial deployment of this low-latency AI inference technology following the company’s approximately $20 billion acquisition of Groq-related assets last year. The first batch of systems will be deployed by AI cloud computing provider NEBIUS (NBIS.US) and is scheduled to go online later this year.
Nvidia Senior Director Dion Harris told the media that Groq 3 LPX will be deployed alongside Nvidia Vera CPUs and Rubin GPUs in Nebius data centers.
The rapid mass production of Groq products reflects the shift in AI applications from model training to real-world deployment, with low-latency AI inference becoming a new market focus for Nvidia. This is particularly important for applications such as AI agents and AI programming, where models must continuously generate content at higher speeds, reducing user wait times.
Last December, Nvidia acquired assets related to AI chip startup Groq for approximately $20 billion, making it the largest acquisition in the company’s history.
A key feature of Groq’s chip architecture is the integration of 500MB high-speed SRAM within the chip, reducing data transfer bottlenecks caused by traditional memory access and thus increasing AI model inference response speed.
Unlike Nvidia’s main GPUs, which are manufactured by TSMC (TSM.US), Groq chips are manufactured by Samsung.
Currently, Nvidia integrates 256 Groq 3 chips into one LPX rack. According to benchmark data from Artificial Analysis cited by Nvidia, the Groq 3 LPX system is capable of generating approximately 3,400 tokens per second.
In generative AI, a token can be understood as the basic unit for processing and generating text. A higher token generation speed per second means the AI can respond to user requests faster, which is especially important for latency-sensitive applications like AI programming and real-time intelligent agents.
Harris stated that for cloud computing companies providing AI inference services, lower latency means they can offer higher-priced premium services to clients who require faster response times.
However, Groq chips are not intended to replace Nvidia’s traditional GPUs.
GPUs are still the core of today’s AI computing infrastructure. They not only perform AI model training but also inference tasks, offering greater versatility and adaptability for different models and technical architectures.
Low-latency chips like Groq focus more on specific stages of the AI inference process, especially the “decoding” phase when the model is generating content.
Harris said: “This isn’t about replacing GPUs, but about using the most cost-effective and high-performance processor for each part of the workload.”
This means Nvidia is attempting to establish a more segmented AI computing architecture: Vera CPUs and Rubin GPUs will handle broad AI computing tasks, while Groq chips are optimized for latency-sensitive inference workloads.
As the AI industry moves from large-scale model training to rapidly growing inference demand, competition in the low-latency inference market is heating up.
AMD (AMD.US) announced earlier this year that it will integrate its rack-level AI systems with Cerebras (CBRS.US) chips, also targeting low-latency AI inference.
The importance of this field is growing with the proliferation of applications such as AI programming. Reports mentioned that OpenAI’s newly released Ultrafast mode promises 750 tokens per second, powered by underlying compute from Cerebras.
By comparison, benchmark tests cited by Nvidia show Groq 3 LPX achieving up to 3,400 tokens per second. However, test conditions and application scenarios may differ between systems, so direct performance comparisons should be interpreted cautiously.
Meanwhile, Nvidia is accelerating shipments of the Vera Rubin system, which entered production earlier this year.
Nvidia CEO Jensen Huang, when unveiling Vera Rubin and Groq 3 LPX in March, predicted that cumulative sales from the current Blackwell chips to the next-generation Vera Rubin systems could reach $1 trillion by 2027.
Huang also revealed at the time that, in data center spaces dedicated to AI programming applications, he plans to allocate about a quarter to Groq chips, with the rest entirely deploying Vera Rubin systems.
This allocation further demonstrates Nvidia’s emphasis on the low-latency inference market. As AI agents, AI programming, and real-time generative AI applications rapidly develop, inference speed is becoming a key competitive metric for cloud computing companies and AI developers.
The official mass production of Groq 3 LPX also means Nvidia is expanding beyond GPU-centric AI chip supply, moving towards providing specialized computing architectures for different AI workloads.
The market will next focus on Nvidia’s latest financial performance. The company is scheduled to release its earnings report this Wednesday. In addition to demand for Blackwell and Vera Rubin, AI inference business and the commercialization progress of Groq may also become new focal points for investors.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like

Tesla (TSLA.US) Officially Ends Its Decade-Long "Solar Roof Tile Dream": Production Halted Due to Financial Unsustainability, Energy Strategy Shifts Entirely to Traditional Photovoltaics and Energy Storage
Tesla stops selling solar roofs ten years after their launch.

Mystery 49,000 ETH Transfer Moves $121.9M as Ethereum Price Holds Steady
