Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Nvidia begins mass production of Groq 3 LPX AI inference chips
00

Nvidia begins mass production of Groq 3 LPX AI inference chips

Aug 24, 2026

Nvidia has commenced mass production of its Groq 3 LPX AI inference accelerator racks, which are built with 256 language processing units manufactured on Samsung's 4-nanometer process. Designed to resolve compounding decode latency in agentic AI, the SRAM-based system achieved a record 3,400 tokens per second in independent benchmarks. Neocloud provider Nebius will be the first to deploy the hardware before December 31, 2026, as Nvidia competes against low-latency offerings from AMD and Cerebras.

Groq 3 LPX production launch

  • ▪Amsterdam-headquartered neocloud Nebius will be the first AI cloud to deploy the Groq 3 LPX hardware through its Token Factory platform before December 31, 2026.
  • ▪The Groq 3 LPX rack integrates 256 language processing units in a liquid-cooled chassis, delivering 128 gigabytes of aggregate on-chip SRAM.
  • ▪Nvidia announced at the Hot Chips 2026 conference on August 24, 2026, that its Groq 3 LPX inference accelerator rack has entered full production.

SRAM decode architecture advantages

  • ▪The Groq 3 LPX system achieved 3,400 output tokens per second running Gemma 4 31B in benchmarks conducted by Artificial Analysis.
  • ▪Each Groq 3 LPU carries 500 megabytes of on-chip SRAM, which is approximately 576 times less than the HBM4 capacity of a single Nvidia Vera Rubin GPU.
  • ▪The Groq 3 LPU stores model weights in on-chip static random-access memory (SRAM) to deliver 150 terabytes per second of bandwidth per chip.

Agentic AI inference bottlenecks

  • ▪The decode phase of large language model inference is memory-bandwidth-bound because the model must load its full weight tensor from memory for every token generated.
  • ▪Nvidia claims the Groq 3 LPX delivers a 4x reduction in decode latency compared to the nearest GPU-based alternative.

Samsung foundry manufacturing partnership

  • ▪Samsung Electronics is the sole manufacturer of the Groq 3 LPX's language processing units, producing them on its 4-nanometer foundry lines.
  • ▪KB Securities projects that the ramp-up of 4-nanometer LPU production and a 15 percent price increase could return Samsung's foundry business to profitability in Q3 2026.
  • ▪Nvidia acquired the assets and technology of chip startup Groq in December 2025 for a reported $20 billion licensing agreement.

AI inference market competition

  • ▪Cerebras Systems announced its CS-4 system on August 18, 2026, claiming it delivers up to 30 times faster inference than GPU systems on trillion-parameter models.
  • ▪Advanced Micro Devices announced plans to integrate its rack-scale systems with chips from Cerebras to focus on low-latency AI inference.

4 sources

Cnbc
Nvidia says Groq racks will be online this year following $20 billion purchase
View source article
Techtimes
Groq 3 LPX Hits Full Production: SRAM Decode Chip Reaches 3,400 Tokens Per Second
View source article
Firstpost
Nvidia Puts Groq 3 LPX Into Full Production, Racks Set to Go Online This Year
View source article
Koreaherald
Nvidia’s Groq chip ramp brings Samsung foundry closer to profit
View source article

Featured stories

View more in Compute, chips & AI infrastructure

Marvell raises annual forecast but shares fall on Google AI deal timing concerns

Aug 27, 2026 · 3 sources

Nvidia pauses revenue-sharing deals with AI cloud companies

Aug 27, 2026 · 3 sources

Glean unveils Tau desktop workspace, claims token-cost edge over Claude

Aug 26, 2026 · 1 source

Anthropic planned $7 billion acquisition of MatX, then abandoned deal

Aug 26, 2026 · 4 sources

Story comments

Loading comments…

Related Projects

NvidiaGroqSamsung

Topics

Compute, chips & AI infrastructureAI tokensAI startupsHardware manufacturingAI inference (scaling)

Featured stories

View more in Compute, chips & AI infrastructure

Marvell raises annual forecast but shares fall on Google AI deal timing concerns

Aug 27, 2026 · 3 sources

Nvidia pauses revenue-sharing deals with AI cloud companies

Aug 27, 2026 · 3 sources

Glean unveils Tau desktop workspace, claims token-cost edge over Claude

Aug 26, 2026 · 1 source

Anthropic planned $7 billion acquisition of MatX, then abandoned deal

Aug 26, 2026 · 4 sources