Nvidia has commenced mass production of its Groq 3 LPX AI inference accelerator racks, which are built with 256 language processing units manufactured on Samsung's 4-nanometer process. Designed to resolve compounding decode latency in agentic AI, the SRAM-based system achieved a record 3,400 tokens per second in independent benchmarks. Neocloud provider Nebius will be the first to deploy the hardware before December 31, 2026, as Nvidia competes against low-latency offerings from AMD and Cerebras.
Aug 27, 2026 · 3 sources
Aug 27, 2026 · 3 sources
Aug 26, 2026 · 1 source
Aug 26, 2026 · 4 sources
Story comments
Loading comments…