AIThe Register1h ago

Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second

Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence

Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second

TL;DRNvidia's Groq hardware achieved record token generation speeds in a benchmark test.

Why it matters: Faster inference speeds could reshape AI economics by reducing latency and compute costs for LLM applications.

Nvidia's $20 billion bet on Groq's LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse …

Read full article

Source: The Register · Opens in new tab