Groq raises $650M to scale its inference-only cloud
Groq — the company behind the LPU, a chip purpose-built for inference rather than training — raised $650 million, led by Disruptive and Infinitum, to expand its inference cloud. It says it runs 13 data centers across four continents, serves more than 5 million developers, and processes trillions of tokens a week, targeting 200 MW of capacity by the end of 2027. Notably, the raise comes roughly seven months after Nvidia paid about $20 billion to license Groq's LPU architecture.
Why it matters: The subtext here is that inference and training are splitting into two different hardware markets with different physics. GPUs were designed to chew through the massive parallel math of training; generating tokens fast and cheap for always-on production traffic is a distinct problem, and dedicated silicon can win on latency and cost-per-token where general-purpose GPUs leave money on the table. The most revealing detail is that Groq is scaling a rival cloud even after Nvidia licensed its core architecture — a sign the demand for specialized inference is large enough to support both the incumbent and the challenger rather than a winner-take-all outcome. For builders, more competition at the inference layer is straightforwardly good: it pushes token prices down and pressures latency, which is exactly what turns AI from a costly demo into something you can afford to run in the hot path of a real product.