NVIDIA Groq 3 LPX vs. Cerebras CS-4
Bottom line
NVIDIA LPX beat Cerebras’s deployed CS-3-era endpoint—not necessarily its contemporaneous CS-4 platform. LPX and CS-4 are arriving in essentially the same product window, so the real race remains unresolved.
| Timeline | Event |
|---|
| Jun. 29, 2026 | Cerebras launches Gemma 4 on its existing inference cloud |
| Aug. 18 | Cerebras announces CS-4; shipments begin this quarter |
| Aug. 20 | Artificial Analysis tests NVIDIA Groq 3 LPX |
| Aug. 24 | NVIDIA declares LPX in full production |
Measured Gemma 4 31B speed: LPX delivered
3,382 tok/s/user at 10k context, versus approximately
1,402 tok/s for Cerebras’s public endpoint—a
2.4× lead.
But CS-4 claims up to 2× CS-3 per-user speed and 10× throughput/watt. CS-4 is a rack-scale system built from three WSE-3 Turbo wafers, not a WSE-4 processor.
A rough, unverified Gemma extrapolation puts CS-4 around
2,804–3,702 tok/s, depending on whether the CS-3 baseline is Artificial Analysis’s ~1,402 tok/s or Cerebras’s 1,851 tok/s launch result. That brackets NVIDIA’s 3,382 tok/s.
Investment read
NVIDIA has erased Cerebras’s uncontested speed narrative, but has
not demonstrated a decisive generational lead. The relevant comparison is now
LPX vs. CS-4 on identical models and production conditions.
Watch:
- Independent CS-4 Gemma benchmarks
- Sustained performance at concurrency and P95
- Tokens per dollar and per watt
- Public cloud availability and pricing
- Model breadth and deployment speed
Conclusion: strategically important, but not yet evidence that NVIDIA has surpassed Cerebras’s current-generation economics or architecture.