NVIDIA Groq 3 LPX vs. Cerebras CS-4

Created by@hypertonxvia MCP
August 25, 2026 at 6:58 AM

NVIDIA Groq 3 LPX vs. Cerebras CS-4

Bottom line

NVIDIA LPX beat Cerebras’s deployed CS-3-era endpoint—not necessarily its contemporaneous CS-4 platform. LPX and CS-4 are arriving in essentially the same product window, so the real race remains unresolved.
TimelineEvent
Jun. 29, 2026Cerebras launches Gemma 4 on its existing inference cloud
Aug. 18Cerebras announces CS-4; shipments begin this quarter
Aug. 20Artificial Analysis tests NVIDIA Groq 3 LPX
Aug. 24NVIDIA declares LPX in full production
Measured Gemma 4 31B speed: LPX delivered 3,382 tok/s/user at 10k context, versus approximately 1,402 tok/s for Cerebras’s public endpoint—a 2.4× lead.
But CS-4 claims up to 2× CS-3 per-user speed and 10× throughput/watt. CS-4 is a rack-scale system built from three WSE-3 Turbo wafers, not a WSE-4 processor.
A rough, unverified Gemma extrapolation puts CS-4 around 2,804–3,702 tok/s, depending on whether the CS-3 baseline is Artificial Analysis’s ~1,402 tok/s or Cerebras’s 1,851 tok/s launch result. That brackets NVIDIA’s 3,382 tok/s.

Investment read

NVIDIA has erased Cerebras’s uncontested speed narrative, but has not demonstrated a decisive generational lead. The relevant comparison is now LPX vs. CS-4 on identical models and production conditions.
Watch:
  • Independent CS-4 Gemma benchmarks
  • Sustained performance at concurrency and P95
  • Tokens per dollar and per watt
  • Public cloud availability and pricing
  • Model breadth and deployment speed
Conclusion: strategically important, but not yet evidence that NVIDIA has surpassed Cerebras’s current-generation economics or architecture.
Show actions

Start a new chat seeded with this artifact.

Built with Modeledge MCP

Connect your MCP client to research filings and earnings calls, then build a research note like this one.

Connect your agent