OpenAI on Cerebras: Explicit Use Cases (Ro Varma / Big Chip Club)
Source
Cerebras
Big Chip Club interview with
Ro Varma, OpenAI product lead for Codex, enterprise, and the Cerebras work. Background: engineer/founder, Palantir, Cursor; joined OpenAI in 2026. Quotes below are from that conversation, lightly cleaned of filler ("um," repeated "like") but not paraphrased. This is a product interview, not a contract or capacity disclosure.
What this is: OpenAI using Cerebras as the inference backend for an ultrafast Codex tier (
Codex Spark), described as a step-function faster frontier model and "the first of many" joint launches.
What this is not: Training, ChatGPT, API mix, live MW, $/MW, or the 750 MW MRA. Ro never discusses those. Internal OpenAI Codex anecdotes (Sites, flight booking, Slack-to-Slack agents, seating charts) are
Codex usage, not evidence those workloads run on Cerebras.
Explicit Cerebras use cases
| Use case | Why it is Cerebras-specific | Revenue relevance |
|---|
| Ultrafast Codex inference (Codex Spark) | Named as the Cerebras-hosted model vs GPU Fast Mode | Primary disclosed product |
| Interactive / "jamming" iteration (frontend, exploratory) | 20s loop vs ~10 min walk-away | Changes how the product is used, not just wait time |
| Creative tasks | Only worth doing if the model is fast enough to iterate | New usage, not just faster existing usage |
| Sub-agent swarms (review, test) | "Cerebras has a sub-agent model"; spawn many and get them back in time | More tokens / more parallel jobs per user |
| Computer use at human speed | Slow frontier models feel clunky; Cerebras makes takeover feel human | Unlocks a product category that needs both intelligence and latency |
| Time-critical async (incident-response agent) | 20 min vs 5 vs 2 is a business difference | Willingness-to-pay for latency, not just intelligence |
| Customer AI products redesigned around speed | CTOs asking when they can get Cerebras | Demand signal outside OpenAI's own Codex surface |
Revenue-relevant quotes
1. Economic break vs GPU Fast Mode (the core claim)> Fast mode is a little bit different because we're essentially able to serve — by using more capacity, we can effectively serve the model faster. This is the main way that most speed increases you see in the market are served. If you apply more compute you do get faster results, but it exponentially backs off in both directions. We offer fast mode at 1.5x speed and it costs about 2.5x. If we wanted to offer 2x or 3x speed with the same mechanism, it doesn't linearly increase the compute cost, it would exponentially increase. What's interesting is with Cerebras, it kind of completely blows up that equation because by using different hardware, we can essentially achieve something that's completely out of, even getting to 10x speed in a completely economically viable way.
> The speed that we unlock with Cerebras is magnitudally faster compared to even something like fast mode in Codex.
2. Product: first of many joint models> Codex Spark was the first of many future models that Cerebras and OpenAI are very excited to launch together, and it's a step function faster in terms of state-of-the-art models than what we've seen before.
3. Named demand: enterprise CTOs> Even just the possibility of the speed that we can offer with it has caused a ton of — I get texts from different CTOs all the time being like, when can we get Cerebras?
> Our customers who are serving AI products have all these ideas of how their products can be fundamentally different if speed is optimized for a lot more. So far in this AI development journey it's just been about intelligence, more and more intelligence, max out intelligence. The product AI product community has always been: if I can get more intelligence I can do more important work for my customers. I don't care about latency because the output is more useful. But now if you say, actually, latency is another variable you can play around with, suddenly it unlocks even different user experiences that I think our customers are really excited about.
4. Capacity ramp (from zero)> OpenAI, we operate at a massive scale in terms of the amount of compute that we need. We're famously very compute-hungry. With Cerebras you all are bringing on these big chips that are very fast and obviously starting from zero capacity and going up from there. There's a ton of optimizations that can be made across the full stack to actually just get the most out of those chips, whether it's the hardware, the way that the model lives on it, the way that the software optimizations work around it.
5. Use cases that need Cerebras-class speedInteractive iteration:
> Speed actually has a big influence on the experience. Frontend iteration is a really good example. If it takes Codex 10 minutes to do the next iteration, then you probably went and did something else. If it takes 20 seconds, then suddenly you're sitting there iterating very quickly with it, jamming.
Sub-agents:
> We'll have sub-agents that get launched to do very targeted tasks, let's say review a thing or test a thing. With Cerebras, [it] has a sub-agent model. It enables spawning a lot more of those things and then those actually coming back with responses in a way that can unlock a lot more of those use cases.
Computer use:
> Computer use is one of those things where the speed actually makes it feel like truly magical. Frontier models today — you kind of need both the intelligence and the speed. The intelligence is great, but if it's really slow it feels kind of clunky. With Cerebras you can watch your computer actually get taken over in a way that it feels like a human is using it.
Incident response:
> There's really interesting async workloads that are kind of really time-critical. For example, an incident response agent: if it takes 20 minutes versus 5 minutes versus 2 minutes, that could be a pretty impactful difference to your business.
6. Partnership depth (not a volume number, but a lock-in hint)> We've been super lucky to partner with Cerebras on bringing these models online. The work that your team has to do to bring these models online is actually very custom to our models, in deep collaboration with our research team, our compute team, our inference team, with our product team. You all have been [involved] both when it comes to making the model work on the chips, also bringing compute online and data centers, and kind of collaboration with us.
What not to attribute to Cerebras
Ro's Codex-at-OpenAI stories (100% of the company on Codex, Sarah Friar / finance, marketing video automation, Sites/AppGarden, event seating apps, Notion→site via computer use, 2 a.m. business-class booking, Slack messages sent by agents) describe
Codex as a daily driver. He does not say those jobs run on Cerebras. Internally he says most people sit on
GPU Fast Mode because 1.5× speed company-wide is ~50% more throughput. Cerebras is framed as the next step-function on top of that, still ramping from zero capacity.
Takeaway
The disclosed OpenAI×Cerebras
product is ultrafast Codex inference (Spark), sold on a speed/cost curve GPUs cannot match (~10× economically vs Fast Mode's 1.5× at 2.5× cost). The disclosed
demand is (i) OpenAI's own interactive Codex surfaces and (ii) enterprise CTOs already asking for access so their AI products can treat latency as a design variable. Capacity is explicitly still ramping from zero. None of this sizes the 750 MW contract; it does say what OpenAI thinks the hardware is
for.