Nebius Token Factory: Estimating the Hidden Managed-Inference Business
Executive conclusion
Nebius does not disclose standalone revenue, annualized revenue, token volume, customer count, or gross margin for Token Factory. The business is reported inside Nebius AI Cloud, and dedicated endpoints are billed per GPU-hour while shared endpoints are billed per token. Any estimate therefore requires reconstructing the business from public traffic, named customers, throughput disclosures, and competitor economics.
My best estimate as of July 28, 2026 is:
| Metric | Central estimate | Defensible range |
|---|
| Current Token Factory annualized product revenue | ~$40M | $10M–$100M |
| Current text-token-equivalent volume | ~5T–7T/month | ~2T–15T/month |
| Average daily text-token-equivalent volume | ~170B–230B/day | ~70B–500B/day |
| Share of Nebius Q1 2026 AI Cloud ARR | ~2% | <1%–5% |
| Estimated year-end 2026 Token Factory ARR | ~$100M | $50M–$200M |
These estimates refer to gross product revenue attributable to workloads sold through Token Factory, including the underlying compute. They should not be added on top of Nebius AI Cloud guidance: Token Factory is a product layer within that segment. The genuinely incremental software-layer economics—better utilization, serving optimization, and a managed-inference price premium—are probably only a fraction of the gross product revenue today.
The central conclusion is that Token Factory is a real production business with credible large customers, but it is still very small relative to Fireworks AI and Together AI. It is probably measured in tens of millions of current ARR rather than hundreds of millions, and certainly not yet at Fireworks's $1B-plus revenue scale.
What is actually disclosed
Financial context
Nebius reported:
- Q1 2026 Nebius AI Cloud revenue of $389.7M.
- AI Cloud ARR of $1.92B at the end of March 2026.
- Full-year 2026 revenue guidance of $3.0B–$3.4B.
- Year-end 2026 ARR guidance of $7B–$9B.
- Token Factory customer wins including Revolut and monday.com.
- No standalone Token Factory revenue, ARR, token volume, customer count, or profitability.
Management described Token Factory as already seeing momentum, but the disclosed 3.5x sequential pipeline increase covers both the core cloud and Token Factory. It cannot be treated as a Token Factory growth rate.
Observable operating evidence
| Evidence | Disclosure | What it proves | What it does not prove |
|---|
| Prosus | Dedicated autoscaling endpoints supporting workloads of up to 200B tokens/day; up to 26x lower cost than proprietary models | The platform can operate at very large production scale for one customer | Average utilization, contract value, model mix, or realized price |
| Revolut | 1.2M support tickets/month on Token Factory; 80% reportedly handled without human intervention | A live enterprise workflow with measurable application volume | Tokens per ticket or revenue |
| monday.com | Selected Token Factory for its AI Work Platform | Direct enterprise customer relationship | Contract size or share of monday's AI traffic |
| monday.com subprocessor list | Nebius provides a data embedding service and AI functionalities, applicable across all monday data regions | Independent confirmation that Nebius is an approved production AI subprocessor | Volume, exclusivity, or which user-facing features use Nebius |
| OpenRouter | Approximately 371.9B Nebius tokens/month, 21.8B/day, and 13–14 models in a late-July snapshot | A directly observable public distribution channel | Total direct Nebius traffic |
| Higgsfield AI | Uses Nebius for on-demand and autoscaling inference | Potentially valuable GPU-intensive multimodal production usage | Whether revenue is recorded under Token Factory or general AI Cloud |
| Hugging Face and Mithril | Public Token Factory relationships | Ecosystem reach and additional production demand | Standalone commercial scale |
| Nebius product disclosure | Customers already run workloads producing hundreds of billions of tokens per day | Aggregate or peak production capacity is at least in the low hundreds of billions per day | Sustained average traffic |
The most important accounting distinction
“Token Factory revenue” can mean two different things:
- Gross product revenue: the entire customer bill for a managed endpoint, including GPUs, networking, storage, and the inference software layer.
- Incremental software value: the additional revenue or gross profit Nebius earns because it manages serving, pooling, autoscaling, routing, post-training, and optimization rather than leasing raw GPUs.
This note estimates the first definition because it is the only practical way to size the product. The second number is economically more important but much smaller and impossible to isolate publicly. If Token Factory currently produces roughly $40M of gross product ARR, a plausible incremental software and utilization benefit might be only $5M–$15M of annualized gross profit or revenue uplift. That is an analytical placeholder, not a disclosed figure.
Dedicated endpoints reinforce this distinction. Nebius bills them per GPU-hour, with capacity isolated for the customer; shared serverless endpoints are billed per token. Therefore, raw token volume and revenue cannot be converted with one universal price.
Sizing method 1: OpenRouter establishes the public floor
The OpenRouter provider table shows approximately 227T total tokens per trailing month across listed providers. Nebius contributes approximately 371.9B, or:
- ~0.16% of OpenRouter monthly token volume
- ~0.30% of current daily token volume
- Approximately 38th by trailing-month tokens and 30th by current-day tokens
This is a real integration but a small one. OpenRouter is unlikely to be financially material to Nebius.
Illustrative annualized gross billings on 371.9B tokens/month are:
| Blended realized price | Annualized billings |
|---|
| $0.10 per 1M tokens | ~$0.45M |
| $0.25 per 1M tokens | ~$1.12M |
| $0.50 per 1M tokens | ~$2.23M |
| $1.00 per 1M tokens | ~$4.46M |
Common Nebius-hosted open models on OpenRouter have input pricing near $0.10–$0.30 per million and output pricing around $0.30–$1.20, although several frontier models are much more expensive. Input-heavy and cached agent traffic normally produces a realized blended rate far below the simple average of input and output list prices.
I therefore estimate that the OpenRouter channel currently represents approximately
$0.5M–$2.5M of Nebius annualized product revenue, with an outer ceiling around $5M under a high-value model mix.
This is the floor, not the whole business. Direct enterprise traffic is clearly larger.
Sizing method 2: Prosus is the principal throughput anchor
Prosus disclosed the ability to handle workloads of up to 200B tokens/day on Token Factory dedicated endpoints. If sustained continuously, that would equal:
- Approximately 6T tokens/month
- Approximately 73T tokens/year
The phrase “up to” is critical. It likely represents peak or designed throughput rather than average billable consumption.
A simple utilization and price sensitivity produces:
| Case | Average utilization of 200B/day capacity | Blended value per 1M tokens | Implied annualized revenue |
|---|
| Low | 20% | $0.15 | ~$2.2M |
| Base | 40% | $0.25 | ~$7.3M |
| High | 60% | $0.40 | ~$17.5M |
| Extreme maximum | 100% | $1.00 | ~$73M |
Actual dedicated-endpoint revenue can differ because the product is billed per GPU-hour rather than per token. Minimum capacity commitments can raise revenue when utilization is low; enterprise discounts can lower it. The best working estimate is therefore
$3M–$20M of Prosus-related annualized Token Factory revenue, centered near $8M–$10M.
Prosus alone could be larger than all public OpenRouter traffic, but the public evidence does not support assuming that its 200B/day peak is continuously utilized.
Sizing method 3: the customer and product build
Public serverless and OpenRouter
This includes OpenRouter plus direct self-service use of public endpoints. OpenRouter contributes only $0.5M–$2.5M of estimated ARR. Direct usage should be larger, but Nebius's modest public catalog traffic and weak OpenRouter share argue against a very large self-service business.
Estimated current ARR:
$1M–$8M, base case
$3M.
Prosus dedicated endpoints
The throughput disclosure makes Prosus the strongest identifiable enterprise revenue contributor.
Estimated current ARR:
$3M–$20M, base case
$8M.
Revolut
The 1.2M-ticket monthly disclosure is operationally impressive but does not necessarily imply high inference revenue. If a support ticket consumes 20,000–200,000 aggregate model tokens across routing, context, repeated calls, and generation, the workload would equal roughly 24B–240B tokens/month. At $0.15–$0.50 per million, annual usage revenue would be only approximately $40,000–$1.4M. A dedicated capacity floor could increase the contract.
Estimated current ARR:
$0.25M–$5M, base case
$1M.
The important lesson is that application volume and inference revenue are not synonymous. A million automated conversations can be strategically valuable to the customer while remaining a modest infrastructure contract.
monday.com
monday.com's own subprocessor list independently identifies Nebius for:
- Platform-level data embeddings using full account data, temporarily processed and deleted.
- AI Work Platform “AI functionalities” using board data, temporarily processed and deleted.
- Availability across all monday.com data regions.
This is stronger evidence than a customer-logo announcement. It indicates real production integration. At the same time, monday publicly uses a multi-provider architecture involving Anthropic, OpenAI, Google, AWS Bedrock, Azure AI, Google Vertex, and self-hosted environments. Nebius is one inference and embedding backend, not the exclusive AI platform.
Embeddings are extremely cheap per token, and no share of monday's generative inference has been disclosed.
Estimated current ARR:
$0.5M–$10M, base case
$3M. Confidence is low.
Higgsfield and multimodal inference
Higgsfield is potentially more valuable in revenue terms than the text-token customers because generative video is GPU intensive. Nebius presents Higgsfield as a Token Factory user for on-demand and autoscaling inference. However, neither GPU count nor commercial classification is disclosed, and part of the bill may be recorded economically as ordinary AI Cloud consumption.
Estimated current ARR attributable to Token Factory:
$2M–$30M, base case
$10M. Confidence is low.
Hugging Face, Mithril, other direct customers, and post-training
These relationships establish a longer tail of production inference, batch usage, model hosting, fine-tuning, and post-training. Token Factory is also increasingly a custom-model platform rather than merely a public model API. That should create revenue not visible on OpenRouter.
Estimated current ARR:
$3M–$30M, base case
$12M.
Combined bottom-up estimate
| Revenue component | Low | Base | High |
|---|
| Public serverless and OpenRouter | $1M | $3M | $8M |
| Prosus | $3M | $8M | $20M |
| Revolut | $0.25M | $1M | $5M |
| monday.com | $0.5M | $3M | $10M |
| Higgsfield / multimodal | $2M | $10M | $30M |
| Other customers and post-training | $3M | $12M | $30M |
| Estimated current annualized revenue | ~$10M | ~$37M | ~$103M |
I round the base case to
approximately $40M and the defensible range to
$10M–$100M.
This table should not be read as account-level forecasts. The customer rows are analytical allocations designed to prevent the estimate from relying on one unsupported aggregate multiplier.
Sizing method 4: token-volume reconstruction
A reasonable reconstruction of current text-token-equivalent volume is:
| Source | Estimated monthly volume |
|---|
| OpenRouter | 0.37T disclosed |
| Prosus average usage | ~1.2T–3.6T, assuming 20%–60% of disclosed peak capacity |
| Revolut | ~0.02T–0.24T under a broad per-ticket assumption |
| monday.com, Hugging Face, Mithril, direct serverless, and other customers | ~1T–6T, highly uncertain |
| Likely current total | ~2T–10T/month |
| Central estimate | ~5T–7T/month |
The wider plausible range is approximately 2T–15T/month, or 70B–500B tokens/day. This is consistent with Nebius saying customers run workloads producing hundreds of billions of tokens per day, while recognizing that such statements often describe peak rather than average throughput.
Cross-check against Fireworks and Together
Fireworks reported more than:
- 40T tokens/day
- $1B annualized revenue
- 95% of tokens from customer-specialized models
- A $17.5B valuation
Together reported more than:
- 400T tokens/month
- Approximately 13.3T tokens/day
At the Token Factory central estimate of 5T–7T/month, Nebius would be approximately:
- 0.4%–0.6% of Fireworks's reported token volume
- 1.3%–1.8% of Together's reported token volume
- Approximately 4% of Fireworks's reported ARR at the $40M central revenue estimate
This order-of-magnitude gap is consistent with all observable evidence. Fireworks and Together publicly emphasize their enormous token scale. Nebius instead emphasizes customer wins, platform breadth, infrastructure ownership, and newly acquired optimization capabilities. If Token Factory were already operating at several hundred million dollars of ARR or tens of trillions of tokens per day, the absence of a disclosed platform-scale metric would be surprising.
OpenRouter is misleading if viewed in isolation. Nebius's 0.37T monthly volume is only moderately below Fireworks's 1.2T and Together's 0.76T on OpenRouter. But OpenRouter represents only approximately 0.1% of Fireworks's company-wide volume and 0.2% of Together's. The channel is disproportionately important to the appearance of smaller providers and does not measure their total business accurately.
Why raw token counts can look impossibly large
A token is not a standardized unit of compute or revenue:
- Prompt tokens are generally cheaper than output tokens.
- Cached prompt tokens can consume a small fraction of the compute of newly processed tokens.
- Agent systems repeatedly resend large system prompts, tool schemas, and context.
- Different model architectures activate radically different numbers of parameters.
- Providers use different tokenizers.
- Dedicated capacity can be billed even when the endpoint is not fully utilized.
- Multimodal workloads do not map cleanly into text-token comparisons.
Together's provisioned-throughput economics illustrate the distortion. One PTU can serve approximately:
- 23,140 output tokens per minute
- 138,840 ordinary input tokens per minute
- 694,200 cached-input tokens per minute
The same capacity therefore processes approximately 30 times as many cached-input tokens as output tokens. Derived cached-input economics are about $0.072 per million tokens, compared with approximately $0.36 for ordinary input and $2.16 for output.
Fireworks's disclosed $1B ARR divided by 40T daily tokens implies only about $0.0685 of revenue per million raw tokens—almost identical to Together's cached-input capacity economics. That does not invalidate the throughput. It demonstrates that provider headline token totals are dominated by low-cost input, cached traffic, optimized specialized models, or measurement scopes that are not directly comparable with billed output tokens.
Probability-weighted current revenue view
My subjective probability distribution is:
| Current annualized Token Factory revenue | Probability |
|---|
| Below $10M | 15% |
| $10M–$25M | 30% |
| $25M–$60M | 35% |
| $60M–$100M | 15% |
| Above $100M | 5% |
Using rough interval midpoints produces a probability-weighted estimate close to
$40M.
The probability above $100M is deliberately low. Reaching that figure would require some combination of:
- Sustained Prosus utilization close to the disclosed maximum;
- A materially larger Higgsfield or undisclosed multimodal contract;
- Direct private traffic more than 100 times the observable OpenRouter channel;
- Meaningful dedicated capacity commitments not visible in token telemetry; or
- Several additional large custom-model customers that Nebius has not yet disclosed.
All are possible, but none is supported strongly enough to make $100M-plus the base case.
Year-end 2026 estimate
A useful sensitivity is Token Factory's share of Nebius's $7B–$9B year-end AI Cloud ARR guidance:
| Token Factory share of AI Cloud ARR | Implied Token Factory ARR |
|---|
| 0.5% | $35M–$45M |
| 1.0% | $70M–$90M |
| 2.0% | $140M–$180M |
| 5.0% | $350M–$450M |
This is not a derivation because strategic hyperscaler capacity and Token Factory have different business models. It is a reasonableness check.
Given the current central estimate near $40M, the product's early stage, enterprise ramps, and the integration of Eigen AI and Clarifai technology, I would carry:
- Conservative year-end 2026 ARR: $50M
- Base year-end 2026 ARR: approximately $100M
- Upside year-end 2026 ARR: $200M
- Speculative breakout case: $350M-plus
The $350M-plus case would require a major new disclosure and should not be embedded in a base Nebius valuation today.
Investment interpretation
Token Factory is currently immaterial to consolidated revenue but potentially material to long-term economics.
The strategic value is not its present $40M estimated run rate. It is the possibility that Nebius can attach managed inference, post-training, data tooling, and agent services to a multibillion-dollar GPU fleet. Even a modest software and utilization improvement across the broader AI Cloud could matter more than standalone Token Factory revenue.
Potential benefits include:
- Higher GPU utilization through pooled and autoscaled inference;
- More revenue per GPU-hour from optimized serving;
- A shift from commodity capacity sales toward custom models and workflows;
- Higher switching costs from fine-tuning, custom weights, observability, data pipelines, and dedicated endpoints;
- Better customer diversification than a small number of hyperscaler contracts;
- A path from Nebius compute to Tavily search and future agentic services.
The risk is that Fireworks, Together, hyperscalers, and low-cost inference specialists establish the developer and enterprise control planes before Nebius's software layer reaches scale. Infrastructure ownership is valuable, but it does not automatically produce inference-platform distribution.
What would change the estimate
The estimate should be revised immediately if Nebius discloses any of the following:
- Total Token Factory tokens served per day or month.
- Standalone Token Factory revenue or ARR.
- Dedicated endpoint count, committed capacity, or average GPU utilization.
- The average mix of shared, batch, and dedicated inference.
- Public versus custom-model token share.
- Realized revenue per million input, cached-input, and output tokens.
- A quantified Higgsfield or monday.com contract.
- Evidence that Prosus's 200B-token daily capacity is sustained rather than peak.
- A large new inference-only customer with disclosed volume.
- Token Factory gross margin or a quantified uplift to AI Cloud margins.
Until then, the appropriate investment-model placeholder is
approximately $40M of current ARR and $100M at year-end 2026, with wide error bars and no incremental addition to reported AI Cloud guidance.
Sources