Nebius has introduced Nebius Token Factory, a production inference platform for open-source and custom AI models. The company says the service supports more than 60 text, code, and vision models, offers OpenAI-compatible APIs, and separates input and output token pricing to make inference costs easier to compare.
Nebius has launched Nebius Token Factory, a production AI inference platform designed to serve open-source and custom models at scale.
In a newsroom post, Nebius said Token Factory is intended for production inference rather than experimentation alone, with support for more than 60 text, code, and vision models. The company describes the platform as offering transparent cost-per-token pricing, volume discounts, and deployment options for open-source and custom models.
Nebius’s Token Factory product page says the service provides OpenAI-compatible APIs, which is intended to make it easier for developers to move applications that already use OpenAI-style interfaces. The same product page also emphasizes a pricing structure that separates input and output tokens, rather than presenting only a single headline price.
That pricing distinction matters because many AI applications have uneven token usage. A document summarization tool, for example, may send a large amount of text into a model and receive a shorter answer. A coding or reasoning assistant may generate comparatively long responses, making output-token pricing more important for total cost.
Nebius Token Factory documentation defines input tokens as the tokens included in an API request and priced per million tokens. It defines output tokens as the tokens generated in the model response, also priced per million tokens.
This separation is common in AI inference pricing because model providers often charge different rates for reading a prompt and generating a response. Output tokens can be priced higher than input tokens because generation is typically more compute-intensive than processing the input.
Independent benchmarking site Artificial Analysis tracks Nebius model performance and price metrics. Its Nebius provider page includes measures such as model intelligence, speed, latency, and pricing. Artificial Analysis also describes a “blended price” approach that combines cached input, uncached input, and output token costs using a selected ratio, allowing models and providers to be compared on a more standardized basis.
Blended pricing can be useful for comparisons, but it does not replace workload-specific cost analysis. Applications that generate long answers, code, or multi-step reasoning traces may be more sensitive to output-token prices. Applications that process large documents with short replies may be more sensitive to input-token prices.
Nebius is positioning Token Factory around production inference, where predictable pricing, model availability, and API compatibility can matter as much as raw benchmark performance. According to Nebius, the service supports more than 60 models across text, code, and vision categories, giving developers a choice of models through a single platform.
The company’s product page also highlights volume discounts and transparent dollar-per-token pricing. For teams moving from prototypes to customer-facing services, these details can affect budgeting and product design. A chatbot, coding assistant, research tool, or document-processing product may have very different ratios of input to output tokens, so separating those prices can help teams estimate cost more accurately.
The OpenAI-compatible API is another practical feature. Nebius says Token Factory supports that interface, which may reduce integration work for developers using existing libraries or application code built around OpenAI-style calls.
Nebius’s announcement comes as more AI infrastructure providers compete not only on training capacity but also on inference economics. As AI products become more widely deployed, customers increasingly need predictable throughput, lower latency, and clear per-token costs.
Artificial Analysis’s coverage of Nebius provides an external reference point for comparing price and performance, while Nebius’s own newsroom post, product page, and documentation describe the company’s intended positioning and pricing mechanics.
The key takeaway is that Nebius Token Factory is a new production inference offering with model variety, OpenAI-compatible APIs, and explicit input/output token pricing. For customers, the most relevant cost metric will depend on the application: some workloads are dominated by prompts, while others are dominated by generated responses.
Nebius has launched Nebius Token Factory, a production AI inference platform designed to serve open source and custom models at scale.
What Nebius announced In a newsroom post, Nebius said Token Factory is intended for production inference rather than experimentation alone, with support for more than 60 text, code, and vision models.
The company describes the platform as offering transparent cost per token pricing, volume discounts, and deployment options for open source and custom models.
Continue reading