Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Nebius launches Token Factory for production AI inference with per-token pricing · News · Kaino
Nebius launches Token Factory for production AI inference with per-token pricing
Kaino
YesterdayJul 27, 2026, 12:00 AM0 views

Nebius launches Token Factory for production AI inference with per-token pricing

Nebius has introduced Nebius Token Factory, a production inference platform for open-source and custom AI models. The company says the service supports more than 60 text, code, and vision models, offers OpenAI-compatible APIs, and separates input and output token pricing to make inference costs easier to compare.

Nebius

Nebius has launched Nebius Token Factory, a production AI inference platform designed to serve open-source and custom models at scale.

What Nebius announced

In a newsroom post, Nebius said Token Factory is intended for production inference rather than experimentation alone, with support for more than 60 text, code, and vision models. The company describes the platform as offering transparent cost-per-token pricing, volume discounts, and deployment options for open-source and custom models.

Nebius’s Token Factory product page says the service provides OpenAI-compatible APIs, which is intended to make it easier for developers to move applications that already use OpenAI-style interfaces. The same product page also emphasizes a pricing structure that separates input and output tokens, rather than presenting only a single headline price.

That pricing distinction matters because many AI applications have uneven token usage. A document summarization tool, for example, may send a large amount of text into a model and receive a shorter answer. A coding or reasoning assistant may generate comparatively long responses, making output-token pricing more important for total cost.

Input tokens, output tokens, and blended pricing

Nebius Token Factory documentation defines input tokens as the tokens included in an API request and priced per million tokens. It defines output tokens as the tokens generated in the model response, also priced per million tokens.

This separation is common in AI inference pricing because model providers often charge different rates for reading a prompt and generating a response. Output tokens can be priced higher than input tokens because generation is typically more compute-intensive than processing the input.

Independent benchmarking site Artificial Analysis tracks Nebius model performance and price metrics. Its Nebius provider page includes measures such as model intelligence, speed, latency, and pricing. Artificial Analysis also describes a “blended price” approach that combines cached input, uncached input, and output token costs using a selected ratio, allowing models and providers to be compared on a more standardized basis.

Blended pricing can be useful for comparisons, but it does not replace workload-specific cost analysis. Applications that generate long answers, code, or multi-step reasoning traces may be more sensitive to output-token prices. Applications that process large documents with short replies may be more sensitive to input-token prices.

Why this matters for developers and enterprise buyers

Nebius is positioning Token Factory around production inference, where predictable pricing, model availability, and API compatibility can matter as much as raw benchmark performance. According to Nebius, the service supports more than 60 models across text, code, and vision categories, giving developers a choice of models through a single platform.

The company’s product page also highlights volume discounts and transparent dollar-per-token pricing. For teams moving from prototypes to customer-facing services, these details can affect budgeting and product design. A chatbot, coding assistant, research tool, or document-processing product may have very different ratios of input to output tokens, so separating those prices can help teams estimate cost more accurately.

The OpenAI-compatible API is another practical feature. Nebius says Token Factory supports that interface, which may reduce integration work for developers using existing libraries or application code built around OpenAI-style calls.

The broader inference market context

Nebius’s announcement comes as more AI infrastructure providers compete not only on training capacity but also on inference economics. As AI products become more widely deployed, customers increasingly need predictable throughput, lower latency, and clear per-token costs.

Artificial Analysis’s coverage of Nebius provides an external reference point for comparing price and performance, while Nebius’s own newsroom post, product page, and documentation describe the company’s intended positioning and pricing mechanics.

The key takeaway is that Nebius Token Factory is a new production inference offering with model variety, OpenAI-compatible APIs, and explicit input/output token pricing. For customers, the most relevant cost metric will depend on the application: some workloads are dominated by prompts, while others are dominated by generated responses.

Key takeaways
  • 1

    Nebius has launched Nebius Token Factory, a production AI inference platform designed to serve open source and custom models at scale.

  • 2

    What Nebius announced In a newsroom post, Nebius said Token Factory is intended for production inference rather than experimentation alone, with support for more than 60 text, code, and vision models.

  • 3

    The company describes the platform as offering transparent cost per token pricing, volume discounts, and deployment options for open source and custom models.

Continue reading

Latest from Kaino News

Story pulse

Freshness

Yesterday

Views

0

Reading

3 min

Byline

Kainotomic Team

Utilities

Topics

Nebius

Sources

Reference material and original reporting used in this story.

Nebius

Published Jul 27, 2026, 12:00 AM

View source