Tool profile

Nebius

ByNebiusNL

Token Factory inference service: OpenAI-compatible API, serverless and dedicated endpoints, fine-tuning and custom weights on Nebius AI Cloud.

Nebius wordmark in white on a dark gradient background, centered and minimal composition
Country
NL
Availability
Proprietary

Nebius (Netherlands) operates Token Factory, an inference service for open models inside Nebius AI Cloud. Access is through an OpenAI-compatible API, with public serverless endpoints and dedicated endpoints for teams that need a fixed region and reserved capacity.

What it does

  • Runs open-model inference through a managed API
  • Creates dedicated endpoints with reserved GPU for predictable workloads
  • Fine-tunes models and serves custom weights
  • Tracks usage, latency and consumption per project

How it works

  • An account is created in the Nebius console and issues an API key
  • The client points at the Token Factory base URL and selects a model
  • Public endpoints run on shared infrastructure and show a global region
  • Dedicated endpoints reserve GPU and keep the chosen region
  • Usage is metered by tokens on public endpoints and by GPU time on dedicated ones

Models and infrastructure

  • A catalogue of open text and image models
  • Recent NVIDIA GPUs on Nebius-owned infrastructure
  • Supervised fine-tuning, LoRA adapter merging and execution sandboxes
  • Built-in observability through an API

Strengths

  • Combines serverless and dedicated endpoints in one service
  • European company, with region and data residency under customer control on dedicated endpoints
  • OpenAI-compatible API simplifies integration

Points of attention

  • Public endpoints do not guarantee a fixed processing region
  • Nebius also sells cloud GPUs; this entry covers the inference service, not the GPU cloud
  • Post-training features and dedicated endpoints require their own configuration and billing