Nebius (Netherlands) operates Token Factory, an inference service for open models inside Nebius AI Cloud. Access is through an OpenAI-compatible API, with public serverless endpoints and dedicated endpoints for teams that need a fixed region and reserved capacity.
What it does
- Runs open-model inference through a managed API
- Creates dedicated endpoints with reserved GPU for predictable workloads
- Fine-tunes models and serves custom weights
- Tracks usage, latency and consumption per project
How it works
- An account is created in the Nebius console and issues an API key
- The client points at the Token Factory base URL and selects a model
- Public endpoints run on shared infrastructure and show a global region
- Dedicated endpoints reserve GPU and keep the chosen region
- Usage is metered by tokens on public endpoints and by GPU time on dedicated ones
Models and infrastructure
- A catalogue of open text and image models
- Recent NVIDIA GPUs on Nebius-owned infrastructure
- Supervised fine-tuning, LoRA adapter merging and execution sandboxes
- Built-in observability through an API
Strengths
- Combines serverless and dedicated endpoints in one service
- European company, with region and data residency under customer control on dedicated endpoints
- OpenAI-compatible API simplifies integration
Points of attention
- Public endpoints do not guarantee a fixed processing region
- Nebius also sells cloud GPUs; this entry covers the inference service, not the GPU cloud
- Post-training features and dedicated endpoints require their own configuration and billing
