DeepInfra (US) is an inference platform that serves open models over an API. It offers endpoints for text, embeddings, reranking, image, video and audio, plus deployment of private models and rental of GPU instances. The focus is per-token pricing and compatibility with the OpenAI format.
What it does
- Calls open models through an OpenAI-compatible API
- Runs embeddings, reranking, image, video and voice generation
- Deploys private models with your own weights or LoRA adapters
- Rents GPU instances for dedicated workloads
How it works
- You pick a model and get an HTTP endpoint
- Billing is per token processed, with no minimum contract
- Private models live on customer-exclusive endpoints
- LoRA adapters allow fine-tuning without duplicating the base model
Models and infrastructure
- Over 200 open text, vision, image, video and audio models
- Deployment of private models and LoRA adapters
- GPU instances and isolated runtime environments
- API compatible with OpenAI-ecosystem SDKs
Public pricing
- Per-token price published per model, with no platform fee
- Billed per image, audio second or video second
- GPU instances billed per hour of use
- Credit program for qualifying startups, with no permanent free tier
Strengths
- Broad open-model catalog with per-token pricing
- Accepts private models and adapters
- Integrates with tools that use the OpenAI format
Points of attention
- Frontier proprietary models are not part of the offering
- Quality and latency vary by model and region
- Very specific workloads may need a dedicated GPU

