Cerebras (US) runs an inference cloud on its own hardware, the wafer-scale engine. The platform serves text models through an OpenAI-compatible API, with a focus on low latency for interactive workloads and agents. There is a free tier, a per-token plan and an enterprise offering with custom weights.
What it does
- Calls text models through an OpenAI-compatible API
- Runs interactive workloads and agents that depend on low latency
- Tests models with free credits
- Contracts dedicated capacity for larger volumes
How it works
- You create an account and get an API key
- The call follows the market-standard chat completions format
- Inference runs on the company’s own hardware, not third-party GPUs
- Enterprise plans include custom weights and queue priority
Models and infrastructure
- Public tier with gpt-oss-120b and qwen-3.8-27b, both open
- Additional models and custom weights on dedicated endpoints
- Wafer-scale hardware, with a stated latency focus
- Free tier, per-token plan and enterprise offering
Public pricing
- Free trial with initial credit, subject to account rules
- gpt-oss-120b at US$ 0.35 per million input tokens and US$ 0.75 output
- qwen-3.8-27b at US$ 0.99 per million input and US$ 1.49 output
- Enterprise with dedicated capacity, custom weights and support
Strengths
- Purpose-built infrastructure with a low-latency focus
- API compatible with the OpenAI format
- Free tier for testing
Points of attention
- The catalog is smaller than that of general-purpose platforms
- Custom weights sit in the enterprise offering
- Availability depends on the company’s capacity

