Crusoe (United States) runs an AI cloud built on its own data centers and, on top of it, Crusoe Intelligence Foundry, its managed serving layer. The product combines token-billed serverless inference, self-service deployments on dedicated GPUs, tailored environments and serverless fine-tuning with exportable weights. The inference engine is MemoryAlloy, built in-house.
What it does
- Calls hosted models through an API without managing GPUs
- Publishes your own or fine-tuned model in a few steps, with a dedicated endpoint
- Fine-tunes models with LoRA and takes the weights wherever the team prefers
How it works
- Serverless inference answers through an OpenAI-compatible API and bills per token
- Self-service deployments deliver an endpoint on a dedicated GPU billed per hour
- Serverless fine-tuning produces adapters that can be exported in an open format
- New accounts get credits to test the whole flow before scaling
Strengths
- Vertically integrated infrastructure: power, data centers and serving from one vendor
- A short path from fine-tuning to a production endpoint
- A managed layer on top of a well-established GPU cloud
Points of attention
- The per-hour GPU cloud remains the main product; this card covers managed serving
- Tailored environments and larger volumes go through sales
- The serverless catalog is smaller than on platforms dedicated only to models

