RunPod (US) is a GPU cloud for AI workloads. It combines per-second GPU rental, serverless endpoints, clusters and public model endpoints. You can run your own containers or call ready-made models through an API, paying for GPU time used.
What it does
- Rents GPUs by the second for training and inference
- Publishes serverless endpoints with custom containers
- Calls ready-made models through public endpoints
- Assembles clusters for distributed workloads
How it works
- You pick a GPU and start a container, or use a ready model
- Pods stay available per session; serverless endpoints scale with demand
- Billing is per second of use, with no data egress fee
- Network volumes keep data between runs
Models and infrastructure
- Dozens of public image, video, text and audio endpoints
- Custom containers with Docker, vLLM and other stacks
- Recent-generation GPUs across multiple regions
- Community network and secure cloud with isolation tiers
Public pricing
- Per-second GPU billing, with no subscription
- Public endpoints billed per token or per generated output
- Storage and volumes billed per GB/month
- Discounts on three- and six-month commitment plans
Strengths
- Wide variety of GPUs and regions
- Serverless model without managing servers
- No data egress fee
Points of attention
- Requires configuring containers for custom workloads
- Costs depend on GPU type and runtime
- Community capacity can vary by region

