fal (US) is an inference platform focused on media models. It serves image, video, audio and 3D models over an API, and lets you deploy custom models on serverless GPU infrastructure. The catalog ranges from open models to proprietary media-generation models.
What it does
- Generates and edits image, video, audio and 3D through an API
- Tests and compares media models in one panel
- Deploys custom models with managed containers and runtime
- Runs media workloads on on-demand GPUs
How it works
- You pick a model and call an HTTP endpoint
- Each successful call is billed by model and output
- Custom models run in containers published on the platform
- Infrastructure scales capacity with the request queue
Models and infrastructure
- Over a thousand image, video, audio and 3D models
- Open and proprietary models in the same catalog
- Serverless for custom models plus dedicated GPUs by the hour
- Billed per generated output and per runtime
Public pricing
- Prepaid credits, with no mandatory subscription
- Price per output, for example per image or per video second
- Dedicated GPUs billed hourly
- Credit programs for startups and research
Strengths
- Focus on generative media with many models in one place
- Custom model deployment without managing servers
- On-demand scaling for generation spikes
Points of attention
- The media focus differs from text-oriented platforms
- Additional serverless features may require approval
- Cost grows with volume and output size

