Generative media AI inference platform: 4,000+ models via API plus serverless GPUs.

Homepage
1 / 2Overview
What is fal?
Generative media platform for developers
fal is a generative media AI inference platform and serverless GPU infrastructure for developers.
Model APIs provide fast, scalable, reliable access to 4,000+ generative media models — image, video, audio, 3D, and language — through one unified API. fal Serverless offers per-second-billed GPU infrastructure to deploy custom models and apps, scaling automatically from zero. The AI Gateway is an OpenAI-compatible unified gateway for calling third-party and proprietary models through a single integration. Developer tooling includes an online Playground, official Python and TypeScript client libraries, and MCP integration (ChatGPT, Claude, Cursor). Teams and enterprise get organization management, model access control, invoice billing, and dedicated support.
Pricing is pay-as-you-go with no fixed subscription: you pay only for the compute you consume. Serverless GPUs are billed per hour/second (1-minute minimum, then per-second): B300 $12.99/h, GB200 $9.99/h, B200 $7.99/h, H200 $6.00/h, H100 $4.50/h, RTX PRO 6000 $4.00/h, with volume discounts. Model APIs are billed per unit (video per second, images per image/megapixel, audio per 1,000 characters or per second, 3D per run), varying by model. Free credits are available on signup (variable validity); purchased credits expire after 365 days. Invoice-based billing and volume discounts available. See fal.ai/pricing for current details.
Platforms and languages
fal Availability
Platforms
Languages
Capabilities
fal Key Features
Model APIs
Fast, scalable, reliable API access to 4,000+ image, video, audio, 3D, and language models.
Serverless GPUs
Per-second-billed GPUs to deploy custom models and apps, scaling automatically from zero.
AI Gateway
OpenAI-compatible unified gateway for third-party and proprietary models.
Developer tooling
Online Playground, official Python/TypeScript clients, MCP integration (ChatGPT, Claude, Cursor).
Teams & enterprise
Organization management, access control, invoice billing, dedicated support.
Best for
Who uses fal?
Developers and AI teams
Build apps on 4,000+ generative media models via API
Enterprises
Deploy custom models on serverless GPU infrastructure
Plans and access
fal Pricing
Paid
Pay-as-you-go; free credits on signup. GPUs from ~$2.49/h.
Common questions
fal FAQs
What are the rate limits?
Each account has concurrency limits; new accounts default to 2 concurrent requests. Contact sales for higher limits.
Am I charged for failed requests?
Server errors (HTTP 500+) are never charged; infrastructure-caused failures have a free-retry policy.
Do credits expire?
Purchased credits expire 365 days after purchase. Free credits and coupons have variable validity.
What happens when my balance runs out?
When the balance falls below the lock threshold, the account is locked and API requests are rejected; top up from the billing panel to unlock.
Can I deploy my own models?
Yes — fal Serverless lets you deploy your own models and apps on fal's GPU infrastructure (any Python environment).
Reviews
See what the community thinks and share your experience.
—
Based on 0 ratings
Rating distribution
Community reviews
No reviews
No written reviews yet.