Serverless cloud for AI and data teams — run functions, batch jobs, and GPU inference from Python or JavaScript with no infra to manage.
You write a function in Python or JavaScript, and Modal runs it in the cloud. No Dockerfiles, no Kubernetes, no cluster to keep alive. You say what container and hardware you want — GPUs included — in the code itself, and Modal builds it, deploys it, and scales it. It starts containers when there's work and stops them when there isn't, so you only pay for the seconds you use.
It handles the jobs a normal serverless setup can't: LLM and image inference, fine-tuning, batch jobs, crons, web endpoints. You get persistent volumes and secrets too. So it feels like serverless, but on real GPUs.
They also serve Kimi K3 with the best tok/s speed currently available, even better than Fireworks Fast, but at the regular price. Run through Modal's shared API at modal.com/endpoints.