Overview¶
With our Containers service, you can create your own inference endpoints to serve your models while paying only for the compute that is in active use.
We support loading containers from any registry and are quite flexible about how the container is built.
You can deploy your first container by following the guide: Quick: Deploy with vLLM
GPU drivers and CUDA compatibility¶
Serverless Containers uses NVIDIA driver version 580.x.x (R580) across all customer-serving GPU pools. Exact patch versions can vary between nodes. If your image requires a specific patch version, contact support to confirm it for your selected GPU and location.
Verda manages the host NVIDIA driver. Your container image supplies the CUDA runtime and application libraries, so its CUDA version does not need to match the node image's CUDA toolkit version exactly. Choose an image whose driver requirements and supported GPU architectures match your deployment.
NVIDIA's CUDA compatibility guide lists R580 as the minimum driver branch for CUDA 13.x minor-version compatibility. Newer drivers also support applications built with older CUDA toolkits, but compatibility within a CUDA major version has feature and PTX limitations. Check the image release's requirements as well as the CUDA compatibility table.
Serverless Containers pricing¶
You're billed for every minute in which a replica processed any usage, including the time spent spinning up or down. The number of currently running replicas will depend on your scaling settings. Charges are aggregated and displayed on your bill in 10-minute intervals.
Features¶
- Scale to hundreds of GPUs when needed with our battle-tested inference cluster
- Scale to zero when idle, so you only pay while your container is running
- Support for any container registry, using either registry-specific authentication methods or a vanilla Docker
config.json-style auth - Both manual and request queue-based autoscaling, with adjustable scaling sensitivity
- Logging and metrics in the console
- RESTful API for managing your deployments
- Python SDK
- Support for async / polling requests
- Shared storage between the Containers and Cloud GPU instances
- Batch jobs - recommended for long inference durations > 3min
Coming soon¶
- Container Registry