#1: Enable Scale-to-Zero
Enable scale-to-zero in Model Garden to automatically power down your GPU instances when there are no active incoming requests.
Model Garden docs ↓
#2: Leverage Spot VMs
Many developer and agent testing workflows are resilient and don’t require 100% continuous uptime. Choose Spot as your VM provisioning model to tap into spare Google Cloud compute capacity at deep discounts.
#3: Right-Size Your Hardware + Limit Auto-Scaling
Keep your monthly bill predictable by updating your settings to apply strict hardware limits, such as:
- capping your accelerator count at a single GPU
- locking your replica count to 1-1
- opting for no capacity reservations