This directory contains bash scripts to automate the end-to-end deployment of a secure vLLM inference server running google/gemma-3-1b-it inside Google Cloud Confidential Space.
Before running these scripts, ensure you have:
gcloudCLI installed and authenticated.- A Google Cloud project with billing and required quota (A3/H100).
- A Hugging Face account with a User Access Token. You must accept the Gemma 3 license agreement on Hugging Face to access the model weights.
The setup.sh script will enable required APIs, build and push the Docker image, create a Service Account with necessary IAM roles, spawn the Confidential VM, and configure an external TCP Load Balancer.
Since we are using the gated Gemma 3 model, you must export your Hugging Face token first:
export HF_TOKEN="your_hugging_face_token_here"
bash setup.sh --project-id <YOUR_PROJECT_ID>After the setup script finishes, the Confidential VM takes several minutes (usually 4-6 minutes) to boot the hardened OS, install NVIDIA drivers, download the model weights from Hugging Face, and allocate the KV cache on the GPU.
Do not proceed to the next step until the load balancer reports the backend as HEALTHY. You can check the health status by running the following command:
gcloud compute backend-services get-health secure-inference-backend \
--region=us-central1 \
--project=<YOUR_PROJECT_ID>Keep running this command periodically until you see healthState: HEALTHY in the output instead of UNHEALTHY.
Once the server is healthy, run the client script. It will automatically read the Load Balancer IP and Image Hash generated by the setup script, perform the Attested TLS handshake, and securely send a prompt to the Gemma 3 model.
bash run_client.sh <YOUR_PROJECT_ID>To avoid ongoing compute charges for the A3 instance and Load Balancer, ensure you tear down all created resources once you are done experimenting:
bash cleanup.sh --project-id <YOUR_PROJECT_ID>