GPU-Accelerated Google Cloud Platform

Accelerate Innovation

Building on over a decade of deep co-engineering across integrated platforms, open source frameworks, and managed services, NVIDIA and Google Cloud are pushing the boundaries of AI—enabling the next generation of AI that learns, reasons, and takes action to advance agentic AI, robotics, drug discovery, grid optimization, smart cities, and more.

Power New Capabilities With Google Cloud and NVIDIA

Explore Customer Stories

NVIDIA Accelerated Infrastructure on Google Cloud

Accelerate next-generation AI with the latest NVIDIA GPUs on Google Cloud, seamlessly integrated with Google Cloud AI Hypercomputer architecture—enabling demanding workloads at scale like LLM training, real-time inference, and advanced agentic AI applications for autonomous decision-making and physical AI for robotics, autonomous vehicles, and digital twins.

Google A4X and A4X Max VMs With NVIDIA GB200 and GB300 NVL72

Google Cloud’s A4X and A4X Max VMs deliver over one exaFLOP of compute per rack and support seamless scaling to tens of thousands of NVIDIA Blackwell and Blackwell Ultra GPUs, enabled by Google’s Jupiter network fabric and advanced networking with NVIDIA® ConnectX®-7 NICs. Google’s third-generation liquid cooling infrastructure delivers sustained, efficient performance, even for the largest AI workloads.

Google A4 VM With NVIDIA HGX B200

Google Cloud’s A4 VMs, accelerated by NVIDIA HGX™ B200, are now generally available. The A4 VM features eight NVIDIA Blackwell GPUs interconnected by fifth-generation NVIDIA NVLink™. Compared to previous generation A3 VMs, the A4 VM offers a significant performance boost, enabling faster model training, real-time inference, and accelerated data analytics. 

Google G4 VMs With NVIDIA RTX PRO 6000 Blackwell Server Edition

Google G4 VMs with NVIDIA RTX PRO™ 6000 Blackwell deliver breakthrough performance for both agentic and physical AI applications, accelerating everything from cost-efficient inference and generative AI to robotics simulation, hyper-realistic 3D rendering, and next-generation game rendering. Unlock next-generation AI and graphics capabilities.

Unlock the Full Potential of NVIDIA Accelerated Computing on Google Cloud

NVIDIA on Google Cloud Marketplace

NVIDIA offers a comprehensive, performance-optimized software stack directly on Google Cloud Marketplace to unlock the full potential of cutting-edge NVIDIA accelerated infrastructure and reduce the complexity of building accelerated solutions on Google Cloud. This lowers TCO through improved performance, simplified deployment, and streamlined development.

NVIDIA AI Enterprise

NVIDIA AI Enterprise is a cloud native platform that streamlines development and deployment of production-grade AI solutions including generative AI, computer vision, speech AI, and more. Easy-to-use microservices provide optimized model performance with enterprise-grade security, support, and stability to ensure a smooth transition from prototype to production for enterprises that run their businesses on AI.

NVIDIA NIM

NVIDIA NIM™, part of NVIDIA AI Enterprise, is a set of easy-to-use inference microservices for accelerating the deployment of AI applications that require natural language understanding and generation. By offering developers access to industry-standard APIs, NIM enables the creation of powerful copilots, chatbots, and AI assistants, while making it easy for IT and DevOps teams to self-host AI models in their own managed environments. NVIDIA NIM can be deployed on GCE, GKE, or Google Cloud Run.

NVIDIA Omniverse

NVIDIA Omniverse™ is a platform of APIs, SDKs, and services that enable developers to integrate OpenUSD, NVIDIA RTX™ rendering technologies into physical AI applications. Use VMs on Google Cloud to accelerate your application development.

Integrations at Every Layer of the Google Cloud Stack

NVIDIA and Google Cloud collaborate closely on integrations that bring the power of the full-stack NVIDIA AI platform to a broad range of native Google Cloud services, giving developers the flexibility to choose the level of abstraction they need. With these integrations, Google Cloud customers can combine the power of both enterprise-grade NVIDIA AI software and the computational power of NVIDIA GPUs to maximize application performance within the Google Cloud services they’re already familiar with.

Google Kubernetes Engine

Combine the power of the NVIDIA AI platform with the flexibility and scalability of GKE to efficiently manage and scale generative AI training and inference and other compute-intensive workloads. GKE's on-demand provisioning, automated scaling, NVIDIA Multi-Instance GPU (MIG) support, and GPU time-sharing capabilities ensure optimal resource utilization. This minimizes operational costs while delivering the necessary computational power for demanding AI workloads.

Vertex AI

Combine the power of NVIDIA accelerated computing with Google Cloud’s Vertex AI, a fully managed, unified MLOps platform for building, deploying, and scaling AI models in production. Leverage the latest NVIDIA GPUs and NVIDIA AI software, like Triton™ Inference Server, within Vertex AI Training, Prediction, Pipelines, and Notebooks to accelerate generative AI development and deployment without the complexities of infrastructure management.

Google Dataproc

Leverage the NVIDIA RAPIDS™ Accelerator for Spark to accelerate Apache Spark and Dask workloads on Dataproc, Google Cloud’s fully managed data processing service—without code changes. This enables faster data processing, extract, transform, and load (ETL) operations, and machine learning pipelines while substantially lowering infrastructure costs. With the RAPIDS Accelerator for Spark, users can also speed up batch workloads within Dataproc Serverless without provisioning clusters.

Google Dataflow

Accelerate machine learning inference with NVIDIA AI on Google Cloud Dataflow, a managed service for executing a wide variety of data processing patterns, including both streaming and batch analytics. Users can optimize the inference performance of AI models using NVIDIA TensorRT’s integration with Apache Beam SDK and speed up complex inference scenarios within a data processing pipeline using NVIDIA GPUs supported in Dataflow.

Cloud Run

DAccelerate the path to deploy generative AI faster with NVIDIA NIM on Google Cloud Run, a fully managed, serverless compute platform for deploying containers on Google Cloud’s infrastructure. With support for NVIDIA GPUs in Cloud Run, users can leverage NIM to optimize performance and accelerate deployment of gen AI models into production in a serverless environment that abstracts away infrastructure management.

Dynamic Workload Scheduler

Get easy access to NVIDIA GPU capacity on Google Cloud for short-duration workloads like AI training, fine-tuning, and experimentation using Dynamic Workload Scheduler. With flexible scheduling and atomic provisioning, users can get access to the compute resources they need within services like GKE, Vertex AI, and Batch while enhancing resource utilization and optimizing costs associated with running AI workloads. 

Google Distributed Cloud

With the NVIDIA Blackwell platform coming to Google Distributed Cloud, enterprises can now securely deploy advanced agentic AI—including Google Gemini models—directly in their own data centers on premises. This integration empowers organizations to harness breakthrough AI performance and scalability for sensitive, regulated workloads while ensuring data privacy, sovereignty, and compliance. By combining the strengths of Google Distributed Cloud and NVIDIA Blackwell, businesses can accelerate innovation with next-generation AI, while maintaining full control over their data and operations.

Google AI Hypercomputer With NVIDIA Dynamo Recipe

Google Cloud developed a recipe for disaggregated inferencing with NVIDIA Dynamo, a high-performance, low-latency platform for frontier AI models. This recipe makes it easy to deploy NVIDIA Dynamo on Google Cloud’s AI Hypercomputer, including Google Kubernetes Engine (GKE), vLLM inference engine, and A3 Ultra GPU-accelerated instances powered by NVIDIA H200 GPUs.

Google Vertex AI Model Garden With NVIDIA Nemotron

The NVIDIA Nemotron family of open models, including Nemotron 3 Nano, will be available on Google Vertex AI Model Garden as NVIDIA NIM microservices. This integration will provide developers and enterprises with access to NVIDIA's leading open-weight models. With a Vertex AI managed deployment, organizations can rapidly develop and deploy custom AI agents powered by Nemotron models while maintaining control over performance, cost, and compliance.

Google Cloud Storage for NVIDIA Run:ai Model Streamer

NVIDIA Run:ai Model Streamer comes with native Google Cloud Storage support, supercharging vLLM inference workloads on Google Kubernetes Engine (GKE). This collaboration accelerates loading LLMs from “cold” storage to memory by nearly 5x for faster AI inference—including an example using a 141 GB Meta Llama 3.3-70B model.

Google Cloud Cluster Director

Cluster Director is a Google Cloud service designed to optimize the deployment and management of large-scale AI and HPC clusters, accelerated by NVIDIA. This includes full support for large-scale AI systems, including Google Cloud’s A4X and A4X Max VMs powered by NVIDIA GB200 and GB300 NVL72 systems. Google Cloud announced the Preview of Cluster Director support for Slurm on GKE, utilizing SchedMD’s Slinky offering—a company recently acquired by NVIDIA.

Google Cloud and NVIDIA Developer Community

Google Cloud and NVIDIA have partnered to create this community for developers, data scientists, AI/ML engineers, and technical practitioners focused on leveraging NVIDIA and Google Cloud technologies for their development.

Additional Resources

Gemma

NVIDIA is collaborating with Google to launch Gemma, a newly optimized family of open models built from the same research and technology used to create the Gemini models. An optimized release with TensorRT-LLM enables users to develop with LLMs using only a desktop with an NVIDIA RTX™ GPU.

RAPIDS cuDF on Google Colab

RAPIDS cuDF is now integrated into Google Colab. Developers can instantly accelerate pandas code up to 50X on Google Colab GPU instances and continue using pandas as data grows—without sacrificing performance.

Accelerate Your Startup

The NVIDIA Inception program helps startups accelerate innovation with developer resources and training, access to cloud credits, exclusive pricing on NVIDIA software and hardware, and opportunities for exposure to the VC community.

Latest News

NVIDIA and Google Cloud Collaborate to Advance Agentic and Physical AI

Advancements in building AI factories will power the next frontier of agentic and physical AI. These include new NVIDIA Vera Rubin-powered A5X VMs, general availability of Google Gemini on Google Distributed Cloud running on NVIDIA Blackwell and NVIDIA Blackwell Ultra, confidential VMs with NVIDIA Blackwell GPUs, and agentic AI on Gemini Enterprise Agent Platform with NVIDIA Nemotron and NVIDIA NeMo™.

Co-Engineered AI Infrastructure for the Agentic AI Era

Google Cloud and NVIDIA provide the foundation and full-stack experience that technology leaders need to scale their agentic AI workloads.

Together, they are innovating every layer of the AI stack—driven by continued momentum for G4 VMs, support for the NVIDIA Vera Rubin NVL72 platform, NVIDIA Dynamo integration with Inference Gateway, enhanced NVIDIA support across Vertex AI Training and Model Garden, and a new AI startup accelerator program for the public sector.

NVIDIA Kicks Off the Next Generation of AI With Rubin—Six New Chips, One Incredible AI Supercomputer

Building on its decade-long partnership with NVIDIA, Google plans to bring the capabilities of the NVIDIA Rubin platform to our customers, offering them the scale and performance needed to advance the boundaries of AI. Google Cloud will be among the first cloud providers to deploy NVIDIA Vera Rubin-based instances in 2026.

Google Cloud Scales MoE Inference on A4X With NVIDIA Dynamo

Google Cloud is scaling frontier mixture-of-experts models faster with A4X machines, powered by NVIDIA GB200 NVL72 systems, and NVIDIA Dynamo. This validated reference architecture achieved over 6,000 total tokens/sec/GPU with 10ms inter-token latency by delivering distributed runtime KV cache management and kernel scheduling with Wide Expert Parallelism (WideEP) across the full 72-GPU compute domain. Explore the technical details and get started with deployment recipes.

Accelerate Model Downloads on GKE With NVIDIA Run:Ai Model Streamer

Google Cloud and NVIDIA are supercharging AI workloads with native Google Cloud Storage support for the NVIDIA Run:ai Model Streamer. This integration reduces model loading "cold start" times on GKE by streaming tensors directly to GPU memory, enabling faster auto-scaling and improved GPU efficiency for high-performance AI inference.

NVIDIA and Google Cloud Accelerate Enterprise AI and Industrial Digitalization

NVIDIA and Google Cloud are accelerating industrial digitalization with the general availability of G4 VMs, powered by NVIDIA Blackwell GPUs. By bringing NVIDIA Omniverse and NVIDIA Isaac Sim™ to Google Cloud Marketplace, this partnership empowers enterprises to scale physical AI and digital twins for manufacturing, automotive, and logistics workloads.

Access the Power of Google Cloud and NVIDIA

Contact Sales