Run multi-host RL training for Qwen3-30b-a3b on TPU v6e

This tutorial shows you how to run multi-host reinforcement learning (RL) training on a Tensor Processing Unit (TPU) v6e-64 cluster by using MaxText and Cluster Toolkit. You use Cluster Toolkit to execute a multi-host training workload and export the results back to Hugging Face format for serving.

Objectives

  • Install Cluster Toolkit and its dependencies.
  • Deploy a Cluster Toolkit cluster.
  • Convert a Hugging Face model to MaxText format.
  • Run an RL training workload on the TPU v6e cluster.
  • Convert the fine-tuned model back to Hugging Face format for serving.

Costs

In this document, you use the following billable components of Google Cloud:

To generate a cost estimate based on your projected usage, use the pricing calculator.

New Google Cloud users might be eligible for a free trial.

When you finish the tasks that are described in this document, you can avoid continued billing by deleting the resources that you created. For more information, see Clean up.

Before you begin

  • You need a Hugging Face access token to use this tutorial. You can sign up for a free account at Hugging Face. After you have an account, generate an access token:

    1. On the Welcome to Hugging Face page, click your account avatar and select Access tokens.
    2. On the Access tokens page, click Create new token.
    3. Select the Read token type and enter a name for your token.
    4. Your access token is displayed. Save the token in a safe place.

  • On the Hugging Face website, accept the license agreement for the model that you plan to train. This tutorial uses the model qwen3-30b-a3b.

To get the permissions that you need to complete this tutorial, ask your administrator to grant you the following IAM roles on your project:

For more information about granting roles, see Manage access to projects, folders, and organizations.

You might also be able to get the required permissions through custom roles or other predefined roles.

Set up your environment variables

Set up your environment variables:

export PROJECT="YOUR_PROJECT_ID"
export REGION="YOUR_REGION"
export ZONE="YOUR_ZONE"
export CLUSTER_NAME="YOUR_CLUSTER_NAME"
export GCS_BUCKET="YOUR_BUCKET_NAME"
export CLOUD_IMAGE_NAME="us-docker.pkg.dev/cloud-tpu-images/maxtext-images/tpu_post_training:0.2.4"
export COMPUTE_TYPE="ct6e-standard-4t"
export TPU_TYPE="v6e-64"
export TOPOLOGY="8x8"
export CLUSTER_NODEPOOL_COUNT=1
export PW_CPU_MACHINE_TYPE="c4d-standard-96"
export RESERVATION="YOUR_RESERVATION_NAME"
export MODEL_NAME="qwen3-30b-a3b"
export CLUSTER_TOOLKIT_VERSION="v1.103.0"
export HF_TOKEN="YOUR_HF_TOKEN"

Replace the following:

  • YOUR_PROJECT_ID: the ID of your Google Cloud project.
  • YOUR_REGION: the region where you want to deploy your cluster.
  • YOUR_ZONE: the zone where you want to deploy your cluster.
  • YOUR_CLUSTER_NAME: the name of your Google Kubernetes Engine cluster (up to 20 characters).
  • YOUR_BUCKET_NAME: a globally unique name for a Cloud Storage bucket.
  • YOUR_RESERVATION_NAME: the name of your reservation.
  • YOUR_HF_TOKEN: your Hugging Face access token.

Install Cluster Toolkit dependencies

To complete this tutorial from a Linux or macOS client or workstation, follow the relevant steps in Install dependencies in the Cluster Toolkit documentation.

If you're using Cloud Shell, then you can skip this section.

Install Cluster Toolkit

Install the prebuilt bundle for Cluster Toolkit in your current working directory by following the instructions at Install Cluster Toolkit.

For example, you can download and extract the bundle into your current working directory as follows:

wget -qO- "https://github.com/GoogleCloudPlatform/cluster-toolkit/releases/download/${CLUSTER_TOOLKIT_VERSION:-v1.103.0}/gcluster_bundle_linux_amd64.tgz" | tar -xz

Extracting the bundle into your current working directory provides the gcluster binary and the examples/ blueprints that you use in the steps that follow.

Create your Cluster Toolkit cluster

To create and deploy a Cluster Toolkit cluster with 64 v6e TPU chips, complete the following steps:

  1. Create a Cloud Storage bucket:

    gcloud storage buckets create "gs://${GCS_BUCKET}" --project="${PROJECT}" --location="${REGION}" || true
  2. Copy the Cluster Toolkit blueprint to your current working directory:

    cp examples/gke-tpu-v6e/gke-tpu-v6e-advanced.yaml .
  3. By default, your cluster node pool service account doesn't have the required permissions to write to your Cloud Storage bucket. To allow the node pool service account to write to your Cloud Storage bucket, you must grant it the Storage Admin role. To grant this role, edit the file gke-tpu-v6e-advanced.yaml by updating the service-account module named node_pool_service_account:

    - id: node_pool_service_account
      source: modules/project/service-account
      settings:
        name: gke-np-sa
        project_roles:
        - logging.logWriter
        - monitoring.metricWriter
        - monitoring.viewer
        - stackdriver.resourceMetadata.writer
        - storage.admin
        - artifactregistry.reader
  4. Use the gcluster deploy command to deploy your Cluster Toolkit cluster by using the blueprint gke-tpu-v6e-advanced.yaml and passing the required variables by using the --vars flag:

    ./gcluster deploy gke-tpu-v6e-advanced.yaml \
        --vars project_id="${PROJECT}" \
        --vars deployment_name="${CLUSTER_NAME}" \
        --vars region="${REGION}" \
        --vars zone="${ZONE}" \
        --vars num_slices="${CLUSTER_NODEPOOL_COUNT}" \
        --vars tpu_topology="${TOPOLOGY}" \
        --vars authorized_cidr="0.0.0.0/0" \
        --vars reservation="${RESERVATION:-}" \
        -l IGNORE --auto-approve -w

Convert the model to MaxText format

To train the model in MaxText format, you must convert it from Hugging Face format to MaxText format.

  1. To simplify subsequent commands, use the gcluster job config command to configure your default project, cluster, and location:

    # Configure gcluster Defaults
    ./gcluster job config set project "${PROJECT}"
    ./gcluster job config set cluster "${CLUSTER_NAME}"
    ./gcluster job config set location "${REGION}"
  2. Use the gcluster job submit command to convert the model from Hugging Face format to MaxText format and store it in your Cloud Storage bucket:

    ./gcluster job submit \
      --name qwen-hf-to-mt \
      --num-slices 1 \
      --image "${CLOUD_IMAGE_NAME}" \
      --compute-type "${COMPUTE_TYPE}" \
      --topology "${TOPOLOGY}" \
      --await-job-completion \
      --command "[ \"\$JOB_COMPLETION_INDEX\" != \"0\" ] || \
      python3 -m maxtext.checkpoint_conversion.to_maxtext \
      model_name=${MODEL_NAME} \
      hf_access_token=${HF_TOKEN} \
      --hf_model_path='Qwen/Qwen3-30B-A3B-Instruct-2507' \
      base_output_directory=gs://${GCS_BUCKET}/${MODEL_NAME}/max-text-format/ \
      scan_layers=True \
      use_multimodal=False \
      skip_jax_distributed_system=true \
      checkpoint_storage_use_zarr3=0 \
      checkpoint_storage_use_ocdbt=0 \
      hardware=cpu \
      --lazy_load_tensors=True"
  3. Use the gcluster job logs command to check the status of the conversion job:

    # Use the list command to check status
    ./gcluster job list
    
    # Check progress of the job (--main-only targets the coordinator pod (Job Index 0, Pod Index 0) to avoid duplicate logs from other workers)
    ./gcluster job logs qwen-hf-to-mt --main-only -f
  4. Verify that the converted model files are available in your Cloud Storage bucket:

    gcloud storage ls "gs://${GCS_BUCKET}/${MODEL_NAME}/max-text-format/"

Start the training workload

After the conversion process has completed, start the RL training workload:

./gcluster job submit \
  --name="qwen-training" \
  --num-slices=1 \
  --image="${CLOUD_IMAGE_NAME}" \
  --compute-type="${COMPUTE_TYPE}" \
  --topology="${TOPOLOGY}" \
  --pathways \
  --pathways-gcs-location="gs://${GCS_BUCKET}/pathways/" \
  --env="GRPC_DNS_RESOLVER=native" \
  --pathways-proxy-env="GRPC_DNS_RESOLVER=native" \
  --pathways-server-env="GRPC_DNS_RESOLVER=native" \
  --pathways-worker-env="GRPC_DNS_RESOLVER=native" \
  --command="export VLLM_HOST_IP=\$(hostname -I | awk '{print \$1}'); \
      JAX_PLATFORMS=proxy,cpu ENABLE_PATHWAYS_PERSISTENCE=1 \
      python3 -m maxtext.trainers.post_train.rl.train_rl \
      run_name=rl \
      base_output_directory=gs://${GCS_BUCKET}/${MODEL_NAME}/trained/ \
      model_name=${MODEL_NAME} \
      load_parameters_path=gs://${GCS_BUCKET}/${MODEL_NAME}/max-text-format/0/items/ \
      data_template_path=maxtext/examples/chat_templates/openmathinstruct2_rl.json \
      reshard_chunk_size=8 \
      hf_access_token=${HF_TOKEN} \
      num_batches=50 \
      batch_size=4 \
      rollout_tensor_parallelism=8 \
      rollout_expert_parallelism=1 \
      trainer_devices_fraction=0.5 \
      sampler_devices_fraction=0.5 \
      tokenizer_path='Qwen/Qwen3-30B-A3B-Instruct-2507' \
      ici_tensor_parallelism=4 \
      ici_expert_parallelism=4 \
      hbm_utilization_vllm=0.65 \
      remat_policy=full \
      async_scheduling=False \
      allow_split_physical_axes=true \
      ragged_gather_reduce_fallback=True \
      vllm_hf_overrides='{architectures: [\"MaxTextForCausalLM\"]}' \
      vllm_additional_config=\"{'maxtext_config': {'model_name': '${MODEL_NAME}', 'allow_split_physical_axes': 'true', 'use_ragged_sort': 'false', 'ragged_gather_reduce_fallback': 'true', 'prefuse_moe_weights': 'true', 'weight_dtype': 'bfloat16'}}\""
  • To check the status of the training job:

    # Use the list command to check status
    ./gcluster job list
    
    # Ensure kubectl credentials are configured
    gcloud container clusters get-credentials "${CLUSTER_NAME}" \
        --location="${REGION}" \
        --project="${PROJECT}"
    
    # Check progress of the job
    kubectl logs -f \
        -l jobset.sigs.k8s.io/replicatedjob-name=pathways-head \
        -c workload-container
  • To verify that the training checkpoints are generated in your Cloud Storage bucket:

    gcloud storage ls "gs://${GCS_BUCKET}/${MODEL_NAME}/trained/rl/checkpoints/actor/"

Convert the trained model back into Hugging Face format

After the training workload has completed, convert the model back to Hugging Face format:

./gcluster job submit \
  --name="qwen-mt-to-hf" \
  --num-slices=1 \
  --image="${CLOUD_IMAGE_NAME}" \
  --compute-type="${COMPUTE_TYPE}" \
  --topology="${TOPOLOGY}" \
  --await-job-completion \
  --command="[ \"\$JOB_COMPLETION_INDEX\" != \"0\" ] || \
  python3 -m maxtext.checkpoint_conversion.to_huggingface \
  model_name=${MODEL_NAME} \
  hf_access_token=${HF_TOKEN} \
  load_parameters_path=gs://${GCS_BUCKET}/${MODEL_NAME}/trained/rl/checkpoints/actor/50/model_params/ \
  base_output_directory=gs://${GCS_BUCKET}/${MODEL_NAME}/hf-trained/ \
  skip_jax_distributed_system=true \
  hardware=cpu \
  scan_layers=True \
  use_multimodal=False \
  weight_dtype=bfloat16 \
  --override_model_architecture"
  • To check the status of the conversion job:

    # Use the list command to check status
    ./gcluster job list
    
    # Check progress of the job (--main-only targets the coordinator pod (Job Index 0, Pod Index 0) to avoid duplicate logs from other workers)
    ./gcluster job logs qwen-mt-to-hf --main-only -f
    # The trained model is now available in gs://${GCS_BUCKET}/${MODEL_NAME}/hf-trained/
  • To verify that the trained Hugging Face model weights and configuration files are present in your Cloud Storage bucket:

    gcloud storage ls -l --readable-sizes "gs://${GCS_BUCKET}/${MODEL_NAME}/hf-trained/"

Clean up

To avoid incurring additional charges, use the gcluster destroy command to delete the resources created during this tutorial:

./gcluster destroy "${CLUSTER_NAME}"
gcloud storage rm -r "gs://${GCS_BUCKET}"

# To delete the local deployment folder and copied blueprint
rm -rf .ghpc "${CLUSTER_NAME}" gke-tpu-v6e-advanced.yaml

What's next