This page lists the metrics emitted by a reinforcement learning fine-tuning job and describes how to view job status and results.
Metrics
The following metrics are emitted at every training step (on the current training batch) and at every evaluation interval (on the entire validation dataset).
Reward metrics
/train_mean_reward— The mean reward on the batch of samples in the current training step./eval_mean_reward— The mean reward on the entire validation dataset.
Generation length metrics
/train_generation_length— The mean generation length, in number of tokens (including thinking tokens and non-thinking tokens), on the batch of samples in the current training step./eval_generation_length— The mean generation length, in number of tokens (including thinking tokens and non-thinking tokens), on the entire validation dataset./train_thinking_token_length— The mean thinking token length on the batch of samples in the current training step./eval_thinking_token_length— The mean thinking token length on the entire validation dataset.
Reward latency metrics
/train_mean_reward_latency— The mean latency for computing rewards on the batch of samples in the current training step./eval_mean_reward_latency— The mean latency for computing rewards on the entire validation dataset./train_p95_reward_latency— The P95 latency for computing rewards on the batch of samples in the current training step./eval_p95_reward_latency— The P95 latency for computing rewards on the entire validation dataset.
Sampling latency metrics
/train_mean_sampling_latency— The mean latency for sampling on the batch of samples in the current training step./eval_mean_sampling_latency— The mean latency for sampling on the entire validation dataset./train_p95_sampling_latency— The P95 latency for sampling on the batch of samples in the current training step./eval_p95_sampling_latency— The P95 latency for sampling on the entire validation dataset.
Batch composition metrics
/learnable_prompt_ratio— The actual batch size divided by all samples (including samples filtered out) for the current training step. Training metric only.
Per-reward metrics (for composite rewards)
For composite rewards, each individual reward emits its own group of metrics
prefixed with ${reward_name}. Metrics with the same reward name are grouped
together in the monitoring UI, where you can fold and unfold them.
${reward_name}/train_mean_reward— The mean reward for${reward_name}on the batch of samples in the current training step.${reward_name}/eval_mean_reward— The mean reward for${reward_name}on the entire validation dataset.${reward_name}/train_processing_time— The mean latency for computing the rewards for${reward_name}on the batch of samples in the current training step.${reward_name}/eval_processing_time— The mean latency for computing the rewards for${reward_name}on the entire validation dataset.${reward_name}/train_p95_processing_time— The P95 latency for computing rewards for${reward_name}on the batch of samples in the current training step.${reward_name}/eval_p95_processing_time— The P95 latency for computing rewards for${reward_name}on the entire validation dataset.${reward_name}/train_executing_code_failure_ratio— The mean failure ratio for computing Code Execution rewards named${reward_name}on the batch of samples in the current training step. Code execution rewards only.${reward_name}/eval_executing_code_failure_ratio— The mean failure ratio for computing Code Execution rewards named${reward_name}on the entire validation dataset. Code execution rewards only.${reward_name}/train_rpc_error_ratio— The mean RPC failure ratio for computing rewards named${reward_name}on the batch of samples in the current training step. Applies to Code execution, Autorater, and Cloud Run rewards; does not apply to string matching.${reward_name}/eval_rpc_error_ratio— The mean RPC failure ratio for computing rewards named${reward_name}on the entire validation dataset. Applies to Code execution, Autorater, and Cloud Run rewards; does not apply to string matching.${reward_name}/train_invalid_rpc_response_ratio— The mean ratio of invalid RPC responses for rewards named${reward_name}on the batch of samples in the current training step. Applies to Code execution, Autorater, and Cloud Run rewards; does not apply to string matching.${reward_name}/eval_invalid_rpc_response_ratio— The mean ratio of invalid RPC responses for rewards named${reward_name}on the entire validation dataset. Applies to Code execution, Autorater, and Cloud Run rewards; does not apply to string matching.${reward_name}/train_clipping_rewards_ratio— The mean clipping ratio for${reward_name}on the batch of samples in the current training step. Rewards are clipped to the range[-1, 1].${reward_name}/eval_clipping_rewards_ratio— The mean clipping ratio for${reward_name}on the entire validation dataset. Rewards are clipped to the range[-1, 1].
View job status and results
You can view a reinforcement learning fine-tuning job's status and results in the Google Cloud console or by using the REST API.
Console
Go to Models > Tuning in the Google Cloud console and select your tuning job. The job details page provides three tabs for monitoring and inspection:
- Monitor tab: Displays real-time tuning progress and
interactive charts for the training and evaluation metrics listed on
this page:
- Use the sticky Filter bar (
Filter metrics by name) or the Show Category selector to filter charts by metric name or category. - Expand or collapse chart groups for general tuning metrics and individual reward functions in composite rewards.
- Inspect checkpoint annotations directly on the metric charts.
- Use the Checkpoints table and its column selector to compare tuning and reward metrics across saved checkpoints, copy a checkpoint endpoint ID, or click Test to evaluate a checkpoint in Agent Studio.
- Use the sticky Filter bar (
- Dataset tab: Inspect training and validation dataset samples,
conversation examples,
referencesfields, and dataset distribution charts. - Details tab: View the base model, tuning method, hyperparameter values, and reward configuration. Click View details on a reward configuration to open a side drawer displaying non-default settings.
REST
Issue a
tuningJobs.get
request to retrieve the job's state, tunedModel, error
details, and metadata.
curl -X GET \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ "https://LOCATION_ID-aiplatform.googleapis.com/v1beta1/projects/PROJECT_ID/locations/LOCATION_ID/tuningJobs/TUNING_JOB_ID"
Replace the following:
PROJECT_ID: your Google Cloud project ID.TUNING_JOB_ID: the ID of the tuning job.LOCATION_ID: the ID of the location.
Intermediate checkpoints
Reinforcement learning fine-tuning produces intermediate checkpoints at the
frequency set by the checkpointInterval hyperparameter. You can deploy and
evaluate these checkpoints independently before the job completes. For
details on the checkpointInterval hyperparameter, see the
Hyperparameters
page.
What's next
- Follow the Google Cloud console quick start or API quick start to create your first reinforcement learning fine-tuning job.
- Define reward functions.
- Configure hyperparameters.