Skip to content

EPP reads one engine's metrics on vLLM data-parallel pods #3084

Description

@mtparet

What happened:

A vLLM pod running data parallelism with internal load balancing (--data-parallel-size N) exposes one series per engine in each metric family:

vllm:num_requests_waiting{engine="0",model_name="m"} 6.0
vllm:num_requests_waiting{engine="1",model_name="m"} 10.0
vllm:num_requests_waiting{engine="2",model_name="m"} 11.0
vllm:num_requests_waiting{engine="3",model_name="m"} 7.0

The core metrics extractor reads one series per family. Spec.getLatestMetric (pkg/epp/framework/plugins/datalayer/extractor/metrics/spec.go:104, main @ 32c11b7) keeps the matching series with the greatest timestamp:

for _, metric := range family.GetMetric() {
	if spec.labelsMatch(metric.GetLabel()) {
		ts := metric.GetTimestampMs()
		if ts > recent {
			recent = ts
			latest = metric
		}
	}
}

vLLM exposes no timestamps, so this is always the first series. WaitingQueueSize, RunningRequestsSize and KVCacheUsagePercent (extractor.go:131-154) describe engine 0 only. The pod above reads a queue of 6.

Every consumer of these fields sees one engine's load: the queue-depth, running-requests, KV-cache-utilization and load-aware scorers, the utilization filter, and the flow-control utilization saturation detector. A pod whose engine 0 is idle scores as idle while its other engines are saturated.

What you expected to happen:

Pod-level values cover every engine:

  • WaitingQueueSize: max over engines. Engines running DP with expert parallelism step in lockstep, so the most blocked engine sets the pod's pace. A mean reads 2 for waiting 7,0,0,0.
  • RunningRequestsSize: mean over engines, rounded up. Scorers compare it across pods without knowing each pod's engine count, and a sum makes a four-engine pod score four times busier at equal per-engine load.
  • KVCacheUsagePercent: max over engines. The value stays a fraction.

A single-series family reads the same value as today. LoRA, cache-info, custom and multi-cluster metrics are unchanged.

I have a fix ready that adds an aggregating read on Spec, used for these three fields, plus unit tests.

How to reproduce it (as minimally and precisely as possible):

Unit level: pass the extractor a PrometheusMetricMap whose vllm:num_requests_waiting family holds four series with different engine labels and no timestamps. WaitingQueueSize equals the first series' value.

Cluster: deploy vLLM with --data-parallel-size 4 behind EPP under load. Compare EPP's llm_d_epp_per_endpoint_queue_size for the pod with the pod's own vllm:num_requests_waiting series. EPP's value matches engine="0".

Anything else we need to know?:

Reported against GAIE as kubernetes-sigs/gateway-api-inference-extension#2950 (closed, not fixed). I found no open llm-d issue covering it. #2306 and #2598 cover SGLang DP ranks for P/D, which is a different problem.

Environment:

  • Kubernetes version:
  • llm-d-router version: 0.11.00
  • Cloud provider or hardware configuration:
  • Install tools:
  • Others: vLLM with data parallelism and internal load balancing

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageIndicates an issue or PR lacks a triage label and requires one.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions