What happened:
A vLLM pod running data parallelism with internal load balancing (--data-parallel-size N) exposes one series per engine in each metric family:
vllm:num_requests_waiting{engine="0",model_name="m"} 6.0
vllm:num_requests_waiting{engine="1",model_name="m"} 10.0
vllm:num_requests_waiting{engine="2",model_name="m"} 11.0
vllm:num_requests_waiting{engine="3",model_name="m"} 7.0
The core metrics extractor reads one series per family. Spec.getLatestMetric (pkg/epp/framework/plugins/datalayer/extractor/metrics/spec.go:104, main @ 32c11b7) keeps the matching series with the greatest timestamp:
for _, metric := range family.GetMetric() {
if spec.labelsMatch(metric.GetLabel()) {
ts := metric.GetTimestampMs()
if ts > recent {
recent = ts
latest = metric
}
}
}
vLLM exposes no timestamps, so this is always the first series. WaitingQueueSize, RunningRequestsSize and KVCacheUsagePercent (extractor.go:131-154) describe engine 0 only. The pod above reads a queue of 6.
Every consumer of these fields sees one engine's load: the queue-depth, running-requests, KV-cache-utilization and load-aware scorers, the utilization filter, and the flow-control utilization saturation detector. A pod whose engine 0 is idle scores as idle while its other engines are saturated.
What you expected to happen:
Pod-level values cover every engine:
WaitingQueueSize: max over engines. Engines running DP with expert parallelism step in lockstep, so the most blocked engine sets the pod's pace. A mean reads 2 for waiting 7,0,0,0.
RunningRequestsSize: mean over engines, rounded up. Scorers compare it across pods without knowing each pod's engine count, and a sum makes a four-engine pod score four times busier at equal per-engine load.
KVCacheUsagePercent: max over engines. The value stays a fraction.
A single-series family reads the same value as today. LoRA, cache-info, custom and multi-cluster metrics are unchanged.
I have a fix ready that adds an aggregating read on Spec, used for these three fields, plus unit tests.
How to reproduce it (as minimally and precisely as possible):
Unit level: pass the extractor a PrometheusMetricMap whose vllm:num_requests_waiting family holds four series with different engine labels and no timestamps. WaitingQueueSize equals the first series' value.
Cluster: deploy vLLM with --data-parallel-size 4 behind EPP under load. Compare EPP's llm_d_epp_per_endpoint_queue_size for the pod with the pod's own vllm:num_requests_waiting series. EPP's value matches engine="0".
Anything else we need to know?:
Reported against GAIE as kubernetes-sigs/gateway-api-inference-extension#2950 (closed, not fixed). I found no open llm-d issue covering it. #2306 and #2598 cover SGLang DP ranks for P/D, which is a different problem.
Environment:
- Kubernetes version:
- llm-d-router version: 0.11.00
- Cloud provider or hardware configuration:
- Install tools:
- Others: vLLM with data parallelism and internal load balancing
What happened:
A vLLM pod running data parallelism with internal load balancing (
--data-parallel-size N) exposes one series per engine in each metric family:The core metrics extractor reads one series per family.
Spec.getLatestMetric(pkg/epp/framework/plugins/datalayer/extractor/metrics/spec.go:104, main @ 32c11b7) keeps the matching series with the greatest timestamp:vLLM exposes no timestamps, so this is always the first series.
WaitingQueueSize,RunningRequestsSizeandKVCacheUsagePercent(extractor.go:131-154) describe engine 0 only. The pod above reads a queue of 6.Every consumer of these fields sees one engine's load: the queue-depth, running-requests, KV-cache-utilization and load-aware scorers, the utilization filter, and the flow-control utilization saturation detector. A pod whose engine 0 is idle scores as idle while its other engines are saturated.
What you expected to happen:
Pod-level values cover every engine:
WaitingQueueSize: max over engines. Engines running DP with expert parallelism step in lockstep, so the most blocked engine sets the pod's pace. A mean reads 2 for waiting7,0,0,0.RunningRequestsSize: mean over engines, rounded up. Scorers compare it across pods without knowing each pod's engine count, and a sum makes a four-engine pod score four times busier at equal per-engine load.KVCacheUsagePercent: max over engines. The value stays a fraction.A single-series family reads the same value as today. LoRA, cache-info, custom and multi-cluster metrics are unchanged.
I have a fix ready that adds an aggregating read on
Spec, used for these three fields, plus unit tests.How to reproduce it (as minimally and precisely as possible):
Unit level: pass the extractor a
PrometheusMetricMapwhosevllm:num_requests_waitingfamily holds four series with differentenginelabels and no timestamps.WaitingQueueSizeequals the first series' value.Cluster: deploy vLLM with
--data-parallel-size 4behind EPP under load. Compare EPP'sllm_d_epp_per_endpoint_queue_sizefor the pod with the pod's ownvllm:num_requests_waitingseries. EPP's value matchesengine="0".Anything else we need to know?:
Reported against GAIE as kubernetes-sigs/gateway-api-inference-extension#2950 (closed, not fixed). I found no open llm-d issue covering it. #2306 and #2598 cover SGLang DP ranks for P/D, which is a different problem.
Environment: