Summary
Add support for collecting AMD GPU telemetry from the AMD Device Metrics Exporter (DME) Prometheus HTTP endpoint. This enables remote AMD GPU monitoring in AIPerf without requiring ROCm on the AIPerf host.
Motivation
AIPerf already supports NVIDIA DCGM (HTTP) and local AMD amdsmi. AMD DME is the standard Kubernetes exporter for AMD GPUs and exposes all relevant metrics (power, energy, utilization, memory, temperature, ECC, clocks) over HTTP. This addition closes the multi-vendor remote-telemetry gap.
Proposed Implementation
New AMDDMETelemetryCollector class extending BaseMetricsCollectorMixin, mirroring the existing DCGMTelemetryCollector pattern. It uses prometheus_client.parser to parse DME Prometheus exposition format and emit AMD-namespaced telemetry fields:
amd_power, amd_energy_consumption, amd_gfx_activity, amd_umc_activity, amd_memory_*, amd_temperature, amd_memory_temperature, amd_ecc_uncorrectable, amd_sm_clock, amd_mem_clock.
New TelemetryMetrics fields include Pydantic ge=0 constraints to satisfy the existing numeric-bounds invariants.
Usage Example
aiperf profile \
--model meta-llama/Llama-3.1-8B-Instruct \
--endpoint-type chat \
--url http://<inference-server-ip>:8000 \
--gpu-telemetry http://<amd-exporter-ip>:5000/metrics
Files Changed
src/aiperf/gpu_telemetry/amd_dme_collector.py (new)
src/aiperf/config/flags/_converter_telemetry.py (auto-detection)
src/aiperf/common/models/telemetry_models.py
src/aiperf/gpu_telemetry/constants.py
src/aiperf/plugin/enums.py
src/aiperf/plugin/plugins.yaml
docs/tutorials/gpu-telemetry.md
Notes
Auto-detection relies on a synchronous httpx probe during CLI parsing. This is documented as a known trade-off and does not affect benchmark results.
No new dependencies are needed (prometheus_client and httpx are already declared in pyproject.toml).
uv run pytest tests/unit/ -n auto passes (12,679 passed; remaining 6 failures are pre-existing environment/permission issues unrelated to this change)
Summary
Add support for collecting AMD GPU telemetry from the AMD Device Metrics Exporter (DME) Prometheus HTTP endpoint. This enables remote AMD GPU monitoring in AIPerf without requiring ROCm on the AIPerf host.
Motivation
AIPerf already supports NVIDIA DCGM (HTTP) and local AMD
amdsmi. AMD DME is the standard Kubernetes exporter for AMD GPUs and exposes all relevant metrics (power, energy, utilization, memory, temperature, ECC, clocks) over HTTP. This addition closes the multi-vendor remote-telemetry gap.Proposed Implementation
New
AMDDMETelemetryCollectorclass extendingBaseMetricsCollectorMixin, mirroring the existingDCGMTelemetryCollectorpattern. It usesprometheus_client.parserto parse DME Prometheus exposition format and emit AMD-namespaced telemetry fields:amd_power,amd_energy_consumption,amd_gfx_activity,amd_umc_activity,amd_memory_*,amd_temperature,amd_memory_temperature,amd_ecc_uncorrectable,amd_sm_clock,amd_mem_clock.New
TelemetryMetricsfields include Pydanticge=0constraints to satisfy the existing numeric-bounds invariants.Usage Example
aiperf profile \ --model meta-llama/Llama-3.1-8B-Instruct \ --endpoint-type chat \ --url http://<inference-server-ip>:8000 \ --gpu-telemetry http://<amd-exporter-ip>:5000/metricsFiles Changed
src/aiperf/gpu_telemetry/amd_dme_collector.py (new)
src/aiperf/config/flags/_converter_telemetry.py (auto-detection)
src/aiperf/common/models/telemetry_models.py
src/aiperf/gpu_telemetry/constants.py
src/aiperf/plugin/enums.py
src/aiperf/plugin/plugins.yaml
docs/tutorials/gpu-telemetry.md
Notes
Auto-detection relies on a synchronous httpx probe during CLI parsing. This is documented as a known trade-off and does not affect benchmark results.
No new dependencies are needed (prometheus_client and httpx are already declared in pyproject.toml).
uv run pytest tests/unit/ -n auto passes (12,679 passed; remaining 6 failures are pre-existing environment/permission issues unrelated to this change)