Ray Serve LLM now offers 4.4x higher request throughput on prefill-heavy workloads, and 24.8x higher request throughput on decode-heavy workloads!
🚀Three major optimizations:
- Direct streaming, bypassing an intermediate Ray Serve deployment on the response path with a new, control plane-only endpoint picker
- A new, Ray V2 executor backend in vLLM, enabling optimizations such as async scheduling
- HAProxy ingress, for ingress request routing at the speed of C
All available in Ray 2.56. This is awesome work with @googlecloud and @vllm_project!
Today we are excited to announce, in partnership with the GKE team at Google Cloud (@googlecloud), a major milestone in Ray Serve LLM’s production serving capability. Ray Serve LLM now matches high performance, rust-based routing frameworks such as vllm-router (@vllm_project) in

