We just released Ray 2.56! This includes
- Ray Data stability improvements: reduced object store spilling, automatic batch size selection
- Ray Serve LLM re-architecture: decoupling request handling from the token streaming response path, LLM serving performance improvements, new routing policies like session-sticky routing via consistent hashing
- Ray Core GPU-domain-aware placement groups: enables placement groups to pack bundles onto nodes that share a ray.io/gpu-domain label instead of only packing at the single-node level
- Kubernetes integration: initial Kubernetes in-place pod resizing support for Autoscaler v2

