vLLM is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, you’ll find vLLM-related work throughout the program, from Simon Mo’s keynote to technical sessions and posters covering LLM inference and serving.
Simon will present the keynote “vLLM Update: Scaling Open Frontier Inference Infrastructure,” covering improvements to vLLM’s core architecture and major optimizations in KV cache management and GPU kernels.
Across #PyTorchCon, technical sessions dig into vLLM-related inference, including attention, KV cache management and transfer, disaggregated serving, elastic expert parallelism, scaling across hardware without forks, and more.
"Meet the Developers of vLLM" will feature George Novack and Nick Hill. Posters cover additional vLLM-related work on custom accelerators, Trainium, expert parallelism and RDMA KV-cache transfer, multi-chip KV-cache transfer, tier-aware routing, and more.
Explore the vLLM sessions: https://lnkd.in/eeD4c_xh
PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register for PyTorch Conference North America, October 20–21 in San Jose: https://hubs.la/Q04vK-VQ0