[Compressed Tensors] Add XPU wNa16 support - #29484
Conversation
Signed-off-by: yiliu30 <yi4.liu@intel.com>
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
There was a problem hiding this comment.
Code Review
This pull request introduces support for wNa16 compressed tensors on XPU by adding a new IPEXwNa16LinearKernel. The changes are mostly self-contained in a new file and registration of the new kernel. However, I've found a critical issue in the implementation of the new kernel within vllm/model_executor/layers/quantization/kernels/mixed_precision/ipex.py. The input feature size for the underlying IPEX linear layer is calculated incorrectly, which will likely lead to runtime failures or incorrect computations. I have also pointed out a confusing and redundant variable assignment that should be cleaned up. Please see the detailed comments for suggestions on how to fix these issues.
|
@robertgshaw2-redhat @mgoin Please help review. The PR is to support the quantized models with compressed tensor format (e.g., quantized by LLM-C/AutoRound) for Intel GPUs. |
Signed-off-by: yiliu30 <yi4.liu@intel.com>
Signed-off-by: yiliu30 <yi4.liu@intel.com>
|
@mgoin @robertgshaw2-redhat , may you help to take a look? |
Signed-off-by: yiliu30 <yi4.liu@intel.com>
|
Great work! Sorry for missing this |
Signed-off-by: yiliu30 <yi4.liu@intel.com>
Signed-off-by: yiliu30 <yi4.liu@intel.com>
Signed-off-by: yiliu30 <yi4.liu@intel.com>
Signed-off-by: yiliu30 <yi4.liu@intel.com>
Purpose
cd vllm python examples/offline_inference/basic/generate.py \ --model Intel/Qwen3-8B-W4A16-G128-AutoRound-LLMC-TEST-ONLY \ --gpu_memory_utilization 0.75 \ --enforce-eagerEssential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.