🚀 The feature, motivation and pitch
Currently, when DP is enabled, we assume the model uses EP or TP for expert layers and as such enable additional synchronization so that dummy forward passes are done in idle ranks when necessary, and certain metadata is shared before each forward pass.
However, this is also done for non-MoE models unnecessarily, which results in unnecessary additional overhead. We should ensure that the extra DP-related sync isn't done in this case.
We do probably still want the DPCoordinator to run, since we'll still want it to propagate the load-balancing stats.
Alternatives
We could consider hard-failing in this situation to avoid folks unintentionally using vLLM with sub-par performance.
🚀 The feature, motivation and pitch
Currently, when DP is enabled, we assume the model uses EP or TP for expert layers and as such enable additional synchronization so that dummy forward passes are done in idle ranks when necessary, and certain metadata is shared before each forward pass.
However, this is also done for non-MoE models unnecessarily, which results in unnecessary additional overhead. We should ensure that the extra DP-related sync isn't done in this case.
We do probably still want the DPCoordinator to run, since we'll still want it to propagate the load-balancing stats.
Alternatives
We could consider hard-failing in this situation to avoid folks unintentionally using vLLM with sub-par performance.