forked from Karylab-cklius/vllm
Port SGLang's push-based 2-buffer allreduce protocol into vLLM as a new communicator backend for small-message reductions. The push protocol eliminates the two explicit cross-GPU NVLink barrier round-trips used by the existing barrier-based CustomAllreduce, replacing them with a sentinel-based data arrival detection mechanism and double-buffered epoch alternation. Key advantages over the barrier-based approach: - Zero barriers: data arrival IS the synchronization (positive-zero sentinel) - Single NVLink round-trip instead of two barrier exchanges + remote reads - All SMs active (SM_count CTAs vs 2 CTAs) for higher NVLink bandwidth - No cudaMemcpy to IPC staging buffer in eager mode - PDL (griddepcontrol) support for kernel overlap on sm_90+ The new PushAllReduce is inserted in the CudaCommunicator dispatch chain above the existing CustomAllreduce for messages below a size threshold (~720 KB at TP=8). Larger messages continue to use the barrier-based path. The existing CustomAllreduce code is not modified. Measured results on DeepSeek-V4-Pro (61 layers, TP=8, 8x NVIDIA B200, BS=1, decode with ISL=4, OSL=33024): - Throughput: +2.14% (84.06 vs 82.30 tokens/s) - TPOT: -2.09% (11.90 vs 12.15 ms/token) Correctness verified via lm_eval gsm8k 5-shot with no regression (exact_match delta within statistical noise). The feature can be disabled at runtime via VLLM_DISABLE_PUSH_ALLREDUCE=1 to fall back to the barrier-based path. Signed-off-by: Alexander Matveev <amatveev@redhat.com>