Alexander Matveev and Alexander Matveev
e3b4fdaf5d
perf: add push-based allreduce for small tensor reductions
...
Port SGLang's push-based 2-buffer allreduce protocol into vLLM as a new
communicator backend for small-message reductions. The push protocol
eliminates the two explicit cross-GPU NVLink barrier round-trips used by
the existing barrier-based CustomAllreduce, replacing them with a
sentinel-based data arrival detection mechanism and double-buffered epoch
alternation.
Key advantages over the barrier-based approach:
- Zero barriers: data arrival IS the synchronization (positive-zero sentinel)
- Single NVLink round-trip instead of two barrier exchanges + remote reads
- All SMs active (SM_count CTAs vs 2 CTAs) for higher NVLink bandwidth
- No cudaMemcpy to IPC staging buffer in eager mode
- PDL (griddepcontrol) support for kernel overlap on sm_90+
The new PushAllReduce is inserted in the CudaCommunicator dispatch chain
above the existing CustomAllreduce for messages below a size threshold
(~720 KB at TP=8). Larger messages continue to use the barrier-based
path. The existing CustomAllreduce code is not modified.
Measured results on DeepSeek-V4-Pro (61 layers, TP=8, 8x NVIDIA B200,
BS=1, decode with ISL=4, OSL=33024):
- Throughput: +2.14% (84.06 vs 82.30 tokens/s)
- TPOT: -2.09% (11.90 vs 12.15 ms/token)
Correctness verified via lm_eval gsm8k 5-shot with no regression
(exact_match delta within statistical noise).
The feature can be disabled at runtime via VLLM_DISABLE_PUSH_ALLREDUCE=1
to fall back to the barrier-based path.
Signed-off-by: Alexander Matveev <amatveev@redhat.com >
2026-06-15 16:23:12 -04:00
Flora Feng and GitHub
cd9078fe59
[Frontend] Skip structural tags for auto tool_choice without strict mode ( #45600 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 19:55:31 +00:00
Wentao Ye and GitHub
e18fe932ca
[Perf] Optimize DSv4 prefill chunk planning, 4.0% E2E Throughput Improvement ( #45061 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-15 19:50:21 +00:00
51ec5cf08f
[Bugfix] Chat Completions Harmony Refactor Clean up ( #45464 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-15 14:45:19 -04:00
7e612a0f06
[KV Offloading] Implement reset_cache for TieringOffloadingManager ( #44541 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 18:42:53 +00:00
+1
0a1c5034f5
[Model] Add MiniMax M3 support ( #45381 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-16 01:01:25 +08:00
RoyWang and GitHub
a3195fab7b
[AMD][Bugfix][Quantization] Honor fused-name match in is_layer_skipped ( #43981 )
2026-06-15 09:37:52 -07:00
Flora Feng and GitHub
0d80979644
[Chore] Consolidate reasoning/tool parser attributes into unified Parser in chat serving ( #45548 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 11:16:45 -04:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
5ed15f42b9
Fix the E8M0 scale computation in the MXFP4 (W4A4) MOE CUTLASS kernel ( #43557 )
...
Signed-off-by: Xin He <xin3.he@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-15 06:04:54 -07:00
Juan Pérez de Algaba and GitHub
b997071ec4
(security) Enforce audio upload size limit before full file materialization ( #45510 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-15 10:25:24 +00:00
Yejing Lai and GitHub
9872921c5f
[XPU] skip UT test_with_ngram_gpu_spec_decoding ( #44423 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-06-15 08:46:30 +00:00
Ting SUN and GitHub
48df95c43e
[Feature][Frontend] Report multimodal token counts in usage.prompt_tokens_details ( #45458 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-15 05:20:58 +00:00
7df4fe1bd7
[Model] Remove XverseForCausalLM ( #45638 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-14 22:09:00 -07:00
c4a3f9d137
[Frontend] Add Streaming Parser Engine and new Qwen3 Parser ( #45413 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 11:59:05 +08:00
Li, Jiang and GitHub
8760f972ca
[CPU] Refine CPU attention frontend ( #45391 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-14 19:26:54 -07:00
Chaojun Zhang and GitHub
2725c84aae
[XPU] Enable sequence parallel support for XPU ( #38608 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-06-14 19:26:46 -07:00
Ting SUN and GitHub
3d6ce816f0
[Bugfix][Model] Validate runai_streamer model_loader_extra_config ( #45291 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-14 19:23:30 -07:00
Taneem Ibrahim and GitHub
2c764c089a
Added real /v1/embeddings support for messages + chat_template_kw ( #45173 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-15 09:08:10 +08:00
Amanzhol Salykov and GitHub
725c3bc808
[ROCm][Perf] Enable W4A16 FlyDSL MoE ( #44400 )
...
Signed-off-by: amd-asalykov <asalykov@amd.com >
Signed-off-by: Amanzhol Salykov <asalykov@amd.com >
2026-06-14 00:14:39 -07:00
4ef4492e9b
[V1][Spec Decode] Add Dynamic SD ( #32374 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-14 00:14:27 -07:00
54bbf51668
[Bugfix] nightly Docker images crash with ImportError: AnthropicOutputConfig since May 28 ( #44795 )
...
Signed-off-by: achyuthan.s <113010327+Achyuthan-S@users.noreply.github.com >
Signed-off-by: Achyuthan S <achyuthan.sivasankar@gmail.com >
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-13 21:45:29 -07:00
71b961dd35
[Perf] SM90 cutlass fp8 mm supports odd M by swap_ab, 180~290% kernel performance improvement ( #44572 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 12:05:45 -07:00
521b88c29e
[Bugfix] Reject structured outputs for diffusion decoders with a clear error ( #45468 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-13 12:04:01 -07:00
Juan Pérez de Algaba and GitHub
470229c37e
[Security] Fix DoS via prompt_embeds on M-RoPE models ( #45252 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-13 10:17:38 +00:00
2b3006076c
[Security] Add timeout guard for regex compilation in structured outp… ( #45118 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 09:52:56 +00:00
WEI CHENG CHIU and GitHub
5b2943f5a6
[Bugfix] Return the tokenizer from maybe_make_thread_pool so it survives pickling ( #45460 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 06:01:35 +00:00
43f0e024bc
[Render] Add /derender endpoints for disaggregated postprocessing ( #43606 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-13 13:55:33 +08:00
Andreas Karatzas and GitHub
1033ffac2e
[CI] Wait for SSL cert refresher events in the test ( #45489 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 04:57:18 +00:00
WEI CHENG CHIU and GitHub
17ee5b1ac5
[Bugfix] Set type/role explicitly in streaming message_start event ( #45376 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 01:40:50 +00:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
badddd254f
[ROCm][DSV4][Perf] Fuse inverse-RoPE and cache bf16 wo_a in o-projection ( #45103 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 15:57:09 -05:00
c90650088d
Add the QuantizedActivation linear-kernel contract ( #44260 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-12 13:48:15 -07:00
Michael Goin and GitHub
9eaacb23ec
[Kernel] Consolidate Marlin thread-tile padding across all dense Marlin paths ( #45295 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-12 13:46:21 -07:00
78739c1946
[Model Runner v2] Migration from v1 to v2, with Qwen and DSv2 MOE models [3/N] ( #42667 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 20:44:52 +00:00
Flora Feng and GitHub
6e4a547176
[Refactor] Deprecate ResponsesParser wrapper, inline parsing into ParsableContext ( #45431 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-12 16:15:41 -04:00
aab639c705
[Core][AMD] Propagate shutdown timeout to MultiprocExecutor ( #43154 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-12 15:13:31 -05:00
Isotr0py and GitHub
6635279d8a
[Migration] Migrate GGUF quantization support to plugin ( #39612 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 12:02:21 -07:00
Jonas I. Liechti and GitHub
d6fd7ce8da
[Model][Dflash] Enable Dflash support for Qwen3NextForCausalLM targets ( #45319 )
...
Signed-off-by: Jonas I. Liechti <j-i-l@t4d.ch >
2026-06-12 10:30:09 -07:00
272c16953e
[Kernel][Helion][1/N] Add Helion kernel for dynamic_per_token_scaled_fp8_quant ( #33790 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-12 12:50:06 -04:00
Chauncey and GitHub
3b8fc3fe6d
[Frontend] Support strict mode for tool calling with ResponsesAPI ( #45396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-12 10:59:59 -04:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
Guan-Ming (Wesley) Chiu and GitHub
c7aa3d2630
[Core] Support structured outputs for beam search ( #35022 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <guanmingchiu@gmail.com >
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-06-12 06:56:25 -07:00
4171ae406c
[V1][Metrics] Add MLA attention metrics for DeepSeek MFU estimation ( #39457 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-06-12 14:28:40 +01:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
f1e13f7df9
[Model] Remove Mono-InternVL (InternLM2VEForCausalLM) ( #45129 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 10:41:09 +00:00
88ed636218
[KV Connector]: Support KV push from Prefill to Decode node using Nixl KV Connector ( #35264 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-06-12 10:38:41 +00:00
2043258dec
[Frontend] Support strict mode for tool calling ( #45003 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: cjackal <44624812+cjackal@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 07:51:48 +00:00
Yuwen Zhou and GitHub
0cd9b7af25
[CPU] Support CPU W4A16 INT4 MoE ( #43409 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-12 07:12:37 +00:00
39dee1114a
[MM][Perf][CG] Support ViT full cudagraphs for mllama4 ( #40660 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-11 22:17:55 -07:00