bnellnm and GitHub
6427603ae8
[MoE Refactor] Move remaining experts classes to experts directory ( #42334 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-12 09:19:46 -04:00
206eaed08d
[MoE Refactor] Move expert map related code into ExpertMapManager class ( #41046 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-05-12 09:18:27 -04:00
8f89381fc6
[Hybrid] Warmup Mamba2 SSD kernel ( #39822 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-12 12:46:22 +00:00
Dipika Sikka and GitHub
a7b801e26d
[MXFP4] Support for linear layers + compressed-tensors integration ( #41664 )
2026-05-12 07:49:33 -04:00
Kunshang Ji and GitHub
4df1be9547
[XPU] bump up vllm-xpu-kernels to v0.1.8 ( #42410 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-05-12 11:47:37 +00:00
bc03f280c8
[XPU] keep generator state of sycl kernel align with pytorch ( #41771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Qiming Zhang <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 11:44:47 +00:00
997132911e
[Doc] Fix typo in llm-d documentation link ( #42397 )
...
Signed-off-by: Florian Woerner <florian.woerner@onmyown.io >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-05-12 04:26:46 -07:00
haosdent and GitHub
fc8bf6eedb
[CI] De-flake Language Models Test (Extended Generation) test_models(False-False-5-32-bigcode/starcoder2-3b) ( #42392 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-12 10:46:48 +00:00
liuzhenwei and GitHub
07a40ede19
[UT][XPU] fix test_parallel_sampling due to global random state ( #42388 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-12 18:03:23 +08:00
Kevin H. Luu and GitHub
e1c8776e90
[CI] Move DockerHub and PyPI publish steps to end of release pipeline ( #42355 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 09:17:42 +00:00
Kevin H. Luu and GitHub
1ff9d33535
[CI] Migrate remaining B200 jobs to b200-k8s with test fixes ( #42387 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 02:00:37 -07:00
7f65f84428
[Bugfix] Fix empty channel/recipient in harmony for /v1/responses ( #35540 )
...
Signed-off-by: kg6-sleipnir <christopherhazen42@gmail.com >
Signed-off-by: chazen <45186108+kg6-sleipnir@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-12 08:45:51 +00:00
amitz-nv and GitHub
ef34592a1a
[Bugfix] Fix double reduce in flashinfer_nvlink_two_sided and flashinfer_nvlink_one_sided backends ( #41382 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-12 07:47:47 +00:00
Kevin H. Luu and GitHub
f69644caf8
[CI] Migrate more B200 jobs to b200-k8s queue ( #42356 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 00:38:31 -07:00
d37e25ffbe
[Frontend] Consolidate Speech to Text entrypoints. ( #42370 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 07:06:57 +00:00
8517cdaf90
[XPU] update dp rank w/o env-var isolation ( #39856 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 14:49:54 +08:00
Lucas Kabela and GitHub
4e498b5e5c
[Bugfix][Performance Improvement] Improve penalties triton kernel performance ( #40657 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-12 05:47:20 +00:00
28ee78af54
Implement custom dataset class for ASR benchmarking ( #41576 )
...
Signed-off-by: Yasmin Moslem <48152713+ymoslem@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 12:17:58 +08:00
ZiTian Zhao and GitHub
630492da30
[Fix] Gemma4 Mixed-Resolution Image Co-Batching Crash ( #42217 )
...
Signed-off-by: zitian.zhao <zitian.zhao@tencentmusic.com >
2026-05-12 03:13:03 +00:00
Chauncey and GitHub
920bf3ec84
[Bugifx] [Qwen3CoderTool] Restore supports_required_and_named for required tool_choice ( #42292 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-12 02:09:56 +00:00
pschlan-amd and GitHub
39dff5ff39
Add VLLM_USE_SPINLOOP_EXT to use more efficient busy polling ( #36517 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 16:11:49 -07:00
d7af6b34d8
[Model Runner V2] Bug fix: logprob dtype int64/int32 issue ( #41761 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 21:55:43 +00:00
bbee532988
[Perf][1/n] Eliminate various GPU<->CPU syncs ( #41429 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 20:36:03 +00:00
53181384e0
[Bugfix] Fix DSV4 swiglu_limit on marlin backend ( #42287 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 13:03:56 -07:00
wang.yuqi and GitHub
a0dc7a0f36
[CI] Consolidate Speech to Text tests ( #42274 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-11 19:50:17 +00:00
56e5810ff1
[BugFix] Prevent orphaned process on NCCL destroy ( #39846 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-11 15:25:26 -04:00
Flora Feng and GitHub
639cbfd274
[CI] Add tests/parser to CI coverage ( #41877 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-11 19:08:54 +00:00
a721315488
[ROCm][Perf] Fix RMSNorm+Quant fusion for gfx950 (non-fnuz) ( #41825 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Chuan Li <chuali@amd.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: Frida Andersson <frida-andersson@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-11 15:00:51 -04:00
6fdb49392e
[Bugfix] Fix int32 overflow in DeepGEMM SiLU/mul FP8 Triton kernel ( #42201 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Signed-off-by: Vensen <vensenmu@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 14:52:31 -04:00
cf0d279142
[Docs] Add Apple Silicon documentation for vLLM-Metal GPU support ( #41987 )
...
Signed-off-by: alexagriffith <agriffith96@gmail.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-05-11 11:34:25 -07:00
5497ffbf7c
Add documentation about vLLM FIPS compliance ( #42190 )
...
Signed-off-by: Vinay Damodaran <vrdn@hey.com >
Signed-off-by: Vinay R Damodaran <vrdn@hey.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
2026-05-11 18:17:02 +00:00
Nick Hill and GitHub
9af6a5ed75
[Model Runner V2] Fix seq_lens_cpu_upper_bound ( #42202 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-11 10:37:50 -07:00
Hexiang Wang and GitHub
7863fff6e5
[ROCm][DSv4] implement flash sparse mla with triton kernels ( #41812 )
...
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
2026-05-11 09:27:11 -07:00
Wentao Ye and GitHub
0d453e2336
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement ( #40408 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-11 12:20:58 -04:00
Wentao Ye and GitHub
3f9c0c25b3
[Bug] Fix kimi dtype issue with mm_projector_forward ( #42081 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 11:45:24 -04:00
Vadim Gimpelson and GitHub
a2e776d716
[Bugfix] Accept canonicalized modelopt_* quant_method in _extract_modelopt_quant_algo ( #42181 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-11 11:10:57 -04:00
4955990f1b
[kv_offload] Move FilterReusedOffloadingManager logic to CPUOffloadingManager ( #41727 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 18:09:29 +03:00
Wentao Ye and GitHub
4b64fc2cbf
[Refactor] Cleanup batch invariant dead code ( #41993 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 10:48:39 -04:00
pschlan-amd and GitHub
5f1b313900
[ROCm] Clean up a bit the AITER FA backend ( #41942 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 22:45:18 +08:00
724ed2fc35
[DSv4] Improved dequant gather K cache kernel ( #42236 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 10:41:12 -04:00
a51376b3f0
[Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA ( #40392 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
Co-authored-by: ElizaWszola <ewszola@redhat.com >
2026-05-11 14:10:50 +00:00
Martin Hickey and GitHub
8415bf2cdb
[kv_offload] Set offloading connector to prefer HND layout ( #41928 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-11 15:05:41 +03:00
Noa Neria and GitHub
ac062147fa
Avoid silent weights corruption when loading Nemotron Nano VL with reusable-buffer loaders like runai distributed streaming ( #42244 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-05-11 12:03:14 +00:00
Chauncey and GitHub
617239b70c
[Frontend]Responses API supports chat_template_kwargs ( #42272 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-11 11:59:39 +00:00
Kyungmin Lee and GitHub
27ae676364
Fix EXAONE-4.5 to align with Transformers update ( #42246 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
2026-05-11 10:25:31 +00:00
haosdent and GitHub
17ed5e61f5
[CI] Make Python-only Installation optional ( #42293 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-11 09:47:16 +00:00
Nicolò Lucchesi and GitHub
5672d100ed
[KV Connector][NIXL][Bugfix] Fix NIXL handshake failures not honoring kv_load_failure_policy ( #40364 )
...
When NIXL handshake fails (e.g., due to compatibility hash mismatch
between prefill and decode instances), requests fail with "engine dead"
error instead of gracefully falling back to local recomputation as configured
by kv_load_failure_policy='recompute'.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:37:21 +00:00
Nicolò Lucchesi and GitHub
770e9bd6b3
[Nixl][PD] Lease renewal TTL KV blocks on P ( #41383 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:27:30 +00:00
Cyrus Leung and GitHub
9efdddca28
[Model] Fix missing maybe_prefix ( #42280 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-11 09:04:06 +00:00
Qiu and GitHub
b1b59720b2
bugfix(flashinfer,dcp): remove kv_cache_layout for BatchDCPPrefillWrapper._new_tokens. ( #38895 )
...
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com >
2026-05-11 08:11:49 +00:00