53181384e0
[Bugfix] Fix DSV4 swiglu_limit on marlin backend ( #42287 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 13:03:56 -07:00
wang.yuqi and GitHub
a0dc7a0f36
[CI] Consolidate Speech to Text tests ( #42274 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-11 19:50:17 +00:00
56e5810ff1
[BugFix] Prevent orphaned process on NCCL destroy ( #39846 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-11 15:25:26 -04:00
Flora Feng and GitHub
639cbfd274
[CI] Add tests/parser to CI coverage ( #41877 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-11 19:08:54 +00:00
a721315488
[ROCm][Perf] Fix RMSNorm+Quant fusion for gfx950 (non-fnuz) ( #41825 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Chuan Li <chuali@amd.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: Frida Andersson <frida-andersson@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-11 15:00:51 -04:00
6fdb49392e
[Bugfix] Fix int32 overflow in DeepGEMM SiLU/mul FP8 Triton kernel ( #42201 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Signed-off-by: Vensen <vensenmu@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 14:52:31 -04:00
cf0d279142
[Docs] Add Apple Silicon documentation for vLLM-Metal GPU support ( #41987 )
...
Signed-off-by: alexagriffith <agriffith96@gmail.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-05-11 11:34:25 -07:00
5497ffbf7c
Add documentation about vLLM FIPS compliance ( #42190 )
...
Signed-off-by: Vinay Damodaran <vrdn@hey.com >
Signed-off-by: Vinay R Damodaran <vrdn@hey.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
2026-05-11 18:17:02 +00:00
Nick Hill and GitHub
9af6a5ed75
[Model Runner V2] Fix seq_lens_cpu_upper_bound ( #42202 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-11 10:37:50 -07:00
Hexiang Wang and GitHub
7863fff6e5
[ROCm][DSv4] implement flash sparse mla with triton kernels ( #41812 )
...
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
2026-05-11 09:27:11 -07:00
Wentao Ye and GitHub
0d453e2336
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement ( #40408 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-11 12:20:58 -04:00
Wentao Ye and GitHub
3f9c0c25b3
[Bug] Fix kimi dtype issue with mm_projector_forward ( #42081 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 11:45:24 -04:00
Vadim Gimpelson and GitHub
a2e776d716
[Bugfix] Accept canonicalized modelopt_* quant_method in _extract_modelopt_quant_algo ( #42181 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-11 11:10:57 -04:00
4955990f1b
[kv_offload] Move FilterReusedOffloadingManager logic to CPUOffloadingManager ( #41727 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 18:09:29 +03:00
Wentao Ye and GitHub
4b64fc2cbf
[Refactor] Cleanup batch invariant dead code ( #41993 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 10:48:39 -04:00
pschlan-amd and GitHub
5f1b313900
[ROCm] Clean up a bit the AITER FA backend ( #41942 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 22:45:18 +08:00
724ed2fc35
[DSv4] Improved dequant gather K cache kernel ( #42236 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 10:41:12 -04:00
a51376b3f0
[Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA ( #40392 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
Co-authored-by: ElizaWszola <ewszola@redhat.com >
2026-05-11 14:10:50 +00:00
Martin Hickey and GitHub
8415bf2cdb
[kv_offload] Set offloading connector to prefer HND layout ( #41928 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-11 15:05:41 +03:00
Noa Neria and GitHub
ac062147fa
Avoid silent weights corruption when loading Nemotron Nano VL with reusable-buffer loaders like runai distributed streaming ( #42244 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-05-11 12:03:14 +00:00
Chauncey and GitHub
617239b70c
[Frontend]Responses API supports chat_template_kwargs ( #42272 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-11 11:59:39 +00:00
Kyungmin Lee and GitHub
27ae676364
Fix EXAONE-4.5 to align with Transformers update ( #42246 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
2026-05-11 10:25:31 +00:00
haosdent and GitHub
17ed5e61f5
[CI] Make Python-only Installation optional ( #42293 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-11 09:47:16 +00:00
Nicolò Lucchesi and GitHub
5672d100ed
[KV Connector][NIXL][Bugfix] Fix NIXL handshake failures not honoring kv_load_failure_policy ( #40364 )
...
When NIXL handshake fails (e.g., due to compatibility hash mismatch
between prefill and decode instances), requests fail with "engine dead"
error instead of gracefully falling back to local recomputation as configured
by kv_load_failure_policy='recompute'.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:37:21 +00:00
Nicolò Lucchesi and GitHub
770e9bd6b3
[Nixl][PD] Lease renewal TTL KV blocks on P ( #41383 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:27:30 +00:00
Cyrus Leung and GitHub
9efdddca28
[Model] Fix missing maybe_prefix ( #42280 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-11 09:04:06 +00:00
Qiu and GitHub
b1b59720b2
bugfix(flashinfer,dcp): remove kv_cache_layout for BatchDCPPrefillWrapper._new_tokens. ( #38895 )
...
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com >
2026-05-11 08:11:49 +00:00
f9f770ca0b
fix nixl side-channel host selection ( #41806 )
...
Signed-off-by: Shahar Mor <smor@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 07:40:37 +00:00
Haoqing Wang and GitHub
5cba6839e6
Document MolmoWeb hf_overrides ( #42163 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-10 23:58:22 -07:00
Jee Jee Li and GitHub
05d610e5cd
[CI/Build] Reduce LoRA model tests. ( #42266 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-11 14:49:08 +08:00
581b5e9afc
[Frontend] Return rendered prompt text in chat completion response ( #42052 )
...
Signed-off-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Cursor <cursor@cursor.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-11 13:53:39 +08:00
wangxiyuan and GitHub
5536fc0c01
[Misc] Replace mamba_type string literals with MambaAttentionBackendEnum ( #41188 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-05-11 03:59:36 +00:00
vllmellm and GitHub
7f95e66a11
[ROCm][Bugfix]: dynamically align BLOCK_DMODEL with Lv in MLA decode kernel ( #41119 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-11 11:14:19 +08:00
yzong-rh and GitHub
b1687527b8
[Bugfix] Gemma 4 chat template crash with missing tool name and tool id ( #42188 )
...
Signed-off-by: Yifan <yzong@redhat.com >
2026-05-11 03:07:45 +00:00
171019ab19
add fused mhc_post_pre kernel ( #41536 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-10 19:56:52 -07:00
Haoqing Wang and GitHub
879a8c3180
Fix Molmo2 image token metadata ( #42162 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-11 01:19:21 +00:00
1b57eb41f2
[MoE] Move various experts classes to fused_moe/experts/ ( #41979 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
2026-05-11 07:54:33 +08:00
Mohammad Miadh Angkad and GitHub
21943d4c25
[Performance] Make safetensors checkpoint prefetch settings configurable ( #41499 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 15:55:15 +00:00
f396bee56f
[DSV4] Add PP support for deepseek-v4 ( #41694 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: qizixi <22851944+zixi-qi@users.noreply.github.com >
2026-05-10 15:47:26 +00:00
215e2f7990
[Bugfix][Mamba] IMA in causal_conv1d kernel for long sequences ( #41617 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 12:38:28 +00:00
Ronen Schaffer and GitHub
e175192d33
[KV Offload] Pass ReqContext to touch(), complete_load(), and complete_store() ( #41366 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-10 15:09:25 +03:00
a54f0d1049
[CPU] Fix spec decode kernel signatures for synthetic mode compatibility ( #41932 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-05-10 12:07:15 +00:00
Isotr0py and GitHub
48698b1b9b
[Bugfix] Fuse Qwen3.5 in_qkvz_proj forwarding with LoRA enabled ( #37912 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-05-10 10:59:02 +00:00
Andreas Karatzas and GitHub
0a309b5ee9
[ROCm] Cap Triton paged attention block size to fix ROCm shared memory OOM ( #38502 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 10:03:00 +00:00
Jee Jee Li and GitHub
84f7a55340
[CI] Trigger LoRA test when changing MoE code. ( #42196 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-10 01:26:09 -07:00
Ethan Feng and GitHub
a2c9d548d7
[Docs] Fix broken local links ( #42160 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-10 01:15:38 -07:00
Yongye Zhu and GitHub
301305c093
Add @zyongye to CODEOWNERS ( #42200 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-10 16:07:32 +08:00
Mohammad Miadh Angkad and GitHub
efd0e7789d
Fix mypy failure on main ( #42197 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 07:55:57 +00:00
a5d0a5afba
[Frontend][Bugfix] Abort ASR engine requests on cancellation ( #41266 )
...
Signed-off-by: abdulrahman-cohere <abdulrahman.abdulrazzag@cohere.com >
Signed-off-by: <>
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-09 23:51:11 -07:00
Andreas Karatzas and GitHub
f2840120f6
[ROCm][CI] Fix NIXL spec-decode acceptance startup and diagnostics ( #41313 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 14:50:16 +08:00