Jee Jee Li
931b3f6110
Done
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-25 13:23:05 +00:00
Andreas Karatzas and GitHub
95995bbef8
[ROCm][Engine] Fix GPU memory leaks in engine shutdown and test workaround for async KV prefix cache reset ( #38503 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 05:25:20 +00:00
07351e0883
[Feature] Warm up readonly multimodal processor during renderer startup ( #40797 )
...
Signed-off-by: Chenguang ZHENG <645327136@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-25 03:57:41 +00:00
Andreas Karatzas and GitHub
428b988c98
[ROCm][CI] Fix trust_remote_code AttributeError in EAGLE3 acceptance length test ( #40306 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 02:59:31 +00:00
Andreas Karatzas and GitHub
e54894fc85
[ROCm][CI] Fix TestSiluMulGroupFp8QuantModel after W8A8 block linear refactor ( #39799 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 11:20:59 +09:00
Angela Yi and GitHub
bc2ae5a3d6
[Test] Increase qwen2_vl num_logprobs to fix torch 2.12 update ( #40818 )
...
Signed-off-by: Angela Yi <angelayi@meta.com >
2026-04-25 00:59:20 +00:00
Wentao Ye and GitHub
a474da2813
[Refactor] Remove unused dead code ( #40640 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-25 07:28:18 +08:00
Lucas Kabela and GitHub
ce6a199ecc
[BE][Bugfix] Respect TORCH_COMPILE_DISABLE env var at the vLLM config level for torch 2.12 ( #40715 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-24 16:25:03 -07:00
Ignacio Sica and GitHub
f88763efc3
[Bugfix] add seq_lens_cpu_upper_bound to CommonAttentionMetadata in mla_runner.py ( #40844 )
...
Signed-off-by: ignaciosica <mignacio.sica@gmail.com >
2026-04-24 23:13:52 +00:00
Artem Perevedentsev and GitHub
333529deae
[EPLB] Fix replica selection bias in fused_moe router ( #40810 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-04-24 22:06:41 +00:00
Zhang Jian and GitHub
8825608205
[Bugfix][CI] Fix wrong residual shape in TestFusedAddRMSNorm.example_inputs that causes flaky test ( #40629 )
...
Signed-off-by: Zhang Jian <jianmusings@gmail.com >
2026-04-24 16:40:07 -04:00
qli88 and GitHub
095d2f87e8
[Bug] Fix GLM-5.1 running error on ROCm platform ( #40763 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-04-24 19:54:40 +00:00
21792520e7
[Build] Add Python 3.14 to supported version list. ( #34770 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-04-24 10:24:05 -07:00
Alex Brooks and GitHub
5e11b40365
[Frontend] Delegate to vLLM Omni When --omni Passed ( #40744 )
...
Signed-off-by: Alex Brooks <albrooks@redhat.com >
2026-04-24 12:30:00 -04:00
f768b4473e
[Docs] Add docs for context extension using the yarn method ( #37430 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
Signed-off-by: labAxiaoming <34019940+labAxiaoming@users.noreply.github.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-24 08:26:09 -07:00
JartX and GitHub
914d0464c1
[Refactor] Unify 2D/3D kernels in triton_unified_attention ( #40631 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-04-24 17:18:06 +02:00
Jinzhen Lin and GitHub
9f771b3ab9
[Quantization] add humming quantization kernel ( #34556 )
2026-04-24 09:29:44 -04:00
Itay Alroy and GitHub
c9d3c6e6af
fused_moe: treat NIXL EP as batched experts ( #40412 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-04-24 08:05:31 -05:00
Or Ozeri and GitHub
51adca74e6
[kv_offload+HMA][9/N]: Support lookup with multiple KV groups ( #39401 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-24 15:32:29 +03:00
Netanel Haber and GitHub
e8eb0490ce
[Bugfix][MoE] Unpad routed output before shared expert add [ Fixes #35949 ] ( #40794 )
...
Signed-off-by: Netanel Haber <nhaber@nvidia.com >
2026-04-24 11:53:23 +00:00
Jiangyun Zhu and GitHub
e8ee2a78db
[Attention] use diff kv backend for mimo v2 flash ( #40045 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-04-24 11:25:55 +00:00
2ec18f5df4
[Bugfix][Parser] Fix Mistral tool parser for HF tokenizers ( #39294 )
...
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-24 19:01:56 +08:00
Dmitry Tokarev and GitHub
6dec49f27e
[Build] Bump CUDA to 13.0.2 to match PyTorch 2.11.0 ( #40669 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com >
2026-04-24 10:27:11 +00:00
Shanshan Shen and GitHub
b5587e1013
[CI/Build] Add e2e test for ViT CUDA graph ( #40780 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-24 18:12:14 +08:00
milesial and GitHub
9ad5abe772
Fix Nano Nemotron VL static image inputs ( #40724 )
...
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com >
2026-04-24 09:18:55 +00:00
Woosuk Kwon and GitHub
7d3195ea9f
[Bugfix] Fix IMA in DSA + MTP ( #40772 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-04-24 01:40:20 -07:00
512f522192
[Model] Gemma4: add bidirectional vision attention for sliding layers with window guard ( #40534 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Luciano Martins <lucianomartins@google.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-24 08:27:46 +00:00
4c34b2f6fc
[XPU] Enable torch.compile for XPU GDN attention ( #39466 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-24 16:26:16 +08:00
Xin Yang and GitHub
cf8a613a87
Support only half types for concat_mla_q kernel ( #37892 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-04-23 23:51:05 -07:00
xiangdong and GitHub
01acf96c6f
[XPU][CI] Fix Docker cleanup races on Intel CI runners ( #40761 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-24 14:08:45 +08:00
079a4cf399
[MoE] Move cutlass moe to fused_moe/experts/ ( #40574 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-24 06:05:49 +00:00
9744b699ba
[Deprecate] Deprecate LLM.reward offline api, use LLM.encode instead. ( #40688 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-24 05:37:50 +00:00
c662b4359e
[Bugfix] Avoid mutating chat_template_kwargs in HYV3ReasoningParser initialization ( #40713 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 13:08:58 +08:00
lyd1992 and GitHub
100c7b65e7
[Platform] Fix RISC-V platform detection (lscpu parsing + non-NUMA meminfo) ( #40427 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-04-24 04:33:05 +00:00
Neil Schemenauer and GitHub
56bdf85e10
[Feature] Avoid eager import of the "mistral_common" package. ( #40043 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
2026-04-24 02:49:16 +00:00
Vinayak Kumar and GitHub
eba73068ea
[Doc] fix capitalization consistency in README (vLLM, Hugging Face) ( #40729 )
...
Signed-off-by: Vinayak Mishra <vinayakmishra448@gmail.com >
2026-04-24 02:23:54 +00:00
Nick Hill and GitHub
e9f331d72e
[MRV2] Ensure warmup covers prefill path ( #40746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-24 01:33:26 +00:00
c9bf77df92
[BUG]: fix HF tokenizer concurrent borrow in tool parsers ( #40059 )
...
Signed-off-by: Yifan <yzong@redhat.com >
Co-authored-by: timon0305 <timon0305@outlook.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-04-23 18:20:30 -07:00
3041344287
[Misc] Added curl retries in install_python_libraries.sh ( #36700 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 01:19:30 +00:00
Doug Campos and GitHub
92762edc53
[Bugfix] Treat <tool_call> as implicit reasoning end in Qwen3 parser ( #35687 )
...
Signed-off-by: Doug Campos <qmx@qmx.me >
2026-04-24 09:10:04 +08:00
626daa2076
[Feat] Unified Synthetic Acceptance Rate for V1 and V2 ( #40662 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 00:48:08 +00:00
Nick Hill and GitHub
fe85a92e86
[Core] Avoid seq_lens_cpu GPU->CPU sync ( #40654 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-24 00:35:55 +00:00
Sage Moore and GitHub
62b1bbe470
[EPLB] Remove asyncio infrastructure from Async EPLB ( #40730 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-24 00:21:15 +00:00
Hemanth Acharya and GitHub
fa4b70555b
[ROCm] Cast score correction bias tensor during model construction for DeepSeek/Kimi-K2 ( #39999 )
...
Signed-off-by: Hemanth Acharya <heachary@amd.com >
2026-04-24 09:02:12 +09:00
447c372ac5
[MoE] Move remaining PrepareAndFinalize to prepare finalize folder ( #39009 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-23 20:00:53 -04:00
ff2c2bd80a
[Docs]Add documentation for bench serve visualization arguments ( #40539 )
...
Signed-off-by: Sophie du Couédic <sop@zurich.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-23 15:48:29 -07:00
Matthew Bonanni and GitHub
cde8d24710
[Spec Decode] Move SpecDecodeBaseProposer out of eagle.py ( #40732 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-23 22:28:27 +00:00
bnellnm and GitHub
4a6dd1c3cc
[Bugfix] Fix DeepSeek V2-Lite Accuracy drop ( #40673 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-23 18:11:37 -04:00
7ff65b1900
[Bugfix] Fix workspace resize leaking reserved GPU memory ( #39226 )
...
Signed-off-by: root <conway.zhu@cohere.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 20:50:05 +00:00
Johnny and GitHub
7f95a66cbf
[NVIDIA] Add sm_110 (Jetson Thor) to CUDA 13.0 build targets ( #39233 )
2026-04-23 15:42:14 -04:00
1b1c01de39
[MoE] Move xpu moe to fused_moe/experts/ ( #40568 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 13:38:10 -04:00
e9ba519f45
[DP][Ray] Pin DP control bundle to same node as first GPU bundle ( #39167 )
...
Signed-off-by: Shahar Mor <smor@nvidia.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 17:21:13 +00:00
Or Ozeri and GitHub
5ef33ab250
[kv_offload+HMA][10/N]: Support load with multiple KV groups ( #39402 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-23 20:00:45 +03:00
bnellnm and GitHub
1c2c1eb8b9
[MoE Refactor] Rename FusedMoE.make_expert_params_mapping to fused_moe_make_expert_params_mapping ( #40671 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-23 11:22:34 -04:00
Nicolò Lucchesi and GitHub
8824f50f1f
[CI] Split disaggregated tests into own test-area ( #40623 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-23 23:20:12 +08:00
0098db9ec1
[ROCm] Implement GPU-to-NUMA-node detection ( #40015 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-23 10:08:48 -05:00
Kunshang Ji and GitHub
53ecc807c0
[XPU] Upgrade torch 2.11 for xpu ( #37947 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 10:07:35 -05:00
b7a2605020
[Bugfix] Make Attention Backend Auto-Selection Batch-Invariance-Aware ( #40193 )
...
Signed-off-by: Srreyansh Sethi <srreyansh.sethi@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 14:57:03 +00:00
d0009ddb0b
[Model] Support Hy3 preview ( #40681 )
...
Signed-off-by: stevenkuang <stevenkuang@tencent.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-23 22:08:26 +08:00
Richard Zou and GitHub
424033f4fc
[Bugfix] Include inductor and functorch configs in compilation cache key ( #40627 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-04-23 09:52:59 -04:00
Isotr0py and GitHub
da1e7311ca
[Misc] use model arch converter for bidi models identification ( #40701 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-23 13:42:52 +00:00
xiangdong and GitHub
01cb41dcf5
[XPU][CI]Temporary disable 3 cases on Intel GPU in CI ( #40683 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-23 21:42:22 +08:00
2f314bc5e6
[CPU] Added faster exp routine for lower precision data types. ( #38112 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Co-authored-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-04-23 13:14:44 +00:00
BadrBasowid and GitHub
2196bac135
[Compilation] Refactor SiluMul activation+quant Fusion Pass ( #39684 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-04-23 09:10:36 -04:00
Matthias Gehre and GitHub
4b7869d6bc
[ROCm] Add gfx1102/gfx1103 support ( #40037 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-23 01:32:04 -07:00
liuzhenwei and GitHub
4a79262e0f
[UT][Hardware] let torchrun example tests use the default backend ( #39879 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-04-23 16:22:28 +08:00
3ed5231c6a
[Build] Switch default CUDA to 13.0, update CUDA architecture lists, clean up stale build-args ( #39878 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-23 15:51:28 +08:00
Nicolò Lucchesi and GitHub
9c2492e501
[Misc] Support Human-readable (k/K/m/M..) json cli arg ( #40473 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-23 09:42:23 +02:00
Shanshan Shen and GitHub
fe57be7809
[MM][CG] Support --enable-vit-cuda-graph option for VLM examples ( #40580 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-22 22:46:14 -07:00
8317cedc77
[Responses] Add tool_choice/tools validation to match OpenAI behavior ( #40399 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 22:46:10 -07:00
Zhengxu Chen and GitHub
98a242ff61
[compile] Skip FX graph deserialiaztion on loading, further reducing warm compile time. ( #40151 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-04-23 13:43:18 +08:00
e4ee48da2d
[MoE refactor] refactor GPTQMarlinMoEMethod with MK ( #37990 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-23 05:21:47 +00:00
Kunshang Ji and GitHub
342c58bc54
[BugFix]fix Qwen3 MoE call gate twice ( #40664 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 05:04:41 +00:00
fe9c3d6c5f
[TurboQuant] enable FA3/FA4 for prefill paths ( #40092 )
...
Signed-off-by: 墨楼 <huangzhilin.hzl@antgroup.com >
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Codex <codex@openai.com >
2026-04-23 07:35:24 +03:00
ccaf5ffaa3
[XPU] disable fusion pattern support on XPU platform ( #39789 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 10:07:45 +08:00
Lucas Kabela and GitHub
0283f303d8
[BE] Fix compile time message to be consistent (use monitoring) ( #40641 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-23 00:12:08 +00:00
ac58e2a170
[Fix][MoRI] Align MoRI-IO message format with P2pNcclConnector and vllm-router ( #39565 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: Matvei Pashkovskii <mpashkov@amd.com >
2026-04-23 08:06:31 +09:00
Lucas Kabela and GitHub
b8401a9bf4
[Bugfix] Fix RMS norm + quant fusion on DeepGEMM UE8M0 path for B200 ( #40552 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-22 22:04:42 +00:00
Honglin Cao and GitHub
9c271f9403
[gRPC] Add standard gRPC health checking (grpc.health.v1) for Kubernetes native probes ( #38016 )
...
Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com >
2026-04-22 21:31:00 +00:00
22fa63cfe8
[Bugfix][Torch 2.12] Fix batch_invariant test with allow_override for torch 2.12 upgrade ( #40562 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 13:48:55 -07:00
8f87eb4622
[Refactor] Clean up log once scope="local" ( #40540 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-22 16:42:43 -04:00
cfa49213d7
[Bugfix][Parser] Fix Mistral pre-v11 tool parser failing on trailing model output ( #40531 )
...
Signed-off-by: dougbtv <dosmith@redhat.com >
Signed-off-by: Doug Smith <dougbtv@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-22 16:35:00 -04:00
29f64c5f5e
FlexAttention non-causal support ( #40394 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 13:22:57 -07:00
Angela Yi and GitHub
eb6661d522
Fix test_startup.py for torch 2.12 ( #40636 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-04-22 19:31:41 +00:00
d622e27d2b
[NVFP4] NVFP4 MOE emulation fallback for H100/MI300/MI350, standardize TritonExperts usage for OCP MX emulation ( #35737 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Signed-off-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-04-22 08:58:54 -07:00
5f76b3fb30
[MoE] Convert CT W8A8 To Oracle Structure ( #39187 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-22 14:53:30 +00:00
bnellnm and GitHub
809d83c2dc
[MoE Refactor] Combine MoERunnerBase + DefaultMoERunner ( #40560 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-22 14:43:17 +00:00
Nicolò Lucchesi and GitHub
33ef1941e2
[Bugfix][CI] Fix v1/kv_connector/unit/test_nixl_connector_hma.py::test_fewer_blocks_with_hma ( #40597 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-22 14:21:02 +01:00
Hank_ and GitHub
a4905133f3
[xpu][rocm] Update current_platform.supports_fp8() for TritonExperts ( #40132 )
...
Signed-off-by: Hank <hcc.mayday@gmail.com >
2026-04-22 13:39:40 +02:00
ecbe42e991
[Doc] Clarify supported keys for --speculative-config ( #40455 )
...
Signed-off-by: Wangxiaoxiaoa <Wangxiaoxiaoa@users.noreply.github.com >
Co-authored-by: Wangxiaoxiaoa <Wangxiaoxiaoa@users.noreply.github.com >
2026-04-22 04:36:17 -07:00
a250f1bd5f
[Bugfix] LoRA for DeepSeek V3.2 ( #35077 )
...
Signed-off-by: Hollow Man <hollowman@opensuse.org >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-22 19:33:50 +08:00
lyd1992 and GitHub
04eac6ba24
[Bugfix][CPU][RISC-V] Clamp exp() input to prevent NaN ( #40428 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-04-22 09:38:18 +00:00
9047288b68
support hotwords for FunASR model ( #39674 )
...
Signed-off-by: zixiao <shunli.dsl@alibaba-inc.com >
Co-authored-by: zixiao <shunli.dsl@alibaba-inc.com >
2026-04-22 02:25:06 -07:00
Johnny Yang and GitHub
ed6d30377d
upgrade tpu-inference to v0.18.0 ( #40395 )
2026-04-22 01:33:45 -07:00
6aa057c9d7
[Multimodal] Support custom video metadata for pre-extracted frame sequences ( #40133 )
...
Signed-off-by: storyicon <storyicon@foxmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 15:50:04 +08:00
Chauncey and GitHub
a2bd09c960
[Bugfix] [Reasoning] Add reasoning_start_str/reasoning_end_str properties to reasoning parsers ( #40566 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-22 07:27:44 +00:00
philip-essential and GitHub
123674879e
[Model] Add block-local attention and YaRN for local layers to Gemma3 ( #39823 )
...
Signed-off-by: Philip Monk <169196560+philip-essential@users.noreply.github.com >
2026-04-21 23:34:50 -07:00
Carl Y and GitHub
4254aeb56f
[fix] flaky test_mla_attn_quant_fusion.py ( #40530 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-04-22 06:29:58 +00:00
aad88f8486
[kv_offload+HMA][8/N]: Support multi-group worker transfer ( #38453 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 08:44:00 +03:00
Bugen Zhao and GitHub
0210024ae7
[Bugfix] Pass effective chat template kwargs to reasoning parsers ( #40460 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-04-21 22:17:51 -07:00
4eafc72928
[Audio] Bundle get_generation_prompt() params into SpeechToTextParams ( #36268 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-22 12:24:18 +08:00
Micah Williamson and GitHub
6d09769700
[ROCm] Support non-causal attention in ROCM_ATTN ( #40176 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-22 12:57:12 +09:00
4506319a28
[compile] mla + group fp8 fusion ( #38877 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 23:16:58 -04:00
9b60e2ffaa
[Bugfix] Fix quantized model initialization failure with prefetch offloading ( #40432 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 20:15:58 -07:00
Martin Hickey and GitHub
3951d3eacd
[MyPy] Enable mypy for vllm/model_executor/layers/ ( #40159 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-04-21 20:15:02 -07:00
6f2c71be8f
[Multimodal] Add PyAV video backend for concurrent video decoding ( #39986 )
...
Signed-off-by: Jaseel Muhammad <jaseel.muhammad@mbzuai.ac.ae >
Signed-off-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-21 20:14:57 -07:00
rasmith and GitHub
2463f00fb6
[AMD][CI][BugFix] Override normalize_e4m3fn_to_e4m3fnuz for fnuz machines in test_moe_layer_no_parallel ( #40550 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-22 02:21:02 +00:00
f946659fff
[Bugfix] Fix W4A8_FP8 MoE tp>1 correctness and view() TypeError ( #40310 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 21:58:33 -04:00
Soila Kavulya and GitHub
f90aa44662
[NIXL][XPU]Fix nixl import on XPU ( #40430 )
...
Signed-off-by: Soila Kavulya <soila.p.kavulya.intel.com>
2026-04-22 09:26:33 +08:00
rasmith and GitHub
cefa5281a7
[ROCm][P/D][MORI][BugFix] Ensure correct api is used when making requests to prefill / decode nodes ( #39835 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-22 09:48:25 +09:00
Jhao-Ting Chen and GitHub
46794958f0
test: add nan/inf clamp regression test for fused_topk_bias ( #40553 )
...
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
2026-04-22 00:46:53 +00:00
Khushali Desai and GitHub
6ff8dea075
[Bugfix] avoid warmup if text only expectation in multi_modal run ( #40409 )
...
Signed-off-by: khushali9 <khushali.desai9@gmail.com >
2026-04-22 00:19:50 +00:00
TJian and GitHub
583e6f2226
[ROCm] [Wheel] [Bugfix] [Critical] Remove any packages installed from github from rocm.txt e.g fastsafetensors as it is incompatible with uv pip ( #40461 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-04-22 00:18:07 +00:00
96a85c5750
[Startup][UX] Enable CUDAGraph memory profiling by default ( #38284 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-04-21 18:16:59 -04:00
9db4650e5e
[MoE Refactor] Add more MoE layer tests ( #39349 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-21 18:12:36 -04:00
bnellnm and GitHub
5e584ce9ec
[MoE Refactor] Remove SharedFusedMoE class ( #35782 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-21 18:12:12 -04:00
Wentao Ye and GitHub
1842447c09
[Refactor] Remove unused param ( #39750 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 14:59:20 -07:00
Wentao Ye and GitHub
16688b26a6
[Perf] Optimize batch invariant with fused rms norm, 2.1% E2E latency improvement ( #40413 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 19:51:03 +00:00
Jakub Zakrzewski and GitHub
6fbec8ed47
[Bugfix][Kernel] nvfp4 cutlass MoE: fix nvfp4 experts quant out-of-bounds read for expert counts not divisible by 4 or 16 ( #40351 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
2026-04-21 19:06:09 +00:00
5544f8c18b
[Performance] Add is_reasoning_end_streaming() override to GptOssReasoningParser ( #35745 )
...
Signed-off-by: Fergus <fergus.barratt00@gmail.com >
Signed-off-by: fergus barratt <fergus.barratt00@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-21 18:31:27 +00:00
9f39b380d0
[Bugfix] Fix spec decode test failures on Blackwell (SM100+) ( #39546 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-21 18:21:19 +00:00
Zijing Liu and GitHub
9a6a66f3b8
[MRv2]fix: model accuracy regression caused by reusing the stale last_sampled_tokens and draft_tokens ( #39833 )
...
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com >
2026-04-21 16:30:32 +00:00
67eb6083e3
Revert "[Misc] Move pyav and soundfile to common requirements" ( #40276 )
...
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-21 09:08:06 -07:00
Harry Mellor and GitHub
6ee081d1d0
Add new tp plan styles to the Transformers modelling backend ( #40467 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-21 08:51:30 -07:00
66cc3fa559
[Model Runner V2] Multiple prompt logprobs support ( #39937 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-21 15:49:05 +00:00
Vadim Gimpelson and GitHub
6d85b36a9f
Revert #38730 and #38791 ( #40032 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-04-21 11:44:11 -04:00
Matthew Bonanni and GitHub
ab5666eb7c
[UX] Bump version in CG memory profiling log message ( #40465 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-21 15:26:06 +00:00
roikoren755 and GitHub
f819265a4a
Default to 'align' mamba cache mode for Mamba-based models when speculative decoding is enabled ( #40454 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-21 14:51:43 +00:00
Shanshan Shen and GitHub
936e0b79aa
[MM][CG] Optimize default max_frames_per_batch auto-infer for ViT CUDA graph video inference ( #40445 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 14:47:53 +00:00
b2a5518679
[XPU][CI] Add misc, engine and lora cases on Intel GPU in CI ( #39887 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-21 22:30:46 +08:00
ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 and GitHub
908a713488
[Bugfix] LoRA: extend expert base_layer loading to Qwen3.5 and Step3.x ( #37114 )
...
Signed-off-by: Hollow Man <hollowman@opensuse.org >
2026-04-21 14:17:03 +00:00
ec5ef0ac73
[Doc] Add Qwen3 AWQ models to documentation ( #40034 )
...
Signed-off-by: Yusuf <yusufmohammad@live.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 09:37:41 -04:00
7b1e0b07d0
[Bugfix] Fix dataset name and path argument validation bug in vllm bench serve ( #40288 )
...
Signed-off-by: talora <talora@nvidia.com >
Signed-off-by: Talor Abramovich <talor19@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-21 06:14:28 -07:00
d249a9e90e
Add Granite 4.1 Vision as built-in multimodal model ( #40282 )
...
Signed-off-by: Artem Spector <artems@il.ibm.com >
Signed-off-by: artemspector <artems@il.ibm.com >
Co-authored-by: artemspector <artems@il.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-21 05:43:39 -07:00
d2e2e856ad
[Frontend] Remove frontend pooling multi task support. ( #37861 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 12:27:44 +00:00
766cb65d00
feat(multimodal): support externally processed mm_kwargs with cache injection ( #39502 )
...
Signed-off-by: Krish Hung <krishung5@gmail.com >
Signed-off-by: krishung5 <krish@nvidia.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-21 11:31:09 +00:00
Jhao-Ting Chen and GitHub
28c222157b
fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs ( #39391 )
...
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
2026-04-21 15:04:41 +04:00
wang.yuqi and GitHub
3975eb6de6
Revert "[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm" ( #40438 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-21 08:47:18 +00:00
Zeyu Zhang and GitHub
5a94a19824
[Bugfix] Normalize malformed dict prompts that carry token IDs in prompt ( #40339 )
...
Signed-off-by: Alchuang22-dev <2584829494@qq.com >
2026-04-21 07:44:36 +00:00
hangy-amd and GitHub
f95c11a848
[Feat] dflash support for ROCm ( #39703 )
...
Signed-off-by: Hang Yang <hangy@amd.com >
2026-04-21 14:58:20 +08:00
milesial and GitHub
257015d5e5
[MoE] Triton MoE Perf regression - restore low latency path ( #39016 )
2026-04-21 02:37:11 -04:00
Shanshan Shen and GitHub
b47840019e
[MM][Misc] Support image+video mixed inputs (per prompt) for VLM examples ( #40335 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 03:43:25 +00:00
SeongJun Lee and GitHub
989cc12d88
[Fix] Add missing space in IP fallback warning ( #40359 )
...
Signed-off-by: lesj0610 <lesj0610@gmail.com >
2026-04-20 20:26:06 -07:00
Wentao Ye and GitHub
301024aa9c
[Deprecation] Deprecate cprofile and cprofile_context ( #39100 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 11:25:22 +08:00
Simon Mo and GitHub
8256833fe6
[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm ( #40331 )
...
Signed-off-by: simon-mo <simon@inferact.ai >
2026-04-21 10:49:32 +08:00
Shanshan Shen and GitHub
8097591286
[Doc] Update ViT CUDA graph doc for mixed (image+video) inputs ( #40355 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 02:31:09 +00:00
20d3743491
[Bugfix] Gemma4: fix multimodal embedder norm order to match HF reference ( #40411 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-04-21 02:28:26 +00:00
Chauncey and GitHub
18563f2072
[Misc] Reduce attention logging levels ( #40086 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-21 02:09:25 +00:00
0e884fe638
[Bugfix] Fix _CONFIG_REGISTRY types getting wrong config class when on-disk model_type differs ( #39554 )
...
Signed-off-by: Misa <misaAle@users.noreply.github.com >
Signed-off-by: Misael Casarez <misacasa@amazon.com >
Co-authored-by: Misael Casarez <misacasa@amazon.com >
2026-04-20 19:04:48 -07:00
fe5c115ee4
[vLLM IR] Add IR op testing and benchmarking infrastructure ( #40167 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Theresa Shan <Theresa.Shan@amd.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-21 00:23:03 +00:00
6867bcd076
[Bugfix] Replace code that disabled shared expert overlap ( #39222 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-20 19:36:16 -04:00
c075702eae
[Misc][UX] Suppress confusing num_gpu_blocks log lines ( #40402 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-20 22:32:46 +00:00
Rita Brugarolas and GitHub
21b086d0aa
[ROCm] Hotfix: guard MLA dual RMS norm fusion against older AITer versions ( #40386 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-04-20 16:20:05 -05:00
Sage Moore and GitHub
3173441b0f
[EPLB] Consolidate is_unchanged/is_received_locally into TransferMetadata ( #37341 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-20 21:12:42 +00:00
Cao Qian and GitHub
8b1f3bebca
[LMCache MP Connector] Add num_lmcache_extra_cached_token in KVTransferParams ( #39843 )
...
Signed-off-by: aeon-x <talexcao@gmail.com >
2026-04-20 20:42:49 +00:00
2390caf157
Enable building MoRI with AMD AINIC stack ( #38371 )
...
Signed-off-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com >
Signed-off-by: Theresa Shan <theresa.shan@amd.com >
Co-authored-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-20 11:17:59 -07:00
Frederik Gossen and GitHub
87805fa11e
[Core] Cache InductorPass.hash_source with functools.cache ( #39328 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 14:06:15 -04:00
Nicolò Lucchesi and GitHub
304d5ba1a0
[Bugfix][CI] Fix tests/distributed/test_torchrun_example_moe.py ( #40349 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-20 11:05:44 -07:00
Tyler Michael Smith and GitHub
81d954f454
[WideEP] Remove naive all2all. Use allgather_reducescatter instead ( #33728 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-04-20 17:53:55 +00:00
Frederik Gossen and GitHub
47fcb8ca68
[Core] Pass donate_graph_module=True to standalone_compile ( #39733 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 17:40:52 +00:00
bai and GitHub
191e3fdaa1
Update flashinfer to 0.6.8 ( #39959 )
...
Signed-off-by: bai <v@gor.io >
2026-04-20 10:37:23 -07:00
Frederik Gossen and GitHub
b9cf629bd0
[Core] Label torch trace logging overhead with dynamo_timed ( #39329 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 17:31:03 +00:00
3461c8b027
[EPLB] Refactor Async EPLB synchronization logic ( #37601 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-04-20 17:05:41 +00:00
726efe177b
[MoE Refactor] Move the shared/fused expert output sum into MoERunnerBase ( #35949 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-20 12:28:46 -04:00
Yan Ma and GitHub
595562651a
[XPU] fix MoE triton backend in online fp8 quantization ( #40109 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-20 11:31:39 -04:00
Hashem Hashemi and GitHub
3a30eaa1d7
Properly enable wvSplitK fp8 path for RDNA ( #37712 )
...
Signed-off-by: Hashem Hashemi <hashem.hashemi@amd.com >
2026-04-20 10:09:24 -05:00
Rita Brugarolas and GitHub
fb5635d3f9
[ROCm] Add MLA dual RMS norm fusion (Q, KV) pass for DeepSeek/Kimi-K2 ( #39242 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-04-20 14:56:27 +00:00
Wentao Ye and GitHub
b42e878ec0
[Bug] Fix dcp error message ( #40053 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-20 10:52:32 -04:00
7243e02aa1
[ROCm][Feature] Enable AITER MLA attention backend to work with Eagle3 speculative decoding on ROCm ( #39616 )
...
Signed-off-by: larryli2-amd <larryli2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-20 09:44:43 -05:00
Sage Moore and GitHub
def8f52200
[CI][EPLB] Add Async EPLB end-to-end integration test to CI ( #40168 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-20 10:22:54 -04:00
Vasiliy Kuznetsov and GitHub
38fa87caca
mxfp8 online quant move to new frontend ( #40152 )
...
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com >
2026-04-20 06:26:12 -07:00
a023edfa5b
[bugfix] Use only onlines CPUs in lscpu ( #40161 )
...
Signed-off-by: kse <kevin.sejourne@cloud-temple.com >
Co-authored-by: kse <kevin.sejourne@cloud-temple.com >
2026-04-20 13:19:57 +00:00
b82fc1364d
[Anthropic][Frontend] Added chat_template_kwargs to /v1/messages ( #40125 )
...
Signed-off-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
2026-04-20 06:10:45 -07:00
Yan Ma and GitHub
e06de7f005
[XPU] enable triton attention test on XPU by removing cuda device binding ( #39627 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-20 20:57:11 +08:00
zhanqiuhu and GitHub
cc3993b05d
nixl refactor [2/N]: unify TpKVTopology + HeteroTPTransferConfig into TransferTopology ( #39529 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-20 12:39:08 +02:00
Ilya Markov and GitHub
50dd4cb427
[EPLB] Add nixl-based eplb communicator ( #36276 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
2026-04-20 10:24:23 +00:00
f774ba028a
[kv_offload+HMA][4/N]: Support sliding window lookup ( #36645 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-20 12:53:51 +03:00
Fadi Arafeh and GitHub
2aab9acf48
[CPU][BugFix] Fix inter-node pipeline parallel ( #40150 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-04-20 17:21:12 +08:00
nemanjaudovic and GitHub
58631d7c3f
[Bugfix] Fix scaled_mm output narrowing for 3D input tensors ( #38093 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-04-20 16:58:39 +08:00
Andreas Karatzas and GitHub
a943839e9a
[ROCm][CI] Introducing new MI300 nodes ( #39531 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-20 16:09:58 +08:00
milesial and GitHub
6d8b80802b
[Docs] Fix thinking_token_budget docs ( #40316 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
2026-04-20 08:09:44 +00:00
wuyingjun and GitHub
77fd2c8631
[Bugfix] Forward mm_processor_kwargs in offline generate APIs ( #40251 )
...
Signed-off-by: wuyingjun <wuyingjun_yewu@cmss.chinamobile.com >
2026-04-20 00:56:56 -07:00
San-Nguyen and GitHub
e729cc823d
[Fix] Add Spacing when Requesting Output Token > max_model_len ( #40324 )
...
Signed-off-by: San-Nguyen <san.nguyen@ibm.com >
2026-04-20 00:25:06 -07:00
velonica0 and GitHub
ec7aafc02a
[CPU][RISC-V] Support multiple RVV VLEN targets via compile-time dispatch ( #39478 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-04-20 14:36:59 +08:00
Julien Denize and GitHub
6097afb9bd
[BUGFIX] Fix Pixtral consolidated format vision weight loading ( #39916 )
...
Signed-off-by: Julien Denize <julien.denize@mistral.ai >
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-19 22:25:03 -07:00
4f4713f96e
[XPU] [torch.compile] Skipping CUDA graph memory estimation to avoid startup errors. ( #39977 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-20 13:04:39 +08:00
Tao He and GitHub
8936118134
[Qwen][Bugfix] Fixes sigmoid activation in torch impl of RMSNormGated. ( #40245 )
...
Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com >
2026-04-20 04:28:19 +00:00
Yuan Tang and GitHub
67ed01c353
fix: Do not make function calls when request has no tools for /v1/responses ( #40314 )
...
Signed-off-by: Yuan Tang <terrytangyuan@gmail.com >
2026-04-20 04:17:30 +00:00
6e10cb54f6
[Bugfix][Responses API] Fix streaming tool calls on /v1/responses ( #39892 )
...
Signed-off-by: Hoang Nguyen <118159510+hnt2601@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-20 11:24:52 +08:00
fcb31c1ac3
[Bugfix] Properly initialize PerTensorScaleParameter for fused-on-disk checkpoints ( #39765 )
...
Signed-off-by: Hemmi Shinichi <shemmi@preferred.jp >
Signed-off-by: Shinichi Hemmi <50256998+Alnusjaponica@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-20 02:53:04 +00:00
Lxx and GitHub
d886c26d4d
[Doc] Fix typos in token_embed pooling documentation ( #40266 )
...
Signed-off-by: YifanLi3 <lyfqlx3@gmail.com >
2026-04-19 19:27:32 -07:00
898beca5a8
[BugFix][XPU] fix lora ops bgmv_expand size not match ( #39989 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-20 08:24:50 +08:00
Kevin H. Luu and GitHub
629d45eacb
[ci] Make ecr authenticate non blocking ( #40305 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-04-19 15:37:53 -07:00
Andrew Barnes and GitHub
f150107efd
[ROCm] Fix cu_seqlens_q off-by-one in AITER FA speculative decode path ( #39120 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-19 18:34:33 +00:00
danisereb and GitHub
d1135a5087
Fix MoE backend selection for LoRA (unquantized MoE) ( #40273 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-04-19 17:18:40 +00:00
982beae809
Optimize nemotron VL image/video preprocessing ( #40283 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: milesial <milesial@users.noreply.github.com >
2026-04-19 15:06:20 +00:00
TJian and GitHub
45232a454e
[FEAT] [Perf] [Gemma4] Fused Gemma4 Routing Function Triton ( #39083 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-04-19 09:57:39 +00:00
Flora Feng and GitHub
03ce1c6ed9
[Bugfix] Kimi-K2 tool parser streaming - fix token leakage, argument truncation, and content dropping ( #38579 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-19 01:30:27 -07:00
omerpaz95 and GitHub
4353c9cb4a
[KV Offload] Pass request context ( #39185 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
2026-04-19 08:54:59 +03:00
4b7f5ea1a0
[KV Connector] Allow metrics of multiple connectors of same types in multi connector. ( #40010 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-19 07:49:10 +03:00
38907e4391
[Frontend] Preserve structured output special tokens in offline LLM.chat ( #39352 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-18 19:46:07 -04:00
d0359f3e04
[Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100 ( #40191 )
...
Signed-off-by: ultranationalism <www913363043@gmail.com >
Signed-off-by: mgoin <mike.goin12@gmail.com >
Co-authored-by: mgoin <mike.goin12@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-04-18 13:58:46 -07:00
Dan Alistarh and GitHub
ed0622e3a8
[Attention] TurboQuant: remove redundant random signs, add prior art attribution ( #40194 )
...
Signed-off-by: Dan Alistarh <d.alistarh@gmail.com >
2026-04-18 14:31:59 -04:00
Yusuf Mohammad and GitHub
b5f6c5f834
Added general ND x ND matmul and unit test for it ( #39909 )
...
Signed-off-by: Yusuf <yusufmohammad@live.com >
2026-04-18 10:05:21 -04:00
Jee Jee Li and GitHub
bfde49e287
[DOC] Add fuse_minimax_qk_norm ( #39782 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-18 00:41:37 -07:00
153ba7f0f3
[Refactor] Drop direct dependency on librosa ( #39079 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-18 06:55:38 +00:00
Chinmay-Kulkarni-AMD and GitHub
87518c3027
[ZenCPU] AMD Zen CPU Backend with supported dtypes via zentorch weekly ( #39967 )
...
Signed-off-by: Chinmay Kulkarni <Chinmay.Kulkarni@amd.com >
2026-04-18 06:22:37 +00:00
Rishapveer Singh and GitHub
aeee7ef939
[Bugfix] Fix k_proj's bias for GLM-ASR ( #40160 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-04-17 22:34:33 -07:00
z1ying and GitHub
cda19ecf4d
[Doc] Fix outdated source reference comment in anthropic/serving.py ( #40189 )
...
Signed-off-by: z1ying <tzzying@outlook.com >
2026-04-17 22:31:13 -07:00
80b18230e0
[Frontend] Add multimodal support to /inference/v1/generate endpoint ( #38405 )
...
Signed-off-by: Nithin Chalapathi <nithin.ch10@gmail.com >
Signed-off-by: Nithin Chalapathi <nithinc@berkeley.edu >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-17 20:31:56 -07:00
z1ying and GitHub
d0697cc7b6
[Doc] Add Realtime Transcription section to supported_models.md ( #39845 )
...
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-18 03:26:14 +00:00
b0755523dc
[Core] Reduce mm scheduler, get_num_embed overhead ( #40143 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-18 11:25:49 +08:00
993859ceb0
[XPU] fix all_reduce all-zero accuracy issue under torch.compile ( #39844 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-18 02:33:07 +00:00
Michael Goin and GitHub
48a65ccb02
[CI] Speed up test_fused_marlin_moe ( #40178 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-17 19:26:21 -07:00
55842a8d69
[XPU]fake impl for xpu fp8_gemm ( #39984 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-18 08:53:56 +08:00
Michael Goin and GitHub
1f45e83756
Remove outdated tests test_mixtral_moe and test_duplicated_ignored_sequence_group ( #40175 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-17 16:49:43 -07:00
Michael Goin and GitHub
a8bffaa133
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 ( #37463 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-17 16:42:32 -07:00
5cdddddd4a
[Kernel] [Helion] Force disable HOP path due to performance regression ( #40171 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-04-17 17:36:49 -04:00
aditi-amd and GitHub
6ef1efd51f
[ROCm] Fix TurboQuant on ROCm: backend routing, flash-attn compat, int64 overflow ( #39953 )
...
Signed-off-by: aditi <aditi.rana@amd.com >
2026-04-17 13:08:37 -07:00
Ryan Rock and GitHub
58da4ee047
[AMD][CI] Update DeepEP branch ( #38396 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
2026-04-17 14:30:20 -05:00
Andreas Karatzas and GitHub
1ae11e2bfc
[ROCm][CI] Build fastsafetensors from source so it links against libamdhip64 ( #39978 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-17 14:30:08 -05:00
251c18d1f8
skip fp8e4b15 on xpu ( #39957 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-17 16:55:08 +00:00
512765d52d
[Misc][UX] Map mimo reasoning and tooling parsers ( #40089 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-17 16:49:21 +00:00
640cc9dd7d
feat: Add LoRA support for Gemma4ForConditionalGeneration ( #39291 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-17 09:39:19 -07:00
ceade1952c
[BugFix] Support custom tool parsers when tool_choice is required and named function ( #39870 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-04-17 16:38:10 +00:00
747256bb5d
[Bugfix][Core] Fix stuck chunked pipeline parallelism with async scheduling ( #38726 )
...
Signed-off-by: Jing Wang <jingwang96@qq.com >
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com >
2026-04-17 16:02:50 +00:00
Michael Goin and GitHub
1174723eba
Fix TURBOQUANT backend selection in cuda.py ( #40060 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-17 07:31:41 -07:00
sychen52 and GitHub
6b2b7bd0eb
Add nvfp4 support to reshape_and_cache_flash ( #37332 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-04-17 07:28:00 -07:00
Ben Browning and GitHub
70770268c3
Add @bbrowning to CODEOWNERS ( #40141 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-04-17 09:51:48 -04:00
Chauncey and GitHub
7a51b3e415
[Bugfix] Fix empty delta detection in Qwen3XMLToolParser streaming ( #40090 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-17 13:34:55 +00:00
Li, Jiang and GitHub
d02421a7db
[CPU] Refactor CPU affinity and memory management ( #39781 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-17 21:01:08 +08:00
Lukas Geiger and GitHub
b1dc87a098
[Models][Gemma4] Prevent GPU/CPU sync in embed_input_ids ( #39234 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-04-17 12:37:21 +00:00
Or Ozeri and GitHub
79a5b63253
[kv_offload]: Fix num CPU blocks for UniformTypeKVCacheSpecs ( #39617 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-17 15:13:55 +03:00
Maral and GitHub
c0c98b8b9a
[Bugfix] Add Marlin kernel in block scaled mm kernel selection. ( #40105 )
...
Signed-off-by: maral <maralbahari.98@gmail.com >
2026-04-17 10:20:32 +00:00
wang.yuqi and GitHub
8d2cff8140
[Examples] Resettle Observability examples. ( #40123 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-17 03:13:31 -07:00
Cyrus Leung and GitHub
4f436782af
[Misc] Improve new PR bot trigger condition ( #40114 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-17 16:56:22 +08:00
z1ying and GitHub
bf45e6d0a5
[Doc] Add Gemma 4 to supported models list ( #39607 )
...
Signed-off-by: z1ying <tzzying@outlook.com >
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-17 13:42:52 +08:00
978a4462bb
[CI Failure] Fix Plugin Tests (2 GPUs) Failure ( #40083 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Michele Gazzetti <michele.gazzetti1@ibm.com >
2026-04-17 04:17:39 +00:00
Michael Goin and GitHub
1948d0c467
[UX] Defer some imports on CLI paths to save ~2s ( #40056 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-16 19:48:37 -07:00
Shinichi Hemmi and GitHub
4c47710bf7
[CI/Build] Apply ruff formatter to pass pre-commit ( #40078 )
...
Signed-off-by: Hemmi Shinichi <shemmi@preferred.jp >
2026-04-17 08:54:32 +08:00
Giancarlo Delfin and GitHub
bf9a5ddb24
[MLA] Optimize mla indexer prepare uniform decode for MTP > 1 ( #39458 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-16 16:27:51 -07:00
bnellnm and GitHub
79e799ebbd
[Bugfix] Temporarily disable B200 fp4 MoE layer tests ( #40057 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-16 19:26:55 -04:00
Netanel Haber and GitHub
c4e601c73c
Bugfix: Parakeet: .conv.pointwise/depthwise_conv1/2.bias weigths can exist even if convolution_bias=False ( #40007 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-16 23:22:05 +00:00
BadrBasowid and GitHub
29057d3bee
[Compilation] Add Unit Tests for VllmFusionPatternMatcherPass ( #39692 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-04-16 22:57:16 +00:00
Matthew Bonanni and GitHub
219bb5b8c0
[Misc] Update committers.md ( #40058 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-16 13:48:41 -07:00
Asaf Gardin and GitHub
ad2b1277f9
[Quantization] Consolidate experts_int8 with fp8 online quantization ( #38463 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-04-16 13:12:20 -07:00
roikoren755 and GitHub
b897f00c9c
Gate SSU dispatch setup ( #40039 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-16 13:06:01 -07:00
adf9bb3c57
[CI] Add weight transfer tests to CI ( #39821 )
...
Signed-off-by: SumanthRH <sumanthrh99@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-16 15:51:45 -04:00
Flora Feng and GitHub
b16fda62b7
[Misc] Add @sfeng33 to CODEOWNERS ( #40048 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-16 12:25:29 -07:00
Yufeng He and GitHub
de111f3246
[Bugfix] Fix bench_serve UTF-8 decode crash on split multi-byte chars ( #38732 )
2026-04-16 12:01:25 -07:00
Jared Wen and GitHub
afabb5f45a
[bugfix] Normalize tool message content from array to string format ( #39899 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-04-16 11:54:39 -07:00
Roger Wang and GitHub
3abb7560c0
[Bugfix] Fix audioflamingo test ( #40052 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-04-16 11:53:58 -07:00
Isotr0py and GitHub
617d1c2ff1
[Misc] Move pyav and soundfile to common requirements ( #39997 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-16 08:52:37 -07:00
Nikita Shapovalov and GitHub
692db29cd4
[Bugfix] Fix Ray compiled-DAG SHM channel stalls by detaching zero-copy np.ndarray logprobs buffers ( #35736 )
...
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
2026-04-16 23:49:29 +08:00
Isotr0py and GitHub
82531edbfb
[Refactor] Remove resampy dependency ( #39524 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-16 08:48:17 -07:00
Nicolò Lucchesi and GitHub
3daca38e22
[Misc] toy_proxy_server handle min_tokens ( #39706 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-16 15:08:22 +00:00
daiyu1111 and GitHub
a302a8fd1b
[Bugfix] Fix LLM priority normalization for single-string prompts ( #40011 )
...
Signed-off-by: daiyu1111 <2356690121@qq.com >
2026-04-16 07:56:06 -07:00
4e8c3f1c19
[Frontend][last/5] Improve pooling entrypoints | clean up. ( #39675 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-16 07:53:23 -07:00
Vasiliy Kuznetsov and GitHub
5e5afafa21
[Doc] add docs for online quant frontend ( #39736 )
...
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com >
2026-04-16 07:52:58 -07:00
Li, Jiang and GitHub
324a3d2bd8
[CI/Build] Improve stability of CPU tests ( #39966 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-16 21:50:36 +08:00
4269b79409
[Model] Use mm_features to compute mrope positions for PaddleOCR-VL ( #39888 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
2026-04-16 06:14:00 -07:00
edc3648966
[Kernel][Helion] Fix inductor fusion of Helion HOP ( #39944 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-16 04:41:26 -07:00
Nicolò Lucchesi and GitHub
9965f501a8
[Nixl] Bump Nixl version to 0.10.1 ( #39922 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-16 11:53:21 +01:00
lalit10 and GitHub
17d87168d2
[Model] Use mm_features for Keye-VL and Keye-1.5-VL M-RoPE ( #39869 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
2026-04-16 02:16:06 -07:00
Netanel Haber and GitHub
98700c6105
Fix #33773 : Replace unconditional pandas import with PlaceholderModule ( #39990 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-16 02:06:51 -07:00
Simon Mo and GitHub
10e49d2638
[Docs] Update PR template to remove release notes google docs ( #39982 )
...
Signed-off-by: Simon Mo <simon.mo@hey.com >
2026-04-16 00:22:03 -07:00
Tim Messerschmidt and GitHub
8d7c962833
[Bugfix] Accept **kwargs in MiniMaxM2Parser.__init__() ( #39861 )
...
Signed-off-by: Tim Messerschmidt <timmesserschmidt@gmail.com >
2026-04-16 15:18:32 +08:00
f4ddaf8cf7
[XPU] use spawn multiproc method on xpu ( #39671 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-16 14:42:07 +08:00
2cdf86044d
Add Jina Embeddings v5 model support ( fixes #38633 ) ( #39575 )
...
Signed-off-by: Abhijit <abroy@redhat.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-16 06:37:10 +00:00
realliujiaxu and GitHub
7845379230
[Bugfix] add support for 'num_attention_groups' in ModelArchConfigConvertorBase for Step3p5 ( #39796 )
...
Signed-off-by: realliujiaxu <realliujiaxu@163.com >
2026-04-16 05:48:00 +00:00
R3hankhan and GitHub
4b7ca37bd4
[CPU][IBM Z][Dockefile][Docs] Fix s390x builds for torch 2.11 and update docs for s390x ( #39910 )
...
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com >
2026-04-15 22:26:21 -07:00
445b7093fd
[perf][cpu] Accelerate BF16 GELU with LUT impl on Arm CPUs ( #37469 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-15 22:26:17 -07:00
18013df6ae
[Bugfix] Reject empty tools array with HTTP 400 ( #39780 )
...
Signed-off-by: Jigang Zhou <zjg0907008@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-16 12:08:04 +08:00
Julien Denize and GitHub
c0722f22de
[Mistral Grammar] Fix tool and reasoning parsing ( #39217 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-15 21:05:04 -07:00
Zhengxu Chen and GitHub
951dca8019
[compile] Invoke split FX graph by codegen. ( #38657 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-04-15 21:03:41 -07:00
vllmellm and GitHub
5f7fab881a
[ROCm][FEAT] Integrate aiter gemm w8a8 ptpc ( #33773 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-04-16 09:55:29 +08:00
Giancarlo Delfin and GitHub
343f65234b
[Model Runner V2][BugFix] fix num_sampled dtype for probabilistic rej… ( #39951 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-15 18:09:11 -07:00
Asaf Gardin and GitHub
19fa90ed0d
[Quantization] - Layerwise reloading of Attention/KV quantized models ( #38995 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-04-15 18:03:32 -07:00
03f8d3a548
Update to transformers v5 ( #30566 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Signed-off-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: khluu <khluu000@gmail.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: jiang1.li <jiang1.li@intel.com >
2026-04-15 16:29:15 -07:00
6dc9491406
[Model] Fix Gemma 4 token repetition by dynamic BOS injection for PT models ( #39842 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-04-15 16:13:07 -07:00
Collin McCarthy and GitHub
27c0ca50a0
Update registry for Nemotron-v3 VL Nano/Super ( #39747 )
...
Signed-off-by: Collin McCarthy <cmccarthy@nvidia.com >
2026-04-15 16:09:11 -07:00
Wentao Ye and GitHub
7c636432c6
[CI Bug] fix flaky test ( #39938 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-15 17:20:06 -04:00
Matthew Bonanni and GitHub
c77e596e2e
[FlashAttention] Don't overwrite flash_attn_interface.py when installing precompiled ( #39932 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-15 16:43:15 -04:00
Benjamin Chislett and GitHub
ac3dac545b
[Bugfix][Perf] Indexer upcast WK to BF16 for fusion ( #38928 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-15 20:39:32 +00:00
Wentao Ye and GitHub
39ac640490
[Bug] Fix batch invariant test issue, bs=1 with max_seq_num = 1 ( #39320 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-15 16:28:43 -04:00
zhanqiuhu and GitHub
0b790a2501
[Speculative Decoding] Add DFlash speculators config parsing ( #38300 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-15 16:22:15 -04:00
zhanqiuhu and GitHub
41488f2acd
[Bugfix][NIXL] Fix _logical_to_kernel_block_ids conversion for non-mamba models ( #39724 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-15 20:08:58 +00:00
102d51c9f3
[CI] Only build release Docker images when NIGHTLY=1 ( #39882 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-15 19:01:13 +00:00
55e1a8e103
[Mooncake] Fix mixed MLA+Eagle block-size validation ( #39596 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-15 11:36:47 -07:00
Monishver and GitHub
21e5a9f48e
Bug/test eagle dp v2 ( #39838 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
2026-04-15 17:48:12 +00:00
Mark McLoughlin and GitHub
8ad6ff0037
[Test] Fix @create_new_process_for_each_test("fork") in interactive shell pipeline ( #29130 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-15 12:22:20 -04:00
f2145efcb6
[BugFix] KeyError on scope["method"] for realtime api websocket in AuthenticationMiddleware ( #36934 )
...
Signed-off-by: daniebrill <50454544+daniebrill@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-15 16:15:01 +00:00
Roy Huang and GitHub
ed33310552
[KVConnector][LMCache] Propagate cache_salt through MP connector for per-user cache isolation ( #39837 )
...
Signed-off-by: royyhuang <royyhuang@gmail.com >
Signed-off-by: royyhuang <roy.y.huang@gmail.com >
2026-04-15 09:10:49 -07:00
3cc328a4be
[SpecDecode][Benchmark] Add SPEED-bench support to benchmarking CLI ( #36029 )
...
Signed-off-by: talora <talora@nvidia.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-15 12:00:07 -04:00
3beb57a238
[XPU] properly handle q_descale on XPU as quant query input not supported ( #39676 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-15 21:52:58 +08:00
8b5531933a
FIX: support language_model.backbone naming in NemotronH Nano VL quantization config ( #39901 )
...
Signed-off-by: <>
Co-authored-by: root <root@lyris0144.lyris.clusters.nvidia.com >
2026-04-15 13:49:48 +00:00
Chauncey and GitHub
db8d4a4a06
[BugFix][Graph] fix: handle empty sym_shape_indices in PiecewiseBackend. ( #39395 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-15 09:28:09 -04:00
zofia and GitHub
fc701c8058
[XPU][MXFP4] add mxfp4 quant op for XPU ( #39857 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-04-15 12:28:19 +00:00
Csrayz and GitHub
68be0f853e
[Metrics] Add request_id to FinishedRequestStats to enable correlation between metrics and requests ( #39710 )
...
Enables external `StatLogger` plugins to correlate per-request metrics
with request-level context. Also, this is a pre-requisite for Prometheus
exemplars in #30972 .
Signed-off-by: Csrayz <33659823+Csrayz@users.noreply.github.com >
2026-04-15 11:24:17 +00:00
Zhenzhong Xu and GitHub
60995c05b4
[Quantization][Autoround][CPU] Add W4A16 Support ( #38192 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
2026-04-15 18:38:31 +08:00
Yan Ma and GitHub
29e5d10205
fix online fp8 for MiniCPM models ( #39862 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-15 09:09:20 +00:00
Or Ozeri and GitHub
235e1f930a
[kv_offload+HMA][3/N]: Remove block_size from KVEvents ( #36644 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-15 11:53:19 +03:00
+86
431cea3eea
[Bugfix] Fix tool_calls Iterable consumed when debug logging is enabled ( #34844 )
...
Signed-off-by: Wojciech Wais <wojciech.wais@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Jaebok Lee <jaebok9541@naver.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Signed-off-by: yuwei <yuwei@dev.local >
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Signed-off-by: Li <chuali@amd.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: lkm2835 <lkm2835@gmail.com >
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Signed-off-by: vnadathur <glvikramn@gmail.com >
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com >
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com >
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: jackcfwang <jackcfwang@tencent.com >
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
Signed-off-by: Richard Zou <zou3519@gmail.com >
Signed-off-by: milesial <milesial@users.noreply.github.com >
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Signed-off-by: whx-sjtu <2952154980@qq.com >
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Signed-off-by: jatseng-ai <jatseng@amd.com >
Signed-off-by: jatseng-ai <janet.tseng@amd.com >
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Signed-off-by: xaguilar-amd <xaguilar@amd.com >
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com >
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Peter Nguyen <petern0408@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: manu <fortin.emmanuel@gmail.com >
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Rahul-Tuli <rtuli@redhat.com >
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: leeyongjun <jqueen.astro@gmail.com >
Signed-off-by: Ziying Tao <tzzying@outlook.com >
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Signed-off-by: ShubyM <shubymishra20@gmail.com >
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: r266-tech <r266.tech@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Signed-off-by: Animesh Jain <anijain@umich.edu >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
Signed-off-by: EricccYang <yangyang4991@gmail.com >
Signed-off-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Signed-off-by: sihao.li <sihao.li@intel.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com >
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Santino Ramos <santinor@inferact.ai >
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: George-ao <yuyiao772@gmail.com >
Signed-off-by: Yuyi Ao <yuyiao772@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Mukesh Baphna <mukesh@hippocraticai.com >
Signed-off-by: Pedram Razavi <pedram.razavi@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: Rishi Puri <riship@nvidia.com >
Co-authored-by: zzaebok <44357534+zzaebok@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: yuwei <yuwei@dev.local >
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Co-authored-by: Chuan (Richard) Li <chuali@amd.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Ganesh R <ganesh.r@amd.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Kyungmin Lee <30465912+lkm2835@users.noreply.github.com >
Co-authored-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Co-authored-by: vnadathur <glvikramn@gmail.com >
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Elham <elham.harirpoush@arm.com >
Co-authored-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Chaofan Wang <jackcfwang@tencent.com >
Co-authored-by: Chendi.Xue <chendi.xue@intel.com >
Co-authored-by: Injae Ryou <injaeryou@gmail.com >
Co-authored-by: Richard Zou <zou3519@users.noreply.github.com >
Co-authored-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: Elvir Crnčević <elvircrn@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
Co-authored-by: Hexiang Wang <56632993+whx-sjtu@users.noreply.github.com >
Co-authored-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: PatchyTIS <58251192+PatchouliTIS@users.noreply.github.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: jatseng-ai <janet.tseng@amd.com >
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: xaguilar-amd <xavier.aguilarfruto@amd.com >
Co-authored-by: Ravitez Dondeti <dondetir@users.noreply.github.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Peter Nguyen <petern0408@gmail.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: zhrrr <43847754+izhuhaoran@users.noreply.github.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Jesus Federico <14651+jefp@users.noreply.github.com >
Co-authored-by: Manu <efortin@users.noreply.github.com >
Co-authored-by: zhanqiuhu <49648934+ZhanqiuHu@users.noreply.github.com >
Co-authored-by: yzong-rh <yzong@redhat.com >
Co-authored-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Rahul-Tuli <rtuli@redhat.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: Lee Yongjun <35302114+elwhyjay@users.noreply.github.com >
Co-authored-by: z1ying <55220715+z1ying@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
Co-authored-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com >
Co-authored-by: ShubyM <shubymishra20@gmail.com >
Co-authored-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: r266-tech <r2668940489@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Le Yang <562593859@qq.com >
Co-authored-by: Animesh Jain <anijain@umich.edu >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Zhengxu Chen <zhxchen17@fb.com >
Co-authored-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Co-authored-by: maobaolong <baoloongmao@tencent.com >
Co-authored-by: sihao_li <165983188+1643661061leo@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: Tihomir Elek <tiho.elek@gmail.com >
Co-authored-by: Yi Liu <yi4.liu@intel.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Santino Ramos <51103228+santiramos27@users.noreply.github.com >
Co-authored-by: haosdent <haosdent@gmail.com >
Co-authored-by: JartX <sagformas@epdcenter.es >
Co-authored-by: Yuyi Ao <yuyiao772@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: mukesh-hai <mukesh@hippocraticai.com >
Co-authored-by: Pedram Razavi <pedram@sierra.ai >
2026-04-15 01:32:47 -07:00
zhanqiuhu and GitHub
799973af4e
[CI][NIXL] Fix PD CI breakage: pin nixl-cu{12,13} versions ( #39851 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-14 23:50:23 -07:00
bcc2306cef
[Bugfix] Respect VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY in prefetch offloader ( #37699 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-14 20:43:29 -07:00
wliao2 and GitHub
3abf858443
[Test] Refactor hard coded device string in test files under compile/quantization/models/model_executor folders ( #38901 )
...
Signed-off-by: Liao, Wei <wei.liao@intel.com >
2026-04-15 11:02:35 +08:00
f4b42df048
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity ( #38479 )
...
Signed-off-by: vibhavagarwal5 <vibhavagarwal5@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-14 19:57:13 -07:00
Giancarlo Delfin and GitHub
3bfe55a037
[Model Runner V2] Disable piecewise cudagraph mode fallback for eagle draft decodes ( #39773 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-14 17:47:57 -07:00
Andrey Talman and GitHub
b569620f72
[CI] Add PyTorch nightly build and test pipeline ( #37226 )
...
Signed-off-by: atalman <atalman@fb.com >
2026-04-14 17:13:24 -07:00
65b9808960
[Bugfix] Disable FlashInfer CUTLASS MoE on SM121 (DGX Spark) ( #39825 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-14 16:03:57 -07:00
Francesco Fusco and GitHub
507df79a29
[Hybrid] Simplify accepted token counting in spec decode for hybrid models ( #38372 )
2026-04-14 15:19:09 -07:00
1696c864b9
[Bugfix][Mooncake] Fix thread-local CUDA context for NVLink transfers in _send_blocks ( #39548 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-04-14 14:13:58 -07:00
Wentao Ye and GitHub
2ad1029233
[Bug] Fix batch invariance nvfp4 support ( #39820 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-14 17:08:17 -04:00
maobaolong and GitHub
b2f749dc97
fix(lmcache): correct store for cached requests while enable prefix cache ( #39719 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-04-14 20:51:27 +00:00
70ed01550c
[Reasoning][Frontend] Add model config to adjust_request in reasoning parser ( #37848 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com >
Signed-off-by: Aaron Pham <contact@aarnphm.xyz >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Aaron Pham <contact@aarnphm.xyz >
2026-04-14 20:29:06 +00:00
bnellnm and GitHub
19ec9a0a62
[MoE Refactor] Refactor ZeroExpertFusedMoE into new framework ( #35549 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-14 16:11:20 -04:00
1a9353bb02
[MoE] Move GPT OSS Triton kernel experts into fused_moe/experts/ ( #39007 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-14 19:27:39 +00:00
roikoren755 and GitHub
ecf5ff7ce3
[Mamba] Flashinfer selective_state_update ( #36162 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-14 15:10:58 -04:00
zhanqiuhu and GitHub
30679319e8
[CI][KVConnector][Metrics] Update multi KV connector edge case according to prefill stats changes ( #39808 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-14 18:59:15 +00:00
240f2636ca
[Kernel] Support TRTLLM GEN NVFP4 MoE for non-512-aligned hidden dims via weight padding ( #39510 )
...
Signed-off-by: root <root@lyris0017.lyris.clusters.nvidia.com >
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Co-authored-by: root <root@lyris0017.lyris.clusters.nvidia.com >
2026-04-14 11:49:56 -07:00
dc8df110bc
add warning when FP8 KV cache misses prefill query quantization ( #39752 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-lyris02.lyris.clusters.nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 14:43:05 -04:00
be0c855ebd
[KV Offload] Unified memory layout for offloading workers ( #37206 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-14 21:33:33 +03:00
Andrew Barnes and GitHub
e64b39ea71
[ROCm] Align AiterFlashAttentionImpl attn_type check with backend ( #39119 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-14 10:36:26 -07:00
Alessandro Sangiorgi and GitHub
2faad08362
[compile] Nest inductor cache under AOT compile dir ( #39718 )
...
Signed-off-by: Alessandro Sangiorgi <asangior@redhat.com >
2026-04-14 17:17:54 +00:00
Rohan Potdar and GitHub
23f3760217
[Bugfix][ROCm]: Allow gpt_oss_mxfp4 quantization method on rocm ( #39754 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-14 17:10:04 +00:00
Mark McLoughlin and GitHub
906a8c15d0
[Core][Metrics] Remove unused SchedulerStats.encoder_cache_usage ( #39693 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-14 12:53:57 -04:00
Micah Williamson and GitHub
4f4f8eaa78
[ROCm][CI] Fix condition for test_per_token_group_quant_fp8_packed ( #39730 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-14 16:14:31 +00:00
Netanel Haber and GitHub
b6890a120a
Bugfix: use_existing_torch.py: Glob recursive subdirs in requirements ( fixes #39024 ) ( #39793 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-14 23:11:46 +08:00
Lucas Kabela and GitHub
c08f3b2a62
Measure encoder compile time seperate from llm backbone ( #39240 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-14 10:52:49 -04:00
Hexiang Wang and GitHub
f02b3269e7
[PluggableLayer][3/N] Apply PluggableLayer to moe-related layers. ( #33556 )
...
Signed-off-by: whx-sjtu <2952154980@qq.com >
2026-04-14 09:55:00 -04:00
e1e318af01
[MoE Refactor] Remove MoE DP chunking ( #39107 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-14 09:48:05 -04:00
f7e62e3d66
[Bugfix] Fix mismatch between global and local attention heads in tensor-parallel mode for param2moe model ( #39707 )
...
Signed-off-by: bhargav-patel-29 <bhargav.patel@tihiitb.org >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 20:13:36 +08:00
18b1c77211
fix: handle ImportError in load_audio ( #39473 )
...
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-04-14 19:09:06 +08:00
Matthias Gehre and GitHub
1e4748c66a
[Bugfix] Fix vllm bench serve to count multimodal tokens in "total input tokens" ( #38654 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-14 11:00:40 +00:00
6f786f2c50
[Bugfix][Model] Fix Devstral Small 2 HF format weight loading ( #39293 )
...
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-14 10:11:18 +00:00
fxmarty-amd and GitHub
4eee77b877
[fix][MOE] Fix MOE experts intermediate_size dimension not being narrowed before weight loading ( #39688 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-04-14 09:35:28 +00:00
xiangdong and GitHub
a1993b96fd
[XPU][CI] Remove Arc in label-xpu ( #39776 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-14 02:27:38 -07:00
Julien Debache and GitHub
893b2affff
feat: add TxtSlicesDataset to allow sampling slices from txt file for benchmarking ( #30156 )
...
Signed-off-by: jdebache <jdebache@nvidia.com >
2026-04-14 09:20:03 +00:00
80118853f4
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3-VL video inference ( #38061 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-14 16:49:32 +08:00
c0ecaed950
[Frontend] Offload blocking preprocessing & postprocessing ops to thread pool for pooling entrypoints. ( #39763 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 08:29:25 +00:00
0008729abf
[Model] Use mm_features for Ernie-4.5 VL M-RoPE ( #39753 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-14 01:11:52 -07:00
d3af8c1831
[Core][Metrics][BugFix] Replace num_cached_tokens/num_external_computed_tokens with PrefillStats ( #37460 )
...
Related to `Counters can only be incremented by non-negative amounts`
error with the `vllm:prompt_tokens_by_source_total` metric.
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-04-14 09:00:45 +01:00
noobHappylife and GitHub
25b3242d8b
Fix Responses API streaming for multiple auto tool calls ( #39626 )
...
Signed-off-by: noobhappylife <aratar1991@hotmail.com >
2026-04-14 13:28:43 +08:00
b075604da1
[Bugfix] Fix Gemma4 tool parser converting bare null to string "null" ( #39679 )
...
Signed-off-by: KimuGenie <baby11686@naver.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-14 04:44:46 +00:00
Flora Feng and GitHub
db8a6d66bf
[Refactor][Parser] Migrate chat completion auto-tool/reasoning/plain streaming to parse_delta ( #39446 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-14 04:39:45 +00:00
Chauncey and GitHub
d2130a47bb
[Bugfix]: Fix MinimaxM2ToolParser missing tools parameter ( #39683 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-14 11:16:39 +08:00
c687bf226a
[LMCache][MP] optimize save when mla enabled ( #38810 )
...
Signed-off-by: idellzheng <idellzheng@tencent.com >
Co-authored-by: Yihua Cheng <yihua98@uchicago.edu >
2026-04-13 17:56:43 -07:00
Giancarlo Delfin and GitHub
ccf90ba784
[Model Runner V2] Add full cuda graph support for eagle prefill ( #37588 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-13 16:01:24 -07:00
Netanel Haber and GitHub
6adacfcb65
ParakeetExtractor performance and UX enhancements ( #39423 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-13 21:37:35 +00:00
Flora Feng and GitHub
14cb86c187
[Refactor][Parser] Simplify parse_delta ( #39728 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-13 21:02:13 +00:00
8213e8f880
Bug/test eagle dp v0 ( #38938 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-13 20:50:08 +00:00
Pedram Razavi and GitHub
3693f922ff
[Bugfix][Pooling] Fix silent weight corruption with buffer-reusing iterators ( #39650 )
...
Signed-off-by: Pedram Razavi <pedram.razavi@gmail.com >
2026-04-13 19:37:27 +00:00
5c18b961d6
[Core][Metrics] expose waiting request breakdown via labeled metric (capacity/deferred) ( #38435 )
...
Signed-off-by: Mukesh Baphna <mukesh@hippocraticai.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-04-13 15:30:55 -04:00
f72b20976c
[Bugfix] Reject non-nvfp4 dtypes when using the flashinfer_nvlink_one_sided all2all backend ( #39717 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-13 19:13:51 +00:00
610a3efcaf
[Doc] Fix Python-only build 404 fallback guidance ( #38052 )
...
Signed-off-by: George-ao <yuyiao772@gmail.com >
Signed-off-by: Yuyi Ao <yuyiao772@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-13 12:09:31 -07:00
JartX and GitHub
f414f90601
[Bugfix][Kernel][ROCm] Fix triton_w4a16 scales mismatch when BLOCK_K > group_size ( #39705 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-04-13 14:29:45 -04:00
Nicolò Lucchesi and GitHub
8625ec267b
[Misc] Multi-turn benchmark output performance json ( #39572 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-13 18:15:23 +00:00
995e9a209e
[Bugfix] Use is_integrated to detect UMA GPUs for memory reporting ( #35356 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-13 11:07:40 -07:00
Yongye Zhu and GitHub
739e5945dc
[Quantization] [Refactor] Create special "GptOssMxfp4MoeMethod" ( #39604 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-13 12:53:58 -04:00
Santino Ramos and GitHub
4d042ed85f
[Bugfix] Fix tensor shape mismatch in sparse attention with speculative decoding ( #39542 )
...
Signed-off-by: Santino Ramos <santinor@inferact.ai >
2026-04-13 08:57:38 -07:00
zhanqiuhu and GitHub
10d9872d3a
[CI][Metrics] Fix local_cache_hit assertion after prompt tokens metrics updates ( #39709 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-13 15:16:56 +00:00
ccd0d1d906
[Bug] Fix rocm sparse attn indexer issue ( #39225 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-13 07:53:45 -07:00
Yi Liu and GitHub
d8ddb31644
[Bugfix][CT] Fix KV cache scale handling ( #39418 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
2026-04-13 10:50:16 -04:00
Ekagra Ranjan and GitHub
1ce0318c68
[Bugfix] stream failure when model name not in audio endpoints ( #36679 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-04-13 14:20:07 +00:00
Tihomir Elek and GitHub
8d825b87d6
[Bug] Fix TypeError when hf_config.architectures is None during model loading ( #38849 )
...
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com >
2026-04-13 12:13:21 +01:00
zofia and GitHub
1b19bd7589
[MXFP8] [XPU] add a new compressed tensor schema and add a xpu mxfp8 gemm kernel ( #38707 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-04-13 16:59:20 +08:00
200a727e94
[Bugfix] Fix Responses API instructions leaking through previous_response_id ( #37727 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-13 08:46:33 +00:00
edbc1abd1c
feat: add max_tokens_per_doc in rerank request. ( #38827 )
...
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-13 01:24:09 -07:00
Flora Feng and GitHub
0e39202ca9
[Bugfix] Fix GLM tool parser streaming with MTP or stream interval ( #39253 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-13 05:10:30 +00:00
9dd5ee0117
[XPU]Enhance environment collection for Intel XPU and optimize layout ( #35698 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-13 12:51:46 +08:00
fa6ae31177
feat: rename logit_bias/logit_scale to logit_mean/logit_sigma for affine score calibration ( #39530 )
...
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-13 04:43:44 +00:00
maobaolong and GitHub
2a3c32ce67
fix(lmcache): correct store for cached requests and num_scheduled_tokens in lmcache_mp_connector.py ( #39655 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-04-13 03:29:19 +00:00
4beeb0689c
fused qknorm+rope kernel optimization for SM9.0 ( #37376 )
...
Signed-off-by: EricccYang <yangyang4991@gmail.com >
Signed-off-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-12 19:58:37 -07:00
cae984060f
[compile] Enable AOT compile with batch invariance mode. ( #39201 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-12 19:58:33 -07:00
Jee Jee Li and GitHub
715681c127
[LoRA] Support dual CUDA streams-Linear Layer ( #35721 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-13 10:57:07 +08:00
Kunshang Ji and GitHub
dc02271d76
[XPU] revert torch-xpu to 2.10 ( #39656 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-13 10:50:29 +08:00
Andreas Karatzas and GitHub
4e4ad41d11
[ROCm][CI] Removed stale tests and extended acceptance test ( #39651 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-13 10:40:26 +08:00
Yongye Zhu and GitHub
620e8924d9
[Bugfix] [Tests] Enforce out tensor device in kernel/moe/test_cutedsl_moe.py ( #39644 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-12 17:08:08 -07:00
Animesh Jain and GitHub
f00c5539d7
[compile] Bug fix for _decompose_size_nodes ( #38360 )
...
Signed-off-by: Animesh Jain <anijain@umich.edu >
2026-04-12 20:20:24 +00:00
Le Yang and GitHub
21fab0a3db
fix(moe): fix RoutedExpertsCapturer assertion failure with DP>1 and MK path ( #37879 )
2026-04-12 10:28:17 -04:00
Nicolò Lucchesi and GitHub
3244a2ebf2
[KVConnector][NIXL] Organize NIXL connector into its own directory ( #39354 )
...
The number of features supported by the connector has grown substantially
and the `nixl_connector.py` file has accumulated a lot of code. Creates a separate
directory and isolates connector/scheduler code in the hope of improving clarity
and maintainability.
Further refactor of components aimed at improving clarity and simplifying code
will follow soon.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-12 13:10:50 +00:00
Mark McLoughlin and GitHub
72ff142c37
[Core][Metrics] Remove vllm:prompt_tokens_recomputed metric ( #38709 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-12 12:22:01 +03:00
Nick Hill and GitHub
ee3c0c83db
[Pooling] Disable async scheduling by default for pooling models ( #39592 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-12 07:23:42 +00:00
cc07dad789
[HMA] [KVEvent] Enable GPU-side KV events for HMA ( #37688 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-04-12 10:01:02 +03:00
17e787a779
fix(kimi_k25): resolve media_placeholder_token_id from tokenizer ( #39344 )
...
Signed-off-by: r266-tech <r266.tech@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-11 21:10:24 -07:00
639402f5a2
Support FP8 KVCache on XPU ( #37731 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-12 03:53:32 +00:00
Andreas Karatzas and GitHub
0f7be0f2f7
[ROCm][CI/Build] Fix memory cleanup in MM test ( #39555 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-12 11:13:34 +08:00
Yan Ma and GitHub
394ff86965
[XPU][CT] support per-channel quantization in xpu fp8 linear method ( #38316 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-12 02:46:28 +00:00
EdalatiAli and GitHub
df1e30e74b
[Quant] add CompressedTensorsW8A8Mxfp8 for linear and MoE layers ( #38815 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
2026-04-11 17:21:36 -06:00
bd8bd52308
[Bugfix] Runtime driver check for cuMemcpyBatchAsync in swap_blocks_batch ( #38919 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-11 11:02:34 -06:00
Wei Zhao and GitHub
59b2f7b640
[Perf] Fuse Zero Initializer for FP8 DeepGemm Block Quant Kernel ( #39547 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-04-11 07:16:51 -07:00
ShubyM and GitHub
92feb9991d
[Gemma4][Bugfix]: Enable Gemma4ForCasualLM to load lora adapters correctly ( #38844 )
...
Signed-off-by: ShubyM <shubymishra20@gmail.com >
2026-04-11 09:06:49 +00:00
d4cb783c10
[Bugfix] Fix GDN FLA kernel crashes with NULL_BLOCK_ID=0 CUDA graph padding ( #39064 )
...
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-11 08:35:19 +00:00
Li, Jiang and GitHub
eb92ba740a
[CI/Build] Fix sentence-transformers version in CPU test ( #39557 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-11 07:04:25 +00:00
z1ying and GitHub
a3e750c0a5
[Misc] Update deprecation warning for --model flag ( #39518 )
...
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-10 23:25:20 -07:00
Lee Yongjun and GitHub
da72daced2
[Bugfix] add SupportsMultiModal to Exaone4_5_MTP ( #39526 )
...
Signed-off-by: leeyongjun <jqueen.astro@gmail.com >
2026-04-10 22:57:26 -07:00
8d0aabdde9
Fix the order of _free_encoder_inputs ( #38907 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-10 22:47:48 -07:00
0f3ce4c74b
[XPU] Fix spec-decode UTs under tests/v1/spec_decode ( #38491 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-11 01:31:00 +00:00
Benjamin Chislett and GitHub
af661a182d
Revert "Add nightly b200 test for spec decode eagle correctness ( #38577 )" ( #39512 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-10 20:07:32 -04:00
Michael Goin and GitHub
7f0b8f2020
[Docs] Use --torch-backend=auto for editable install docs ( #39511 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-10 15:27:02 -07:00
Michael Goin and GitHub
11e2375fe2
[Refactor] Move MXFP8 GEMM management into MxFp8LinearKernel ( #39205 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-10 14:02:03 -07:00
fc645f1acc
Add structure to requirements/ directory ( #39024 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-10 13:46:41 -07:00
Fynn Schmitt-Ulms and GitHub
2d80cf9d6e
Fix pre-commit labeled trigger system ( #39523 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-04-10 13:54:49 -06:00
e7cfd7c5b9
Add Gemma4 Eagle3 support ( #39450 )
...
Signed-off-by: Rahul-Tuli <rtuli@redhat.com >
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Rahul-Tuli <rtuli@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-10 12:35:35 -07:00
yzong-rh and GitHub
e816a8811f
[Bugfix] Fix FlashInfer crash with kv_cache_dtype_skip_layers ( #39002 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-04-10 18:50:47 +00:00
zhanqiuhu and GitHub
e281cb721c
[CI] Add MultiConnector (Nixl+Offloading) e2e edge case tests ( #39343 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-10 17:35:03 +00:00
Manu and GitHub
51cfc0e76c
perf(moe): add tuned fused_moe config for RTX PRO 6000 Blackwell Server Edition ( #39183 )
...
Signed-off-by: manu <fortin.emmanuel@gmail.com >
2026-04-10 11:32:42 -06:00
b87575d24b
feat: add logit_scale to PoolerConfig for affine score calibration ( #39435 )
...
Signed-off-by: Jesus Federico <jefp@amazon.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-10 17:21:14 +00:00
TJian and GitHub
42c6bb4b75
[ROCm] [AITER] Revert AITER version to v0.1.10.post3 ( #39509 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-04-10 16:25:52 +00:00
Jee Jee Li and GitHub
ecd1ea1363
[Kernel] Porting the TRTLLM minimax_allreduce_rms kernels ( #37045 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-11 00:20:20 +08:00
zhrrr and GitHub
8f121f7879
[Model Runner V2] support auto resolve cudagraph mode/sizes based on attn backend ( #32936 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-04-10 08:27:15 -07:00
wang.yuqi and GitHub
cb5f7501cb
[New Model]: jinaai/jina-reranker-v3 ( #38800 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-10 15:20:40 +00:00
Peter Nguyen and GitHub
8d0f908b98
[Model] Implement LoRA support for Qwen3ASRForConditionalGeneration ( #37247 )
...
Signed-off-by: Peter Nguyen <petern0408@gmail.com >
2026-04-10 18:34:31 +04:00
Nicolò Lucchesi and GitHub
c9dddc144b
[CI] Add Nixl+OffloadingConnector e2e integration tests ( #39200 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-10 21:40:40 +08:00
c1cc7344fb
[ROCm] Add RDNA 3.5/4 device IDs (gfx1150, gfx1151, gfx1201) ( #38455 )
...
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com >
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-10 11:35:07 +00:00
xaguilar-amd and GitHub
f976e3b98b
[Performance] Remove unnecessary zero-fill of MLA decode output tensor in Aiter backend ( #37539 )
...
Signed-off-by: xaguilar-amd <xaguilar@amd.com >
2026-04-10 11:27:35 +00:00
d468322dc1
[Kernel][Hardware][AMD] Add TritonW4A16LinearKernel for ROCm ( #37352 )
...
Signed-off-by: jatseng-ai <jatseng@amd.com >
Signed-off-by: jatseng-ai <janet.tseng@amd.com >
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-10 10:25:27 +00:00
967146e7bd
[model] support FireRedLID ( #39290 )
...
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
2026-04-10 08:43:58 +00:00
8e8a3becd1
[ZenCPU] Make PT Backport Patch Accessible to vLLM ( #38205 )
...
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-04-10 08:29:35 +00:00
1dfd64c1cc
[PluggableLayer][3/N] Apply PluggableLayer to llm_head and vocab embedding layer ( #33465 )
...
Signed-off-by: whx-sjtu <2952154980@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-10 16:12:59 +08:00
ad720aefe9
[Bugfix] Fix V1 dummy run writing NaN to KV cache null block ( #39444 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-04-10 10:09:46 +02:00
milesial and GitHub
270e8a4102
Nemotron Nano VL: Streamline pixel shuffle ( #37580 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
2026-04-10 07:31:19 +00:00
Richard Zou and GitHub
f44afef6d6
[compile] Allow strings in custom ops without regressing compilation times ( #38123 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-04-10 07:26:37 +00:00
447ce22212
[GGUF] Support non-standard quant types with prefix (e.g. UD-IQ1_S) ( #39471 )
...
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-10 07:22:53 +00:00
Chendi.Xue and GitHub
65e4e46f66
update CODEOWNERS file ( #39439 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
2026-04-10 15:05:31 +08:00
49d20346e4
[Perf] Reduce H2D pageable memory copies ( #38794 )
...
Signed-off-by: jackcfwang <jackcfwang@tencent.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-10 15:03:26 +08:00
Nick Hill and GitHub
ef076c1b73
[Core] Change max_model_len in EngineCoreReadyResponse to be non-None ( #39442 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-10 14:34:57 +08:00
Yan Ma and GitHub
ec68d53b2b
Add platform manual_seed_all API ( #38468 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-10 13:43:50 +08:00
Elham and GitHub
13e6b1b908
[BugFix][CPU] Add CPU profiler summary file output ( #38366 )
...
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com >
2026-04-10 13:41:15 +08:00
Isotr0py and GitHub
58c0a928c9
[Bugfix] Fix broken explicit unquantized kv cache dtype support ( #38922 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-09 22:27:53 -07:00
3dd60971de
[feat]: make DCP error msg clearer ( #28443 )
...
Signed-off-by: vnadathur <glvikramn@gmail.com >
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com >
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Co-authored-by: vnadathur <glvikramn@gmail.com >
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com >
2026-04-10 13:27:22 +08:00
Ronen Schaffer and GitHub
a5b17fba8f
[KV Offload] Implement shutdown() in OffloadingConnector and related classes ( #39182 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-04-10 08:06:22 +03:00
Cyrus Leung and GitHub
c48b2b83bd
[Mergify] Update model vendor auto-label rules ( #39312 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-10 04:25:37 +00:00
e7a1387e73
Add EXAONE-4.5 ( #39388 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-09 20:53:26 -07:00
f83de7196f
[BugFix] Fix OOB read in CUTLASS grouped GEMM with epilogue ( #38571 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-09 23:52:52 -04:00
Ganesh R and GitHub
445a2a4d1a
feat(cpu): add CPU support for draft model speculative decoding ( #32662 )
...
Signed-off-by: R <Ganesh.R@amd.com >
2026-04-10 11:49:52 +08:00
Kunshang Ji and GitHub
55d037e2e5
[CT][FP8][Marlin] refactor CompressedTensorsW8A16Fp8 to use kernel abstraction ( #38244 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-04-10 09:58:35 +08:00
Chauncey and GitHub
ecbfbb8d61
[Feature] Add auto-detection for reasoning_config when only reasoning_parser is set ( #38214 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-10 01:36:26 +00:00
Chuan (Richard) Li and GitHub
e0613702ad
[ROCm] Fix AITER ops fake impl and minor bugs ( #36092 )
...
Signed-off-by: Li <chuali@amd.com >
2026-04-09 17:56:17 -07:00
Ibrahim Arshad and GitHub
9853a3c159
fix(gdn): Align prefill warmup with real prefill path ( #39169 )
...
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
2026-04-10 00:49:50 +00:00
Artem Perevedentsev and GitHub
bb6047db13
[Model][Perf] Enable checkpoints prefetching for Lustre FS by default ( #39422 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-04-10 00:47:59 +00:00
467d3247c3
[LMCache] vLLM Block Allocation Event ( #38856 )
...
Signed-off-by: yuwei <yuwei@dev.local >
Co-authored-by: yuwei <yuwei@dev.local >
2026-04-09 17:30:29 -07:00
Cyrus Leung and GitHub
e5de19ff9a
[CI/Build[ Don't auto-rebase PRs with CI failures ( #39443 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-09 13:57:37 -07:00
edee96519a
[Spec Decode] fix returning size mismatch on extract hidden states proposer ( #38610 )
...
Signed-off-by: Jaebok Lee <jaebok9541@naver.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-09 20:39:39 +00:00
Rishi Puri and GitHub
adaabb8a55
Add nightly b200 test for spec decode eagle correctness ( #38577 )
...
Signed-off-by: Rishi Puri <riship@nvidia.com >
2026-04-09 20:09:09 +00:00
Ekagra Ranjan and GitHub
f7cad67412
[ASR] Fix spacing bw chunks in multi chunk audio transcription ( #39116 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-04-09 12:46:33 -07:00
Xinyu Chen and GitHub
a8134aef4e
[XPU] check is_xccl_available before oneccl warmup ( #39302 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
2026-04-09 12:42:17 -07:00
Michael Goin and GitHub
2800706f06
[Refactor] Move NVFP4 GEMM management into NvFp4LinearKernel ( #39129 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-09 15:05:36 -04:00
Cyrus Leung and GitHub
0d310ffbeb
[CI/Build] Update auto-rebase rule ( #39429 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-09 10:59:56 -07:00
Micah Williamson and GitHub
d5f75fdf50
[ROCm] Correctly guard fused_silu_mul_block_quant on ROCm ( #39387 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-09 17:59:03 +00:00
PikaPikachu and GitHub
827268e98d
[Quantization] Support Quark W8A8 INT8 MoE inference ( #36320 )
...
Signed-off-by: kangletian <Letian.Kang@amd.com >
2026-04-09 17:24:43 +00:00
Wentao Ye and GitHub
56e19d7ee2
[Model Runner V2] Fix flex attention kv blocks calculation issue ( #39353 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-09 13:07:43 -04:00
Andreas Karatzas and GitHub
9036d4c464
[ROCm][CI] Resolved nvidia package deps issue ( #39421 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-10 00:06:06 +08:00
a8c6ee9b78
[Performance Improvement] Update batched_count_greater_than to handle batch size 1 without recompile ( #38933 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-04-09 23:51:31 +08:00
Cyrus Leung and GitHub
3b1d9c3156
[CI/Build] Fix memory cleanup in MM test ( #39411 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-09 08:50:45 -07:00
Cyrus Leung and GitHub
54d244f28f
[UX] Improve error message for MM input too long ( #39409 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-09 13:20:19 +00:00
Richard Zou and GitHub
6c749399b7
[BugFix] fix tests/kernels/moe/test_moe_layer.py ( #39404 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-04-09 08:48:59 -04:00
91eea72330
[Tests] Add Qwen3-VL multimodal memory leak check ( #39268 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-09 04:54:46 -07:00
df2503e125
nemotron-nano-vl: Allow use_audio_in_video to be passed at vllm serve time ( #38538 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-04-09 11:44:39 +00:00
Nick Hill and GitHub
c8d98f81f6
[Core] Simplify API server handshake ( #39364 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-09 18:56:15 +08:00
Harry Mellor and GitHub
d87fb264df
[Docs] Bring README updates into docs README ( #39397 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-09 10:35:00 +00:00
wang.yuqi and GitHub
66c079ae83
[Frontend][4/n] Improve pooling entrypoints | pooling. ( #39153 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-09 10:09:45 +00:00
Shengqi Chen and GitHub
b6c9be509e
[CI] fix possible user permission issues in nightly index generation ( #39390 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-04-09 08:14:07 +00:00
ed733802f0
Fix NUMA binding on non-CDMM Grace-Blackwell systems ( #39361 )
...
Signed-off-by: Qidong Su <soodoshll@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-09 07:36:51 +00:00
Andrew Barnes and GitHub
8a34c5087a
[ROCm] Remove unnecessary fp8 roundtrip in gather cache NHD dequant ( #39122 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-09 15:12:22 +08:00
Wentao Ye and GitHub
ed2f282bc8
[Perf] Optimize redundant sync for pooling model, 3.7% Throughput Improvement ( #39113 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-08 23:12:23 -07:00
Zhewen Li and GitHub
9e78555743
[Docker] Add fastsafetensors to NVIDIA Dockerfile ( #38950 )
2026-04-08 22:21:37 -07:00
e80e633927
[XPU] Skip VLLM_BATCH_INVARIANT for XPU in EAGLE DP test ( #39164 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-09 12:45:16 +08:00
490f17d0c7
[Multimodal] Fix nested_tensors_equal: add length check for lists and tuple support ( #38388 )
...
Signed-off-by: khairulkabir1661 <khairulkabir1661@users.noreply.github.com >
Co-authored-by: khairulkabir1661 <khairulkabir1661@users.noreply.github.com >
2026-04-09 04:40:37 +00:00
Yongye Zhu and GitHub
2e98406048
[Refactor] Improve indexer decode path metadata preparation ( #38865 )
2026-04-08 20:49:15 -07:00
Chendi.Xue and GitHub
ef5a226819
[PD][HeteroArch]Fix accuracy issue with CPU_ATTN as Decoder and Flash_ATTN as prefiller ( #38935 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
2026-04-09 11:19:07 +08:00
Wentao Ye and GitHub
aec18492d0
[CI] Fix mypy for vllm/v1/ops ( #39219 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-09 11:06:34 +08:00
2a49284c8a
Fix Responses JSON schema alias serialization ( #38519 )
...
Signed-off-by: noobhappylife <aratar1991@hotmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-09 10:50:16 +08:00
Ilya Boytsov and GitHub
d37b378762
[Model] Update ColModernVBERT to support latest HF checkpoint ( #39307 )
...
Signed-off-by: Ilya Boytsov <ilyaboytsov1805@gmail.com >
2026-04-09 10:48:51 +08:00
92fbec391b
[Bug] Fix routing bias dtype for trtllm per-block fp8 moe ( #38989 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-08 19:42:43 -07:00
2f41d6c063
[Bugfix] Fix cpu-offload-gb assertion with non-default block sizes ( #36461 )
...
Signed-off-by: AjAnubolu <anuboluajay@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-08 19:42:16 -07:00
Dipika Sikka and GitHub
3aecdf08b4
[Gemma4] Support quantized MoE ( #39045 )
...
Signed-off-by: Dipika Sikka <dipikasikka1@gmail.com >
2026-04-08 21:57:53 -04:00
eb4205fee5
[UX] Integrate DeepGEMM into vLLM wheel via CMake ( #37980 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-08 18:56:32 -07:00
83aea2147f
[XPU][UT] update UTs in CI ( #39296 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <jikunshang95@gmail.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-09 09:38:16 +08:00
Maral and GitHub
2e9034c998
[W8A8 Block Linear Refactor][2/N] Remove W8A8Fp8BlockLinearOp and adopt Fp8 block linear kernel selections. ( #33892 )
...
Signed-off-by: maral <maralbahari.98@gmail.com >
Signed-off-by: Maral <maralbahari.98@gmail.com >
2026-04-09 08:50:39 +08:00
8332078cfd
[Bugfix] FlashInfer MXINT4 MoE crashes, missing do_finalize ( #39315 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-08 20:36:33 -04:00
Richard Zou and GitHub
ba4a78eb5d
[torch.compile] Allow usage of Opaque Objects in PyTorch 2.11 ( #39286 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-04-08 23:21:10 +00:00
Kai Song and GitHub
f3c7941ec8
[Bugfix]Fix EP precision for Qwen3.5, Qwen3-Next ( #39181 )
...
Signed-off-by: Song Kai <songkai05@baidu.com >
2026-04-09 01:47:48 +04:00
Wentao Ye and GitHub
3352bf8b03
[CI Bug] Fix pre-commit issue in main ( #39347 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-08 14:10:05 -07:00
7c94ae16c6
[BugFix] --max-model-len=-1 causes over-limit requests to hang and starve the entire service ( #39102 )
...
Signed-off-by: triangle14 <y1019026570@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-04-08 14:03:17 -07:00
ad05edfbca
tests/v1/e2e/spec_decode: assert async scheduling is used (#39206 )
...
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-08 20:30:03 +00:00
Wentao Ye and GitHub
2018137242
[Feature] Batch invariant nvfp4 linear support ( #39322 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-08 16:29:13 -04:00
a776a48b1c
[MoE] Move DEEP_GEMM into experts/ subdirectory ( #39005 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-08 19:23:08 +00:00
8477fe427d
[Tool] adjust_request to reasoning parser, and Gemma4 fixes ( #39027 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-04-08 19:04:04 +00:00
e24e0a43a4
[Attention] relax the head dim 512 and paged kv for sm90+FA4 ( #38835 )
...
Signed-off-by: Siyuan Fu <siyuanf@nvidia.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-04-08 18:23:18 +00:00
b55d830ec7
[Perf][Kernel] Persistent TopK scheduler: unified CUDAGraph-safe kernel with dynamic per-row dispatch - DeepSeek-V3.2 DSA decode ( #37421 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-04-08 13:35:57 -04:00
75e01a39a1
[Feature] NUMA binding support for GPU workers ( #38635 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Jason Li <jasonlizhengjian@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-08 09:55:24 -07:00
Or Ozeri and GitHub
512c5eb455
[kv_offload+HMA][5/N]: Track group block hashes and block IDs ( #37109 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-08 19:50:28 +03:00
Flora Feng and GitHub
13151a4df4
[Bugfix] Fix Gemma4 streaming tool call corruption for split boolean/number values ( #39114 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-08 16:46:27 +00:00
Gregory Shtrasberg and GitHub
56c976c1b5
[ROCm] Enable fused_silu_mul_block_quant on ROCm ( #38817 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
2026-04-08 11:23:32 -05:00
d74a306c4b
[Core] Use tuple_return in split_module for tuple-conformant subgraphs ( #38752 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
Co-authored-by: Boyuan Feng <boyuan@meta.com >
2026-04-08 09:09:58 -07:00
Gregory Shtrasberg and GitHub
0e9f0a516c
[ROCm][CI-Build] Cherry pick triton BUFFER_OPS fix and update AITER ( #38580 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
2026-04-08 10:38:03 -05:00
haosdent and GitHub
8904fc4d19
[Bugfix] Fix V1 logprobs empty strings for multi-byte UTF-8 tokens when logprobs > 0 ( #34875 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-08 15:30:00 +00:00
nemanjaudovic and GitHub
1a2c17634e
[Bugfix] Add missing ASRDataset import and CLI args in benchmarks/throughput.py ( #38114 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-04-08 13:53:53 +00:00
Matthew Bonanni and GitHub
308cec5864
[FlashAttention] Symlink FA4 instead of copying when using VLLM_FLASH_ATTN_SRC_DIR ( #38814 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-08 12:04:34 +00:00
wang.yuqi and GitHub
4e2ab1861d
[CI Failure] pin nomic-embed-text-v1 revision ( #39292 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-08 11:43:06 +00:00
140cbb1186
[Bugfix] Cuda Clean up scales Kvcache fp8/int8_per_token_head ( #39224 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-08 04:08:04 -07:00
6155bbd1dd
[Bugfix][Docs] Fix ReadTheDocs build crash from mocked torch decorator ( #39284 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-08 09:43:01 +00:00
rasmith and GitHub
78434b923c
[CI][AMD][BugFix][Kernel] Cast induction variable to int64 on MI350 for chunk_gated_delta_rule_fwd_kernel_h_blockdim64 to avoid illegal memory access ( #39087 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-08 16:57:18 +08:00
Michael Goin and GitHub
2488d1dca2
[Docs] Update README ( #39251 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-08 11:34:07 +08:00
yoke and GitHub
d734445fcd
[Bugfix][Frontend] Fix Gemma4 streaming HTML duplication after tool calls ( #38909 )
...
Signed-off-by: yoke233 <yoke2012@gmail.com >
2026-04-08 11:03:54 +08:00
Flora Feng and GitHub
927975ead8
[Parser] Migrate response api streaming to unified parser ( #38755 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Signed-off-by: Andrew Xia <axia@meta.com >
2026-04-08 10:09:00 +08:00
Flora Feng and GitHub
9ea7d670d8
[Bugfix] Fix Qwen3 tool parser for Responses API tools ( #38848 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-08 10:08:51 +08:00
7b80cd8ac3
[Docs] Add Phi-4-reasoning-vision to supported models + examples ( #39232 )
...
Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com >
Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com >
2026-04-08 02:02:26 +00:00
Andrey Talman and GitHub
2111997f96
[release 2.11] Update to torch 2.11 ( #34644 )
2026-04-07 18:55:48 -07:00
Flora Feng and GitHub
5af684c319
[CI] Add reasoning parser tests to CI ( #37025 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-08 00:57:36 +00:00
Md. Mekayel Anik and GitHub
d521dcdbcc
docs: clarify SMT and OMP acronyms in CpuPlatform ( #39085 )
2026-04-07 17:42:07 -07:00
Giancarlo Delfin and GitHub
5daf62271d
[Model Runner V2] Fuse probabilistic rejection sample kernels ( #38496 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-07 17:37:37 -07:00
ad3304425b
[XPU] add xpu backend implementation of mxfp8 quant ( #38682 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-08 08:30:35 +08:00
Lucas Wilkinson and GitHub
70406eb1dc
[Attention][V0 Deprecation] Deprecate accept output buffer ( #39125 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-04-07 17:14:58 -04:00
Yubo Wang and GitHub
08bfedc152
[Bugfix] Fix extract_hidden_states crash with quantized KV cache dtype ( #39160 )
...
Signed-off-by: Yubo Wang <yubowang2019@gmail.com >
2026-04-07 11:18:33 -07:00
Flora Feng and GitHub
0102bd2f4c
[Parser] Pass request.tools to tool parser ( #38860 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-08 01:36:21 +08:00
rasmith and GitHub
83d09d36b5
[CI][Bugfix][AMD][ Ensure weights created when using emulating OCP MXFP4 ( #36993 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-08 00:37:16 +08:00
92b9afeecd
[XPU] Quick fix for TritonMLA to remove cuda hardcode ( #39088 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-08 00:17:58 +08:00
Jinzhen Lin and GitHub
7310555482
[Bugfix] Fix marlin nvfp4 rescaling ( #37502 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
2026-04-07 08:57:17 -07:00
96b5004b71
[KVConnector] Support 3FS KVConnector ( #37636 )
...
Signed-off-by: wuchenxin <wuchenxin.wcx@alibaba-inc.com >
Signed-off-by: ibifrost <47308427+ibifrost@users.noreply.github.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-04-07 15:46:00 +00:00
kkyyxhll and GitHub
98e1a43af7
[Bugfix][Quantization] Fix PerTensorScale loading with tuple shard_id in MergedColumnParallelLinear ( #38517 )
...
Signed-off-by: loukang <loukang@xiaohongshu.com >
2026-04-07 11:16:26 -04:00
729eb59f60
[KVConnector]: prioritize external connector over internal registry ( #38301 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-07 15:03:11 +00:00
Ilya Boytsov and GitHub
6e1100889e
fix(test): recompute Jina ColBERT rotary inv_freq cleared by transformers v5 weight loader ( #39176 )
...
Signed-off-by: Ilya Boytsov <ilyaboytsov1805@gmail.com >
2026-04-07 22:40:55 +08:00
edcc37a8ce
Fix Mistral yarn warning in Transformers v5 ( #37292 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-04-07 13:23:33 +00:00
Harry Mellor and GitHub
79df4a794d
Automatically add links to API docs for matching strings in docs ( #37434 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-07 21:21:18 +08:00
Ronen Schaffer and GitHub
7c139ab23f
[KV Offload] Clean up ARC/LRU refactoring leftovers: group ARC tests and fix stale comment ( #38217 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-04-07 15:14:45 +03:00
Wei Zhao and GitHub
0be9516ea4
[Bug] Fix Trtllm Fp8 MoE Weight Shuffle Memory Fragamentation ( #39054 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-04-07 08:04:08 -04:00
Kyle Mylonakis and GitHub
7b9de7c892
[Bugfix] Correct mistake in chained comparison in static assert logic ( #38699 )
...
Signed-off-by: Kyle Mylonakis <kyle@protopia.ai >
2026-04-07 18:24:39 +08:00
Rohan Potdar and GitHub
dd9342e6bc
only patch runtime_env for torch >= 2.10 ( #38763 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-07 09:29:23 +00:00
8060bb0333
[vLLM IR] rework gemma_rms_norm ( #39014 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-04-07 01:37:00 -07:00
Rishapveer Singh and GitHub
da4c0e4db9
[Model] Use AutoWeightsLoader for FalconH1 ( #39092 )
...
Signed-off-by: Rishapveer Singh <215205492+rishaps@users.noreply.github.com >
2026-04-07 16:25:17 +08:00
Netanel Haber and GitHub
a9a0e0551f
nano-nemotron-vl: get_mm_max_tokens_per_item for audio, video, image == seq_len ( #38727 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-07 00:23:29 -07:00
Andrew Barnes and GitHub
5c35517a3e
[ROCm] Remove unused IS_FNUZ parameter from reshape_and_cache_shuffle_kernel ( #39123 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-07 07:17:59 +00:00
Andreas Karatzas and GitHub
a435e3108d
[ROCm][CI] Fix test repo-root assumptions ( #39053 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-07 13:36:21 +08:00
Andreas Karatzas and GitHub
2df2c85be4
[Kernels][MoE] Fix legacy_routing to use bitmatrix-based routing path ( #38504 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-07 10:57:09 +08:00
Nick Hill and GitHub
62095e82c1
[BugFix][MRV2] Fix cuda event reuse race ( #39115 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-07 00:21:09 +00:00
bnellnm and GitHub
b2b2c5239e
[MoE Refactor] Split up compressed_tensors_moe.py ( #38960 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-06 20:07:54 -04:00
00d7b497b3
[NVFP4] Support NVFP4 dense models from modelopt and compressed-tensors on AMD Instinct MI300, MI355X and Hopper through emulation ( #35733 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Signed-off-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-04-06 16:18:27 -06:00
Matthew Bonanni and GitHub
9c81f35b1a
[Attention][MLA] Re-enable FA4 as default MLA prefill backend ( #38819 )
2026-04-06 17:51:46 -04:00
Woosuk Kwon and GitHub
f186cfe75e
[MRV2] Fix hanging issue with DeepSeek V3.2 by setting skip_attn=False ( #39098 )
...
Signed-off-by: WoosukKwon <woosuk.kwon@berkeley.edu >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-04-06 12:55:13 -07:00
Netanel Haber and GitHub
dfa5062a8f
NemotronH default mamba_ssm_cache_dtype=float32; enable auto-hook for NemotronHNanoVLV2Config ( #39032 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-06 19:47:46 +00:00
e8ebbdde83
[Quantization] Add FlashInfer CuteDSL batched experts backend for NVFP4 MoE ( #38251 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-06 11:57:53 -07:00
namgyu-youn and GitHub
94fbb09894
[EASY] Drop duplicate KV-cache initialization ( #38799 )
...
Signed-off-by: namgyu-youn <namgyu.dev@gmail.com >
2026-04-06 18:05:39 +00:00
Wentao Ye and GitHub
419e73cdfa
[Bug] Fix mistral version dependency ( #39086 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-06 13:31:19 -04:00
f01482408c
[MoE Refactor][Test] FusedMoE layer test ( #24675 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-06 17:17:23 +00:00
zhanqiuhu and GitHub
bfdc0a3a99
[NIXL][Mamba][3/N] Heterogeneous TP: 3-read conv state transfer ( #37635 )
2026-04-06 19:07:02 +02:00
93bada494f
[MoE Refactor] Split of DefaultMoERunner class ( #35326 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-06 12:41:59 -04:00
Frederik Gossen and GitHub
608914de30
[Core] Re-enable Inductor pre-grad passes in standalone compile (torch>=2.12) ( #38944 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-06 09:37:13 -07:00
Wentao Ye and GitHub
4ae218c122
[Refactor] Remove unused dead code ( #38842 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-06 11:52:05 -04:00
Lukas Geiger and GitHub
f40d9879f2
[Models][GDN] Remove GPU/CPU syncs in GDNAttentionMetadata.build during speculative decoding ( #38047 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-04-06 15:39:37 +00:00
Lucas Wilkinson and GitHub
47e605092b
[Gemma4] Enable Fast Prefill Optimization ( #38879 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-04-06 11:19:39 -04:00
Walter Beller-Morales and GitHub
e69a265135
[Feat][Core] safely abort requests when FSM fails to advance ( #38663 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-04-06 08:00:16 -07:00
Julien Denize and GitHub
fef56c1855
[Mistral Grammar] Support Grammar Factory ( #38150 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-06 10:28:51 -04:00
c5e3454e5a
[Model] Add support for BharatGen's Param2MoE model ( #38000 )
...
Signed-off-by: bhargav-patel-29 <bhargav.patel@tihiitb.org >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-06 16:19:56 +08:00
f6983f01de
MiniMax-M2: add Eagle3 speculative decoding support ( #37512 )
...
Signed-off-by: liuchenbing <chenliumail@163.com >
Signed-off-by: liucb <liuchengbao_work@163.com >
Co-authored-by: liuchenbing <chenliumail@163.com >
2026-04-05 19:50:18 -07:00
Andreas Karatzas and GitHub
780ba37458
[ROCm][Quantization] Add asymmetric INT8 quantization support to TritonInt8ScaledMMLinearKernel ( #38501 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-06 09:42:10 +08:00
Micah Williamson and GitHub
9570654c6d
[ROCm][CI] Run Kernels Core Operation Test On MI325 and mitigate flakiness ( #38184 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-06 09:42:02 +08:00
Netanel Haber and GitHub
d56e952239
nano_nemotron_vl: fix tensor device mismatch exception when video profiling ( #39029 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-05 22:23:45 +00:00
Kevin H. Luu and GitHub
56de443db1
[ci] Switch some CI jobs to H200 MIG slices ( #38956 )
2026-04-05 13:26:11 -07:00
4dd49b06f8
[Bug] Fix Import paths for encoder_cudagraph modules ( #38997 )
...
Signed-off-by: greg pereira <grpereir@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-05 19:11:58 +00:00
f53fa26e05
[Bugfix] Fix invalid JSON in Gemma 4 streaming tool calls by stripping partial delimiters ( #38992 )
...
Signed-off-by: greg pereira <grpereir@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-05 17:11:18 +00:00
1af6f78ae5
[Perf] Change Trtllm fp8 MoE to use Shuffled Weights and BlockMajorK Layout ( #38993 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-05 10:54:31 -04:00
228023b3a5
[Bugfix][MoE] Fix 6-8% decode regression: prefer multi-stream shared expert overlap ( #38990 )
...
Signed-off-by: Martin Vit <martin@voipmonitor.org >
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-05 10:28:31 -04:00
Aaron Batilo and GitHub
9a528260ef
[Bugfix][Spec Decode] Fix extract_hidden_states for VLM models ( #38987 )
...
Signed-off-by: Aaron Batilo <abatilo@coreweave.com >
2026-04-05 02:41:54 -07:00
968ed02ace
[Quantization][Deprecation] Remove Petit NVFP4 ( #32694 )
...
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
2026-04-05 00:07:45 +00:00
Robert Shaw and GitHub
7d266abb22
Revert "[vLLM IR] gemma_rms_norm" ( #38998 )
2026-04-04 17:48:08 -04:00
Xiaoshuang Wang and GitHub
156405d243
[vLLM IR] gemma_rms_norm ( #38780 )
...
Signed-off-by: Icey <1790571317@qq.com >
2026-04-04 13:55:52 -04:00
99e5539a67
[Perf][GDN] Align TMA usage with upstream FLA ( #38981 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-05 00:38:02 +08:00
a88ce94bbb
[IR][RmsNorm] pass None if not has_weight ( #38961 )
...
Signed-off-by: Linkun Chen <github@lkchen.net >
Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-04-04 11:02:30 -04:00
Ziming Qi and GitHub
2a36d8fb72
[Bugfix][CPU] Fix macOS compatibility broken by #36487 ( #38970 )
...
Signed-off-by: Ziming (2imi9) <148090931+2imi9@users.noreply.github.com >
2026-04-04 14:05:58 +00:00
93726b2a1c
Refactor Arctic loading to use AutoWeightsLoader ( #38955 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
Co-authored-by: Lalit Laxminarayan Bangad <lalitbangad@meta.com >
2026-04-04 05:01:09 +00:00
Yongye Zhu and GitHub
8617f8676b
[Bugfix] Fix DSV32 weight loading ( #38870 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-03 19:57:52 -07:00
Andreas Karatzas and GitHub
06fd9ffcc4
[ROCm][CI] Fix ROCm Dockerfile conftest generation for older Docker parsers ( #38959 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-04 10:41:41 +08:00
Wentao Ye and GitHub
cab4064cd5
[Bug] Fix workspace manager _current_workspaces size ( #38853 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-04 01:29:45 +00:00
Wentao Ye and GitHub
062f1a2d70
[Bug] Fix compile error for swap_blocks_batch in CUDA 13 ( #38915 )
2026-04-03 16:56:38 -07:00
elenalil-aws and GitHub
81994e1d0e
[Bugfix][LoRA] Fix missing in_proj_z in Qwen3_5ForConditionalGenerati… ( #38927 )
...
Signed-off-by: elenalil-aws <elenalil@amazon.com >
2026-04-03 23:30:09 +00:00
Andreas Karatzas and GitHub
4b506ff90a
[ROCm][CI] Minor missing import patch ( #38951 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-03 23:01:20 +00:00
Andreas Karatzas and GitHub
5875bb2e9c
[ROCm][CI] Added back missing common deps ( #38937 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-03 15:58:57 -07:00
Kevin H. Luu and GitHub
f0d3ad9f3e
[ci] Remove soft fail for AMD image build job ( #38941 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-04-03 20:42:33 +00:00
Divin Honnappa and GitHub
121ea5a21f
Removed GPU state confirmation and cleanup steps. ( #38238 )
...
Signed-off-by: Divin Honnappa <divin.honnappa@amd.com >
2026-04-03 13:11:08 -07:00
Jeffrey Wang and GitHub
ab79863e6c
Remove MQ multi-node tests ( #38934 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
2026-04-03 20:00:08 +00:00
Nick Hill and GitHub
5f1de2b14b
[Model Runner V2] Add config validation for not-yet-supported features ( #38758 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-03 12:08:08 -07:00
yzong-rh and GitHub
a5a623d961
[Bugfix] Re-enable Renormalize routing for TRT-LLM MoE experts ( #38859 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-04-04 01:48:17 +08:00
Xiaoshuang Wang and GitHub
f8c3af2d85
[vLLM IR] add import_ir_kernels() to support OOT platforms ( #38807 )
...
Signed-off-by: Icey <1790571317@qq.com >
2026-04-03 17:25:19 +00:00
danisereb and GitHub
50cd5674b3
Fix invalid logprobs with MTP enabled and sync scheduling ( #38711 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-04-03 12:24:37 -04:00
Vasiliy Kuznetsov and GitHub
7b1a7423be
[Frontend] new online quantization frontend ( #38138 )
...
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com >
2026-04-03 11:58:39 -04:00
Nicolò Lucchesi and GitHub
97f92c6b47
[KVConnector] Skip register_kv_caches on profiling ( #38558 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-03 15:40:16 +00:00
46f02e00f2
[Bugfix] Fix AWQ models batch invariance issues ( #38670 )
...
Signed-off-by: yusuf <yusuf@deeplearningmachine.mynet >
Signed-off-by: <>
Co-authored-by: yusuf <yusuf@deeplearningmachine.mynet >
2026-04-03 14:54:15 +00:00
6b4872240f
[XPU] bump up xpu-kernel v0.1.5, transpose moe weights ( #38342 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Signed-off-by: Qiming Zhang <qiming1.zhang@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-03 14:10:02 +00:00
Necofish and GitHub
580090db6b
[Kernel] Add swapAB support for SM120 CUTLASS blockwise FP8 GEMM ( #38325 )
2026-04-03 15:49:59 +02:00
Artem Perevedentsev and GitHub
cb10b7e80b
[GDN] Eliminate GPU->CPU sync in prepare_chunk_indices during prefill ( #38361 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-04-03 13:38:02 +00:00
bf8b022e60
[Intel][Triton] Support round_int8 for Intel backend ( #38825 )
...
Signed-off-by: Mieszko Dziadowiec <mdziadowiec@habana.ai >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-03 20:47:35 +08:00
xiangdong and GitHub
40ee64c00e
[XPU][CI] Skip test_topp_only and test_topk_and_topp cases on Intel GPU in CI ( #38904 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-03 20:44:52 +08:00
1b117cb0ac
[ROCm] Fix aiter persistent mode mla with q/o nhead<16 for kimi-k2.5 tp8 ( #38615 )
...
Signed-off-by: wufann <36477220+wufann@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-03 03:54:00 -07:00
Anton Ivanov and GitHub
abebd9323d
[CPU] Replace OMP initialization ( #36487 )
...
Signed-off-by: Anton Ivanov <anton.ivanov@cambridgegreys.com >
2026-04-03 18:42:43 +08:00
Hyeonki Hong and GitHub
25f2b55319
[Frontend] feat: add streaming support for token generation endpoint ( #37171 )
...
Signed-off-by: Hyeonki Hong <hyeonki.hong@moreh.io >
2026-04-03 10:20:32 +00:00
xiangdong and GitHub
cb4ff07f8b
[XPU][CI] Skip test_topk_only cases on Intel GPU in CI ( #38899 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-03 09:50:41 +00:00
Gregory Shtrasberg and GitHub
a7d79fa133
[ROCm][CI/Build] Fix the pytest hook to properly print out the summary ( #38585 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
2026-04-03 17:24:26 +08:00
Netanel Haber and GitHub
fa9e68022d
Fix Nano Nemotron VL regressions ( #38655 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-03 15:22:06 +08:00
Isotr0py and GitHub
5506435419
[Misc] Clean up Gemma4 implementation ( #38872 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-03 05:47:02 +00:00
Yifan Qiao and GitHub
311c981647
[MRV2][KVConnector] Fix missing build_connector_worker_meta ( #38698 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-04-03 08:42:52 +03:00
Li, Jiang and GitHub
21d7ecc5b0
[CI/Build] Add audio deps in Dockerfile.cpu ( #38876 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-03 05:05:14 +00:00
Aaron Hao and GitHub
4729b90838
[Bug] Add e_score_correction_bias to SKIP_TENSORS ( #38746 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
2026-04-02 21:15:05 -07:00
shunting314 and GitHub
8b141ed8c3
full cudagraph for flex-attn ( #36298 )
...
Signed-off-by: shunting314 <shunting@meta.com >
2026-04-02 21:15:01 -07:00
2ad7c0335f
[Model] Add Phi4ForCausalLMV for microsoft/Phi-4-reasoning-vision-15B ( #38306 )
...
Signed-off-by: Varun Sundar Rabindranath <vsundarr@redhat.com >
Co-authored-by: Varun Sundar Rabindranath <vsundarr@redhat.com >
2026-04-02 21:14:57 -07:00
Bowen Bao and GitHub
201d2ea5bf
[CI][ROCm] Add Qwen3.5-35B-A3B-MXFP4 model eval into CI ( #38664 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-04-03 04:05:45 +00:00
103f0de565
[ROCm][Quantization][1/N] Refactor quark_moe w_mxfp4 w/ oracle ( #38774 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-03 03:29:57 +00:00
wliao2 and GitHub
32e0c0bfa2
refactor hard coded device string in test files under tests/v1 and tests/lora ( #37566 )
...
Signed-off-by: Liao, Wei <wei.liao@intel.com >
2026-04-03 11:21:47 +08:00
4a06e1246e
[Perf] Batch KV cache swap copies via cuMemcpyBatchAsync ( #38460 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-03 03:13:23 +00:00
Carl Y and GitHub
3bc2734dd0
[Kernel] Fuse FP8 output quantization into merge_attn_states ( #36518 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-04-03 01:47:04 +00:00
1f5ec2889c
[mla] Support fused FP8/NVFP4 output quantization in MLA attention ( #35792 ) ( #36205 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
Signed-off-by: Carl Y <4531192+carlyou@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-02 21:16:11 -04:00
ee3cf45739
[XPU] Initial support for GDN attention on Qwen3-next/Qwen3.5 ( #33657 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Co-authored-by: Chendi Xue <chendi.xue@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-03 08:59:11 +08:00
Matthew Bonanni and GitHub
05e68e1f81
[CI] Fix test_nixl_connector ( #38838 )
2026-04-02 17:52:13 -07:00
Vadim Gimpelson and GitHub
771913e4a0
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3.5 ( #38832 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-04-03 04:45:57 +04:00
71a9125c67
[New Model]: add support for telechat3 ( #38510 )
...
Signed-off-by: xiayongqiang <xiayq1@chinatelecom.cn >
Co-authored-by: xiayongqiang <xiayq1@chinatelecom.cn >
2026-04-03 08:26:22 +08:00
Nicolò Lucchesi and GitHub
66e86f1dbd
[Kernel] Mamba support different layout for Conv state ( #37416 )
2026-04-03 01:50:09 +02:00
Michael and GitHub
bb39382b2b
[Bugfix]: Fix Gemma4ToolParser.__init__() missing tools parameter ( #38847 )
...
Signed-off-by: Michael Hospedales <hospedales@me.com >
2026-04-02 14:35:19 -07:00
zhanqiuhu and GitHub
7b743ba953
[CI] Fix: pass string cache_dtype in test_register_kv_caches ( #38836 )
2026-04-02 19:42:09 +00:00
188defbd0b
[CI] Add flashinfer.py to attention test source deps ( #38792 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-04-02 19:24:29 +00:00
08ed2b9688
feat(models): implement Google Gemma 4 architecture support (MoE, Multimodal, Reasoning, Tool-Use) ( #38826 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Luciano Martins <lucianomartins@google.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
2026-04-02 11:13:28 -07:00
ecd5443dbc
Bump helion dependency from 0.3.2 to 0.3.3 ( #38062 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-02 10:59:33 -07:00
58262dec6e
[Bugfix] Fix test mocks after SM100 restriction in #38730 ( #38791 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-02 13:12:58 -04:00
Lucas Wilkinson and GitHub
cb3935a8fc
[FA4] Update flash-attention to latest upstream FA4 ( #38690 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-04-02 17:02:37 +00:00
Bowen Bao and GitHub
82a006beeb
[CI][ROCm] Add gpt-oss w4a8 in CI ( #38292 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-04-03 00:06:01 +08:00
wang.yuqi and GitHub
a9b4f07ba2
[Frontend] Re-enable running MaxSim on GPU ( #38620 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-03 00:03:13 +08:00
d9408ffba3
Triton MLA perf fixes ( #33529 )
...
Signed-off-by: Koushik Dutta <koushd@gmail.com >
Co-authored-by: root <root@ubuntu-nvidia.localdomain >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-02 09:40:01 -04:00
16a65e4173
[Bugfix] Enable batch-invariant Triton matmul on all Ampere GPUs (SM 8x) ( #38427 )
...
Signed-off-by: yusuf <yusufmohammad@live.com >
Signed-off-by: yusuf <yusuf@deeplearningmachine.mynet >
Signed-off-by: Yusuf Mohammad <79484377+YM2132@users.noreply.github.com >
Signed-off-by: <>
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: yusuf <yusuf@deeplearningmachine.mynet >
2026-04-02 09:29:58 -04:00
bsliu and GitHub
c0817e4d39
[Model] Add support for Cheers multimodal model ( #38788 )
...
Signed-off-by: bsliu <1187291748@qq.com >
Signed-off-by: 吴炳贤 <wubingxian24@mails.ucas.ac.cn >
2026-04-02 21:01:40 +08:00
Harry Mellor and GitHub
dfe5e31689
Don't compile vision encoder for Transformers backend ( #30518 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-02 12:42:29 +00:00
2ce3d0ce36
[Feature] KV cache per-token-head INT8/FP8 quantization ( #38378 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: yangyang4991 <yangyang4991@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
2026-04-02 08:13:26 -04:00
Jiangyun Zhu and GitHub
4eefbf9609
[Perf] fuse kernels in gdn ( #37813 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-04-02 11:52:18 +00:00
vllmellm and GitHub
551b3fb39f
[ROCm] Enable VLLM triton FP8 moe for gfx1201, tuned for Qwen3-30B-A3B-FP8 tp=2 and Qwen/Qwen3.5-35B-A3B-FP8 tp=2 ( #38086 )
...
Signed-off-by: big-yellow-duck <jeffaw99@hotmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-04-02 08:13:42 +00:00
Li, Jiang and GitHub
c6f722b93e
[CPU] Support gelu act in cpu_fused_moe ( #38770 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-02 14:14:32 +08:00
Xin Yang and GitHub
9bd7231106
Revert "[Kernel] Add gpt-oss Router GEMM kernel ( #37205 )" ( #38778 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-04-01 22:02:32 -07:00
73f48ce559
[Kernel] [Helion] Use warning_once in get_gpu_name to prevent log spam ( #38743 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-04-01 21:30:31 -07:00
3aab680e3e
[ROCm][Bugfix] Fix ROCm runtime failure due to missing symbol ( #38750 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
Signed-off-by: Gregory Shtrasberg <156009573+gshtras@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: tjtanaavllm <tunjian.tan@amd.com >
2026-04-01 21:30:11 -07:00
Sergey Zinchenko and GitHub
5a2d420c17
[Bugfix] Use dedicated MM processor cache in /tokenize to prevent sender-cache pollution ( #38545 )
...
Signed-off-by: Sergey Zinchenko <sergey.zinchenko.rnd@gmail.com >
2026-04-01 21:14:49 -07:00
Benjamin Chislett and GitHub
5f96f9aff1
[Perf] DSV3.2 Indexer Fused Weights Projection ( #38684 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-02 03:34:49 +00:00
Luka Govedič and GitHub
694449050f
Fix multiline-format string for python 3.10 ( #38739 )
...
Signed-off-by: Luka Govedic <luka.govedic@gmail.com >
2026-04-02 03:19:35 +00:00
Nick Hill and GitHub
6241521dd2
[BugFix] Fix precommit breakage due to conflicting in-flight merges ( #38759 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-01 15:35:06 -07:00
Kevin H. Luu and GitHub
1785dc5501
Revert "[Bugfix] Fix Qwen3CoderToolParser anyOf/oneOf type resolution for nullable params ( #37831 )" ( #38751 )
2026-04-02 06:34:28 +08:00
Chang Su and GitHub
54500546ac
[Bugfix] Preserve original ImportError in gRPC server entrypoint ( #38673 )
...
Signed-off-by: Chang Su <chang.s.su@oracle.com >
2026-04-01 22:16:44 +00:00
Jeffrey Wang and GitHub
de5e6c44c6
[Feat][Executor] Introduce RayExecutorV2 ( #36836 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
2026-04-01 14:34:29 -07:00
yzong-rh and GitHub
cb268e4e55
[Refactor] Simplify FutureWrapper in MultiprocExecutor ( #38644 )
...
Signed-off-by: Yifan <yzong@redhat.com >
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-04-01 21:28:26 +00:00
Stefano Castagnetta and GitHub
6183cae1bd
[Bugfix] Restrict TRTLLM attention to SM100, fixing GB300 (SM103) hang ( #38730 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-01 12:08:40 -07:00
Monishver and GitHub
c09ad767cd
Feature/silu block quant fusion v1 ( #32996 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
2026-04-01 18:50:43 +00:00
Wentao Ye and GitHub
c9a9db0e02
[Compile] Fix nvfp4 compile warning ( #38573 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-01 18:28:57 +00:00
Chauncey and GitHub
cbe7d18096
[Misc] Rename think_start_str/think_end_str to reasoning_start_str/reasoning_end_str ( #38242 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-01 09:56:45 -07:00
Michael Goin and GitHub
db5d0719e1
[Kernel] Add MXFP8 to Marlin GEMM/MoE and refactor Mxfp8LinearOp ( #34664 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-01 09:41:42 -07:00
dc0428ebb8
[NIXL][BUG] Fix Triton heterogeneous TP ( #37940 )
...
Signed-off-by: Yifan <yzong@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-01 17:23:15 +02:00
Jesus Talavera and GitHub
148c2072ec
Add ibm-granite/granite-vision-3.3-2b to supported models documentation ( #38714 )
...
Signed-off-by: Jesus Talavera <jesus.talavera@ibm.com >
2026-04-01 08:22:25 -07:00
2f5c3c1ec0
[Misc] Fix docstring typo: buildin -> builtin ( #38722 )
...
Co-authored-by: majianhan <majianhan@kylinos.cn >
2026-04-01 07:39:46 -07:00
Fynn Schmitt-Ulms and GitHub
fa246d5231
Fix shape comment in extract_hidden_states example ( #38723 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-04-01 07:29:33 -07:00
bnellnm and GitHub
7cf56a59a2
[MoE Refactor] Make SharedExperts class for use with DefaultMoERunner ( #35153 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-01 09:44:08 -04:00
5e30e9b9a9
[Bugfix] Revert "Zero-init MLA attention output buffers to prevent NaN from CUDA graph padding" ( #38359 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-04-01 09:11:10 -04:00
582340f273
[Bugfix] Fix Qwen3CoderToolParser anyOf/oneOf type resolution for nullable params ( #37831 )
...
Signed-off-by: AAISSJ <maze0717@g.skku.edu >
Signed-off-by: <>
Co-authored-by: 세덩 <saison@sedeong-ui-MacBookAir.local >
2026-04-01 20:22:29 +08:00
992368522f
[KVTransfer] Fix TpKVTopology.is_kv_replicated equality case ( #38179 )
...
Signed-off-by: JianDan0212 <zhangyj0212@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-01 12:41:49 +02:00
Juan Pérez de Algaba and GitHub
58ee614221
(security) Enforce frame limit in VideoMediaIO ( #38636 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-04-01 10:23:45 +00:00
Harry Mellor and GitHub
f9f6a9097a
Add verified label to trigger pre-commit ( #38708 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-01 02:31:02 -07:00
Zhanda Zhu and GitHub
c75a313824
[Perf] triton bilinear_pos_embed kernel for ViT ( #37948 )
...
Signed-off-by: Zhanda Zhu <zhandazhu@gmail.com >
2026-04-01 01:52:02 -07:00
Lukas Geiger and GitHub
4f6eed3bd4
[Core] Simplify multimodal masking ( #34246 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-04-01 01:18:22 -07:00
Li, Jiang and GitHub
36d7f19897
[CPU] Support head_size 512 in cpu_attn ( #38676 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-01 05:42:27 +00:00
Jeffrey Wang and GitHub
2d725b89c5
[Bugfix] Lazy import diskcache to avoid sqlite3/libstdc++ ImportError at startup ( #38649 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
2026-04-01 05:31:20 +00:00
ef53395e2c
[bugfix] do not add extra linebreak for score/rerank with chat template ( #38617 )
...
Signed-off-by: augusto.yjh <augusto.yjh@antgroup.com >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: wang.yuqi <noooop@126.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-01 04:50:07 +00:00
Lucas Wilkinson and GitHub
eb47454987
[Bugfix][MLA] Add logits size budget to sparse indexer prefill chunking ( #36178 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-04-01 00:15:53 -04:00
Matthew Bonanni and GitHub
116f4be405
[1/N][Cleanup] Standardize on use of is_quantized_kv_cache ( #38659 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-01 04:08:01 +00:00
Wentao Ye and GitHub
7b01d97a22
[Perf] Optimize mean pooling using chunks and index_add, 5.9% E2E throughput improvement ( #38559 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-01 03:54:58 +00:00
17b72fd1c8
Fix priority preemption regression test in scheduler ( #37051 )
...
Signed-off-by: HarshRathva <harshrathvaai@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-01 06:36:12 +03:00
c49497726b
[ROCm][perf] Shuffle KV cache to use paged_attention_common ( #32914 )
...
Signed-off-by: Samu Tamminen <stammine@amd.com >
Co-authored-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-04-01 03:30:19 +00:00
cb0b443274
[Misc] Add 20 regression tests for 11 tool parser bug fixes ( #38172 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-01 03:00:31 +00:00
40bb175027
[vLLM IR] 1/N Implement IR skeleton and rms_norm op ( #33825 )
...
Signed-off-by: Luka Govedič <lgovedic@redhat.com >
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Signed-off-by: chzhang <chaojun.zhang@intel.com >
Signed-off-by: Luka Govedic <luka.govedic@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Luka Govedič <ProExpertProg@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-03-31 22:15:05 -04:00
0fab52f0aa
Fix NaN from stale FP4 scale padding in create_fp4_scale_tensor ( #38148 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-03-31 19:14:59 -07:00
Yifan Qiao and GitHub
91e4521f9f
[Feat][v1] Simple yet General CPU KV Cache Offloading ( #37160 )
...
Signed-off-by: Yifan Qiao <yifanqiao@berkeley.edu >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-03-31 17:58:37 -07:00
31a719bcd3
[ROCm][perf] fix Aiter sparse MLA with MTP>1 ( #37887 )
...
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Signed-off-by: Stig-Arne Grönroos <sgronroo@amd.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-31 19:22:23 -04:00
2e56975657
Generative Scoring ( #34539 )
...
Signed-off-by: Vedant Jhaveri <vjhaveri@linkedin.com >
Co-authored-by: Vedant Jhaveri <vjhaveri@linkedin.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-03-31 16:02:11 -07:00
Chang Su and GitHub
36f1dc19ae
feat(grpc): add periodic stats logging and servicer log forwarding ( #38333 )
...
Signed-off-by: Chang Su <chang.s.su@oracle.com >
2026-03-31 15:50:07 -07:00
Asaf Gardin and GitHub
3dc01ef352
[Quantization] Consolidate dummy format logic into DummyModelLoader ( #38637 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-03-31 22:20:45 +00:00
cc671cb110
[Kernel] [Helion] [17/N] Add Helion kernel torch.compile support ( #38592 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-03-31 17:06:42 -04:00
Wentao Ye and GitHub
856589ed9a
[Refactor] Remove dead code in kv connector and model runner ( #38383 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-31 17:05:23 -04:00
czhu-cohere and GitHub
517b769b58
[Perf] Fix DBO overlap: capture DeepEP event before yield ( #38451 )
...
Signed-off-by: root <conway.zhu@cohere.com >
2026-03-31 20:38:59 +00:00
d9b90a07ac
[MoE Refactor] Migrate Unquantized to Full Oracle Flow ( #36286 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Signed-off-by: yzong-rh <yzong@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
2026-03-31 15:43:33 -04:00
Olya Kozlova and GitHub
598190aac3
[fix] Remove trtllm ragged mla prefills ( #36540 )
...
Signed-off-by: Olya Kozlova <okozlova@nvidia.com >
2026-03-31 12:30:27 -07:00
b779eb3363
[Model] Sync upstream BT=chunk_size fix for GDN chunk_fwd_kernel_o, simplify warmup to single pass ( #38343 )
...
Signed-off-by: AuYang <459461160@qq.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-03-31 23:03:24 +04:00
077a9a8e37
[torch.compile] Refactor Attention Quant Fusion Pass and Remove Boilerplate ( #37373 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-03-31 14:15:50 -04:00
Run Yu and GitHub
07edd551cc
[CI/Build] Resolve a dependency deadlock when installing the test dependencies used in CI ( #37766 )
...
Signed-off-by: Run Yu <yurun00@gmail.com >
2026-03-31 18:05:14 +00:00
mikaylagawarecki and GitHub
7c080dd3c5
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI ( #37503 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-03-31 10:21:13 -07:00
Yi Liu and GitHub
0dd25a44ea
[Quantization][Autoround][XPU] Add W4A16 Support ( #37986 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
2026-03-31 16:48:24 +00:00
SandishKumarHN and GitHub
3896e021a0
[Bugfix] Fix FusedMoE weight loading with padded hidden dimensions ( #37010 )
...
Signed-off-by: SandishKumarHN <sandish@fb.com >
2026-03-31 12:22:26 -04:00
zhang-prog and GitHub
b6e636c12c
[Fix] handle PaddleOCR-VL image processor max_pixels across Transformers v4/v5 ( #38629 )
...
Signed-off-by: zhangyue66 <zhangyue66@baidu.com >
2026-03-31 15:50:41 +00:00
Jingu Kang and GitHub
f1ff50c86c
[Bugfix] clamp dA_cumsum differences to prevent Inf in Mamba2 SSD kernels ( #37501 )
...
Signed-off-by: Jingu Kang <jg.k@navercorp.com >
2026-03-31 17:35:51 +02:00
757068dc65
[Bugfix][Async] Fix async spec decoding with hybrid models ( #38556 )
...
Signed-off-by: SandishKumarHN <sandishkumarhn@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: SandishKumarHN <sandishkumarhn@gmail.com >
2026-03-31 11:08:54 -04:00
Nicolò Lucchesi and GitHub
7337ff7f03
[Docs] PD with Nixl compat matrix ( #38628 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-03-31 15:01:21 +00:00
Kyle Sayers and GitHub
5869f69c5f
[Online Quant] [QeRL] Minor code cleanup ( #38574 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-03-31 14:56:43 +00:00
4dfad17ed1
replace cuda_device_count_stateless() to current_platform.device_count() ( #37841 )
...
Signed-off-by: Liao, Wei <wei.liao@intel.com >
Signed-off-by: wliao2 <wei.liao@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-31 22:32:54 +08:00
wenjun liu and GitHub
e8057c00bc
[CI] Avoid concurrent docker pull in intel XPU CI runners to prevent rate limit issues ( #38594 )
...
Signed-off-by: wendyliu235 <wenjun.liu@intel.com >
2026-03-31 22:23:18 +08:00
Nicolò Lucchesi and GitHub
7430389669
[Bugfix][CI] Skip flaky test_eagle test ( #38566 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-03-31 09:42:37 -04:00
ElizaWszola and GitHub
202f147cf2
Fix MLA runs when use_inductor_graph_partition=True ( #38631 )
...
Signed-off-by: ElizaWszola <ewszola@redhat.com >
2026-03-31 13:37:43 +00:00
Jiangyun Zhu and GitHub
ea7bfde6e4
[CI] fix LM Eval Qwen3.5 Models (B200) ( #38632 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-03-31 13:20:08 +00:00
d71a15041f
[XPU]move testing dependencies from Dockerfile to xpu-test.in ( #38596 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-31 12:49:43 +00:00
abdbb68386
[EPLB] Add alternative communication for EPLB weight exchange ( #33176 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-03-31 08:17:12 -04:00
liuzhenwei and GitHub
0c63739135
[EPD] update EPD script arguments ( #36742 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-03-31 12:02:09 +00:00
719735d6c5
[CI Failure] pin colmodernvbert revision ( #38612 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-31 10:54:54 +00:00
Maosheng Liao and GitHub
aae3e688f8
Fix document of torchrun_example.py ( #31113 )
2026-03-31 10:54:23 +00:00
Matthew Bonanni and GitHub
7d65463528
[WIP][CI][Bugfix] Fix test_run_eagle_dp ( #38584 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-31 12:30:25 +02:00
Mateusz Sokół and GitHub
8278825b57
DOC: TPU mention fix ( #38129 )
...
Signed-off-by: Mateusz Sokół <mat646@gmail.com >
2026-03-31 03:27:56 -07:00
Chang Su and GitHub
acf7292bf2
[Misc] Move --grpc CLI argument into make_arg_parser ( #38570 )
...
Signed-off-by: Chang Su <chang.s.su@oracle.com >
2026-03-31 03:24:05 -07:00
Chauncey and GitHub
ce884756f0
[Feature]: add presence_penalty and frequency_penalty fields to Responses API ( #38613 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-31 08:45:57 +00:00
wang.yuqi and GitHub
d9d21eb8e3
[Frontend][3/n] Improve pooling entrypoints | scoring. ( #28631 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-31 07:52:00 +00:00
Yintong Lu and GitHub
f09daea261
[CPU] Support int8 compute mode in CPU AWQ ( #35697 )
...
Signed-off-by: Yintong Lu <yintong.lu@intel.com >
2026-03-31 15:27:37 +08:00
Kevin H. Luu and GitHub
42318c840b
[ci] Remove benchmarks job ( #38611 )
2026-03-31 06:46:21 +00:00
zhangyiming and GitHub
1ac6694297
[OOT] Add OOT support for linear kernel. ( #37989 )
...
Signed-off-by: menogrey <1299267905@qq.com >
2026-03-31 14:33:21 +08:00
6cc7abdc66
[kv_offload+HMA] Fix num_blocks with different per-layer page sizes and improve assert message ( #38554 )
...
Signed-off-by: Kfir Toledo <kfir.toledo@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-03-31 06:00:40 +00:00
Flora Feng and GitHub
d53cb9cb8e
[Tool Parser][2/3] Use self.tools instead of request.tools in tool parsers ( #38189 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-31 13:41:36 +08:00
Louie Tsai and GitHub
44eef0ca1e
vLLM Benchmark Suite perf regression after PR#32723 ( #38576 )
...
Signed-off-by: louie-tsai <louie.tsai@intel.com >
2026-03-31 05:23:17 +00:00
Andreas Karatzas and GitHub
b9cdc85207
[ROCm][CI] Fix Whisper translation test attention backend selection ( #38508 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-31 13:21:49 +08:00
Flora Feng and GitHub
3e802e8786
[Mypy] Fix adjust_request typing ( #38264 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-31 04:21:18 +00:00
Martin Hickey and GitHub
350af48e14
[KVConnector] Remove redundant method KVConnectorOutput::merge() ( #38546 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-03-31 07:11:02 +03:00
Lucas Kabela and GitHub
e31915063d
[Bugfix] Fix for builtins (forward fix of pytorch/177558) ( #37234 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-03-31 01:08:11 +00:00
Flora Feng and GitHub
29e48707e8
[Refactor] Consolidate Tool type alias in tool_parsers/utils.py ( #38265 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-31 00:55:51 +00:00
4ac227222f
[Bugfix][DCP] Fix CUDA graph capture for Decode Context Parallelism ( #36070 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-30 20:20:43 -04:00
Vadim Gimpelson and GitHub
bb51d5b40d
Add @vadiklyutiy as committer ( #38589 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-03-31 07:50:04 +08:00
Prathmesh Bhatt and GitHub
93b3ec1585
feat(attention): extract KV-cache update from FlashAttentionDiffKV ba… ( #36466 )
...
Signed-off-by: Prathmesh Bhatt <71340361+Prathmesh234@users.noreply.github.com >
2026-03-30 23:16:09 +00:00
e812bf70bd
Restore non-hf processor path for Nano-Nemotron-VL (bypass call_hf_processor_mm_only) - fixes #38018 ( #38567 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-03-30 21:56:52 +00:00
bcc6f67447
[Bugfix] Use null block (0) for padded block table entries ( #35431 )
...
Signed-off-by: SandishKumarHN <sandish@fb.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-30 14:02:51 -07:00
Asaf Gardin and GitHub
1fc69f59bb
[Bug fix][Quantization] Fix dummy weight loading ( #38478 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-03-30 16:38:02 -04:00
Micah Williamson and GitHub
d9c7db18da
[ROCm][CI] Pin test_hybrid test to TRITON_ATTN on ROCm ( #38381 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-03-30 20:26:46 +00:00
Ilya Markov and GitHub
12701e8af2
[EPLB] Optmize eplb mapping and record in router for prefill ( #36261 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
2026-03-30 19:48:33 +00:00
Benjamin Chislett and GitHub
494636b29d
[Feat][Spec Decode] DFlash ( #36847 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-03-30 15:03:15 -04:00
mikaylagawarecki and GitHub
ab1a6a43fa
[3/n] Migrate cutlass/scaled_mm_entry.cu torch stable ABI ( #37221 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-03-30 11:20:13 -07:00
b5e608258e
[Refactor] Unify engine process monitoring in engine manager and add Ray backend support ( #35862 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-03-30 10:16:09 -07:00
Matthew Bonanni and GitHub
2c734ed0e0
[Bugfix][MLA] Change default SM100 MLA prefill backend back to TRT-LLM ( #38562 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-30 09:51:24 -07:00
3b1dbaad4e
[HMA]Fix corner case when hybrid page_size can not be evenly divided issue (blk_size=64,tp=4) ( #37467 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Chendi.Xue <chendi.xue@intel.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-03-30 16:47:30 +00:00
Johnny and GitHub
b4a2f3ac36
[NVIDIA] Bugfix NVFP4 DGX Spark and RTX50 ( #38423 )
...
Signed-off-by: johnnynunez <johnnynuca14@gmail.com >
Signed-off-by: Johnny <johnnynuca14@gmail.com >
2026-03-30 09:36:18 -07:00
roikoren755 and GitHub
8e6293e838
[Mamba] Add stochastic rounding support ( #35753 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-03-30 12:33:49 -04:00
dbdd9ae067
[ROCm][Bugfix] fix exception related to trust_remote_code for MiniMax-M2.1-MXFP4 ( #37698 )
...
Signed-off-by: Hongxia Yang <hongxiay.yang@amd.com >
Co-authored-by: Hongxia Yang <hongxiay.yang@amd.com >
2026-03-30 15:49:23 +00:00
e8b055a5ac
[Bugfix] Handle ParallelLMHead in compressed-tensors get_quant_method ( #37291 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-03-30 07:30:52 -07:00
tomeras91 and GitHub
246dc7d864
[Misc] Add @tomeras91 as a maintainer of Nemotron related code + mamba block ( #38547 )
...
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com >
2026-03-30 21:12:17 +08:00
Thomas Parnell and GitHub
7c3f88b2a8
[Bugfix] Remove false-positive format mismatch warnings in FLA ops ( #38255 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-03-30 12:32:26 +00:00
Li, Jiang and GitHub
6557f4937f
[Bugfix][CPU] Skip set_num_threads after thread binding ( #38535 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-03-30 20:13:00 +08:00
Andreas Karatzas and GitHub
677424c7ac
[Core][CI] Add opt-in media URL caching via VLLM_MEDIA_CACHE ( #37123 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-30 04:58:53 -07:00
1031c84c36
Fix ambiguous num_blocks for hybrid attn mamba ( #37236 )
...
Signed-off-by: Collin McCarthy <cmccarthy@nvidia.com >
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-03-30 11:09:45 +00:00
aliialsaeedii and GitHub
7e76af14fa
[Bugfix][Frontend] Return 400 for corrupt/truncated image inputs instead of 500 ( #38253 )
...
Signed-off-by: aliialsaeedii <ali.al-saeedi@nscale.com >
2026-03-30 10:26:46 +00:00
3683fe6c06
[Bugfix] Fix shared-object aliasing in n>1 streaming with tool calls ( #38158 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Yifan <yzong@redhat.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-03-30 10:12:13 +00:00
Nicolò Lucchesi and GitHub
cc06b4e86b
[Mamba][Bugfix] Raise on insufficient cache blocks instead of silently capping cudagraph sizes ( #38270 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-03-30 09:41:50 +00:00
TJian and GitHub
03ac6ca895
[ROCm] [DOC] Update the Documentation to include ROCm Nightly Wheel support ( #38457 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-03-30 02:25:46 -07:00
haosdent and GitHub
a08b7733fd
[CI] Fix SPLADE pooler test broken by #38139 ( #38495 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-03-30 07:48:33 +00:00
Tan Pin Siang and GitHub
85c0950b1f
[ROCm] Enable MORI EP for unquantized MoE with AITER backend ( #37529 )
...
Signed-off-by: Tan Pin Siang <pinsiang.tan@amd.com >
2026-03-30 15:19:33 +08:00
Juan Pérez de Algaba and GitHub
57861ae48d
(security) Fix SSRF in batch runner download_bytes_from_url ( #38482 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-03-30 07:10:01 +00:00
ac30a8311e
[Bugfix][Model] Fix PixtralForConditionalGeneration LoRA ( #36963 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-03-29 23:59:42 -07:00
PikaPikachu and GitHub
63babd17f1
[Model][Quantization] Add GGUF support for MiniMax-M2.1 ( #36965 )
...
Signed-off-by: kangletian <Letian.Kang@amd.com >
2026-03-30 14:24:06 +08:00
Kevin H. Luu and GitHub
fec5aeca12
[ci] Soft fail and disable retry for AMD build image job ( #38505 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-03-29 23:05:26 -07:00
Jaewon and GitHub
d816834c1a
[MoE] Add RoutingMethodType.Simulated to TRT-LLM FP8/NVFP4 kernel allowlists ( #38329 )
...
Signed-off-by: Jaewon Lee <jaewon@meta.com >
2026-03-29 22:53:43 -07:00
Roger Wang and GitHub
92f0db57a8
[Misc] Always use forward_mulmat for Conv3d on newer versions of torch. ( #38487 )
2026-03-30 05:39:41 +00:00
Andreas Karatzas and GitHub
bea23536f6
[CI] Add temperature=0.0, reduce max_tokens, and add debug prints to audio_in_video tests ( #38492 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-30 05:36:45 +00:00
Jiangyun Zhu and GitHub
c133f33746
Add @ZJY0516 to CODEOWNERS ( #38497 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-03-29 21:10:00 -07:00
a6db99ba02
[Bugfix] Support multi-type params parsing for DeepSeek v3.2 ( #33703 )
...
Signed-off-by: Stanislav Kirillov <stas@nebius.com >
Co-authored-by: Stanislav Kirillov <stas@nebius.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-03-30 04:07:28 +00:00
Andreas Karatzas and GitHub
4f2ed5fddb
[ROCm][CI] Enable hybrid chunked prefill test ( #38317 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-30 10:30:26 +08:00
Kyle Sayers and GitHub
d28d86e8a3
[QeRL] Fix online quantized reloading ( #38442 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-03-29 14:56:41 -06:00
Wentao Ye and GitHub
995dea1354
[Perf] Remove redundant device copies for CPU-only pooling token IDs, 48.9% E2E throughput improvement ( #38139 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-29 18:12:50 +00:00
allgather and GitHub
8c0b6267d7
[Transformers v5] fix missing pixtral/voxtral multimodal dispatch ( #38410 )
...
Signed-off-by: allgather <all2allops@gmail.com >
2026-03-29 09:59:06 +00:00
Andreas Karatzas and GitHub
43cc5138e5
[ROCm][CI] Fix cross-attention dispatch for encoder-decoder models ( #38450 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-28 22:08:03 -07:00
Shubhra Pandit and GitHub
5b8c30d62b
[Spec Decode, BugFix] Propagate norm_before_fc from Eagle3 speculator ( #38111 )
...
Signed-off-by: Shubhra Pandit <shubhra.pandit@gmail.com >
2026-03-29 00:42:06 +00:00
haosdent and GitHub
d39b8daf5f
[Feature] Add Qwen3-ForcedAligner support via token classification pooling ( #35367 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-03-29 00:27:52 +00:00
Walter Beller-Morales and GitHub
fafca38adc
[BugFix][Frontend] apply task instruction as system prompt in cohere v2/embed ( #38362 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-03-28 18:30:54 +00:00
aa4eb0db78
[CI]revert initialize_model context manager ( #38426 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-28 16:56:50 +00:00
Andreas Karatzas and GitHub
af89140efc
[ROCm][CI] Fix UV install in Dockerfile.rocm to detect curl failures and retry ( #38415 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-29 00:47:42 +08:00
haosdent and GitHub
b2bc736b12
[CI] Fix Ernie4.5-VL initialization test ( #38429 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-03-28 22:43:24 +08:00
whyiug and GitHub
58c959a767
[Misc]: clean up non-core lint issues ( #37049 )
...
Signed-off-by: whyiug <whyiug@hotmail.com >
2026-03-28 10:28:16 -04:00
Bvicii and GitHub
bda3eda82d
[Bugfix] Disallow renderer_num_workers > 1 with mm processor cache ( #38418 )
...
Signed-off-by: Bvicii <yizhanhuang2002@gmail.com >
2026-03-28 06:32:52 -07:00
Michael Goin and GitHub
2bf5b70ae8
[CI Bugfix] Pre-download missing FlashInfer headers in Docker build ( #38391 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-03-28 06:09:00 -07:00
yzong-rh and GitHub
6dad4c5722
[Test] Fix flaky race condition in test_abort_final_step ( #38414 )
...
Signed-off-by: Yifan <yzong@redhat.com >
2026-03-28 09:06:56 +00:00
171775f306
Fix Device Index for ROCm Ray Workers in MoE Benchmark ( #38108 )
...
Signed-off-by: Liwen <53441624+li-liwen@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-28 08:27:11 +00:00
TJian and GitHub
58a249bc61
[ROCm] [Release] Update ROCm variant from rocm700 to rocm721 ( #38413 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-03-28 06:07:03 +00:00
IriKa and GitHub
148a5c1226
[Bugfix]fix output Nan/Inf in marlin if dtype=float16 ( #33972 )
...
Signed-off-by: IriKa Qiu <qiujie.jq@gmail.com >
2026-03-27 16:36:08 -07:00
Wei Zhao and GitHub
b69bf2f0b1
[Perf] Use torch compile to fuse pack topk in trtllm moe ( #37695 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
2026-03-27 17:30:46 -06:00
rongfu.leng and GitHub
88149b635e
Add nvidia h800 moe config ( #31201 )
...
Signed-off-by: rongfu.leng <rongfu.leng@daocloud.io >
2026-03-27 16:28:48 -07:00
83a4df049d
[ROCm][Documentation] update quickstart and installation to include rocm nightly docker tips ( #38367 )
...
Signed-off-by: Hongxia Yang <hongxiay.yang@amd.com >
Co-authored-by: Hongxia Yang <hongxiay.yang@amd.com >
2026-03-27 23:20:19 +00:00
Gregory Shtrasberg and GitHub
731285c939
[ROCm][CI/Build] ROCm 7.2.1 release version; torch 2.10; triton 3.6 ( #38252 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
2026-03-27 18:03:12 -05:00
+3
97d19197bc
[NVIDIA] Fix DGX Spark logic ( #38126 )
...
Signed-off-by: johnnynunez <johnnynuca14@gmail.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Signed-off-by: Sathish Sanjeevi <sathish.krishnan.p.s@gmail.com >
Signed-off-by: guillaume_guy <guillaume.guy@airbnb.com >
Signed-off-by: Guillaume Guy <guillaume.c.guy@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Sathish Sanjeevi <SKPsanjeevi@users.noreply.github.com >
Co-authored-by: Guillaume Guy <guillaume.c.guy@gmail.com >
Co-authored-by: guillaume_guy <guillaume.guy@airbnb.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-03-27 15:26:07 -07:00
Giancarlo Delfin and GitHub
384e4d5f48
[Model Runner V2] Rebuild attention metadata before eagle decode full… ( #38311 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-27 13:46:42 -07:00
Nicolò Lucchesi and GitHub
44a6528028
[CI] Skip failing test ( #38369 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-03-27 13:25:19 -07:00
Kyle Sayers and GitHub
648edcf729
[QeRL] Compose online quantization with quantized reloading ( #38032 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-03-27 13:22:33 -07:00
7ba425e916
Add short flag -sc for --speculative-config argument ( #38380 )
...
Co-authored-by: Claude <noreply@anthropic.com >
2026-03-27 12:04:22 -07:00
Gregory Shtrasberg and GitHub
b8665383df
[ROCm] Fix GPT-OSS import for triton 3.6 ( #37453 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
2026-03-27 18:00:57 +00:00
0e9358c11d
{ROCm]: gpt-oss fusion/padding fixes ( #38043 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-03-27 12:19:15 -04:00
21d2b53f88
Remove need for explicit \n in docstring lists for --help formatting ( #38350 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-27 08:38:00 -07:00
Jonas M. Kübler and GitHub
98e7f223b9
enable skipping of SW attention layers when using FP8 KV cache ( #33695 )
...
Signed-off-by: Jonas Kuebler <kuebj@amazon.com >
2026-03-27 07:25:02 -06:00
b111f8a61f
fix(security): Add VLLM_MAX_N_SEQUENCES environment variable and enforce limit ( #37952 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Russell Bryant <rbryant@redhat.com >
Co-authored-by: Russell Bryant <rbryant@redhat.com >
2026-03-27 09:02:10 -04:00
Sage Moore and GitHub
497e234d38
[EPLB] Cleanup the transfer logic for the various eplb maps ( #34520 )
...
Signed-off-by: Sage Moore <sagmoore@redhat.com >
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-03-27 10:18:46 +01:00
dtc and GitHub
6287e7fa20
[P/D] Mooncake: Add unit tests and minor fixes for mooncake connector ( #36946 )
...
Signed-off-by: Tianchen Ding <dtcccc@linux.alibaba.com >
2026-03-27 09:26:40 +01:00
84e439a9cb
[CI/Build] Move nightly wheel index generation to a single post-build step ( #38322 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com >
2026-03-27 07:44:18 +00:00
a1746ff9ec
[Doc] Clarify Helm chart location in deployment guide ( #38328 )
...
Signed-off-by: Yuichiro Utsumi <utsumi.yuichiro@fujitsu.com >
Signed-off-by: Yuichiro Utsumi <81412151+utsumi-fj@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-27 15:43:02 +08:00
Flora Feng and GitHub
aee4c14689
[Bugfix] Fix Hermes tool parser when stream interval > 1 ( #38168 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-27 14:42:26 +08:00
Bowen Bao and GitHub
0ae89f18fd
[Refactor] Move FusedMoE hidden_size roundup to quant_method ( #34285 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-03-26 23:38:26 -07:00
wenjun liu and GitHub
c2b17d71af
[CI] Add xpu auto-label rule for Intel GPU/XPU PRs ( #38320 )
...
Signed-off-by: wendyliu235 <wenjun.liu@intel.com >
2026-03-27 14:22:38 +08:00
Li, Jiang and GitHub
becaed6ec8
[CPU] Support CT W4A16 on CPU MP kernel ( #38219 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-03-27 14:15:28 +08:00
Xiaoshuang Wang and GitHub
a8eab8f30d
[Model] Extract GatedDeltaNetAttention into shared layer for Qwen3Next and Qwen3.5 ( #37975 )
...
Signed-off-by: wxsIcey <1790571317@qq.com >
Signed-off-by: Icey <1790571317@qq.com >
2026-03-27 14:13:21 +08:00
cjackal and GitHub
2babac0bed
[frontend] dump openai responses type by alias ( #38262 )
...
Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com >
2026-03-27 05:58:20 +00:00
Or Ozeri and GitHub
7cc302dd87
[kv_offload+HMA][7/N]: Support register_kv_caches for hybrid models ( #37853 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-03-27 08:38:33 +03:00
999dfc1622
[Bugfix] Offload blocking tokenizer ops to shared thread pool to unblock event loop ( #34789 )
...
Signed-off-by: Bvicii <yizhanhuang2002@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-03-26 22:17:00 -07:00
wenjun liu and GitHub
d86060122a
[CI/Build] enable Intel XPU test flow with prebuilt image ( #37447 )
...
Signed-off-by: wendyliu235 <wenjun.liu@intel.com >
2026-03-26 18:16:04 -07:00
Harry Mellor and GitHub
f73bcb1c51
Various Transformers v5 config fixes ( #38247 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-26 23:06:59 +00:00
yzong-rh and GitHub
28048bd6b0
[Bugfix] Add missing f-string prefix in xgrammar choices error message ( #38162 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-03-26 21:43:03 +00:00
Giancarlo Delfin and GitHub
c32e97602d
[Model Runner V2] Enable forcing a specific acceptance rate during rejection sampling ( #38045 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-26 13:38:12 -07:00
0904b6550d
Fix multi-node allreduce fusion ( #38136 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: root <root@theia0053.lyris.clusters.nvidia.com >
2026-03-26 20:24:36 +00:00
Stig-Arne Grönroos and GitHub
f26fcdfb9e
[Bugfix][ROCm] Fix lru_cache on paged_mqa_logits_module ( #37547 )
...
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
2026-03-26 19:01:05 +00:00
TJian and GitHub
bc9c6fbbe6
[ROCm] [Bugfix] [Release] Fix nightly rocm release pipeline ( #38263 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-03-26 18:47:10 +00:00
Andreas Karatzas and GitHub
bff9a1c266
[ROCm][CI] Override PYTORCH_ROCM_ARCH with detected GPU arch in test containers ( #38165 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 18:33:45 +00:00
Andreas Karatzas and GitHub
db01535e2b
[ROCm][CI] Add uv pip compile workflow for rocm-test.txt lockfile ( #37930 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 12:44:01 -05:00
a4cf9b22ba
[ROCM][Bugfix] Use correct stride in cp_mha_gather_cache_kernel for hybrid model ( #37228 ) ( #37228 )
...
Signed-off-by: jennyyyyzhen <yzhen@hmc.edu >
Co-authored-by: yZhen <yZhen@fb.com >
2026-03-26 10:33:39 -07:00
Andreas Karatzas and GitHub
9c3ae04bfe
[ROCm][CI] Add LM Eval Qwen3.5 Models test for MI355 ( #38155 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 16:51:18 +00:00
Andreas Karatzas and GitHub
a8e48a7b85
[CI] Fix conch kernel crash on 3D input by reshaping to 2D before GEMM ( #38178 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 11:46:03 -05:00
Divakar Verma and GitHub
b9dbc5c4ab
[Mamba][APC] Add test case to compare apc outputs ( #34977 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-03-26 16:40:35 +00:00
60af7b967b
[Releases] [ROCm] Enable Nightly Docker Image and Wheel Releases for ROCm ( #37283 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxiay.yang@amd.com >
2026-03-26 16:32:25 +00:00
Andreas Karatzas and GitHub
bdc1719eb9
[ROCm][CI] Fix AITER state leak in shared_fused_moe_routed_transform test ( #38137 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 09:26:46 -07:00
0aac2048bf
[Bugfix] Restore CUDA graph persistent buffers for FP8 FlashMLA decode ( #35175 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-26 16:13:39 +00:00
Chuan (Richard) Li and GitHub
cb2263218e
[Bugfix][Minor] Fix potential NameError in mamba backend selector and misc typos ( #35886 )
...
Signed-off-by: Li <chuali@amd.com >
2026-03-26 11:59:24 -04:00
Wentao Ye and GitHub
e054f152fa
[CI] Add batch invariant test for b200 ( #38014 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-26 11:54:54 -04:00
zhang-prog and GitHub
0f5b526040
[Fix] Remove unused packing_position_embedding from PaddleOCRVL for better checkpoint compatibility ( #38232 )
...
Signed-off-by: zhangyue66 <zhangyue66@baidu.com >
2026-03-26 15:34:49 +00:00
be1a85b7a2
Revert "[MoE Kernel] Flashinfer nvfp4 cutedsl moe kernel integration" ( #38050 ) ( #38169 )
...
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-03-26 07:59:09 -07:00
Cyrus Leung and GitHub
2e225f7bd2
[Renderer] Consolidate factory methods ( #38218 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-26 12:19:22 +00:00
Jared Wen and GitHub
757eafcf37
[bug-fix] GLM OCR Patch Merger context_dim ( #37962 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-03-26 05:11:21 -07:00
wang.yuqi and GitHub
dcdc145893
[CI] Reorganize scoring tests ( #38207 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-26 12:07:01 +00:00
Andreas Karatzas and GitHub
f2d16207c7
[ROCm][CI] Fix flaky GPTQ compile correctness test ( #38161 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 19:57:00 +08:00
Andreas Karatzas and GitHub
37a83007fe
[ROCm][CI] Fix wvSplitKrc mock argument order in test_rocm_unquantized_gemm ( #38167 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 19:54:59 +08:00
Wentao Ye and GitHub
bf5eec638d
[Refactor] Remove unused utils ( #38153 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-26 17:08:19 +08:00
Mateusz Sokół and GitHub
b1cb1d3d2c
DOC: Documentation pages fixes ( #38125 )
...
Signed-off-by: Mateusz Sokół <mat646@gmail.com >
2026-03-26 16:55:42 +08:00
Kunshang Ji and GitHub
6ae8bbd0c2
[XPU] Disable xpu graph by default ( #38193 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-26 01:53:45 -07:00
Cyrus Leung and GitHub
a9213c0ffe
[Doc] Fix outdated reference to CUDAGraphManager ( #38209 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-26 01:52:38 -07:00
Cyrus Leung and GitHub
502c41a8f6
[Model] Use helper function to run MM processors with token inputs (where applicable) ( #38018 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-26 16:44:04 +08:00
Vadim Gimpelson and GitHub
52069012fe
[Bugfix] Fix DeepGemm E8M0 accuracy degradation for Qwen3.5 FP8 on Blackwell ( #38083 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-03-26 01:21:47 -07:00
Fadi Arafeh and GitHub
71161e8b63
[cpu][ci] remove soft-fail for Arm CI and add quant model tests ( #37691 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-03-26 07:03:31 +00:00
Terry Gao and GitHub
38de822310
[Model] Add torch.compile support for InternVL vision encoder ( #38049 )
...
Signed-off-by: tianrengao <terrygao87@gmail.com >
2026-03-25 23:52:29 -07:00
Jee Jee Li and GitHub
2bfbdca23c
[Bugfix] Fix benchmark_fused_collective.py ( #38082 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-25 23:51:00 -07:00
Matej Rojec and GitHub
2908094567
Add /v1/chat/completions/batch endpoint for batched chat completions ( #38011 )
...
Signed-off-by: Matej Rojec <64556640+MatejRojec@users.noreply.github.com >
2026-03-26 12:13:33 +08:00
e6bf9f15ec
[Bugfix][CI] Fix Marlin FP8 Linear Kernel for Compressed Tensors Format ( #38092 )
...
Signed-off-by: BadrBasowid <Badr.Basowid@gmail.com >
Signed-off-by: BadrBasowid <61441185+BadrBasowid@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-25 21:11:43 -07:00
144030c84e
Relocate Encoder CUDA graph manager ( #38116 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-03-25 20:52:12 -07:00
Flora Feng and GitHub
e2db2b4234
[Tool Parser][1/3] Pass tools to ToolParser constructor ( #38029 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-26 10:29:06 +08:00
Chauncey and GitHub
87f05d6880
[Revert] Remove DeepGEMM availability check in DeepseekV32IndexerMetadataBuilder ( #38076 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-26 01:43:51 +00:00
Andreas Karatzas and GitHub
36f6aede23
[Misc] Optimized check to encapsulate both CUDA and ROCm platforms ( #34549 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-26 09:43:07 +08:00
Xin Yang and GitHub
9704a5c310
Disable dual stream execution of input projection for Qwen3 ( #38152 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-03-26 01:20:39 +00:00
Wei Zhao and GitHub
74056039b7
Fix minimax m2.5 nvfp4 kv scales weight loading ( #37214 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-03-26 00:48:06 +00:00
Jacob Platin and GitHub
d7d51a7ee5
[Bugfix] Fix Qwen3.5-FP8 Weight Loading Error on TPU ( #37348 )
...
Signed-off-by: Jacob Platin <jacobplatin@google.com >
2026-03-26 00:46:01 +00:00
Harry Mellor and GitHub
3c3c084240
Various Transformers v5 fixes ( #38127 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-26 00:10:08 +00:00
Ekagra Ranjan and GitHub
7b54f60db0
[Cohere] Enable Cohere-Transcribe ( #38120 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-03-25 16:13:51 -07:00
Rohan Potdar and GitHub
a0e8c74005
[ROCm]: Update rope+kvcache fusion conditions and disable custom op by default ( #36716 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-03-25 20:58:44 +00:00
70a2152830
[MultiModal] add support for numpy array embeddings ( #38119 )
...
Signed-off-by: guillaume_guy <guillaume.guy@airbnb.com >
Signed-off-by: Guillaume Guy <guillaume.c.guy@gmail.com >
Co-authored-by: guillaume_guy <guillaume.guy@airbnb.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-03-25 20:13:04 +00:00
Sathish Sanjeevi and GitHub
978fc18bf0
[ROCm] Utilize persistent MLA kernel from AITER ( #36574 )
...
Signed-off-by: Sathish Sanjeevi <sathish.krishnan.p.s@gmail.com >
2026-03-26 03:00:42 +08:00
7d6917bef5
[ROCm] Fix MoE kernel test failures on gfx950 ( #37833 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-03-25 13:46:40 -05:00
Mark McLoughlin and GitHub
e38817fadb
[Core][KV Connector] Remove use of num_cached_tokens in error handling ( #38096 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-03-25 18:20:48 +00:00
Nick Hill and GitHub
72cad44d3c
[Frontend] Move APIServerProcessManager target server fn ( #38115 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-03-25 18:14:41 +00:00
Cyrus Leung and GitHub
ba2f0acc2d
[Misc] Reorganize inputs ( #35182 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-25 10:22:54 -07:00
Yongye Zhu and GitHub
678b3c99e8
[MoE Kernel] Flashinfer nvfp4 cutedsl moe kernel integration ( #38050 )
2026-03-25 10:16:40 -07:00
mikaylagawarecki and GitHub
bf4cc9ed2d
[2/n] Migrate per_token_group_quant to torch stable ABI ( #36058 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-03-25 10:15:13 -07:00
1ac2ef2e53
[CI/Docs] Improve aarch64/DGX Spark support for dev setup ( #38057 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-25 09:24:42 -07:00
Richard Zou and GitHub
6e37c46b35
[compile] Add some more startup tests for top models ( #38046 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-03-25 12:02:22 -04:00
Wentao Ye and GitHub
1bf2ddd0ee
[Refactor] Rename WAITING_FOR_FSM to WAITING_FOR_STRUCTURED_OUTPUT_GRAMMAR ( #38048 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-25 11:41:44 -04:00
e7221180e1
[Kernel] Optimize SM120 CUTLASS blockwise FP8 GEMM ( #37970 )
...
Signed-off-by: Necofish <liuxiangyang@mail.ustc.edu.cn >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-03-25 08:20:04 -07:00
4a76ad12e0
[Bugfix] Preserve CUDA arch suffix (a/f) for SM12x — fixes NVFP4 NaN on desktop Blackwell ( #37725 )
...
Signed-off-by: Rob Tand <robert.tand@icloud.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-03-25 08:18:25 -07:00
d7e93e13fb
[Feature] EPLB Support for GPU Model Runner v2 ( #37488 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-25 08:16:39 -07:00
cd7643015e
[Feature] Support per-draft-model MoE backend via --speculative-config ( #37880 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Signed-off-by: [Andrii Skliar] <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-03-25 14:31:52 +00:00
Ben Browning and GitHub
a1a2566447
[Docs] Add guide for editing agent instruction files ( #37819 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-03-25 13:54:09 +00:00
yjz and GitHub
b745e8b5d3
[KVTransfer][Mooncake] Add heterogeneous TP support for disaggregated P/D in MooncakeConnector ( #36869 )
...
Signed-off-by: JianDan0212 <zhangyj0212@gmail.com >
2026-03-25 14:24:07 +01:00
Harry Mellor and GitHub
d215d1efca
[Mypy] Better fixes for the mypy issues in vllm/config ( #37902 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-25 06:14:43 -07:00
Fadi Arafeh and GitHub
34d317dcec
[CPU][UX][Perf] Enable tcmalloc by default ( #37607 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-03-25 20:39:57 +08:00
7ac48fd357
[Model] Add AutoWeightsLoader support for jais ( #38074 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
2026-03-25 12:38:40 +00:00
Harry Mellor and GitHub
d6bb2a9d9a
Fix Plamo 2/3 & LFM2 for Transformers v5 ( #38090 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-25 12:29:49 +00:00
Harry Mellor and GitHub
1e673a43ce
Better weight tying check for multimodal models ( #38035 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-25 12:07:23 +00:00
Andreas Karatzas and GitHub
04417ecd5f
[ROCm][CI] Rename filepath test to point to correct file ( #38102 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 20:05:46 +08:00
R0CKSTAR and GitHub
242c93f744
[Docs] Adds vllm-musa to custom_op.md ( #37840 )
...
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com >
2026-03-25 11:54:36 +00:00
Matthias Gehre and GitHub
a889b7f584
[Bugfix] Pass drafter quant_config to ParallelLMHead in Eagle3 ( #37280 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-03-25 11:42:58 +00:00
Harry Mellor and GitHub
ba2910f73a
Fix offline mode test for Transformers v5 ( #38095 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-25 11:39:48 +00:00
Andreas Karatzas and GitHub
f262a62aa1
[ROCm][CI] Fix flaky Cohere/OpenAI embedding parity test ( #37616 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 10:55:51 +00:00
Andreas Karatzas and GitHub
9ac2fcafbb
[CI] Fix realtime WebSocket timeout deadlock and unhandled model validation errors ( #37483 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 11:24:33 +01:00
Kunshang Ji and GitHub
e9ae3f8077
[Hardware][XPU] Align memory usage with cuda on xpu ( #37029 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-25 18:14:29 +08:00
Andreas Karatzas and GitHub
04cec4f927
[ROCm][CI] Increase OpenAPI schema test timeouts ( #38088 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 18:06:58 +08:00
Kunshang Ji and GitHub
14771f7150
[XPU] support MLA model on Intel GPU ( #37143 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-25 17:43:42 +08:00
189ddefbfd
[ROCm] Attention selector reordering ( #36702 )
...
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
2026-03-25 17:42:56 +08:00
Chauncey and GitHub
09c3dc9186
[Revert] Remove CUDA torch fallbacks for fp8_mqa_logits/fp8_paged_mqa_logits_torch function ( #37968 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-25 06:19:37 +00:00
vllmellm and GitHub
42e9547976
[ROCm][Test] Fix ROCM_AITER_UNIFIED_ATTN attn+quant fusion test ( #37640 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-03-25 05:06:15 +00:00
Chauncey and GitHub
a32783bb35
[Bugfix] Fix IndexError when accessing prev_tool_call_arr in OpenAIToolParser ( #37958 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-25 12:06:21 +08:00
Baorun (Lauren) Mu and GitHub
9d0351c91d
[Docs] Add Encoder (ViT) CUDA Graphs section to CUDA Graphs design doc ( #37914 )
...
Signed-off-by: Baorun Mu <bmu@nvidia.com >
2026-03-24 19:53:24 -07:00
Artem Perevedentsev and GitHub
a93a53f8a1
[Performance] Auto-enable prefetch on NFS with RAM guard ( #37673 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-03-24 17:31:14 -07:00
Andreas Karatzas and GitHub
679c6a3ecc
[Bugfix][ROCm][MoE] Fix mxfp4 oracle regressions from #37128 ( #37787 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 08:17:33 +08:00
Andreas Karatzas and GitHub
8bbb7c7f20
[ROCm][CI][PD] Add Hybrid SSM integration tests to CI ( #37924 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-25 07:58:39 +08:00
Kevin H. Luu and GitHub
af945615b5
[release] Move the rest of release jobs to release queue ( #38044 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-03-24 16:40:58 -07:00
82580b10ac
[Perf] Disable inductor runtime asserts by default for serving perfor… ( #37485 )
...
Signed-off-by: tianrengao <terrygao87@gmail.com >
Co-authored-by: Tianren Gao <tianren@fb.com >
2026-03-24 19:37:51 -04:00
Netanel Haber and GitHub
a0d487b2e1
nano_nemotron_vl: suppress readonly torch.from_numpy() warning in image and video resize paths ( #37903 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-03-24 23:25:56 +00:00
Junhao and GitHub
b73b5b0629
Make microbatch optimization (DBO) work with general models ( #37926 )
...
Signed-off-by: Junhao Li <junhao@ubicloud.com >
2026-03-24 14:40:08 -07:00
Michael Goin and GitHub
0f0e03890e
[UX] Add flashinfer-cubin as CUDA default dep ( #37233 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-03-24 14:13:08 -07:00
Woosuk Kwon and GitHub
4b53740d7f
[MRV2] Fix for DS v3.2 ( #38030 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-24 14:03:24 -07:00
Nick Hill and GitHub
4e824d1c83
[Model Runner V2][Minor] Simplify PP logic ( #38031 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-03-24 13:57:17 -07:00
amey asgaonkar and GitHub
0c1809c806
Add Ubuntu 24.04 support for Docker builds ( #35386 )
...
Signed-off-by: aasgaonkar <aasgaonkar@nvidia.com >
2026-03-24 13:34:44 -07:00
liangel-02 and GitHub
8c47fdfdb1
[FlexAttention] allow custom mask mod ( #37692 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-03-24 16:03:24 -04:00
Javier De Jesus and GitHub
54b0578ada
[Bugfix] Pass hf_token through config loading paths for gated model support ( #37920 )
...
Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com >
2026-03-24 15:22:05 -04:00
Richard Zou and GitHub
89f572dbc0
[BugFix] fix VLLM_USE_STANDALONE_COMPILE=0 ( #38015 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-03-24 19:08:26 +00:00
Richard Zou and GitHub
71a4a2fbd0
[BugFix] Fix order of compile logging ( #38012 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-03-24 18:58:18 +00:00
Nick Cao and GitHub
935c46dd9b
[Model] Add Granite 4.0 1B speech to supported models ( #38019 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
2026-03-24 18:23:41 +00:00
057fc94cbd
[Bugfix] Fix structured output crash on CPU due to pin_memory=True ( #37706 )
...
Signed-off-by: Willy Hardy <whardy@redhat.com >
Signed-off-by: Will Hardy <whardy@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-24 17:44:17 +00:00
b58c5f28aa
docs: fix broken offline inference paths in documentation ( #37998 )
...
Signed-off-by: Vineeta Tiwari <vineeta.tiwari2@ibm.com >
Signed-off-by: Vineeta Tiwari <vineetatiwari2000@gmail.com >
Co-authored-by: Vineeta Tiwari <vineeta.tiwari2@ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-24 17:35:14 +00:00
Ming Yang and GitHub
c07e2ca6e0
Fix Mamba state corruption from referencing stale block table entries ( #37728 ) ( #37728 ) ( #37728 )
2026-03-24 10:29:59 -07:00
4df5fa7439
[Bugfix] Force continuous usage stats when CLI override is enabled ( #37923 )
...
Signed-off-by: Your Name <you@example.com >
Co-authored-by: Your Name <you@example.com >
Co-authored-by: OpenCode <noreply@openai.com >
2026-03-24 10:29:50 -07:00
sihao_li and GitHub
a5416bc52e
[XPU] Support Intel XPU hardware information collection in usage stats ( #37964 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
2026-03-24 10:29:17 -07:00
Harry Mellor and GitHub
b3601da6e7
[Mypy] Fix mypy for vllm/model_executor (except vllm/model_executor/layers) ( #37904 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-24 17:14:01 +00:00
dc78c2c933
[Core] add option to schedule requests based on full ISL ( #37307 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-03-24 13:01:12 -04:00
4731884796
[Feature] limit thinking tokens (hard limit) ( #20859 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-24 09:53:07 -07:00
Harry Mellor and GitHub
8de5261e69
Update new contributor message ( #37999 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-24 16:01:41 +00:00
1b6cb920e6
[Deprecate] Deprecate pooling multi task support. ( #37956 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-03-24 14:07:47 +00:00
Li, Jiang and GitHub
352b90c4a4
[Bugfix] Add replacement of _compute_slot_mapping_kernel on CPU ( #37987 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-03-24 07:00:20 -07:00
Sage and GitHub
1c0aabdeb0
[Bugfix] Suppress spurious CPU KV cache warning in launch render ( #37911 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-03-24 12:36:18 +00:00
Ilya Markov and GitHub
14acf429ac
[EPLB] Remove main waits in case of slow EPLB ( #36271 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
2026-03-24 11:50:44 +00:00
Harry Mellor and GitHub
ce57fd5557
[Docs] Fix build ( #37991 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-24 03:20:49 -07:00
Flora Feng and GitHub
2e67fa756d
Fix tool_parser_cls type annotation from Callable to type[ToolParser] ( #37957 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-23 22:58:27 -07:00
Ronen Schaffer and GitHub
e3c6c10cad
[KV Offload] Refactor CPU offloading: pluggable CachePolicy, remove Backend abstraction, restructure into cpu/ package ( #37874 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-03-24 07:02:51 +02:00
jetxa and GitHub
16a664df24
[Frontend][Bugfix] Pass default_chat_template_kwargs to AnthropicServingMessages ( #37899 )
...
Signed-off-by: jetxa <jetxzhang@outlook.com >
2026-03-24 05:00:12 +00:00
Kevin H. Luu and GitHub
7281199a8c
[release] Move agent queue to Release cluster queues ( #37783 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-03-23 20:36:47 -07:00
Kevin H. Luu and GitHub
b2dd75eb48
Downsize CPU jobs to use small queue ( #37913 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-03-23 20:36:37 -07:00
Wentao Ye and GitHub
c59a132f96
[V0 Deprecation] Refactor kv cache from list to element ( #37487 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-23 20:10:11 -07:00
Andreas Karatzas and GitHub
de99d91ece
[ROCm][CI] Split Entrypoints Integration (API Server 1) into 3 jobs ( #37906 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-24 09:48:37 +08:00
Wentao Ye and GitHub
83c9d525b6
[CI] Add batch invariant test: Block FP8 + small MOE ( #37895 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-23 21:16:14 -04:00
Giancarlo Delfin and GitHub
8f4824b664
[Model Runner V2] Gather multimodal embeddings before draft model postprocess ( #37932 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-23 18:14:13 -07:00
roikoren755 and GitHub
56777b5c89
[Test] E2E Nemotron-3-Super tests ( #36803 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-03-23 17:49:56 -07:00
2488a82f89
[CI] Split V1 Others into 3 separate jobs ( #37016 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-24 06:44:38 +08:00
dc6908ac6a
[Bugfix] Register VLLM_BATCH_INVARIANT in envs.py to fix spurious unknown env var warning ( #35007 )
...
Signed-off-by: Ranran <1012869439@qq.com >
Signed-off-by: Ranran <hzz5361@psu.edu >
Signed-off-by: ran <hzz5361@psu.edu >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-03-23 18:31:14 -04:00
yzong-rh and GitHub
e85f8f0932
[Bug][MoE] Strengthen _supports_current_device() checks in the TRTLLM FP8, NVFP4, and FlashInfer CuteDSL MoE experts ( #36728 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-03-23 17:02:57 -04:00
5bf3c42d4c
[Bug][MoE] Fix TRTLLM NVFP4 Routing Kernel Precision ( #36725 )
...
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
2026-03-23 20:19:06 +00:00
Kyle Sayers and GitHub
38364a7e32
[Sparse24] [Deprecation] Remove Sparse24 CT integration and kernels ( #36799 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-03-23 16:03:29 -04:00
fafe76b4af
[Async][Spec Decoding] Zero-bubble async scheduling + spec decoding ( #32951 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
Co-authored-by: zhrrr <43847754+izhuhaoran@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-03-23 15:37:22 -04:00
ffb5b32b5f
[MRV2] Consider spec decoding in warmup ( #37812 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-03-23 17:45:43 +00:00
Kunshang Ji and GitHub
91fd695b75
[CI] split Entrypoints Integration (API Server 1) into 3 jobs ( #37882 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-23 10:37:56 -07:00
Nicolò Lucchesi and GitHub
1cbbcfe8a3
[CI][PD] Add Hybrid SSM integration tests to CI ( #37657 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-03-23 23:58:19 +08:00
Angela Yi and GitHub
aceadb5ee1
Use lazy graph module during split_module to defer recompile() ( #37609 )
...
Signed-off-by: angelayi <yiangela7@gmail.com >
2026-03-23 11:21:29 -04:00
Yufeng He and GitHub
ec2280611a
[Bugfix] Fix RoBERTa position_ids accumulation on CUDA graph padding ( #37884 )
2026-03-23 15:15:12 +00:00
yanghui1-arch and GitHub
7151ae6528
[Bugfix] RoBERTa position_id accumulation in CUDA graph padding region ( #37873 )
...
Signed-off-by: dass90 <3053034939@qq.com >
2026-03-23 14:59:21 +00:00
Wentao Ye and GitHub
45bd5c8e75
[Mypy] Fix mypy for vllm/config ( #37808 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-23 14:33:59 +00:00
10a1018c12
[ROCm] fix sleep mode not releasing GPU memory problem on ROCm ( #37533 )
...
Signed-off-by: bingzhaodong <aaab8b@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-03-23 06:07:19 -07:00
Jee Jee Li and GitHub
aec2dc6c0d
[Bugfix][LoRA] Fix incorrect LoRA Log ( #37877 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-23 11:42:52 +00:00
DorBernsohn and GitHub
7938d12119
[Bugfix] Fix CPU backend crash in KV cache block zeroing ( #37550 )
...
Signed-off-by: DorBernsohn <dor.bernsohn@gmail.com >
2026-03-23 11:35:45 +00:00
Kunshang Ji and GitHub
debd6e768c
[XPU][MoE Refactor] Refactor xpu mxfp4 support into oracle ( #37784 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-23 11:10:41 +00:00
Andrew Xia and GitHub
9ace378a63
[Frontend][Responses API] Fix arrival_time recording for TTFT on initial request ( #37498 )
...
Signed-off-by: Andrew Xia <axia@meta.com >
2026-03-23 09:58:08 +00:00
Kunshang Ji and GitHub
27d5ee3e6f
[FP8]add FP8 WoQ kernel abstraction. ( #32929 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-03-23 09:47:47 +00:00
wangxiyuan and GitHub
35141a7eed
[Misc]Update gitignore ( #37863 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-03-23 01:14:10 -07:00
Chuan (Richard) Li and GitHub
e99fb98867
[ROCm] Fix fused_moe_fake signature mismatch and other AITER bugs ( #36100 )
...
Signed-off-by: Li <chuali@amd.com >
2026-03-23 15:48:31 +08:00
Artem Perevedentsev and GitHub
a16133a0f1
[Perf] [Bugfix] Fix Triton autotuning in inference for Qwen3.5 ( #37338 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-03-23 00:37:58 -07:00
54ab804e87
[Bugfix] Store Qwen3Next A_log in fp32 ( #37810 )
...
Signed-off-by: effortprogrammer <yhjhoward7@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-03-23 15:36:57 +08:00
02e6efe56d
[Bugfix] JAIS: Only apply ALiBi when position_embedding_type='alibi' ( #37820 )
...
Co-authored-by: r266-tech <r266-tech@users.noreply.github.com >
2026-03-23 07:36:34 +00:00
410d300893
[ROCm][Refactor] Enable AWQMarlinConfig on ROCm to use choose_mp_linear_kernel ( #36505 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-03-23 15:36:08 +08:00
Yan Ma and GitHub
d3fe857135
update doc for online fp8 quantization ( #37851 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-03-23 05:19:03 +00:00
Baorun (Lauren) Mu and GitHub
f85e479e66
[Feature] ViT Full CUDA Graph ( #35963 )
...
Signed-off-by: Baorun Mu <bmu@nvidia.com >
2026-03-23 13:01:10 +08:00
Jee Jee Li and GitHub
1f0d210641
[CI/Build][LoRA] Update Qwen35 LoRA testing ( #37816 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-23 12:55:49 +08:00
Ben Browning and GitHub
3bbe2e1e6e
[Test] Consolidate tool parser unit tests to tests/tool_parsers ( #37834 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-03-23 04:24:25 +00:00
6e04e79326
always use embed&token_classify for bge-m3 ( #37632 )
...
Signed-off-by: augusto.yjh <augusto.yjh@antgroup.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-23 03:10:57 +00:00
Lasha Koroshinadze and GitHub
e7767eccae
Fix AudioFlamingo3/MusicFlamingo HF parity and RoTE handling ( #37643 )
...
Signed-off-by: Lasha <26011196+lashahub@users.noreply.github.com >
2026-03-23 10:29:07 +08:00
Woosuk Kwon and GitHub
43877a620b
[MRV2] Enable PP CUDA graph test ( #37830 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-22 16:30:25 -07:00
63f49b8bd4
[Model Runner V2] Enable piecewise CUDA graphs for pipeline parallelism ( #35162 )
...
Signed-off-by: Zhanqiu Hu <zh338@cornell.edu >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-22 20:48:25 +00:00
Woosuk Kwon and GitHub
a5e9d511de
[MRV2] Use FP64 for Gumbel noise ( #37798 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-22 12:28:10 -07:00
Yongye Zhu and GitHub
c058ff44d4
[Bigfix]fix lora test by pass padded size back to the layer ( #37811 )
2026-03-22 13:20:13 -06:00
Woosuk Kwon and GitHub
ce9b1d76cf
[MRV2] Skip hidden states allocation for PW CUDA graphs ( #37818 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-22 11:47:21 -07:00
Netanel Haber and GitHub
e74c17e153
Enable NemotronHPuzzle + NemotronHMTP ( #37803 )
2026-03-22 15:13:58 +00:00
Wentao Ye and GitHub
eaf4978621
[Test] Only Run MLA model when user explicitly set for batch invariance ( #37719 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-22 09:09:12 -04:00
Wentao Ye and GitHub
77d24c4bfe
[Bug] Fix fp8 deepgemm batch invariant ( #37718 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-22 08:57:20 -04:00
b3e846017d
[Model Runner V2] Support multi-modal embeddings for spec decode model ( #36097 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-22 02:48:43 -07:00
Andreas Karatzas and GitHub
cd1242d82a
[ROCm][CI] Stabilize ROCm speech-to-text translation test with lower min acc threshold ( #37723 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 17:32:08 +08:00
Robert Shaw and GitHub
4383f1532e
[MoE] Move PF Methods to Folder ( #35927 )
2026-03-22 02:42:59 -06:00
Andreas Karatzas and GitHub
6eedec6e36
[ROCm][CI] Make some duplicated tests optional so that they are only evaluated in our nightly ( #37780 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 16:03:18 +08:00
Andreas Karatzas and GitHub
ffc8531524
[ROCm][CI] Added missing resampy dependency for MM audio tests ( #37778 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 16:02:41 +08:00
Andreas Karatzas and GitHub
6ecba840d7
[ROCm][CI] get_cu_count was renamed to num_compute_units in #35042 ( #37764 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 16:02:21 +08:00
Andreas Karatzas and GitHub
3b06c55c78
[ROCm][CI] Fix MEGA_AOT_ARTIFACT fallback when PyTorch < 2.10.0 lacks AOT support ( #37763 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 16:02:03 +08:00
Yang Liu and GitHub
b050700462
[Perf] Optimize glm4.xv VIT ( #37779 )
...
Signed-off-by: Yang <lymailforjob@gmail.com >
2026-03-22 06:12:34 +00:00
Andreas Karatzas and GitHub
5dac719b2b
[Bugfix] Handle libsndfile sf_error(NULL) race condition in audio fallback ( #37782 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 13:37:29 +08:00
Andreas Karatzas and GitHub
c862481c02
[CI] Skip ISAAC multimodal tests due to broken upstream HF model weights ( #37781 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 13:23:32 +08:00
Andreas Karatzas and GitHub
c86b17cfe6
[ROCm][CI] Add large_gpu_mark to test_max_tokens_none for ROCm ( #37717 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 12:25:16 +08:00
Andreas Karatzas and GitHub
66f927f205
[Bugfix] Fix pooling non-determinism from pinned prompt_lens aliasing ( #37775 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 03:22:24 +00:00
Andreas Karatzas and GitHub
e78bc74268
[ROCm][CI] close missing quote in kernels/moe block in run-amd-test.sh ( #37774 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-22 09:42:34 +08:00
Robert Shaw and GitHub
6b2fa3a762
[MoE] Move FlashInfer CuteDSL experts into fused_moe/experts/ ( #37759 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
2026-03-21 19:15:16 -04:00
eeee5b262d
[Quantization][Deprecation] Remove PTPC FP8 ( #32700 )
...
Signed-off-by: Robert Shaw <robshaw@redhat.com >
Co-authored-by: Robert Shaw <robshaw@redhat.com >
2026-03-21 22:10:16 +00:00
Robert Shaw and GitHub
5ad0446572
Revert "Consolidate AWQ quantization into single awq_marlin.py file" ( #37768 )
2026-03-21 17:20:41 -04:00
Robert Shaw and Claude
8cc700dd6a
Consolidate AWQ quantization into single awq_marlin.py file
...
Merge awq.py and awq_marlin.py into a single file, eliminating the
circular import between them. awq.py becomes a backward-compat shim.
Follows the same structure as gptq_marlin.py.
Co-authored-by: Claude
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
2026-03-21 17:09:17 -04:00
80b70884eb
Add tensor IPC transfer mechanism for multimodal data ( #32104 )
...
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com >
Signed-off-by: Brandon Pelfrey <brandonpelfrey@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-03-21 20:10:20 +00:00
Mohammad Miadh Angkad and GitHub
61e381dcf0
[Perf] Add SM 10.3 (B300/GB300) all-reduce communicator tuning ( #37756 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-03-21 19:43:47 +00:00
Mohammad Miadh Angkad and GitHub
88f1b374f5
[Core] Enable allreduce fusion by default for SM 10.3 (B300/GB300) ( #37755 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-03-21 19:40:37 +00:00
Francesco Fusco and GitHub
298e510848
[Hybrid] calling get_mamba_groups() once at MambaCopyBuffers.create() ( #37318 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
2026-03-21 09:29:43 +00:00
3982bc2cd0
[ROCm] Enable DeepEP ROCm as all2allbackend for AMD GPUs. ( #34692 )
...
Signed-off-by: Tej Kiran <vpolamre@amd.com >
Co-authored-by: Tej Kiran <vpolamre@amd.com >
2026-03-21 00:32:31 -07:00
Andreas Karatzas and GitHub
02eec7ecbe
[ROCm][CI] Update GSM8K eval config to use fp8-and-mixed models list (MI355) ( #37721 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-21 15:27:12 +08:00
17ee641c45
[Responses API] Add kv_transfer_params for PD disaggregation ( #37424 )
...
Signed-off-by: bongwoobak <bongwoobak@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-03-21 13:48:54 +08:00
Andreas Karatzas and GitHub
0d50fa1db6
[ROCm][CI] Mark gemma3 as large GPU test to avoid OOM on MI250 ( #37610 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-21 12:57:25 +08:00
Simon Mo and GitHub
1fa1e53a73
Revert "[compile] Initialize passes at VllmBackend init" ( #37733 )
2026-03-20 21:35:49 -07:00
Andreas Karatzas and GitHub
3ffa52009f
[ROCm][CI] Guard CudaPlatform/RocmPlatform imports to fix test collection on cross-platform builds ( #37617 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-21 11:58:58 +08:00
87bd91892f
[MoE Refactor] Mxfp4 oracle rebased ( #37128 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-21 03:37:04 +00:00
Isotr0py and GitHub
c7f98b4d0a
[Frontend] Remove librosa from audio dependency ( #37058 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-03-21 11:36:15 +08:00
tmm77 and GitHub
1c472f8fe1
Add get_device_uuid for rocm ( #37694 )
...
Signed-off-by: Tiffany Mintz <Tiffany.Mintz@amd.com >
2026-03-21 11:33:16 +08:00
c57d38d603
elastic_ep: Fix issues with repeated scale up/down cycles ( #37131 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
Co-authored-by: Ron Tourgeman <rtourgeman@nvidia.com >
2026-03-20 23:13:02 +00:00
Kaihang Jiang and GitHub
e5ed6c6c13
[BugFix] Allow qk_nope_head_dim=192 in FlashInfer MLA backend checks ( #37475 )
...
Signed-off-by: Kaihang Jiang <kaihangj@nvidia.com >
2026-03-20 16:14:55 -06:00
Wentao Ye and GitHub
b3d0b37908
[Refactor] Remove unused dead code ( #36171 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-20 16:12:51 -06:00
Santino Ramos and GitHub
85f671b8e1
[Model Runner V2] Support Streaming Inputs ( #37028 )
...
Signed-off-by: Santino Ramos <elsantinoramos@gmail.com >
2026-03-20 20:42:25 +00:00
Andreas Karatzas and GitHub
8bc6b5cdb0
[ROCm][CI] Setting some mi325_4 tests back to optional (in parity with upstream) ( #37711 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 12:25:08 -07:00
Vadim Gimpelson and GitHub
4f16ebbbd3
[Bugfix] Disable monolithic TRTLLM MoE for Renormalize routing ( #37591 ) ( #37605 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-03-20 12:19:26 -07:00
Angela Yi and GitHub
12fd17eb51
[compile] Initialize passes at VllmBackend init ( #35216 )
...
Signed-off-by: angelayi <yiangela7@gmail.com >
2026-03-20 11:40:33 -07:00
Cyrus Leung and GitHub
37aadf6237
[Model] Update Kimi-K25 and Isaac processors to fit HF-style ( #37693 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-20 18:30:22 +00:00
Le Yang and GitHub
d7d2b5e405
[Bugfix] Disable --calculate-kv-scales for hybrid GDN/Mamba+Attention… ( #37565 )
...
Signed-off-by: Young-Leo <562593859@qq.com >
2026-03-20 18:28:34 +00:00
6ec5e9fd37
refactor: abstract deepgemm support into platform ( #37519 )
...
Co-authored-by: sherryC41 <sherry.c.c41@gmail.com >
2026-03-20 17:54:08 +00:00
Lucas Wilkinson and GitHub
e1d85e5c24
[Attention] Support distinguishing between short extends and decodes ( #37303 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-03-20 10:49:36 -07:00
79eb9369c5
fix CUDAGraph memory being counted twice ( #37426 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
Signed-off-by: Peter Pan <peter.pan@daocloud.io >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-20 17:36:32 +00:00
Woosuk Kwon and GitHub
e80cfe575d
[MRV2] Avoid recompilation of _gather_block_tables_kernel ( #37645 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-20 10:31:45 -07:00
Xin Yang and GitHub
d0532bf38d
[Perf] Eliminate redundant SparseMatrix creation in gpt_oss_triton_kernels ( #37683 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-03-20 11:28:41 -06:00
Andreas Karatzas and GitHub
fb4e8bf442
[ROCm][CI] Fix accuracy for llama-nemotron-vl pooling tests ( #37613 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 10:16:59 -07:00
Harry Mellor and GitHub
6ade4bc5a5
Fix various config related issues for Transformers v5 ( #37681 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-20 16:30:12 +00:00
Zhengxu Chen and GitHub
2e089b96a8
[compile] Add compiled artifact counter for VLLM_USE_MEGA_AOT_ARTIFACT=1. ( #37589 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-03-20 16:22:46 +00:00
Martin Hickey and GitHub
880be2b1b8
[Metrics] Some small refactoring for better maintainability ( #33898 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-03-20 16:11:34 +00:00
Zhengxu Chen and GitHub
c0f5fae601
[compile] Fix aot test failures with torch 2.12. ( #37604 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-03-20 16:06:29 +00:00
Rémi Delacourt and GitHub
aa84e43ccb
[Pixtral] Enable Pixtral language model support Eagle3 ( #37182 )
...
Signed-off-by: remi <remi@mistral.ai >
2026-03-20 15:50:15 +00:00
Matthias Gehre and GitHub
5e806bcf54
[Bugfix] Fix ConchLinearKernel channelwise quantization (group_size=-1) ( #37329 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-03-20 10:32:21 -05:00
Matthias Gehre and GitHub
56a62c310c
[Bugfix] Reject channelwise quantization (group_size <= 0) in ExllamaLinearKernel ( #37331 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-03-20 10:31:57 -05:00
1779c09898
[ROCm] Enable wvSplitK skinny GEMM kernel for RDNA4/gfx1x decode ( #34709 )
...
Signed-off-by: L.B.R. <lbr@mmonad.com >
Co-authored-by: L.B.R. <lbr@mmonad.com >
2026-03-20 10:11:23 -05:00
xuebwang-amd and GitHub
44eea10f68
[ROCm][Quantization] make quark ocp mx dtype parser robust for weight-only quantization ( #36232 )
...
Signed-off-by: xuebwang-amd <xuebwang@amd.com >
2026-03-20 10:10:03 -05:00
Ilya Boytsov and GitHub
8b6c6b9505
[Model] Add LFM2-ColBERT-350M support ( #37528 )
...
Signed-off-by: Ilya Boytsov <ilyaboytsov1805@gmail.com >
2026-03-20 14:57:57 +00:00
Harry Mellor and GitHub
9f6d9dd371
Fix attribute error in isaac_patch_hf_runner ( #37685 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-20 14:49:40 +00:00
Jee Jee Li and GitHub
dd20ee4e3e
[UX] Enable torch_profiler_with_stack ( #37571 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-20 11:17:26 +00:00
Chauncey and GitHub
0523449c9c
[Misc] Use logger.info_once for auto tool choice log message ( #37661 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-20 10:40:36 +00:00
Flora Feng and GitHub
b4c1aef21c
[Refactor] Relocate tests from tests/v1/entrypoints/ to tests/entrypoints/ ( #37500 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-20 02:50:34 -07:00
Flora Feng and GitHub
6050b93bed
[Refactor] Move serve entrypoint tests under tests/entrypoints/serve/ ( #37595 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-20 02:10:47 -07:00
Andreas Karatzas and GitHub
5a4a179591
[ROCm][CI] Fix granite_speech test for gfx90a by selecting compatible attention backend ( #37611 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 17:07:26 +08:00
Andreas Karatzas and GitHub
37cd9fc107
[ROCm][CI] Remove deepep DBO tests on gfx90a ( #37614 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 17:07:07 +08:00
Andreas Karatzas and GitHub
9cfd4ebb5e
[ROCm][CI] Update GSM8K eval config to use fp8-and-mixed models list ( #37619 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 17:06:53 +08:00
wang.yuqi and GitHub
ed359c497a
[Model] Deprecate the score task (this will not affect users). ( #37537 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-20 08:07:56 +00:00
Giancarlo Delfin and GitHub
dcee9be95a
[Model Runner V2] Fix draft logits not populated during cudagraph replay ( #37639 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-20 07:43:47 +00:00
Andreas Karatzas and GitHub
bd8c4c0752
[CI] Removing deprecated rlhf examples reference ( #37585 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-20 15:20:33 +08:00
0140eafb15
[Bug] Fix FlashInfer allreduce fusion workspace uninitialized error ( #37461 )
...
Signed-off-by: root <root@prenyx0169.a51.clusters.nvidia.com >
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: <>
Co-authored-by: root <root@prenyx0169.a51.clusters.nvidia.com >
Co-authored-by: root <root@prenyx0042.a51.clusters.nvidia.com >
2026-03-20 03:09:21 -04:00
Kunshang Ji and GitHub
bdf6a0a57b
[XPU] bump vllm-xpu-kernels to v0.1.4 ( #37641 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-20 15:04:38 +08:00
0674d1fee7
[PluggableLayer][MM] Add PluggableLayer for CustomQwen2Decoder ( #37293 )
...
Signed-off-by: Wangbei25 <wangbei41@huawie.com >
Signed-off-by: Wangbei25 <wangbei41@huawei.com >
Co-authored-by: Wangbei25 <wangbei41@huawie.com >
2026-03-20 06:24:07 +00:00
Cyrus Leung and GitHub
30108fc8b0
[Model] Refactor Step3-VL processor to HF style ( #37579 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-20 06:05:08 +00:00
Flora Feng and GitHub
e2d1c8b5e8
[Refactor] Relocate entrypoint tests to match serving code structure ( #37593 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-20 05:31:23 +00:00
Huanxing and GitHub
6951fcd44f
[XPU] Automatically detect target platform as XPU in build. ( #37634 )
...
Signed-off-by: huanxing <huanxing.shen@intel.com >
2026-03-20 13:30:15 +08:00
Giancarlo Delfin and GitHub
39474513f6
[Model Runner V2] fix draft attention metadata generation ( #37364 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-19 21:05:15 -07:00
638a872d77
fix(xpu): Re-compute compile ranges after platform-specific config updates ( #37523 )
...
Signed-off-by: Yuxiang Liang <yuxiang.liang@intel.com >
Signed-off-by: Yuxiang Liang <yuliang@habana.ai >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-20 03:52:35 +00:00
Flora Feng and GitHub
9040151fe1
[V0 Deprecation] Deprecate --disable-frontend-multiprocessing ( #37612 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-20 11:31:43 +08:00
Jee Jee Li and GitHub
8fbe3f303f
[Bugfix][LoRA] Fix Qwen35 LoRA ( #36976 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-20 11:09:32 +08:00
Xiao and GitHub
ea2c148fa7
[compile][graph_partition]Add tensor size handling ( #36038 )
...
Signed-off-by: Xiao Fu <xiaofu@meta.com >
2026-03-19 19:55:25 -07:00
Tianmu Li and GitHub
47b7af0d87
[Feat] Enable CompressedTensorW4A8Int for XPU ( #37207 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
2026-03-20 02:34:28 +00:00
tianshu-Michael-yu and GitHub
269bf46d99
fix: disambiguate multimodal prefix cache keys ( #36708 )
...
Signed-off-by: tianshu.yu <tianshuyu.formal@gmail.com >
2026-03-20 10:33:20 +08:00
Flora Feng and GitHub
e5a77a5015
[CI] Update mergify tool-calling label paths ( #37478 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-20 02:22:23 +00:00
Itay Alroy and GitHub
ca1ac1a4b4
Fix DP coordinator ZMQ port TOCTOU ( #37452 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-03-20 00:58:31 +00:00
Divakar Verma and GitHub
4ca3fa6bb4
[ROCm][Bugfix] fix cache block size mismatch for aiter unified attention ( #37606 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-03-20 00:00:08 +00:00
Flora Feng and GitHub
be12afd284
[Bugfix] Fix Deepseekv32 tool parser when stream interval > 1 ( #36056 )
2026-03-19 19:51:25 -04:00
Wentao Ye and GitHub
df3c0291a3
[Bug] Fix EmbedIOprocessor "classify" <-> "embed" ( #37573 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-20 07:40:10 +08:00
Wentao Ye and GitHub
2be1a0f74b
[Refactor] Remove dead code in pooling model ( #37572 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-20 07:39:43 +08:00
4120a05ff1
Fix AttributeError in Qwen3.5 GDN layers with quantized models ( #37448 )
...
Signed-off-by: Jim Smith <jim@joshua8.ai >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Xin Yang <105740670+xyang16@users.noreply.github.com >
2026-03-19 19:21:14 -04:00
rasmith and GitHub
98ff042917
[CI][BugFix][AMD] Don't set VLLM_ROCM_USE_AITER anymore in test_rocm_aiter_topk since its not necessary ( #36996 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-03-20 07:12:45 +08:00
Artem Perevedentsev and GitHub
b55156eae9
[Performance] Enable Triton autotuning disk cache by default ( #37188 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-03-19 17:36:28 -04:00
Laith Sakka and GitHub
112944fab9
test Qwen/Qwen3-4B-Instruct-2507 for unbacked ( #36064 )
...
Signed-off-by: Laith Sakka <lsakka@meta.com >
2026-03-19 17:28:45 -04:00
bnellnm and GitHub
91be5f9be3
[MoE Refactor] Rename "naive" all2all backend ( #36294 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-03-19 15:50:34 -04:00
Aaron Hao and GitHub
4ee847e400
Comment fix for async rl example ( #35244 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
2026-03-19 19:46:07 +00:00
Andreas Karatzas and GitHub
040a505ff5
[ROCm][CI] Cleaning and restructuring amd-ci legacy pipeline ( #34839 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-19 14:30:58 -05:00
bnellnm and GitHub
9279c59a0e
[MoE Refactor] DefaultMoERunner simplifcation ( #33049 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-03-19 15:07:44 -04:00
Wentao Ye and GitHub
7454096199
[Log] Log once in local node by default ( #37568 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-19 12:04:59 -07:00
Andreas Karatzas and GitHub
fb8b5e05fc
[CI] Add retry with 4x backoff to HTTP fetches for transient failures ( #37218 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-19 19:00:20 +00:00
Harry Mellor and GitHub
e5d96dc8fc
Fix SpeculatorsConfig now that PreTrainedConfig is a dataclass in Transformers ( #37574 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 18:04:40 +00:00
daa05bf340
[Bugfix] Fix AttributeError when serving MXFP8 models with DeepGEMM installed ( #37358 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-19 17:58:33 +00:00
Lucas Kabela and GitHub
7769b58307
[torch.compile][BE][Multimodal] Remove requirement to set_model_tag to avoid cache conflict ( #37345 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-03-19 17:26:12 +00:00
Chauncey and GitHub
2f9f946b22
[P/D] AnthropicMessages add kv_transfer_params for PD disaggregation ( #37535 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-19 16:41:20 +00:00
Fadi Arafeh and GitHub
2890aecce5
[CPU][UX] Do not crash when tcmalloc/libiomp are not ldpreloaded ( #37561 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-03-19 16:35:45 +00:00
Harry Mellor and GitHub
34f093b417
[CI] Gate pre-commit on ready label or number of contributions ( #37544 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 16:21:57 +00:00
Harry Mellor and GitHub
4dce8321a9
Run MacOS smoke test on daily cron job instead of every commit ( #37567 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 16:19:50 +00:00
Cyrus Leung and GitHub
657855ab41
[Misc] Cleanup more configs and processors ( #37560 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-19 15:45:23 +00:00
Wei Zhao and GitHub
e27b8ba3d1
[Bug] Fix fp8 trtllm MoE modular kernel supported routing methods ( #37346 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-03-19 11:43:06 -04:00
Woosuk Kwon and GitHub
40b8363b45
[MRV2] Use fp32 for draft logits ( #37526 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-19 08:41:21 -07:00
mikaylagawarecki and GitHub
8b10e4fb31
[1/n] Migrate permute_cols to libtorch stable ABI ( #31509 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-03-19 11:27:26 -04:00
+3
104605cbf2
Remove deprecated reasoning_content message field(part-2) ( #37480 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: Ifta Khairul Alam Adil <ikaadil007@gmail.com >
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Philip Ottesen <phiott256@gmail.com >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Andy Lo <andy@mistral.ai >
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Signed-off-by: sihao.li <sihao.li@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: JartX <sagformas@epdcenter.es >
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Philip Ottesen <phiott256@gmail.com >
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Andy Lo <andy@mistral.ai >
Co-authored-by: Thillai Chithambaram <79466435+thillai-c@users.noreply.github.com >
Co-authored-by: sihao_li <165983188+1643661061leo@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 15:20:08 +00:00
96266f119b
[LoRA] Minor improvements to LoRA log ( #37557 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-03-19 15:18:06 +00:00
Sage Moore and GitHub
7c0cf3bcd0
Cap the number of API servers to 1 when using Elastic EP. ( #37466 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-03-19 10:42:57 -04:00
Harry Mellor and GitHub
572b432913
Stop bench CLI from recursively casting all configs to dict ( #37559 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 14:04:03 +00:00
Cyrus Leung and GitHub
9515c20868
[Misc] Clean up processing logic ( #37541 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-19 13:30:20 +00:00
DorBernsohn and GitHub
c63ca2b2e6
[Bugfix] Add Kimi-K2.5 reasoning/tool parser aliases and tool_call_id support ( #37438 )
...
Signed-off-by: DorBernsohn <dor.bernsohn@gmail.com >
2026-03-19 21:08:00 +08:00
Harry Mellor and GitHub
a32eaf5bb2
[CI] Merge cleanup_pr_body.yml and reminder_comment.yml ( #37552 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 12:55:07 +00:00
e390742c59
Fix KV Offloading + MLA AssertionError by using num_kv_heads=1 in cpu… ( #37536 )
...
Signed-off-by: xueliangyang-oeuler <yxl546827391@gmail.com >
Co-authored-by: xueliangyang-oeuler <yxl546827391@gmail.com >
2026-03-19 12:05:07 +00:00
Cyrus Leung and GitHub
7a6ebcbfcf
[Model] Remove unnecessary get_language_model ( #37545 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-19 20:00:36 +08:00
Cyrus Leung and GitHub
c7bc12c20f
[CI/Build] Split out MM pooling tests ( #37542 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-19 11:36:11 +00:00
f9e2a38386
[Docs] Reorganize pooling docs. ( #35592 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 11:25:47 +00:00
Harry Mellor and GitHub
4426447bba
Don't log exc_info when vLLM tries to doenload a file that doesn't exist ( #37458 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-19 10:38:29 +00:00
Li, Jiang and GitHub
3322e26420
[Bugfix] Avoid more OpenMP thread reallocation in CPU torch compile ( #37538 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-03-19 10:24:39 +00:00
Cyrus Leung and GitHub
765e461065
[Bugfix] Fix Nemotron Parse loading ( #37407 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-19 09:55:29 +00:00
Duyi-Wang and GitHub
6a9cceb219
[Bugfix][ROCm] Fix MoRI + AITER FP8 dispatch compatibility for defer_input_quant ( #37418 )
...
Signed-off-by: Duyi-Wang <duyi.wang@amd.com >
2026-03-19 09:49:27 +00:00
yassha and GitHub
199f914183
fix(cpu): add null check for aligned_alloc in ScratchPadManager ( #37369 )
...
Signed-off-by: yassha <50112520+yassha@users.noreply.github.com >
2026-03-19 17:45:06 +08:00
Kunshang Ji and GitHub
ca21483bf9
[MISC] fix pin_memory=torch.cuda.is_available(), use is_pin_memory_available ( #37415 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-19 09:23:24 +00:00
TJian and GitHub
da70c87e81
[CI] Fix wrong path test file, missing rlhf_async_new_apis.py ( #37532 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-03-19 02:21:55 -07:00
Collin McCarthy and GitHub
0b6d52629f
Support temporal compression for Nemotron-3-VL videos ( #36808 )
...
Signed-off-by: Collin McCarthy <cmccarthy@nvidia.com >
2026-03-19 08:02:19 +00:00
Ziming Huang and GitHub
d3cc379567
[Perf] Fix slow hasattr in CUDAGraphWrapper.__getattr__ ( #37425 )
...
Signed-off-by: 智鸣 <hzm414167@alibaba-inc.com >
2026-03-19 15:43:48 +08:00
cdpath and GitHub
354cd580d5
fix(anthropic): remove non-standard 'data: [DONE]' from Anthropic streaming ( #37510 )
...
Signed-off-by: cdpath <cdpath@outlook.com >
2026-03-19 07:23:35 +00:00
zhanqiuhu and GitHub
d49f273144
[SSM/Mamba] Follow-up: N-1 prefill for P/D disaggregation ( #37310 )
2026-03-19 08:22:00 +01:00
Flora Feng and GitHub
b21d384304
[Refactor] Relocate endpoint tests to mirror serving code directory structure ( #37504 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-19 07:19:36 +00:00
Hongxia Yang and GitHub
e3126cd107
[ROCm] issue management - request information for bug issues on ROCm ( #37009 )
...
Signed-off-by: Hongxia Yang <hongxiay.yang@amd.com >
2026-03-19 03:51:29 +00:00
Wentao Ye and GitHub
e37ff5b5c8
[Perf] Optimize token_embed for pooling models, 1.0% token throughput improvement ( #37347 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-19 10:27:51 +08:00
Aaron Hao and GitHub
6accb21f2a
[bug] Fix deadlock with pause resume and collective_rpc ( #37024 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
2026-03-19 01:49:02 +00:00
Giancarlo Delfin and GitHub
053f3b6309
[Model Runner V2] Spec decode rejection sampler logprobs support ( #37237 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-19 01:36:27 +00:00
Aaron Hao and GitHub
5f82706a21
[BUG] Exclude SKIP_TENSORS from get_layer_size() + new weight sync example for dpep ( #37334 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
2026-03-19 00:45:10 +00:00
c32a58cc2a
[EPLB] Simplify EPLB rearrange by only returning one map ( #36267 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-03-18 20:34:00 -04:00
ef2c4f778d
[Bugfix] Zero-init MLA attention output buffers to prevent NaN from CUDA graph padding ( #37442 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-03-19 00:28:37 +00:00
sihao_li and GitHub
9dade5da3a
[XPU]Unify xpu test dependencies in dockerfile.xpu ( #36477 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
2026-03-19 08:12:07 +08:00
Thillai Chithambaram and GitHub
828f862acb
[Bugfix] Expand quantization method support in perf metrics ( #37231 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
2026-03-18 23:54:19 +00:00
Andy Lo and GitHub
577df69b26
[Bugfix] Fix KV scales inconsistency in fp8 MLA & FlashInfer kv_cache_dtype "auto" leading to gibberish ( #37054 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-03-18 23:07:29 +00:00
Giancarlo Delfin and GitHub
04244fd0e1
[Model Runner V2] Spec decode rejection sampler greedy support ( #37238 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-03-18 15:59:03 -07:00
Michael Goin and GitHub
9482b0b085
[Bugfix] Remove assertion for NVFP4 scale dynamic range ( #37465 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-03-18 15:37:49 -07:00
Woosuk Kwon and GitHub
5bc1da147f
[LoRA][BugFix] Fix skipped LoRA adapters for Mistral3 ( #36928 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-03-18 22:34:19 +00:00
Philip Ottesen and GitHub
0091017188
fix(worker): optimize swap_states to copy only active token prefixes ( #34733 )
...
Signed-off-by: Philip Ottesen <phiott256@gmail.com >
2026-03-18 14:59:27 -07:00
Wentao Ye and GitHub
0d81a1fe61
[V0 Deprecation] Deprecate virtual engine ( #37195 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-18 14:30:14 -07:00
Netanel Haber and GitHub
6ae4c8d6fc
chunk parakeet into 30s clips to prevent OOMs on long audios ( #36671 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-03-18 14:22:24 -07:00
JartX and GitHub
a913b612d8
[Bugfix] Fix ROCm crash in qwen3_next multi-stream events ( #36795 ) ( #37427 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-03-18 16:06:31 -04:00
Harry Mellor and GitHub
5ce2d10e4a
Fix models which use layer_type_validation for Transformers v5 ( #37398 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-18 18:41:51 +00:00
Chengyu Fang and GitHub
738d0a281f
[Bugfix] Fix incorrect use of merge_size in Qwen3-VL video timestamp calculation ( #37439 )
...
Signed-off-by: chengyufang <cnyvfang@outlook.com >
2026-03-18 11:36:34 -07:00
youkaichao and GitHub
70b81c4f3d
[bugfix][async scheduling] fix extra cuda context in device 0 with EP/DP ( #37449 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
2026-03-18 18:32:30 +00:00
Cyrus Leung and GitHub
7476d148db
[Model] Remove unnecessary processor definition for Nemotron Parse ( #37456 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-18 18:25:13 +00:00
Cyrus Leung and GitHub
f3732bd931
[Misc] Clean up model registry ( #37457 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-18 18:24:44 +00:00
Wentao Ye and GitHub
0ef7f79054
[Perf] Add tuned triton moe config for Qwen3.5 H200, 9.9% E2E throughput improvement ( #37340 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-18 14:18:34 -04:00
Or Ozeri and GitHub
5dd8df0701
[kv_offload+HMA][2/N]: Support multiple KV groups in GPULoadStoreSpec ( #36642 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-03-18 19:26:40 +02:00
Harry Mellor and GitHub
39bfb57b7c
Add API docs link if the CLI arg is a config class ( #37432 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-18 17:19:35 +00:00
RonaldBXu and GitHub
c9d838fc33
Adding deterministic lora benchmarking to vLLM Bench ( #36057 )
...
Signed-off-by: Ubuntu <ubuntu@ip-172-31-43-201.ap-northeast-1.compute.internal >
Signed-off-by: Ronald Xu <ronaldxu@amazon.com >
2026-03-18 16:02:03 +00:00
Xin Yang and GitHub
b1169d7be8
[Kernel] Add gpt-oss Router GEMM kernel ( #37205 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-03-18 08:15:56 -07:00
17808394bc
standardize load_weights using AutoWeightsLoader for kimi_linear and minimax_text_01 ( #37371 )
...
Signed-off-by: XuLiu <xuliu40@gmail.com >
Co-authored-by: XuLiu <xuliu40@gmail.com >
2026-03-18 15:05:37 +00:00
elvischenv and GitHub
296839a1b0
[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE ( #30647 )
...
Signed-off-by: elvischenv <219235043+elvischenv@users.noreply.github.com >
2026-03-18 15:01:26 +00:00
Wentao Ye and GitHub
c373b5c00d
[Log] Reduce duplicate log ( #37313 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-18 10:57:44 -04:00
Itay Alroy and GitHub
de1a86b7de
elastic_ep: Fix stateless group port races ( #36330 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-03-18 14:36:18 +00:00
Cyrus Leung and GitHub
99267c23ca
[2/3] Refactor InternVL-based processors ( #37324 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-18 22:22:19 +08:00
Or Ozeri and GitHub
525f2eeb0b
[kv_offload+HMA][6/N]: Split offloading_connector.py ( #37405 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-03-18 14:42:46 +01:00
918b7890a1
[Bugfix] Fix base64 JPEG video frames returning empty metadata ( #37301 )
...
Signed-off-by: Yufeng He <40085740+universeplayer@users.noreply.github.com >
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Yufeng He <40085740+universeplayer@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-03-18 13:40:03 +00:00
Andy Lo and GitHub
98b09ddc27
[NIXL][Bugfix] metrics & testing minor bug ( #36051 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-03-18 14:39:14 +01:00
Shwetha Poojary and GitHub
cef1f302d2
[Model] Enable LoRA support for tower and connector in H2OVL ( #31696 )
...
Signed-off-by: shwetha-s-poojary <shwetha.s-poojary@ibm.com >
2026-03-18 13:26:47 +00:00
17c47fb869
[Bugfix] Fix EP weight filter breaking EPLB and NVFP4 accuracy ( #37322 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Kevin H. Luu <khluu000@gmail.com >
2026-03-18 18:30:29 +08:00
Chauncey and GitHub
b322b197f1
[Build] Bump python openai version ( #32316 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-18 18:20:10 +08:00
Andreas Karatzas and GitHub
eaf7c9b976
[CI] Fix PaddleOCR-VL HF test failure due to create_causal_mask API rename ( #37328 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-18 09:44:12 +00:00
47a1f11bff
[docs] Add docs for new RL flows ( #36188 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-18 09:04:26 +00:00
fad09e8a1f
fix(glm47): improve tool call parsing and content normalization ( #37386 )
...
Signed-off-by: karanb192 <karan@example.com >
Co-authored-by: karanb192 <karan@example.com >
2026-03-18 08:12:21 +00:00
Jee Jee Li and GitHub
8c31f47c63
[LoRA] Make LoRA respect language_model_only ( #37375 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-03-18 07:53:34 +00:00
Li, Jiang and GitHub
261801242f
[Bugfix] Avoid OpenMP thread reallocation in CPU torch compile ( #37391 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-03-18 07:51:39 +00:00
fcf0687b27
[kv_offload+HMA][0/N]: Support block-level preemption handling ( #34805 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-03-18 08:49:53 +02:00
86b7e3c95a
[XPU] skip unsupported ut and update test_nixl_connector ( #37179 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-03-18 13:32:59 +08:00
Andrew Xia and GitHub
0e95916155
[responsesAPI] parser.extract_response_outputs can take in token IDs ( #37130 )
...
Signed-off-by: Andrew Xia <axia@meta.com >
2026-03-18 05:31:31 +00:00
Andreas Karatzas and GitHub
ce2ef42fd3
[CI] Stabilize test_cpu_offloading by waiting for async offload before cache reset ( #37335 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-18 05:26:20 +00:00
Andreas Karatzas and GitHub
8b6325758c
[ROCm][CI] Add ROCM_EXTRA_ARGS to audio_in_video test server fixture ( #37349 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-18 04:55:40 +00:00
gxd3 and GitHub
a0dd1995c7
[Hardware][TPU] Add supports_async_scheduling() method to Executor interface so that it can be extended for Executor implementations. ( #36924 )
...
Signed-off-by: Guangxiang Du <gxd@google.com >
2026-03-18 12:53:28 +08:00
Xin Yang and GitHub
f1740006e4
[Perf] Enable dual stream execution of input projection for Qwen3 ( #36795 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-03-18 11:13:27 +08:00
Andreas Karatzas and GitHub
58cde5c026
[ROCm][CI] Skip trtllm kvfp8 dequant tests on ROCm ( #37330 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-18 11:12:26 +08:00
761e0aa7a0
[Performance] Add --enable-ep-weight-filter CLI option ( #37351 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-18 09:36:55 +08:00
ff9fbc9aff
[Kernel][Helion] [16/N] Refactor register_kernel API to be more Dynamo-friendly ( #36705 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-18 01:23:35 +00:00
Divakar Verma and GitHub
e6c4797704
[ROCm][Quantization] add fp8xfp8 attn support for rocm_aiter_unified_attn ( #36927 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-03-18 08:49:32 +08:00
Michael Goin and GitHub
09e4576f65
[Kernel] Add non-gated support for NVFP4 CUTLASS MoE ( #37320 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-03-17 18:12:04 -04:00
Andreas Karatzas and GitHub
3ed7b1e6e0
[ROCm] Validate block_size for explicitly selected attention backends ( #36846 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-17 17:04:40 -05:00
JartX and GitHub
e8f9dbc369
[Bugfix][ROCm] Fix worker startup OOM on ROCm by skipping unreliable cudagraph memory profiling ( #36720 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-03-17 17:55:34 -04:00
Yong Hoon Shin and GitHub
de35c06c66
Make KV connector metadata build overridable via plugin ( #37336 )
...
Signed-off-by: Yong Hoon Shin <yhshin@meta.com >
2026-03-17 21:29:06 +00:00
c0745a851a
[Model] Add ColQwen3.5 4.5B support ( #36887 )
...
Signed-off-by: Athrael Soju <athrael.soju@gmail.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-03-17 21:17:02 +00:00
Ekagra Ranjan and GitHub
b5ca9c3557
[Models] Cohere ASR ( #35809 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-03-17 21:04:17 +00:00
245758992e
[Bugfix] Rescale NVFP4 weight scales to fix BF16 dequant underflow ( #34577 )
...
Signed-off-by: ricky-chaoju <ricky.chen@infinirc.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-03-17 20:48:42 +00:00
1204cf0a9d
[Bugfix] Fix mock.patch resolution failure for standalone_compile.FakeTensorMode on Python <= 3.10 ( #37158 )
...
Signed-off-by: Dimitrios Bariamis <12195802+dbari@users.noreply.github.com >
Co-authored-by: Dimitrios Bariamis <12195802+dbari@users.noreply.github.com >
2026-03-17 20:13:06 +00:00
Wei Zhao and GitHub
b36adfa349
[Perf] Set Flashinfer sparse MLA as default backend for FP8 kv cache ( #37252 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-03-17 20:09:20 +00:00
Michael Goin and GitHub
e78821b438
[Deprecation] Deprecate --calculate-kv-scales option ( #37201 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-03-17 19:57:24 +00:00
Cyrus Leung and GitHub
51f0acda79
[Model] Remove unused handle_oov_mm_token ( #37321 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-17 19:44:52 +00:00
fa75204b16
bump compressed-tensors version to 0.14.0.1 ( #36988 )
...
Signed-off-by: Brian Dellabetta <bdellabe@redhat.com >
Co-authored-by: Dipika Sikka <dipikasikka1@gmail.com >
2026-03-17 15:36:19 -04:00
Wentao Ye and GitHub
bdb903bb5f
[Bug] Fix FlashInfer MNNVL socket collisions under concurrent vLLM jobs ( #36674 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-03-17 15:19:52 -04:00
Andrey Talman and GitHub
68f783a727
[Torch 2.11] Guard torch._C._cpu attribute checks for forward compatibility ( #35673 )
...
Signed-off-by: atalman <atalman@fb.com >
2026-03-17 18:47:59 +00:00
c5030c439d
[CI] Split Distributed Tests (4 GPUs) and Kernel MoE tests ( #37100 )
...
Signed-off-by: Avinash Singh <avinashsingh.rcoem@gmail.com >
Signed-off-by: Avinash Singh <107198269+avinashsingh77@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Kevin H. Luu <khluu000@gmail.com >
2026-03-17 11:44:55 -07:00
Michael Goin and GitHub
51b2333be1
[Perf] Optimize top-k search in apply_top_k_top_p_triton sampler ( #37225 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-03-17 11:35:17 -07:00
Andreas Karatzas and GitHub
4ed51308c8
[CI] Fix GPU memory leak when RemoteOpenAIServer fails to start in __init__ ( #37230 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-17 09:08:08 -07:00
Cyrus Leung and GitHub
c781fbbab3
[Bugfix] Standardize custom HF Processor init ( #37289 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-17 15:38:55 +00:00
Richard Zou and GitHub
979ff44cea
[BugFix] PyTorch Compilation Tests should error if any test fails ( #37300 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-03-17 15:26:38 +00:00
Benjamin Chislett and GitHub
f63ed7b5ac
[Bugfix] Fix DP MTP Dummy Run ( #35243 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-03-17 11:16:48 -04:00
Ning Xie and GitHub
c9e5096256
[openapi] remove redundant exception stack trace[4/N] ( #37157 )
...
Signed-off-by: Andy Xie <andy.xning@gmail.com >
2026-03-17 15:06:25 +00:00
2ff0ad9694
[UltraVox] Fix output type ( #37224 )
...
Signed-off-by: vasqu <antonprogamer@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-17 14:51:17 +00:00
Isotr0py and GitHub
a836524d20
[Chore] Replace all base64 usages with faster pybase64 package ( #37290 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-03-17 14:44:19 +00:00
3717a4dd47
[Misc][LoRA] Add --lora-target-modules to restrict LoRA to specific modules ( #34984 )
...
Signed-off-by: Bhoomit Vasani <bhoomit.2010@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-17 14:36:41 +00:00
Harry Mellor and GitHub
ecfcdd2ce4
Fix Phi3 test that fails with Transformers v5 ( #37298 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-03-17 14:29:24 +00:00
Siew's Capital Jarvis and GitHub
c25dbc2d27
[Bugfix] Fix unclean shutdown crash with AllReduce Fusion workspace ( #36955 )
...
Signed-off-by: Jarvis <brayden.stanley.0127@gmail.com >
2026-03-17 14:22:09 +00:00
Jonas M. Kübler and GitHub
77d2a5f17b
pick up tuned prefill configs for FP8 FA3 ( #36265 )
...
Signed-off-by: Jonas M. Kübler <44084297+jmkuebler@users.noreply.github.com >
Signed-off-by: Jonas Kuebler <kuebj@amazon.com >
2026-03-17 07:00:26 -07:00
Sage and GitHub
59192dfd39
[Frontend] Complete OpenAI render delegation ( #37287 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-03-17 13:53:55 +00:00
Umut Polat and GitHub
56cb1baa66
[Misc] Use VLLMValidationError in batch, pooling, and tokenize protocol validators ( #36256 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
2026-03-17 13:52:30 +00:00
Cyrus Leung and GitHub
f340324335
[1/2] Move InternVL-based processors ( #37260 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-03-17 21:50:56 +08:00
2660b9289c
Bugfix for offloading+prefetch for GLM-4.7-FP8 ( #37178 )
...
Signed-off-by: Benjamin Merkel <benjamin.merkel@tngtech.com >
Co-authored-by: Benjamin Merkel <benjamin.merkel@tngtech.com >
2026-03-17 21:22:09 +08:00
Viacheslav and GitHub
293f036e6d
Add gigachat 3.1 tool parser + fix gigachat3 tool parser ( #36664 )
...
Signed-off-by: Viacheslav Barinov <viacheslav.teh@gmail.com >
2026-03-17 12:03:20 +00:00
youkaichao and GitHub
0fb142a454
[perf][connector] optimize build_connector_meta when host buffer transfer is not used ( #37165 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
2026-03-17 11:59:35 +00:00
Sage and GitHub
00f8e0d211
[Frontend] Delegate tokenization serving preprocessing to OpenAIServingRender ( #37266 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-03-17 11:22:54 +00:00
zhao, zhenhui and GitHub
4af9ed21cb
[Bugfix](xpu): prevent “selected index k out of range” in TP decode path ( #37259 )
...
Signed-off-by: zhenzhao <zhenzhao@habana.ai >
2026-03-17 11:14:07 +00:00
Augusto Yao and GitHub
9c7cab5ebb
[Feature]: Support for multiple embedding types in a single inference call ( #35829 )
...
Signed-off-by: augusto.yjh <augusto.yjh@antgroup.com >
2026-03-17 17:05:42 +08:00
Chauncey and GitHub
132bfd45b6
[Bugfix][ResponsesAPI] Fix crash when tool_choice=required exceeds max_output_tokens ( #37258 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-03-17 08:54:52 +00:00
24b4272a8c
Fix infinite recursive search issue in quark.py ( #32779 )
...
Signed-off-by: Yanwen Lin <lyw1124278064@gmail.com >
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com >
Signed-off-by: kimheesu <wlskaka4@gmail.com >
Co-authored-by: Yanwen Lin <lyw1124278064@gmail.com >
Co-authored-by: Kim Hee Su <wlskaka4@gmail.com >
2026-03-17 07:19:15 +00:00
Benjamin Chislett and GitHub
8a680463fa
[Bugfix] Fix NemotronH MTP + Chunked Prefill ( #35447 )
2026-03-17 07:07:33 +01:00
Nick Cao and GitHub
20b14095a4
[Bugfix] Fix loading Music Flamingo ( #35535 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
2026-03-17 05:24:40 +00:00
17c1bdf371
[Bugfix] dtype mismatch in ngram gpu propose ( #37246 )
...
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
2026-03-17 05:19:55 +00:00
Flora Feng and GitHub
3e3d320c1b
[Refactor] Relocate responses API tests ( #37241 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-17 05:14:52 +00:00
Andreas Karatzas and GitHub
54a62a79f7
[ROCm] Fix AttributeError for torch.compiler.skip_all_guards_unsafe on older PyTorch ( #37219 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-17 11:34:49 +08:00
Flora Feng and GitHub
384dc7f77b
[Refactor] Relocate completion and chat completion tests ( #37125 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-17 11:31:23 +08:00
Flora Feng and GitHub
f04d5226f8
[CI] Fix flaky tool_use chat completion tests with deterministic seed ( #37027 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-03-17 03:24:34 +00:00
Kyuyeun Kim and GitHub
0a0a1a198b
Add ability to replace oot ops when using lora ( #37181 )
...
Signed-off-by: Kyuyeun Kim <kyuyeunk@google.com >
2026-03-16 18:04:15 -07:00
6c1cfbad32
Support non-contiguous KV cache in TRTLLM fp8 dequant kernel ( #36867 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: Pavani Majety <pavanimajety@gmail.com >
2026-03-16 17:48:42 -07:00
Harry Huang and GitHub
45f526d652
[BugFix] Correct max memory usage for multiple KV-cache groups ( #36030 )
...
Signed-off-by: huanghaoyan.hhy <huanghaoyan.hhy@alibaba-inc.com >
2026-03-17 00:38:52 +00:00
Julien Denize and GitHub
5db91f0aaf
Fix some Mistral parser issues ( #37209 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-03-17 00:08:56 +00:00
Walter Beller-Morales and GitHub
061980c36a
[Feature][Frontend] add support for Cohere Embed v2 API ( #37074 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-03-16 19:55:53 -04:00
Ben Browning and GitHub
7a49742b88
[CI/Build] Add common tool call parser test suite ( #27599 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-03-16 19:46:20 -04:00
Terry Gao and GitHub
3e6a1e1686
[Custom Ops] Add functional + out variant for scaled_fp4_quant ( #34389 )
...
Signed-off-by: tianrengao <terrygao87@gmail.com >
2026-03-16 18:51:46 -04:00
Julien Denize and GitHub
7961486a9b
Fix EagleMistralLarge3Model initialization ( #37232 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-03-16 15:41:00 -07:00
Andreas Karatzas and GitHub
4f9b14c21c
[CI] Stabilize multinode DP internal LB completion tests ( #36356 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-03-16 15:40:23 -07:00
Yuchen Fama and GitHub
31a458c091
[Doc] Clarify schema enforcement behavior for tool_choice modes ( #37064 )
...
Signed-off-by: yfama <yuchengu@gmail.com >
2026-03-16 22:27:42 +00:00
Wei Zhao and GitHub
a3a51d20e7
[Benchmark] Improvements to attention benchmark script ( #37115 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-03-16 22:22:40 +00:00
EdalatiAli and GitHub
e5b807607c
[Quant][Feature] Support online MXFP8 quantization for MoE and dense models ( #35448 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
2026-03-16 18:07:39 -04:00
fd4d96302a
Fix eplb nvfp4 experts hook ( #37217 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Signed-off-by: Elvir Crncevic <elvir@anthropic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-03-16 22:03:54 +00:00
Krish Gupta and GitHub
c0f011918d
[Bugfix] opcheck false mutation error in rms_norm_per_block_quant ( #36688 ) ( #36779 )
...
Signed-off-by: Krish Gupta <krishom70@gmail.com >
2026-03-16 21:11:33 +00:00
Zhengxu Chen and GitHub
e6ae4b1be1
[compile] Enable mega aot artifact for torch 2.12+. ( #37198 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-03-16 21:05:51 +00:00