 Jee Jee LiandGitHub
|
8cd174fa35
|
[LoRA] MoE LoRA Refactor (#40338)
|
2026-04-26 01:55:19 +00:00 |
|
  
|
c798593f0d
|
[Bugfix] Fix the DSML token leakage in DSV4/3.2 (#40806)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>
Co-authored-by: Windswithyou 1694599440@qq.com
|
2026-04-26 08:58:50 +08:00 |
|
 
|
12a3f6454b
|
[Bugfix][MoE] Only unpad routed output before shared expert add or routed output transform (#40865)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-04-25 20:50:12 +00:00 |
|
 Or OzeriandGitHub
|
60cd878a3b
|
[kv_offload+HMA][11/N]: Support store with multiple KV groups (#39403)
Signed-off-by: Or Ozeri <oro@il.ibm.com>
|
2026-04-25 20:00:46 +03:00 |
|
 rasmithandGitHub
|
1e9f19ca3f
|
[CI][AMD]BugFix] Fix deadlock occuring in test_moe_layer (#40767)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-04-25 09:34:14 -04:00 |
|
 labAxiaomingandGitHub
|
6646c0c7e0
|
[Opt] Optimize deepstack buffer handling for multimodal Qwen3 models (#40145)
Signed-off-by: xiaoming <1259730330@qq.com>
|
2026-04-25 21:04:26 +08:00 |
|
 Andreas KaratzasandGitHub
|
95995bbef8
|
[ROCm][Engine] Fix GPU memory leaks in engine shutdown and test workaround for async KV prefix cache reset (#38503)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-25 05:25:20 +00:00 |
|
 
|
07351e0883
|
[Feature] Warm up readonly multimodal processor during renderer startup (#40797)
Signed-off-by: Chenguang ZHENG <645327136@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-04-25 03:57:41 +00:00 |
|
 Andreas KaratzasandGitHub
|
428b988c98
|
[ROCm][CI] Fix trust_remote_code AttributeError in EAGLE3 acceptance length test (#40306)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-25 02:59:31 +00:00 |
|
 Andreas KaratzasandGitHub
|
e54894fc85
|
[ROCm][CI] Fix TestSiluMulGroupFp8QuantModel after W8A8 block linear refactor (#39799)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-25 11:20:59 +09:00 |
|
 Angela YiandGitHub
|
bc2ae5a3d6
|
[Test] Increase qwen2_vl num_logprobs to fix torch 2.12 update (#40818)
Signed-off-by: Angela Yi <angelayi@meta.com>
|
2026-04-25 00:59:20 +00:00 |
|
 Wentao YeandGitHub
|
a474da2813
|
[Refactor] Remove unused dead code (#40640)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-25 07:28:18 +08:00 |
|
 Lucas KabelaandGitHub
|
ce6a199ecc
|
[BE][Bugfix] Respect TORCH_COMPILE_DISABLE env var at the vLLM config level for torch 2.12 (#40715)
Signed-off-by: Lucas Kabela <lucaskabela@meta.com>
|
2026-04-24 16:25:03 -07:00 |
|
 Ignacio SicaandGitHub
|
f88763efc3
|
[Bugfix] add seq_lens_cpu_upper_bound to CommonAttentionMetadata in mla_runner.py (#40844)
Signed-off-by: ignaciosica <mignacio.sica@gmail.com>
|
2026-04-24 23:13:52 +00:00 |
|
 Artem PerevedentsevandGitHub
|
333529deae
|
[EPLB] Fix replica selection bias in fused_moe router (#40810)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-24 22:06:41 +00:00 |
|
 Zhang JianandGitHub
|
8825608205
|
[Bugfix][CI] Fix wrong residual shape in TestFusedAddRMSNorm.example_inputs that causes flaky test (#40629)
Signed-off-by: Zhang Jian <jianmusings@gmail.com>
|
2026-04-24 16:40:07 -04:00 |
|
 qli88andGitHub
|
095d2f87e8
|
[Bug] Fix GLM-5.1 running error on ROCm platform (#40763)
Signed-off-by: Qiang Li <qiang.li2@amd.com>
|
2026-04-24 19:54:40 +00:00 |
|
 
|
21792520e7
|
[Build] Add Python 3.14 to supported version list. (#34770)
Signed-off-by: Neil Schemenauer <nas@arctrix.com>
Co-authored-by: Simon Mo <simon.mo@hey.com>
|
2026-04-24 10:24:05 -07:00 |
|
 Alex BrooksandGitHub
|
5e11b40365
|
[Frontend] Delegate to vLLM Omni When --omni Passed (#40744)
Signed-off-by: Alex Brooks <albrooks@redhat.com>
|
2026-04-24 12:30:00 -04:00 |
|
 
|
f768b4473e
|
[Docs] Add docs for context extension using the yarn method (#37430)
Signed-off-by: xiaoming <1259730330@qq.com>
Signed-off-by: labAxiaoming <34019940+labAxiaoming@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2026-04-24 08:26:09 -07:00 |
|
 JartXandGitHub
|
914d0464c1
|
[Refactor] Unify 2D/3D kernels in triton_unified_attention (#40631)
Signed-off-by: JartX <sagformas@epdcenter.es>
|
2026-04-24 17:18:06 +02:00 |
|
 Jinzhen LinandGitHub
|
9f771b3ab9
|
[Quantization] add humming quantization kernel (#34556)
|
2026-04-24 09:29:44 -04:00 |
|
 Itay AlroyandGitHub
|
c9d3c6e6af
|
fused_moe: treat NIXL EP as batched experts (#40412)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-04-24 08:05:31 -05:00 |
|
 Or OzeriandGitHub
|
51adca74e6
|
[kv_offload+HMA][9/N]: Support lookup with multiple KV groups (#39401)
Signed-off-by: Or Ozeri <oro@il.ibm.com>
|
2026-04-24 15:32:29 +03:00 |
|
 Netanel HaberandGitHub
|
e8eb0490ce
|
[Bugfix][MoE] Unpad routed output before shared expert add [Fixes #35949] (#40794)
Signed-off-by: Netanel Haber <nhaber@nvidia.com>
|
2026-04-24 11:53:23 +00:00 |
|
 Jiangyun ZhuandGitHub
|
e8ee2a78db
|
[Attention] use diff kv backend for mimo v2 flash (#40045)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
|
2026-04-24 11:25:55 +00:00 |
|
   
|
2ec18f5df4
|
[Bugfix][Parser] Fix Mistral tool parser for HF tokenizers (#39294)
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com>
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
|
2026-04-24 19:01:56 +08:00 |
|
 Dmitry TokarevandGitHub
|
6dec49f27e
|
[Build] Bump CUDA to 13.0.2 to match PyTorch 2.11.0 (#40669)
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
|
2026-04-24 10:27:11 +00:00 |
|
 Shanshan ShenandGitHub
|
b5587e1013
|
[CI/Build] Add e2e test for ViT CUDA graph (#40780)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-24 18:12:14 +08:00 |
|
 milesialandGitHub
|
9ad5abe772
|
Fix Nano Nemotron VL static image inputs (#40724)
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com>
|
2026-04-24 09:18:55 +00:00 |
|
 Woosuk KwonandGitHub
|
7d3195ea9f
|
[Bugfix] Fix IMA in DSA + MTP (#40772)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-04-24 01:40:20 -07:00 |
|
   
|
512f522192
|
[Model] Gemma4: add bidirectional vision attention for sliding layers with window guard (#40534)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Signed-off-by: Luciano Martins <lucianomartins@google.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-24 08:27:46 +00:00 |
|
 
|
4c34b2f6fc
|
[XPU] Enable torch.compile for XPU GDN attention (#39466)
Signed-off-by: yuwenzho <yuwen.zhou@intel.com>
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-04-24 16:26:16 +08:00 |
|
 Xin YangandGitHub
|
cf8a613a87
|
Support only half types for concat_mla_q kernel (#37892)
Signed-off-by: Xin Yang <xyangx@amazon.com>
|
2026-04-23 23:51:05 -07:00 |
|
 xiangdongandGitHub
|
01acf96c6f
|
[XPU][CI] Fix Docker cleanup races on Intel CI runners (#40761)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
|
2026-04-24 14:08:45 +08:00 |
|
 
|
079a4cf399
|
[MoE] Move cutlass moe to fused_moe/experts/ (#40574)
Signed-off-by: Jackmin801 <ongjackm@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-04-24 06:05:49 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9744b699ba
|
[Deprecate] Deprecate LLM.reward offline api, use LLM.encode instead. (#40688)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
|
2026-04-24 05:37:50 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
c662b4359e
|
[Bugfix] Avoid mutating chat_template_kwargs in HYV3ReasoningParser initialization (#40713)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-24 13:08:58 +08:00 |
|
 lyd1992andGitHub
|
100c7b65e7
|
[Platform] Fix RISC-V platform detection (lscpu parsing + non-NUMA meminfo) (#40427)
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn>
|
2026-04-24 04:33:05 +00:00 |
|
 Neil SchemenauerandGitHub
|
56bdf85e10
|
[Feature] Avoid eager import of the "mistral_common" package. (#40043)
Signed-off-by: Neil Schemenauer <nas@arctrix.com>
|
2026-04-24 02:49:16 +00:00 |
|
 Vinayak KumarandGitHub
|
eba73068ea
|
[Doc] fix capitalization consistency in README (vLLM, Hugging Face) (#40729)
Signed-off-by: Vinayak Mishra <vinayakmishra448@gmail.com>
|
2026-04-24 02:23:54 +00:00 |
|
 Nick HillandGitHub
|
e9f331d72e
|
[MRV2] Ensure warmup covers prefill path (#40746)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-24 01:33:26 +00:00 |
|
  
|
c9bf77df92
|
[BUG]: fix HF tokenizer concurrent borrow in tool parsers (#40059)
Signed-off-by: Yifan <yzong@redhat.com>
Co-authored-by: timon0305 <timon0305@outlook.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>
|
2026-04-23 18:20:30 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
3041344287
|
[Misc] Added curl retries in install_python_libraries.sh (#36700)
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-24 01:19:30 +00:00 |
|
 Doug CamposandGitHub
|
92762edc53
|
[Bugfix] Treat <tool_call> as implicit reasoning end in Qwen3 parser (#35687)
Signed-off-by: Doug Campos <qmx@qmx.me>
|
2026-04-24 09:10:04 +08:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
626daa2076
|
[Feat] Unified Synthetic Acceptance Rate for V1 and V2 (#40662)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-24 00:48:08 +00:00 |
|
 Nick HillandGitHub
|
fe85a92e86
|
[Core] Avoid seq_lens_cpu GPU->CPU sync (#40654)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-24 00:35:55 +00:00 |
|
 Sage MooreandGitHub
|
62b1bbe470
|
[EPLB] Remove asyncio infrastructure from Async EPLB (#40730)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
|
2026-04-24 00:21:15 +00:00 |
|
 Hemanth AcharyaandGitHub
|
fa4b70555b
|
[ROCm] Cast score correction bias tensor during model construction for DeepSeek/Kimi-K2 (#39999)
Signed-off-by: Hemanth Acharya <heachary@amd.com>
|
2026-04-24 09:02:12 +09:00 |
|
 
|
447c372ac5
|
[MoE] Move remaining PrepareAndFinalize to prepare finalize folder (#39009)
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Jackmin801 <ongjackm@gmail.com>
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com>
|
2026-04-23 20:00:53 -04:00 |
|