 Taneem IbrahimandGitHub
|
4a4fdabe28
|
[Misc] Aligning tokwise pooler heads for consistency (#43041)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-05-19 06:16:42 +00:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
f1e3f0e6d6
|
[XPU] Use custom op collective behavior (#41354)
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-19 14:14:59 +08:00 |
|
  
|
9fd8487d2f
|
[Docs] Add SVG images for pooling models. (#42626)
Signed-off-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-18 22:50:38 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
27f4ba9481
|
fix: use keyword arguments for shard_id and expert_id in weight_loade… (#42671)
Signed-off-by: junyanxu <junyanxu5513@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-19 05:29:04 +00:00 |
|
 
|
6e889b582b
|
[ci] Route 28 gpu_1_queue tests to h200_35gb queue (#43030)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-18 21:58:36 -07:00 |
|
  
|
fab07e4d0f
|
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B (#42289)
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist <noreply@google.com>
|
2026-05-18 21:22:33 -07:00 |
|
 
|
3ca8db2ef8
|
add cutedsl dsv4 indexer fp8 kernel (#42899)
Signed-off-by: george <george@inferact.ai>
Co-authored-by: george <george@inferact.ai>
|
2026-05-18 21:17:56 -07:00 |
|
 Woosuk KwonandGitHub
|
87b08c5f64
|
[Model Refactoring] Move DeepSeek V4 layers to models/deepseek_v4/ [2/N] (#43039)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-18 21:00:58 -07:00 |
|
 
|
fba010dd74
|
[Bugfix][MRV2] Fix KVCache tensor explicit kernel_block_size dim (#42766)
Signed-off-by: NickLucche <nlucches@redhat.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-18 20:25:41 -07:00 |
|
 Mohammad Miadh AngkadandGitHub
|
da03e549b3
|
[UX] Add a persistent cache for FlashInfer autotuning (#42537)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-18 20:25:37 -07:00 |
|
 Kunshang JiandGitHub
|
36dcaf25d8
|
[XPU] add gptq(int4) support (#37844)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-19 11:17:09 +08:00 |
|
 Ofir ZafrirandGitHub
|
8f16c4a5c0
|
[BugFix][CPU][Spec Decode] Fix Eagle implementation on CPU backend (#42468)
Signed-off-by: Ofir Zafrir <ofir.zafrir@intel.com>
|
2026-05-19 03:16:07 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
afd7b1dce9
|
[Bugfix] Use platform-agnostic device in example_connector load (#42926)
Signed-off-by: Revital Sur <eres@il.ibm.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-19 03:12:04 +00:00 |
|
 Woosuk KwonandGitHub
|
287471b994
|
[Model Refactoring] Migrate DeepSeek V4 to vllm/models/ [1/N] (#43004)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-18 19:50:02 -07:00 |
|
 
|
239b5ff30c
|
[Frontend] Add --spec-method/--spec-model/--spec-tokens CLI aliases (#42476)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-18 17:22:27 -07:00 |
|
 Artem PerevedentsevandGitHub
|
f85c76d701
|
[CI/Build] Bump nvidia-cutlass-dsl to 4.5.1 (#42991)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-05-18 16:58:15 -07:00 |
|
 shanjiazandGitHub
|
a171e6b52d
|
Add parallel drafting to v2 model runner unsupported features (#43010)
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com>
|
2026-05-18 16:39:09 -07:00 |
|
 Wentao YeandGitHub
|
37ece593c1
|
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement (#42774)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-18 16:38:12 -07:00 |
|
 Flora FengandGitHub
|
57fef4e0bf
|
[Refactor] Extract shared coerce_to_schema_type utility from Minimax M2 tool parser (#43006)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-18 17:55:39 -04:00 |
|
 haosdentandGitHub
|
0191354827
|
[Perf][MLA] Enable FULL cudagraph capture for TRITON_MLA decode (#42885)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-18 14:29:10 -07:00 |
|
 Wentao YeandGitHub
|
cd49a05d5a
|
[Refactor] Remove dead code (#42889)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-18 16:41:22 -04:00 |
|
 Ronen SchafferandGitHub
|
84747489de
|
Tier offload followup (#42529)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-18 19:41:58 +00:00 |
|
 Tuukka SarviandGitHub
|
8fc1c284b9
|
[ROCm] Guard AITER GDN decode fast path by layout (#42880)
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com>
|
2026-05-18 11:56:22 -07:00 |
|
 Amit PortnoyandGitHub
|
ce88f01c9a
|
[Docs] update attribution to reflect EDEN foundation (#41666)
Signed-off-by: amitport <1131991+amitport@users.noreply.github.com>
|
2026-05-18 11:22:56 -07:00 |
|
 Wentao YeandGitHub
|
00e20e76f7
|
[Refactor] Remove dead cuda kernels (#42767)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-18 11:14:21 -07:00 |
|
 czhu-cohereandGitHub
|
9758a6e5c5
|
[BugFix] support PP for Cohere vision model (#42819)
Signed-off-by: <conway.zhu@cohere.com>
Signed-off-by: root <conway.zhu@cohere.com>
|
2026-05-18 11:12:06 -07:00 |
|
 Bowen BaoandGitHub
|
a2c8fc6657
|
[ROCm][Quantization][3/N] Refactor quark_moe w4a4 w/ oracle (#41436)
Signed-off-by: Bowen Bao <bowenbao@amd.com>
|
2026-05-18 13:46:13 -04:00 |
|
 
|
6859ca7615
|
[Bugfix] fix swiglu limit issue for humming backend + deepseek v4 (#42541)
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
|
2026-05-18 17:32:26 +00:00 |
|
 Mohammad Miadh AngkadandGitHub
|
67f58ce23f
|
[Bugfix] Fix DSV4 MTP after ROCm mHC integration (#42930)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-18 17:02:01 +00:00 |
|
 Wei ZhaoandGitHub
|
8c296de63b
|
[Perf] Re-enable flashinfer autotune by default and cleanup (#42857)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-05-18 09:12:27 -07:00 |
|
 Harry MellorandGitHub
|
b12745e4f3
|
Fix --convert passed without --runner on causal models (#42935)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-18 15:56:09 +00:00 |
|
 Wentao YeandGitHub
|
e26736973a
|
[Model Runner V2] Fix prompt logprobs calculation Sizes of tensors must match error (#42778)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-18 08:27:21 -07:00 |
|
 Netanel HaberandGitHub
|
47829b1159
|
[Bugfix] mamba: run single-token extends as decodes (#42430)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
|
2026-05-18 15:26:00 +00:00 |
|
 Blanc SwanandGitHub
|
4a39b4f553
|
[Model] Add Apertus Tool Parser (#41154)
Signed-off-by: Blanc <swan.blanc@infomaniak.com>
|
2026-05-18 11:20:04 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
78e7a7b9b0
|
Refactor AWQ Marlin MoE onto modular WNA16 oracle (#42483)
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com>
Signed-off-by: Siddharth Bedekar <104613085+bedeks@users.noreply.github.com>
Co-authored-by: Robert Shaw <robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-18 08:02:43 -07:00 |
|
 
|
f5d3dc7115
|
[Model Runner v2] Support update_config (#42783)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-18 10:26:07 -04:00 |
|
  
|
1ac10f159a
|
Revert "[torch.compile] Add patch for fullgraph compilation" (#42686) (#42913)
Co-authored-by: Luka Govedič <luka.govedic@gmail.com>
Co-authored-by: Zhewen Li <zhewenli@inferact.ai>
|
2026-05-18 09:02:51 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
e5417657e5
|
[KV Connector][Offloading] Flush all pending jobs on last step (#42611)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Signed-off-by: liranschour <liranschour@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-18 12:59:42 +00:00 |
|
 xiangdongandGitHub
|
2e40faf08b
|
[XPU][CI] Temporarily skip test_moe_lora_align_block_size_mixed_base_and_lora[1] in Intel GPU CI (#42954)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
|
2026-05-18 20:34:48 +08:00 |
|
 Nicolò LucchesiandGitHub
|
69c91d010a
|
[MRv2] Default to MRv1 when a connector is present (#42955)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-18 20:34:16 +08:00 |
|
 roikoren755andGitHub
|
737bfa3a43
|
[Bugfix][Hybrid][NemotronH] Fix mamba_cache_mode=all + speculative decoding crash (#41233)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-05-18 14:54:00 +03:00 |
|
 Kfir ToledoandGitHub
|
e414e1f1c0
|
[Bugfix][KV Offload] count appended GPU blocks in store group_sizes (#42945)
Signed-off-by: Kfir Toledo <kfir.toledo@ibm.com>
|
2026-05-18 11:36:02 +00:00 |
|
 
|
df852ed503
|
fix: remove unused norm for dpskv4 (#41710)
Signed-off-by: inisis <desmond.yao@buaa.edu.cn>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-18 18:33:29 +08:00 |
|
 Yuwen ZhouandGitHub
|
88a860d754
|
[CPU] Add MXFP4 W4A16 MoE support (#41922)
Signed-off-by: yuwenzho <yuwen.zhou@intel.com>
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com>
|
2026-05-18 03:04:45 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
cac81b6eda
|
[CPU Backend] Improve cpu thread utilization (#42666)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-18 03:04:41 -07:00 |
|
 Li, JiangandGitHub
|
b4601ad43f
|
[CPU] Add fused GDN support for AMX CPU platform (#42707)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-05-18 03:04:36 -07:00 |
|
 Jee Jee LiandGitHub
|
2267f70070
|
[Kernel] Pack topk id/weights triton kernel (#42527)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-18 03:04:31 -07:00 |
|
 
|
965d076148
|
[CPU] Specify required KV cache layout for CPU attention backend (#42740)
Signed-off-by: Tony Lin <tony.lin@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-05-18 17:38:54 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
c38bed4248
|
delete xpu ci (#42582)
Signed-off-by: wenjun.liu <wenjun.liu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-18 16:36:45 +08:00 |
|
 Xin YangandGitHub
|
998714b21b
|
[Perf] Add do_not_specialize in fused FP8 RoPE kernel (#42849)
Signed-off-by: Xin Yang <xyangx@amazon.com>
|
2026-05-18 01:32:46 -07:00 |
|