 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
4dcd10eb0d
|
[1/N][KV-Cache Layout Refactor] Refactor DSV4 KV cache config construction (#44454)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-06-07 14:53:37 +00:00 |
|
 Charlie FuandGitHub
|
228bcc436b
|
[ROCm][Kernel] Enable permute_cols for ROCm (#44674)
Signed-off-by: charlifu <charlifu@amd.com>
|
2026-06-07 09:50:03 +00:00 |
|
 
|
3d3ba460a2
|
Modify torch dependency in xpu.txt (#43087)
Signed-off-by: Bram Vanroy <2779410+BramVanroy@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-07 16:33:50 +08:00 |
|
 Mohammad Miadh AngkadandGitHub
|
66ecfd0568
|
[Dependency] Remove stale cuDNN frontend upper bound (#42599)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-07 16:09:25 +08:00 |
|
 Andreas KaratzasandGitHub
|
f0f6805d8a
|
[CI] Stabilize the multi-audio OpenAI server path (#44051)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-07 15:54:32 +08:00 |
|
 
|
15652a6b70
|
[Doc] Fix multimodal torch.compile troubleshooting to not use removed VLLM_TORCH_COMPILE_LEVEL (#44378)
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-06-07 00:34:07 -07:00 |
|
 Yifan QiaoandGitHub
|
51ef688831
|
[Bugfix][Mooncake] Fix per-group block_size/block_hash and group_idx in MooncakeStoreConnector KV events (#44103)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-06-07 07:12:43 +00:00 |
|
 Jared WenandGitHub
|
6ac69203e8
|
[videoloader] implement glm46v video loader (#44417)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
|
2026-06-07 06:27:20 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1505b3d8a1
|
[Cohere] Enable Cohere Mini Code model and update Command A-plus test registry (#44707)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-06 22:44:16 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
32f34d3935
|
[feature] add index share feature for DSA MTP (#44420)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-06 22:04:14 -07:00 |
|
 Qiuyang YueandGitHub
|
9c7f7741d4
|
[Bugfix] Fix benchmark_moe.py after inplace mechanism removal (#44041)
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com>
|
2026-06-07 00:32:00 -04:00 |
|
 
|
6181e80fe0
|
[XPU] add xpu branch in compressed_tensors_moe_w4a4_mxfp4 (#44540)
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>
Co-authored-by: Kunshang Ji <jikunshang95@gmail.com>
|
2026-06-07 12:27:34 +08:00 |
|
 Yan MaandGitHub
|
3bb46975bd
|
[XPU][Feature] transparent sleep mode support for XPU platform (#37149)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-06-07 10:45:31 +08:00 |
|
 Chaojun ZhangandGitHub
|
810966453a
|
[XPU] Support cpu kv offloading and tiering offloading on XPU platform (#36423)
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
|
2026-06-07 09:59:28 +08:00 |
|
 Woosuk KwonandGitHub
|
2a983c79ac
|
[DSV4] Decouple DS V4 Sparse MLA Metadata from DS V3.2 (#44699)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-06 20:37:56 -04:00 |
|
    
|
bc5745a00f
|
[ROCm][MLA] Replace torch.cat in sparse-MLA forward_mqa with fused concat_mla_q (#42838)
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com>
Signed-off-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 18:20:50 -05:00 |
|
 Nick HillandGitHub
|
3b3d5287fa
|
[BugFix] Resolve multiple async kv load deadlock (#44560)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-06 23:05:47 +00:00 |
|
 
|
062b05ff3a
|
[ROCm][Perf] Fused MoE W4A16 HIP kernel for AMD RDNA3 (gfx1100) (#44075)
Signed-off-by: JartX <sagformas@epdcenter.es>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 15:30:39 -05:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
fa27d4e9cf
|
[PERF] [Qwen3.5] Split mixed prefill+decode batches: route decodes to the recurrent kernel (#44700)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-06 22:13:50 +08:00 |
|
 Vadim GimpelsonandGitHub
|
67d3792d99
|
[Bugfix] Fix Qwen3.5-FP8 nightly fail. Guard fused_add_rms_norm input/weight dtype mismatch in RMSNorm + quant fusion (#44694)
|
2026-06-06 08:46:14 -04:00 |
|
  
|
00d1fb7747
|
[Bugfix][ROCm] ApplyRotaryEmb: fall back to native when flash_attn rotary grid would exceed the HIP per-dim limit (#43684)
Signed-off-by: vLLM ROCm fix <noreply@example.com>
Signed-off-by: amd-fuweiy <fuweiy@amd.com>
Co-authored-by: vLLM ROCm fix <noreply@example.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 02:29:16 -07:00 |
|
  
|
c9b4b184b4
|
[Bugfix][Voxtral] Add fetch_audio to MistralCommonFeatureExtractor (transformers>=5.10 compat) (#44559)
Signed-off-by: Yadan Wei <weiyadan@amazon.com>
Co-authored-by: Yadan Wei <weiyadan@amazon.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-06 07:58:09 +00:00 |
|
 
|
f87df1df9e
|
[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig (#44613)
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-05 23:14:41 -07:00 |
|
 Taneem IbrahimandGitHub
|
eafbb06331
|
[Misc] Replaced asserts with proper exceptions to improve UX for pooling (#44593)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-06-06 05:57:26 +00:00 |
|
 
|
ec0a31d4aa
|
[Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 (#44692)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-06 10:44:21 +08:00 |
|
 Devin LaiandGitHub
|
c8beda4cc3
|
[Rust Frontend] Add Phi-4 mini JSON tool parser (#44213)
|
2026-06-06 10:40:00 +08:00 |
|
 
|
2f27c9a150
|
Preserve layout-changing clones (#44574)
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com>
|
2026-06-05 20:45:24 -04:00 |
|
 
|
4765f0f189
|
[Bugfix] Fix sequence_parallel_chunk_impl custom op aliasing its input (#44130)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-05 23:56:36 +00:00 |
|
 Terrence ZhaoandGitHub
|
a50e675b0d
|
[Cohere] fix RoutingMethodType (#44021)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
|
2026-06-05 16:25:53 -07:00 |
|
 Daoyuan LiandGitHub
|
f6a708ab2b
|
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models (#44435)
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>
|
2026-06-05 16:04:32 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4200f62147
|
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path (#42832)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-05 16:22:19 -05:00 |
|
 Walter Beller-MoralesandGitHub
|
c73b0d0db9
|
[Core][Engine] allow DP ray placement groups to be set on specific nodes (#44669)
Signed-off-by: walterbm <walter.beller.morales@gmail.com>
|
2026-06-05 20:07:47 +00:00 |
|
 Harry MellorandGitHub
|
e28e369f78
|
Male Mergify comment less spammy (#44666)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 10:56:52 -07:00 |
|
 yzong-rhandGitHub
|
703fb17b13
|
[Bugfix] GPT-OSS instruction rendering (#44330)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-06-05 13:52:32 -04:00 |
|
 Sting LinandGitHub
|
b593396c7a
|
Upgrade tpu-inference to v0.21.0 (#44621)
Signed-off-by: StingLin <sting.lin@cienet.com>
|
2026-06-05 16:12:49 +00:00 |
|
 FlameandGitHub
|
91e17d4315
|
Fix sarvam forward compatibility with transformers v5 (#38804)
Signed-off-by: vikrantpalle <vikrantpalle@gmail.com>
|
2026-06-05 11:51:44 -04:00 |
|
 TJianandGitHub
|
aa6fb8a329
|
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm (#44648)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-05 15:51:17 +00:00 |
|
 Effi OferandGitHub
|
6a894574bf
|
Add objectstore as a secondary tier to multi-tier kv cache offloading (#41968)
Signed-off-by: Effi Ofer <effi.ofer@gmail.com>
|
2026-06-05 18:05:41 +03:00 |
|
 Yan MaandGitHub
|
7f003a1285
|
Support MiniCPMV batched preprocessing (#44609)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-06-05 15:05:31 +00:00 |
|
 Harry MellorandGitHub
|
ef0df7dbd6
|
[CI] Bump mypy version 1.19.1 -> 1.20.2 (#44647)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:56:27 +00:00 |
|
 Harry MellorandGitHub
|
a80af24356
|
Speed up docs build (#44635)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:51:44 +00:00 |
|
 Harry MellorandGitHub
|
c66b19800b
|
[CI] Bump mistral-common (#44649)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:18:50 +00:00 |
|
 
|
6a11d72df7
|
[Reasoning][Structured Outputs] Add Command A plus tags for structural tags (#44588)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
|
2026-06-05 06:51:20 -07:00 |
|
 Woosuk KwonandGitHub
|
02d2da0748
|
[DSV4] Move more ops out of eager breakpoint (#44561)
|
2026-06-05 06:42:41 -07:00 |
|
 adhithyamulticorewareandGitHub
|
bbb6c274c8
|
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call (#44615)
Signed-off-by: ADHITHYA BALAKRISHNAN <adhithya.balakrishnan@multicorewareinc.com>
|
2026-06-05 12:47:56 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
62215e72c6
|
Remove KV cache scale boilerplate from model weight loading methods (#43167)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-05 05:19:04 -07:00 |
|
  
|
7fe7800fa4
|
[BUG] Fix FP64 Gumbel precision coverage (#43150)
Signed-off-by: tianyu-z <zhangtianyupro@gmail.com>
Signed-off-by: Tianyu Zhang <53099276+tianyu-z@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-06-05 19:04:14 +08:00 |
|
 
|
8a83e6f2d7
|
[Rust Frontend] Batch auto-abort requests by engine (#44591)
Signed-off-by: Hugh Ryan <197298026+HueCodes@users.noreply.github.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-05 02:59:09 -07:00 |
|
 Chunyang WenandGitHub
|
efc347f1b2
|
docs: fix tokenizer optimization typo (#44066)
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com>
|
2026-06-05 02:12:49 -07:00 |
|
 Nicolò LucchesiandGitHub
|
d98b8f371c
|
[NixlConnector] Initiate deprecation cycle for kv_both role (#43874)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-06-05 11:08:17 +02:00 |
|