    
|
bc5745a00f
|
[ROCm][MLA] Replace torch.cat in sparse-MLA forward_mqa with fused concat_mla_q (#42838)
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com>
Signed-off-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 18:20:50 -05:00 |
|
 Nick HillandGitHub
|
3b3d5287fa
|
[BugFix] Resolve multiple async kv load deadlock (#44560)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-06 23:05:47 +00:00 |
|
 
|
062b05ff3a
|
[ROCm][Perf] Fused MoE W4A16 HIP kernel for AMD RDNA3 (gfx1100) (#44075)
Signed-off-by: JartX <sagformas@epdcenter.es>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 15:30:39 -05:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
fa27d4e9cf
|
[PERF] [Qwen3.5] Split mixed prefill+decode batches: route decodes to the recurrent kernel (#44700)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-06 22:13:50 +08:00 |
|
 Vadim GimpelsonandGitHub
|
67d3792d99
|
[Bugfix] Fix Qwen3.5-FP8 nightly fail. Guard fused_add_rms_norm input/weight dtype mismatch in RMSNorm + quant fusion (#44694)
|
2026-06-06 08:46:14 -04:00 |
|
  
|
00d1fb7747
|
[Bugfix][ROCm] ApplyRotaryEmb: fall back to native when flash_attn rotary grid would exceed the HIP per-dim limit (#43684)
Signed-off-by: vLLM ROCm fix <noreply@example.com>
Signed-off-by: amd-fuweiy <fuweiy@amd.com>
Co-authored-by: vLLM ROCm fix <noreply@example.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-06 02:29:16 -07:00 |
|
  
|
c9b4b184b4
|
[Bugfix][Voxtral] Add fetch_audio to MistralCommonFeatureExtractor (transformers>=5.10 compat) (#44559)
Signed-off-by: Yadan Wei <weiyadan@amazon.com>
Co-authored-by: Yadan Wei <weiyadan@amazon.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-06 07:58:09 +00:00 |
|
 
|
f87df1df9e
|
[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig (#44613)
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-05 23:14:41 -07:00 |
|
 Taneem IbrahimandGitHub
|
eafbb06331
|
[Misc] Replaced asserts with proper exceptions to improve UX for pooling (#44593)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-06-06 05:57:26 +00:00 |
|
 
|
ec0a31d4aa
|
[Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 (#44692)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-06 10:44:21 +08:00 |
|
 Devin LaiandGitHub
|
c8beda4cc3
|
[Rust Frontend] Add Phi-4 mini JSON tool parser (#44213)
|
2026-06-06 10:40:00 +08:00 |
|
 
|
2f27c9a150
|
Preserve layout-changing clones (#44574)
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com>
|
2026-06-05 20:45:24 -04:00 |
|
 
|
4765f0f189
|
[Bugfix] Fix sequence_parallel_chunk_impl custom op aliasing its input (#44130)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-06-05 23:56:36 +00:00 |
|
 Terrence ZhaoandGitHub
|
a50e675b0d
|
[Cohere] fix RoutingMethodType (#44021)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
|
2026-06-05 16:25:53 -07:00 |
|
 Daoyuan LiandGitHub
|
f6a708ab2b
|
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models (#44435)
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>
|
2026-06-05 16:04:32 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4200f62147
|
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path (#42832)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-05 16:22:19 -05:00 |
|
 Walter Beller-MoralesandGitHub
|
c73b0d0db9
|
[Core][Engine] allow DP ray placement groups to be set on specific nodes (#44669)
Signed-off-by: walterbm <walter.beller.morales@gmail.com>
|
2026-06-05 20:07:47 +00:00 |
|
 Harry MellorandGitHub
|
e28e369f78
|
Male Mergify comment less spammy (#44666)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 10:56:52 -07:00 |
|
 yzong-rhandGitHub
|
703fb17b13
|
[Bugfix] GPT-OSS instruction rendering (#44330)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-06-05 13:52:32 -04:00 |
|
 Sting LinandGitHub
|
b593396c7a
|
Upgrade tpu-inference to v0.21.0 (#44621)
Signed-off-by: StingLin <sting.lin@cienet.com>
|
2026-06-05 16:12:49 +00:00 |
|
 FlameandGitHub
|
91e17d4315
|
Fix sarvam forward compatibility with transformers v5 (#38804)
Signed-off-by: vikrantpalle <vikrantpalle@gmail.com>
|
2026-06-05 11:51:44 -04:00 |
|
 TJianandGitHub
|
aa6fb8a329
|
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm (#44648)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-05 15:51:17 +00:00 |
|
 Effi OferandGitHub
|
6a894574bf
|
Add objectstore as a secondary tier to multi-tier kv cache offloading (#41968)
Signed-off-by: Effi Ofer <effi.ofer@gmail.com>
|
2026-06-05 18:05:41 +03:00 |
|
 Yan MaandGitHub
|
7f003a1285
|
Support MiniCPMV batched preprocessing (#44609)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-06-05 15:05:31 +00:00 |
|
 Harry MellorandGitHub
|
ef0df7dbd6
|
[CI] Bump mypy version 1.19.1 -> 1.20.2 (#44647)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:56:27 +00:00 |
|
 Harry MellorandGitHub
|
a80af24356
|
Speed up docs build (#44635)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:51:44 +00:00 |
|
 Harry MellorandGitHub
|
c66b19800b
|
[CI] Bump mistral-common (#44649)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:18:50 +00:00 |
|
 
|
6a11d72df7
|
[Reasoning][Structured Outputs] Add Command A plus tags for structural tags (#44588)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
|
2026-06-05 06:51:20 -07:00 |
|
 Woosuk KwonandGitHub
|
02d2da0748
|
[DSV4] Move more ops out of eager breakpoint (#44561)
|
2026-06-05 06:42:41 -07:00 |
|
 adhithyamulticorewareandGitHub
|
bbb6c274c8
|
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call (#44615)
Signed-off-by: ADHITHYA BALAKRISHNAN <adhithya.balakrishnan@multicorewareinc.com>
|
2026-06-05 12:47:56 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
62215e72c6
|
Remove KV cache scale boilerplate from model weight loading methods (#43167)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-05 05:19:04 -07:00 |
|
  
|
7fe7800fa4
|
[BUG] Fix FP64 Gumbel precision coverage (#43150)
Signed-off-by: tianyu-z <zhangtianyupro@gmail.com>
Signed-off-by: Tianyu Zhang <53099276+tianyu-z@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-06-05 19:04:14 +08:00 |
|
 
|
8a83e6f2d7
|
[Rust Frontend] Batch auto-abort requests by engine (#44591)
Signed-off-by: Hugh Ryan <197298026+HueCodes@users.noreply.github.com>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-05 02:59:09 -07:00 |
|
 Chunyang WenandGitHub
|
efc347f1b2
|
docs: fix tokenizer optimization typo (#44066)
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com>
|
2026-06-05 02:12:49 -07:00 |
|
 Nicolò LucchesiandGitHub
|
d98b8f371c
|
[NixlConnector] Initiate deprecation cycle for kv_both role (#43874)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-06-05 11:08:17 +02:00 |
|
 Chao-Ju ChenandGitHub
|
e64237ae82
|
[Rust Frontend] Support include_reasoning=false (#44391)
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com>
|
2026-06-05 16:47:50 +08:00 |
|
 
|
d61d8566ec
|
[Bugfix] Update mistral tokenizer test for continue_final_message fix (#44622)
Signed-off-by: Xu Zhou <xuzhou9417@163.com>
Co-authored-by: Xu Zhou <xuzhou9417@163.com>
|
2026-06-05 16:13:26 +08:00 |
|
 UranusandGitHub
|
d2f70da116
|
fix: pad dummy run query_start_loc (#44603)
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com>
|
2026-06-05 00:43:04 -07:00 |
|
 
|
6542d48964
|
[Bugfix] Fix test_invocations flaky failure with newer openai SDK (#44618)
Signed-off-by: Xu Zhou <xuzhou9417@163.com>
Co-authored-by: Xu Zhou <xuzhou9417@163.com>
|
2026-06-05 07:36:20 +00:00 |
|
 Ting SUNandGitHub
|
ca73293fa6
|
[Bugfix][Rust Frontend] Fix UTF-8 char-boundary panic in incremental detokenizer (#44620)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-06-05 07:36:17 +00:00 |
|
 Vic WenandGitHub
|
ef3af56a97
|
Fix LLM.wait_for_completion output type docstring (#44617)
Signed-off-by: viiccwen <viiccwen@gmail.com>
|
2026-06-05 00:16:38 -07:00 |
|
  
|
b4a6f26c90
|
[ROCm][perf] Use workspace manager for sparse indexer allocations (#41002)
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com>
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com>
Co-authored-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-04 23:46:29 -07:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
165b7864d0
|
[ROCM] [FEAT] Integrate Aiter hipBLASLt GEMM online tuning (#40426)
Signed-off-by: hanlin12 <hanlin12@amd.com>
Signed-off-by: Han Lin <hanlin12@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-04 23:45:36 -07:00 |
|
 Li, JiangandGitHub
|
c505cd93ef
|
[CI/Build] Disable CPU-Compatibility Tests (#44605)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-06-05 13:14:26 +08:00 |
|
 qizixiandGitHub
|
96229fa99e
|
[KVConnector][1/N] PP-aware handshake aggregation and intermediate-PP output plumbing (#43720)
Signed-off-by: zixi-qi <zixi@inferact.ai>
|
2026-06-04 22:04:19 -07:00 |
|
 
|
da1daf40bf
|
[Bugfix] Exclude vision embedder from quantization in Gemma4 Unified (#44571)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
|
2026-06-04 20:47:38 -07:00 |
|
 Woosuk KwonandGitHub
|
4efd6ffde0
|
[DSV4] Refactor DeepseekV4Attention (#44569)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-04 20:23:07 -07:00 |
|
 Chris LeonardandGitHub
|
56aff0dd15
|
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to torch stale ABI. (#44334)
|
2026-06-04 20:14:43 -07:00 |
|
 zofiaandGitHub
|
063ce98fb7
|
[XPU][MoE] support block_fp8_moe on xpu (#42139)
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com>
|
2026-06-05 08:36:58 +08:00 |
|
 Bugen ZhaoandGitHub
|
62d6f06e3d
|
[Rust Frontend] Skip loading multimodal processor if --language-model-only is specified (#44500)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-04 17:02:54 -07:00 |
|