Shengqi Chen and GitHub
a02155c787
Merge branch 'main' into cuda-arch-fixup
2026-07-21 09:33:48 +08:00
Chris Leonard and GitHub
97a98006b0
Update qutlass cmake for stable abi ( #47879 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-20 18:31:10 -07:00
0a684ab0c0
[Bugfix] Fix WSL circular import from pin_memory warning_once ( #48444 )
...
Signed-off-by: AlejandroParedesLT <alejandroparedeslatorre@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-20 18:30:55 -07:00
0d9210a502
Fixes non-coalesced HBM access in marlin_int4_fp8_preprocess_kernel_awq ( #47268 )
...
Signed-off-by: xjx <493337577@qq.com >
Signed-off-by: flutist-alibaba <30485581+flutist@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-20 18:30:38 -07:00
1d874867ea
[Misc][Docs] Fix broken protocol link in speech_to_text doc ( #47212 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-21 01:05:05 +00:00
2e2e626b40
[Bugfix] Count per-group blocks in get_max_concurrency_for_kv_cache_config ( #48317 )
...
Signed-off-by: David Orman <ormandj@corenode.com >
Co-authored-by: Luke Alonso <lalonso@gmail.com >
Co-authored-by: Martin Vit <martin@voipmonitor.org >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-21 00:29:10 +00:00
Nick Hill and GitHub
af91f4b3e4
[Cleanup] Remove unused StructuredOutputRequest.status field ( #49235 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-21 00:06:16 +00:00
2396a61108
[Attention][MLA][DCP] Query replication for MLA decode (DeepSeek-V2/R1 + Kimi-K2.5) ( #45964 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-20 23:51:27 +00:00
97a668152b
[RL Infra][FlashInfer] Enable router replay output from FlashInfer monolithic MoE kernel ( #44214 )
...
Signed-off-by: Xuanyu Zhang <xuanyu.zhang@mistral.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-20 16:45:10 -07:00
58b2012aa2
[copy of #45208 ] CuMem slept-L1 fragmentation accounting ( #49208 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: Justin Wood <justin.m.wood@me.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: haosdent <haosdent@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Justin Wood <jwood@me.com >
2026-07-20 23:08:11 +00:00
Ning Xie and GitHub
b7c20d0cfa
[chore] adjust logo be more friendly to white background terminal ( #48938 )
...
Signed-off-by: Andy Xie <andy.xning@gmail.com >
2026-07-20 15:15:28 -07:00
TJian and GitHub
a2b1f9fc3b
[ROCm] [Release] [Bugfix] Fix the per commit wheel release pipeline. ( #49245 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-20 22:12:38 +00:00
642076d26c
Support loading sample_from_anchor flag from speculators config ( #48639 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-07-20 14:52:43 -07:00
Charlie Fu and GitHub
5feb3950e5
[ROCm][CI] fix test_rocm_quick_reduce.py ( #49234 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-07-20 16:39:33 -05:00
4ec199b66a
[Bugfix][Spec-Decode] Populate draft seq_lens_cpu_upper_bound for spec-decode attention metadata ( #44492 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Signed-off-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-20 20:50:58 +00:00
7ca017778f
[Feat][Perf] Add new warmup infrastructure for JITs ( #47451 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 13:21:55 -07:00
fbfe58133d
[Bugfix][KV Offload] Preserve reachable tails for hybrid SWA groups ( #48911 )
...
Signed-off-by: Colton Ottley <colton@ottleyengineering.com >
Co-authored-by: Colton Ottley <colton@ottleyengineering.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-20 22:12:36 +03:00
9dd62d80ab
Cosmos3 FP8 ModelOpt/Diffusers remapping ( #48952 )
...
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com >
Signed-off-by: wkutak <wkutak@nvidia.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-20 11:31:48 -07:00
f878367898
[Revert][Bugfix] Restore MiniCPM-V 4.6 ViT QKV weight loader ( #49193 )
...
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: wjinxu <1299461899@qq.com >
2026-07-20 18:17:46 +00:00
bd091079cb
[Attention] FlashAttention 4 SM100 FP8 kv cache support ( #42569 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 10:53:27 -07:00
b23bd73f54
[XPU]add sycl path for Mhc ( #47245 )
...
Signed-off-by: root <xiaolong.guo@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:32:54 +00:00
Bugen Zhao and GitHub
e2d7adeb64
[Rust Frontend] Bump xgrammar-structural-tag and enable local extension ( #49161 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-20 16:22:24 +01:00
Isotr0py and GitHub
15cb8e140d
[Multimodal] Allow keeping original image mode for ImageIO ( #49159 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-20 13:42:45 +00:00
f007cceb42
[KV Offload] Support self-describing KV events with TieringOffloadingSpec ( #48679 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 16:41:58 +03:00
0a5069e4e3
[Bugfix][Gemma4] Fix ModelOpt mixed-precision MoE config mapping ( #48563 )
...
Signed-off-by: wangqian <601731555@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 06:39:28 -07:00
8ce53a616e
[Bugfix] Zero new KV blocks for quantized + sliding-window hybrid caches ( #47574 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-20 13:18:17 +00:00
Lena Onyshchenko and GitHub
ae10e855ab
[Misc][Docs] Remove duplicate CodeGeex4 row in XPU model table ( #47210 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
2026-07-20 10:05:36 +00:00
hcl and GitHub
530ee36a0d
fix(openai): reject non-numeric logprobs with 400 instead of 500 ( #49144 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
2026-07-20 10:04:50 +00:00
Salt Sato and GitHub
d835ad572c
[Bugfix][Rust Frontend] Map missing prompt logprobs for single-token prompts in chat and raw generate ( #49111 )
...
Signed-off-by: Feathbow <feathbow@gmail.com >
2026-07-20 10:00:06 +00:00
47d0597ca2
[Misc][Docs] Fix broken csrc kernel links in fusions doc ( #47211 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-20 09:44:25 +00:00
Reid and GitHub
818cf61e91
[Rust Frontend] Fix macro-based content format detection ( #49042 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-20 09:39:13 +00:00
c01618fdc8
[Rust][Benchmark] Integrate vllm-bench to vllm-rs & vllm CLI ( #48930 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-20 09:31:25 +00:00
823eaf667d
[XPU] FP8 o_proj with fp8_bmm and load-time scale transpose ( #48334 )
...
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 16:32:03 +08:00
f1f1259692
[Rust Frontend] Use zero-copy slicing for multimodal tensors ( #48781 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-07-20 16:28:25 +08:00
df13b5aef5
[XPU] [MoE] add quant input when prepare for fusedmoe ( #47122 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:47:26 +08:00
4938d44a3b
[CPU] fixes heterogeneous NIXL KV transfer into CPU_ATTN decode workers ( #47871 )
...
Signed-off-by: Spycsh <sihan.chen@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-20 07:33:13 +00:00
37bf988c2f
[XPU][Bugfix] Fix GroupCoordinator device_index ( #47295 )
...
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:25:56 +08:00
aoshen02 and GitHub
9459fc6471
[Bugfix][RL] Set vLLM config during weight reload ( #45989 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-20 15:02:56 +08:00
5245c80564
[Doc] Document blocks_per_chunk in the KV offloading guide ( #49100 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-07-20 09:48:43 +03:00
9bc266d923
[Bugfix][KV Offload] Propagate EAGLE mode to SimpleCPU coordinator ( #49071 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 06:39:11 +00:00
5c9f6557d7
[Hardware][CPU] Enable granite-4 model on cpu ( #47641 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akashkaothalkar@dhcp-9-123-5-76.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akashkaothalkar@Akashs-MBP.lan >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@dhcp-9-123-5-76.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@Akashs-MBP.lan >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-20 06:15:16 +00:00
dcfebf93f4
[Bugfix] Fix logprobs token-string collision from SentencePiece space… ( #48674 )
...
Signed-off-by: Allen Shen <aoshen@inferact.ai >
Co-authored-by: mvanhorn <mvanhorn@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 12:17:18 +08:00
752bd10647
[ROCm][CI] Fix sparse MLA metadata sync fixture ( #49128 )
...
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-19 23:02:03 -05:00
Thien Tran and GitHub
2730b657c4
[Bugfix] Fix broken NVVM caused by CuteDSL 4.6.0 ( #49108 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-19 19:45:56 -07:00
1dcbbd9cac
[CI] Move compatible 1xL4 jobs to H200 35GB MIG ( #43024 )
...
Signed-off-by: Simon Mo <simon@inferact.ai >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-19 19:21:25 -07:00
ace9fda495
[CI/Build][BugFix][The Rock][AMD] Add spawn method in vision examples to avoid reinitialization ( #47932 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-19 13:41:52 -05:00
TJian and GitHub
ef0aa7ca2f
[ROCm] [Release] [Per-commit] Reenable per commit rocm wheel ( #49044 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-19 13:38:04 -05:00
Taneem Ibrahim and GitHub
e6d1310b2a
[Bugfix] Reject removed pooling parameters ( #48984 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-07-19 05:18:03 -07:00
yzong-rh and GitHub
ac5f38a0f7
[Refactor] Extract StructuredOutputsParams creation logic from Request.to_sampling_params ( #49003 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-19 05:18:00 -07:00
b6ff8a2f50
[Core] Add MRV2 virtual-batch PCP for MLA ( #46570 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Codex <noreply@openai.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-19 02:53:15 +00:00
9243e0124e
[Multimodal] Automatically fallback to ViT DP when TP is unavailable ( #49046 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-18 14:41:04 -07:00
Andreas Karatzas and GitHub
df362b2d6d
[ROCm][CI] Ensure sliding window tests release GPU memory ( #49055 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-18 20:44:05 +00:00
SYLAR and GitHub
7c2acd38b7
[Bugfix] Qwen3-VL/Qwen-Omni: honor max_pixels/min_pixels for video prompts ( #49015 )
2026-07-18 10:29:11 -07:00
yzong-rh and GitHub
a287eb163f
[Front-end] [Messages] Populate num_cache_creation_tokens ( #48535 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-18 13:04:35 -04:00
frida-andersson and GitHub
e94243893d
[ROCm][DSv3.2][Perf] Cap sparse MLA decode KV-splits with a work-per-split heuristic ( #46832 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-07-18 09:39:37 -07:00
29c0ec4d63
[ci] Move 3 entrypoints tests to h200_35gb queue ( #43164 )
...
Signed-off-by: Simon Mo <simon@inferact.ai >
Signed-off-by: Simon Mo <simon@simon-mac-mini-9.local >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-18 08:43:49 -07:00
Michael Goin and GitHub
c7ce03bcbd
[Bugfix] Bump tml-fa4 for cutlass-dsl 4.6 API compatibility ( #48988 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-18 05:59:33 -07:00
Harry Mellor and GitHub
c233d90aa8
Remove even more unnecessary load_weights methods ( #48496 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-18 08:40:27 +00:00
d96aee0951
[Bugfix] Re-sync parameter tp_rank after process_weights_after_loading (fix replicated / disable_tp weight reload) ( #48025 )
...
Signed-off-by: Alex Xu <alexxu@roblox.com >
Co-authored-by: YQ-Wang <yiqingwang@roblox.com >
Co-authored-by: alexhxu <alex.xu1015@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-18 08:40:06 +00:00
Francesco Fusco and GitHub
c71a583aa9
[Perf][Hybrid] Vectorize _copy_mamba_state_block to uint64 for temporal ( #48110 )
2026-07-18 04:43:09 +00:00
xuebwang-amd and GitHub
f12b80c6ef
[ROCm][Bugfix] Fix GPT-OSS Quark MXFP4 MoE loading - emulation buffer not block-aligned ( #43979 )
...
Signed-off-by: xuebwang-amd <xuebwang@amd.com >
2026-07-18 03:49:41 +00:00
Jee Jee Li and GitHub
da64db78b9
[LoRA] Optimize TrtLlmLoRAExperts ( #48759 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-18 10:26:14 +08:00
425c4eafb0
[Sampler] Stop upcasting logits to fp32 in apply_sampling_params ( #48641 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-17 18:15:20 -07:00
Michael Goin and GitHub
02c01f442b
[Model] Use standard ModelOpt config for Inkling NVFP4 ( #48990 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-17 18:13:14 -07:00
fae543015c
[Frontend]Flatten beam-search beams with itertools.chain instead of sum ( #48829 )
...
Signed-off-by: Wang Xingda <wangxingda1993@126.com >
Co-authored-by: 王兴达 <wangxingda@360itdeMacBook-Pro.local >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-17 23:09:14 +01:00
c9be3a8aa1
[Kernel][Helion] Disable warp specialization in rms_norm_per_block_quant B200 configs ( #48797 )
...
Signed-off-by: Shangdi Yu <shangdiy@meta.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-17 21:39:09 +00:00
41ea2dd44a
[Bugfix][V1/V2] Fix prompt_logprobs to respect logprobs_mode ( #47680 )
...
Signed-off-by: Wojciech Wais <wojciech.wais@gmail.com >
Signed-off-by: Federico Kamelhar <209537060+fede-kamel@users.noreply.github.com >
Signed-off-by: Allen Shen <aoshen@inferact.ai >
Co-authored-by: Wojciech Wais <wojciech.wais@gmail.com >
Co-authored-by: Federico Kamelhar <209537060+fede-kamel@users.noreply.github.com >
2026-07-17 21:58:59 +01:00
088c0be268
[CI] Fix macOS wheel release annotation context ( #48771 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-17 13:44:48 -07:00
fcd2255d16
[Hardware][GPU] Profiler config additional to increase it scope and annotation details ( #37524 )
...
Signed-off-by: devalshahamd <deval.shah@amd.com >
Signed-off-by: Deval Shah <devashah@amd.com >
Signed-off-by: Deval Shah <deval.shah@amd.com >
Co-authored-by: Deval Shah <devashah@amd.com >
2026-07-17 13:38:59 -07:00
Wentao Ye and GitHub
b5433b6f50
[Perf] Optimize dsv4 routing using specialized kernel, 2.94% E2E TPOT improvement ( #48660 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-17 13:35:06 -07:00
cc25f028b7
[Loader] Improve InstantTensor loading ( #46868 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-17 16:30:02 -04:00
c4cd2bd544
[Bugfix] MoRIIO toy P/D proxy: fix DP-rank index aliasing + harden for high-concurrency bursts ( #46115 )
...
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: QinPR <1905873179@qq.com >
Co-authored-by: Peiran Qin <66068739+QinPR@users.noreply.github.com >
2026-07-17 12:35:04 -07:00
5784507da4
[Attention] Allow selecting a different attention backend per KV-cache group ( #48012 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-17 15:19:02 -04:00
labAxiaoming and GitHub
bf578e1abd
[Bugfix][GLM4V] Fix video dummy profiling and memory usage ( #48729 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
2026-07-18 01:44:45 +08:00
Fangzhou Ai and GitHub
efed8a1e83
[ROCm][Perf][DSV4] Improve sparse decode reduction occupancy on gfx950 ( #48788 )
...
Signed-off-by: fai <fangzhouai@gmail.com >
2026-07-17 10:24:59 -07:00
11d291511a
[Bugfix][Tool Parser] Preserve whitespace in parameter values (MiniMax M2, Qwen3, MiniCPM5 XML) ( #48846 )
...
Signed-off-by: mosya415 <263250241+mosya415@users.noreply.github.com >
Signed-off-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
Co-authored-by: mosya415 <263250241+mosya415@users.noreply.github.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-17 16:45:41 +00:00
877dae9c68
[Refactor] Remove deepseek dead code ( #48780 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-17 14:57:13 +00:00
passtoor-agi and GitHub
c4dd6d78fd
Fix: Restore data_parallel_size > 1 for use_sequence_parallel_moe ( #48849 )
...
Signed-off-by: passtoor-agi <305788622+passtoor-agi@users.noreply.github.com >
2026-07-17 10:27:50 -04:00
JooHo Lee and GitHub
ce2aecc4dc
[Performance] Use CuTe-DSL for FlashInfer MXFP4 quantization ( #48417 )
...
Signed-off-by: BWAAEEEK <jooho414@gmail.com >
2026-07-17 06:53:48 -07:00
f38f3d11fb
[Bugfix][KV Offloading] Offload last block at request finish and prevent reuse race ( #48596 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-17 16:50:49 +03:00
d4b4562917
[XPU] Bump vllm_xpu_kernels to v0.1.11.1 ( #48942 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-17 20:43:23 +08:00
7b3192523e
[Bugfix]Fix transformer backend failed: AttributeError: 'Parameter' object has no attribute 'weight_loader' ( #48699 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-17 12:43:44 +01:00
Yejing Lai and GitHub
4c6e2e4b30
[XPU][UT]fix _POSSIBLE_KERNELS error on XPU ( #47516 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-17 11:19:05 +00:00
liuzhenwei and GitHub
8502958810
[XPU] support HND layout ( #47975 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-17 10:54:34 +00:00
ce4bdcbda4
[Bugfix] Enable FlashAttention MLA prefill for Mistral Small 4 head dims ( #48855 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-17 18:07:00 +08:00
liuzhenwei and GitHub
d5b1ec2684
[XPU] allow forcing flash attn for mm_prefix ( #48828 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-17 09:44:18 +00:00
867ff69733
[CI] Gate non-default release wheel builds ( #48772 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-17 02:16:44 -07:00
Sage and GitHub
109b736b86
[docs] preserve page path in stable-docs announcement link ( #48839 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-07-17 08:56:45 +00:00
69d4f5ef63
[Bugfix][Multimodal] Fix Qwen3-Omni use_audio_in_video with mixed image/video inputs ( #46213 )
...
Signed-off-by: wendadawen <wendadawen@qq.com >
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
2026-07-17 08:31:16 +00:00
426d48bfa1
[KV Offload] Add optional tier locality to FS/OBJ KV events ( #48281 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-17 10:31:52 +03:00
26c909ed74
[Model] Support TranslateGemma-12b-it ( #41599 )
...
Signed-off-by: Zhang Jian <jianmusings@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-17 07:17:59 +00:00
fb1d8ccaf5
[rl] Stateful Trainer Send: New Abstractions [1/N] ( #48042 )
...
Signed-off-by: haoaaron <ahao@anyscale.com >
Signed-off-by: Aaron Hao <ahao@anyscale.com >
Co-authored-by: Sumanth R Hegde <39546518+SumanthRH@users.noreply.github.com >
2026-07-17 15:11:06 +08:00
9354f22204
[Rust][Benchmark] Port in vllm-bench ( #48107 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-17 14:25:29 +08:00
aoshen02 and GitHub
17fdd42100
[Bugfix][Attention] Preserve post-load tensors across weight reloads ( #48251 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-17 14:15:26 +08:00
472d330c21
Add blocks_per_chunk configuration for KV offloading to support heterogeneous KV cache groups ( #48878 )
...
Signed-off-by: Debasish-87 <22btics06@suiit.ac.in >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-17 09:00:12 +03:00
3b6c96a101
[Bugfix][Pooling] Fix wrong scores for chunked prefill under torch.compile ( #48901 )
...
Signed-off-by: seewoo <seewoo@ucsc.edu >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-17 05:19:12 +00:00
Martin Hickey and GitHub
4d4e04f452
[Render] Add round trip parity test and docs for derender ( #48617 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-07-17 05:02:46 +00:00
Micah Williamson and GitHub
67fe73b2b4
[CI] Extend max-model-len for test_parsable_context to allow reasoning to finish ( #48873 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-17 11:36:44 +08:00
+1
ee8f36d0b3
[Warmup] Show CuTeDSL compilation progress ( #48881 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 20:17:51 -07:00
+1
f3e9497e92
[Model] Add Inkling LoRA support [4/N] ( #48884 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-17 09:52:15 +08:00
Thien Tran and GitHub
fe784ff22e
[M3] Improve indexer for long-context decode (sm100) ( #48582 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-16 18:12:15 -07:00
Daoyuan Li and GitHub
b88abb5036
[Misc] Remove orphaned env vars and stale env-var references ( #44749 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-07-17 00:00:47 +00:00
67f9046e4a
[Bugfix] Sparse MLA: enable fp8_ds_mla dense prefill ( #48642 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 22:44:04 +00:00
f17be06fbe
[Perf] Optimize clamp to clamp_ ( #48143 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 18:41:07 -04:00
2cab53ddee
[Model][Hardware][AMD]: Part 1/2 -> Enable e2e QK Norm + RoPE + KV Cache runtime fusion for Qwen3-30B-A3B on ROCM_AITER_FA, and ROCM_AITER_UNIFIED_ATTN ( #42749 )
...
Signed-off-by: Jack Hu <Jack.Hu@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-07-16 17:39:04 -05:00
ab0a20d151
[Docs] Add Phi-3.5-mini-instruct to batch invariance tested models ( #46396 )
...
Signed-off-by: Yuval Luria <yuvalluria@users.noreply.github.com >
Co-authored-by: Yuval Luria <yuvalluria@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-07-16 18:05:58 -04:00
4a394bfcda
[Spec Decode][DSpark] Add Gemma4-12B DSpark draft model ( #47216 )
...
Signed-off-by: DiegoCao <DiegoCao@users.noreply.github.com >
Co-authored-by: DiegoCao <DiegoCao@users.noreply.github.com >
2026-07-16 21:51:47 +00:00
Michael Goin and GitHub
c95c663049
[Quant] Add nvfp4_per_token online MoE quantization ( #48538 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-16 14:25:27 -07:00
HDCharles and GitHub
ab3c1aedf3
[Bugfix] Fix activation quantization dispatch for WNA4Int/WNA8Int ( #48785 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-07-16 17:13:02 -04:00
+1
fb5ec0dc9e
[Model] Add Inkling MTP=1 support [3/N] ( #48869 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 13:27:21 -07:00
971dac2caa
[Bugfix][KV-transfer] MoRIIO: retry RDMA send-queue-full backpressure instead of failing the read ( #47495 )
...
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io >
Signed-off-by: Edwin Lim <edwinlim0919@gmail.com >
Signed-off-by: harishk-mangoboost <harish.kambhampaty@mangoboost.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: harishk-mangoboost <harish.kambhampaty@mangoboost.io >
2026-07-16 20:02:27 +00:00
Shangdi Yu and GitHub
efa2e424f6
[Helion] Fix degenerate scale_ub in kernel input generators ( #48868 )
...
Signed-off-by: Shangdi Yu <shangdiy@meta.com >
2026-07-16 19:57:45 +00:00
02bf9c7907
Fix Quark mxfp4 quantized model loading issue under mtp ( #46757 )
...
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-16 14:58:47 -04:00
Woosuk Kwon and GitHub
f61163e6c7
[Model] Add Hopper FA4 relative attention for Inkling ( #48858 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-07-16 11:26:52 -07:00
Wentao Ye and GitHub
626c90b2d5
[Refactor] Move fla to third party ( #48500 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-16 19:22:36 +01:00
+1
251f7e478e
[Model] Add PW CUDA graph support for Inkling [2/N] ( #48822 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 10:28:56 -07:00
ce65385618
[KV Offload] Split tiering_lookup_delay into sync/async histograms ( #47679 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-07-16 20:03:45 +03:00
music-dino and GitHub
7d56fe2adc
[ROCm][CI] Avoid HIP init at config time via lazy aiter import in Quark OCP-MX ( #48015 )
...
Signed-off-by: Dino Music <Dino.Music@amd.com >
2026-07-16 16:47:08 +00:00
Zhongdongming Dai and GitHub
75bdad40b5
[Bug][Quantization] Fix humming is_layer_skipped for compressed-tensors "re:" ignore entries ( #48507 )
...
Signed-off-by: Zhongdongming Dai <zhongdongmin@nvidia.com >
2026-07-16 07:47:42 -07:00
d08eebad16
[Perf][MoE] Write FlashInfer combine into final output ( #47156 )
...
Signed-off-by: snordmann <snordmann@nvidia.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-16 17:21:03 +03:00
wang.yuqi and GitHub
3e90d015ba
[Frontend] Overlap preprocessing and computation for pooling models offline inference ( #47699 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-16 14:00:20 +00:00
7cd1d57b74
[CI/Build][Docker] Bump nvidia-cutlass-dsl to 4.6.0 and drop packaging workarounds ( #47442 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-16 13:51:41 +00:00
b8168e33e0
[ROCm][Perf][DSV4] Enable split sparse decode on gfx942 ( #46275 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 12:28:20 +00:00
ovidiusm and GitHub
d803b44dbe
[NIXL] Bump nixl to 1.3.1 ( #47559 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
2026-07-16 14:24:33 +02:00
530852f959
[KV Connector] Fix PD async scheduling race condition for hybrid attn models ( #48481 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: llx-08 <2596671364@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-16 11:42:37 +01:00
Nicolò Lucchesi and GitHub
a317bc5739
[Misc][Nixl] Unify _logical_to_remote_kernel_block_ids ( #48717 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-16 18:36:25 +08:00
a9531edfa6
[KV Offload] Define clean backend configuration boundary ( #48150 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Signed-off-by: Chang Guo <cguo51@asu.edu >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-16 13:27:05 +03:00
Reid and GitHub
8c3393f373
[Bugfix][Rust Frontend] Limit chat top_logprobs in responses ( #48134 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-16 09:41:44 +00:00
Ilia Yastrebov and GitHub
9f8cbfd8eb
Vectorize prep xfer list creation ( #48209 )
...
Signed-off-by: Ilia Yastrebov <iyastrebov@nvidia.com >
2026-07-16 11:39:05 +02:00
ea1d65fe6d
[Rust Frontend] Add Seed-OSS tool parser ( #47741 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-07-16 17:28:02 +08:00
f44f3d6f79
[Rust Frontend] Wait for mock engine endpoints before ZMQ connect ( #47965 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-16 09:20:31 +00:00
cc706b05a5
[Bugfix][Rust Frontend] Detokenizer: avoid leaking prompt on zero-generated-token completions ( #47707 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: xiaguan <751080330@qq.com >
2026-07-16 09:11:21 +00:00
Thien Tran and GitHub
85e296950c
BF16x3 router GEMM ( #47973 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-16 17:04:20 +08:00
Reid and GitHub
dc9f845ddc
[Rust Frontend] Fix mock engine test shutdown race ( #48738 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-16 07:27:43 +00:00
Elvir Crnčević and GitHub
12f2c515a7
[Bugfix] Fix offloading set_ overflow for packed non-uniform KV caches ( #48530 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
2026-07-16 10:04:39 +03:00
+1
6570c9800c
[Model] Add Inkling model support [1/N] ( #48799 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-15 23:40:07 -07:00
8bfd683901
[Spec Decode] Add kv_cache_dtype to speculative_config to control separately from target ( #48787 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-15 23:49:23 -06:00
Micah Williamson and GitHub
7dc2698632
[ROCm][CI] Set "highest" matmul precision for reference hf_runner in test_bert_for_masked_lm ( #48784 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-16 05:49:15 +00:00
ErenAta16 and GitHub
59b964f37d
fix(lora): validate LoRA rank is positive in PEFTHelper ( #48437 )
...
Signed-off-by: ErenAta16 <erena6466@gmail.com >
2026-07-16 05:22:52 +00:00
6a9f24aa8c
[ROCm][CI] Fix cuda graph mem profile issue ( #48764 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-16 04:23:18 +00:00
ba47bb5be1
Bump flashinfer version to 0.6.14 ( #47669 )
...
Signed-off-by: AmeenP <ameenp360@gmail.com >
Signed-off-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Pavani Majety <pmajety@nvidia.com >
2026-07-15 21:00:40 -07:00
df8a0900df
[BugFix] Don't apply weight in batch-invariant RMSNorm when has_weight=False ( #48741 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: Michael Gokhman <michael.gokhman@yahoo.com >
2026-07-16 11:59:03 +08:00
2db39c7049
[Bugfix][Spec Decode] Fix eagle3 first-layer qkv_proj prefix for quantized drafts ( #48068 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-16 11:34:41 +08:00
rongfu.leng and GitHub
3935829f89
[Docs] fix error key name ( #48802 )
...
Signed-off-by: rongfu.leng <lenronfu@gmail.com >
2026-07-16 03:16:16 +00:00
qli88 and GitHub
7746961277
[CI] Fix flaky lora test ( #47375 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Signed-off-by: qli88 <qiang.li2@amd.com >
2026-07-16 02:23:03 +00:00
qli88 and GitHub
5de1add806
[feature]Add int4 quantization support for emulation moe backend ( #48451 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-07-16 02:00:12 +00:00
Mike G and GitHub
915dffaa5f
[Attention] Mirror Triton KV dtype checks in MLA ( #47060 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-07-16 01:54:52 +00:00
nemanjaudovic and GitHub
81e13a0591
[Compilation] Skip x.size(dim) in _decompose_size_nodes ( #42543 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-07-15 17:59:14 -07:00
BRIJ RAJ KISHORE and GitHub
f95e3f0edb
[Tests] Gate Step3VL under Transformers v5 ( #44349 )
...
Signed-off-by: brijrajk <22271048+brijrajk@users.noreply.github.com >
2026-07-15 17:59:10 -07:00
5a65ba5f17
[Refactor] Move iteration logging to the frontend ( #46647 )
...
Signed-off-by: maxyanghu <hyoung2991@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Shang Wang <shangw@nvidia.com >
2026-07-15 17:59:05 -07:00
9d1c695be5
[XPU] Add DSpark speculative decoding support for DeepSeek-V4 ( #47677 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 17:59:02 -07:00
3c1bc1fc0d
[ROCm][Perf] Optimize sparse attention prefill kernel for DeepSeek-V4 ( #48519 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-07-15 17:58:59 -07:00
Michael Goin and GitHub
3a5e88e629
[Bugfix] Fix local speculators with dots in the name from classifying as custom_class ( #48754 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-15 17:58:06 -07:00
0becb7486b
[BugFix][MLA] Support kv_cache_dtype_skip_layers for MLA attention ( #47309 )
...
Signed-off-by: liuruikang <liuruikang.cs@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-16 00:06:11 +00:00
Wentao Ye and GitHub
2dab187f75
[Perf] Optimize fused_topk_bias for DSv4, 1.5~2x kernel performance improvement ( #47463 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-15 23:40:55 +00:00
Giuseppe Grossi and GitHub
015b0320de
Add giuseppegrossi to rocm label auto cc action ( #48643 )
...
Signed-off-by: giuseppegrossi <ggrossi@amd.com >
2026-07-15 16:38:56 -07:00
4238b011a7
[Feature] Migrate moe sp support to non-torch compiled path for GLM5.2 ( #47881 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 23:33:15 +00:00
kliuae and GitHub
eb33ff34dd
[ROCm][Perf] DSv4 two-stage compressor kernel for HCA prefill ( #47718 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
2026-07-15 23:31:51 +00:00
Mike G and GitHub
2bd8957627
[Bugfix][NVFP4 MoE] Pad gated intermediate to 64 for FlashInfer TRT-LLM shuffle (M%128) ( #46880 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-07-15 17:37:15 -04:00
Nicolò Lucchesi and GitHub
3034c8d389
[CI][PD] Add optional/nightly DSv4 Disaggregated eval ( #42310 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-15 21:04:54 +00:00
ecf4aa5ce2
[Bugfix] Fix FlashInfer non-causal draft attention (DFlash/DSpark) on Blackwell ( #48167 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-15 12:44:01 -07:00
49e777cf08
[CI][ROCm] Retry failed Docker build steps once ( #48773 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 14:31:23 -05:00
b7950e798f
[Bugfix] Initialize draft CUDA-graph keys for the native draft_model proposer ( #47460 )
...
Signed-off-by: Alagappan Valliappan <avalliappan@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 15:09:00 -04:00
de100ffb62
[Docs] Document pooling config resolution ( #48497 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 14:24:16 -04:00
Sage and GitHub
43cd340247
[Fix] Align OpenAI vllm_xargs value types across request schemas ( #48252 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
Signed-off-by: Sage <80211083+sagearc@users.noreply.github.com >
2026-07-15 17:48:24 +00:00
1d99f0f421
[ROCm][BugFix] Triton W4A16 handling for GPTQ/AutoGPTQ qzeros layout ( #47770 )
...
Signed-off-by: giuseppegrossi <ggrossi@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 11:55:47 -05:00
Andreas Karatzas and GitHub
0885b51981
[CI][ROCm] Stabilize ci_base hash calculation and image handoff ( #48746 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 10:56:56 -05:00
Xiaohong (Sean) Chen and GitHub
6036bf110a
[Kernel][Helion] Add Helion kernel benchmark script ( #48512 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-07-15 15:43:06 +00:00
Xiaohong (Sean) Chen and GitHub
2fa63e0fff
[Kernel][Helion] Helion kernel lazy registration ( #48264 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-07-15 15:42:46 +00:00
61141ed265
[Hardware][XPU] Register batch-invariant kernels for XPU ( #41934 )
...
Signed-off-by: tzielinski-habana <tomasz.zielinski@intel.com >
Signed-off-by: Tomasz Zielinski <85164140+tzielinski-habana@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Chendi.Xue <chendi.xue@intel.com >
2026-07-15 11:19:44 -04:00
05eed72aec
[ROCm] Re-enable cudagraph memory profiling, captured on the current stream ( #48526 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 10:03:48 -05:00
Gopala-Krishna Char and GitHub
5810e884f1
[Model] Add RobertaForTokenClassification / XLMRobertaForTokenClassification ( #47991 )
...
Signed-off-by: krishy91 <crgkc.r@gmail.com >
2026-07-15 14:30:33 +00:00
615834ee58
[KVOffload][P2P] Well-known default host/port env vars and per-DP-rank control port ( #47636 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 15:22:56 +03:00
Chaojun Zhang and GitHub
5811ed6a05
[Test][kv_offload] Fix flaky drain() helper in test_fs_tier.py ( #48545 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-15 14:53:47 +03:00
Tahsin Tunan and GitHub
1b30ae4ca4
[Rust Frontend] Fix flaky tls_handshake_timeout_drops_silent_client test ( #47873 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-15 11:05:38 +00:00
Tahsin Tunan and GitHub
4e04bcbce6
[Rust Frontend] Tolerate whitespace before the outer brace in JSON tool-call parsers ( #48034 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-15 11:03:37 +00:00
Nicolò Lucchesi and GitHub
66b6c684ab
[PD][Bugfix] Fix validation of cache shape for attn backends enforcing different kernel_block_size ( #48125 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-15 18:26:02 +08:00
c0302d9497
[Bugfix] Fix parallel_tool_calls=null crash in Responses API from_request() ( #48098 )
...
Signed-off-by: mahadrehmann <mahadrehman04@gmail.com >
Signed-off-by: Mahad Rehman <114791389+mahadrehmann@users.noreply.github.com >
Co-authored-by: muhammadfawaz1 <135441198+professorsab@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-07-15 18:01:17 +08:00
Jee Jee Li and GitHub
313fae3e89
[Bugfix] Fix GLM5 config ( #48711 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-15 09:55:39 +00:00
7aab6e2684
[ROCm][Bugfix] Enable the fp32 head_dtype torch.mm fast path on ROCm ( #48688 )
...
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 08:18:22 +00:00
9dd2e72828
fix flaky multi example connector consistency ( #48206 )
...
Signed-off-by: aarushjain29 <aarushi.jain2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 09:20:34 +02:00
Giuseppe Grossi and GitHub
d119beb1b9
[ROCm] Add tuned selective_state_update config for AMD MI350 ( #48159 )
...
Signed-off-by: Giuseppe Grossi <ggrossi@amd.com >
2026-07-15 10:18:09 +03:00
12a8057bfe
[CI/Build] Split release artifact annotations by type ( #48600 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-15 00:00:52 -07:00
e281ac663a
[Rust Frontend] Integrate MM audio support ( #48554 )
...
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-15 15:00:17 +08:00
adce068118
[ROCm][CI] fix test_common.py ( #48676 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 06:42:06 +00:00
b6770d7b54
[ROCm] Run init test engine in-process to avoid KV-cache OOM ( #48527 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 06:39:31 +00:00
3b39fd284a
[Bugfix][Spec Decode] Support heterogeneous QK fusion geometry ( #48671 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-14 22:37:10 -07:00
6472131298
[Bugfix] Set kv_quant_mode on the generic MLA KV-cache spec ( #48379 )
...
Signed-off-by: Mikhail Kostryukov <mike@triptrack.net >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-15 03:36:52 +00:00
37aa52821d
Build with ABI stable FlashMLA ( #48174 )
...
Signed-off-by: Jane Xu <janeyx@meta.com >
Signed-off-by: Shengqi Chen <i@harrychen.xyz >
Co-authored-by: Shengqi Chen <i@harrychen.xyz >
2026-07-14 20:29:28 -07:00
96d2ceda4b
[Security] Replace diskcache to eliminate pickle deserialization ( #44549 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-14 20:29:24 -07:00
Jee Jee Li and GitHub
fdf2cf66d3
[LoRA][1/N] Integrate flashinfer MoE LoRA for BF16 model ( #48632 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-15 10:54:00 +08:00
HDCharles and GitHub
9b2be4e9a5
[Quant] Enable humming w[2-7]a[4,8] inference with compressed-tensors ( #46390 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-07-14 20:22:31 -06:00
Andreas Karatzas and GitHub
3ad85e0de4
[CI][AMD] Configure MI300 tests for native execution without DinD ( #48387 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 02:14:16 +00:00
4f7fffb92f
[Core][LoRA] Support fp32 lm_head (head_dtype) on the LoRA path ( #48525 )
...
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 09:51:09 +08:00
6e073440b1
[ROCm][CI] Remove mxfp4 test skips after amd-quark 0.12 release ( #47330 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 01:25:05 +00:00
gnovack and GitHub
f7aadae5e5
add pad-aware reduce path ( #48385 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-14 18:05:50 -07:00
442c421e79
[Perf] Remove redundant repeat and copy for dsv4, 1.8% E2E TPOT improvement. ( #48137 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 00:48:10 +00:00
0bd6b85a1f
[Bugfix] Preserve unloaded non-persistent buffers during layerwise reload ( #44371 )
...
Signed-off-by: Joan Velja <joan.velja22@gmail.com >
Co-authored-by: Dakai An <77474977+andakai@users.noreply.github.com >
2026-07-14 17:46:29 -07:00
aoshen02 and GitHub
3ca242d1b6
[Bugfix][R3] Exclude draft routers from expert capture ( #48622 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-14 17:45:35 -07:00
Joe Rowell and GitHub
7e950521b3
fix: size FlashInfer prefill workspace to batch head footprint ( #48428 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-07-14 17:18:41 -07:00
Micah Williamson and GitHub
0f0f28b537
[Bugfix][CI] Fix test_head_dtype quant_method test on ROCm ( #48654 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-14 18:31:36 -05:00
520a20ba4e
[Bugfix] MoRIIO toy P/D proxy: add /health ( #45222 )
...
Signed-off-by: Chaemin Lim <chaemin.lim@mangoboost.io >
Signed-off-by: Edwin Lim <edwinlim0919@gmail.com >
Co-authored-by: Edwin Lim <edwin.lim@mangoboost.io >
Co-authored-by: Jaeyoun Kim <jaeyoun.kim@mangoboost.io >
Co-authored-by: Edwin Lim <edwinlim0919@gmail.com >
2026-07-14 22:56:46 +00:00
9182e86971
Log fully resolved pooling config at startup ( #48030 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-14 22:00:46 +00:00
Matthew Bonanni and GitHub
313d01f507
[CI][Bugfix] Fix FlashAttention reported MLA dimension support ( #48631 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-14 21:33:02 +00:00
Divakar Verma and GitHub
05d4f8bba3
[ROCm][CI] fix flashinfer import check ( #48647 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-07-14 20:54:19 +00:00
Michael Goin and GitHub
0b54201a04
[CI] Build macOS arm64 CPU wheel natively on the macmini queue ( #48289 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-14 19:40:26 +00:00
32e632dfeb
[Reasoning] Optimize TPOT for thinking budget when used with speculative decoding ( #46662 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-14 18:55:40 +00:00
7ffb98e248
[ROCm] Retune MI355 selective_state_update float32 config on the unified effective_batch grid ( #48373 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-14 18:26:35 +00:00
cdaa40d2a8
[KV Offload] Split cpu_cache_usage_perc into write/read usage gauges ( #47666 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-14 20:13:41 +03:00
ca3618bc69
[Doc] Sync four function docstrings with their signatures ( #45437 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-14 13:10:13 -04:00
Michael Goin and GitHub
b2f7d2560a
[Bugfix] Make MLA+SWA check the layer's backend, not the model config ( #48520 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-14 09:53:34 -07:00
Wentao Ye and GitHub
1ff9429655
[CI Bug] Fully solve accuracy issue for DSv3.2 + MTP + Sequence Parallel ( #48036 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-14 10:00:24 -04:00
af453e5647
[Bugfix] Gemma4 parser: classify channel-less output consistently in streaming and non-streaming ( #48262 )
...
Signed-off-by: Adhithya Balakrishnan <adhithya.b2004@gmail.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-14 09:30:16 -04:00
32aef44388
[Bugfix] Include inline per-token-head scales in offloaded page transfer width ( #48411 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-07-14 16:07:26 +03:00
7a74a9662b
[NIXL] Avoid reading expired blocks in bidirectional turn-2 read ( #47021 )
...
Signed-off-by: Tomer Gilad <tgilad@nvidia.com >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-14 13:03:41 +00:00
karthik and GitHub
b6754f536e
[Model] Enable LoRA support for tower and connector in LlavaNextVideo ( #48594 )
...
Signed-off-by: gangula-karthik <gkarthik923@gmail.com >
2026-07-14 20:09:38 +08:00
Juan Pérez de Algaba and GitHub
793cf79c89
[Bugfix][Security] Fix concurrent sparse invariant race bypassing CVE remediation ( #48583 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-14 11:08:24 +00:00
50ac1c7bab
[Misc] Rename VLLM_TRITON_ATTN_USE_TD to VLLM_TRITON_USE_TD ( #45781 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-14 10:32:57 +00:00
f04d3f640e
[Test] Enable KV cache events for HMA models in CPU offloading test ( #47754 )
...
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-14 12:22:27 +03:00
xiangdong and GitHub
0a9396a25e
[XPU][CI] Add tests/v1/e2e/general/test_correctness_sliding_window.py in Intel GPU CI ( #47231 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Signed-off-by: xiangdong <40376367+zxd1997066@users.noreply.github.com >
2026-07-14 08:50:16 +00:00
038ec293b1
[Bugfix] Return 400 instead of 500 when multimodal data is sent to a text-only model ( #48473 )
...
Signed-off-by: Hoang Nguyen Tien <hoang.nguyentien.2601@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-14 08:15:43 +00:00
894ebb27f5
Add Cosmos3 Edge Reasoner model ( #48291 )
...
Signed-off-by: Bartosz Stefaniak <bstefaniak@nvidia.com >
Co-authored-by: Bartosz Stefaniak <bstefaniak@nvidia.com >
2026-07-14 08:14:50 +00:00
Juan Pérez de Algaba and GitHub
c9a788eedc
fix(security): guard lm-format-enforcer regex compile with timeout ( #47595 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-14 07:18:11 +00:00
0762f2afeb
[Perf][Feat] Add generic cuteDSL LL BF16 router (GEMM) ( #42562 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-13 23:01:21 -07:00
31be872f55
[ROCm] Retune MI355 selective_state_update float16 config on the unified effective_batch grid ( #48372 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-14 05:16:29 +00:00
wangxiyuan and GitHub
94c0ef3001
[Misc] Clean up "swap_space" ( #48549 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-07-14 04:43:45 +00:00
Matt Woodson and GitHub
af1f036a70
[Bugfix] Skip minimax_m3 tool parser tests when Rust extension is absent ( #48523 )
...
Signed-off-by: Matt Woodson <mwoodson@redhat.com >
2026-07-14 04:43:22 +00:00
95aab66e95
[ROCm][MiniMax-M3][Spec Decode] Support speculative decode with AITER sparse PA ( #47984 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-07-14 04:12:53 +00:00
nemanjaudovic and GitHub
dcf4072da9
[Perf][ROCm] Fix GDN KKT warmup regression on RDNA by avoiding fp32 tl.dot ( #45000 )
...
Signed-off-by: Saeid Rostami <srostami@amd.com >
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-07-13 20:48:54 -07:00
382bbd5144
[ROCm][Kernel] Add HybridW4A16LinearKernel: Triton prefill + HIP skinny decode ( #40977 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-13 20:22:00 -07:00
b50ef9c6ed
[ROCm][MiniMax-M2] Dispatch fused QK-norm + AllReduce via AITER ( #44849 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: Pawel Kowalski <pawel.kowalski@amd.com >
2026-07-14 03:10:16 +00:00
Dan Blanaru and GitHub
9e289c553c
up FI fp8 moe topk to 32 ( #44462 )
2026-07-14 02:58:16 +00:00
c4f5cd60da
[1/N] Add dense MHA path for sparse MLA short sequences ( #47327 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-14 00:29:56 +00:00
0b0ef8d7eb
[Quantization][INC][ARK] Support INT2 XPU WOQ Linear ( #47521 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-14 08:29:45 +08:00
21472f32ea
add pad-aware swiglu limit kernel ( #48287 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 16:48:19 -07:00
fec64fea75
[BugFix] Correct OTEL span start time for Dynamo compilation ( #40698 )
...
Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-07-13 16:25:03 -07:00
8b8af2caf7
[Frontend] Expose logprob_token_ids on Python OpenAI endpoints ( #43463 )
...
Signed-off-by: Lang Zhao <lang.zhao@galileo.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-13 14:40:21 -07:00
Snehlata and GitHub
7738ef35b8
[Feat] Add Support for BertForMaskedLM to vLLM ( #48463 )
...
Signed-off-by: atalhens <sneh.lata@nutanix.com >
2026-07-13 20:56:25 +00:00
9a21f0d1a3
[BugFix] Initialize model_config for Qwen3-VL MoE ( #44863 )
...
Signed-off-by: wenpengw-nv <wenpengw@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-13 13:43:53 -07:00
Nick Hill and GitHub
8ac8375270
[Core] Preserve Marconi caching with selective hybrid cache retention ( #47782 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-13 21:24:20 +01:00
shanjiaz and GitHub
7dc447dda7
Added sliding window attention support for qwen-eagle3 architecture ( #47568 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-07-13 20:20:44 +00:00
7fc97042c3
Add DCP + Eagle support for Tokenspeed MLA backends ( #48180 )
...
Signed-off-by: Pavani Majety <pmajety@nvidia.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-13 11:46:02 -07:00
Micah Williamson and GitHub
18c4067a54
[ROCm][CI] Unblock AMD: Language Models Test (Extended Pooling) ( #48513 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-13 18:38:10 +00:00
550218b136
[Bugfix][Frontend] Flush engine reasoning parser at engine-reasoning → tool streaming boundary ( #47606 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-13 14:06:10 -04:00
Gavin Morris and GitHub
5c342876a6
[Doc] Add DeepseekV32ForCausalLM to supported_models.md ( #48293 )
...
Signed-off-by: Gavin Morris <gmorriscs@gmail.com >
2026-07-13 17:43:59 +00:00
9427c45386
[ROCm][CI] Transformers: pass only one of input_ids/inputs_embeds ( #48258 )
...
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-13 17:28:50 +00:00
43c8cbf79b
[EC Connector] CPU Offloading EC Connector ( #47423 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-13 20:09:41 +03:00
62286308c9
[Misc] Improve Matryoshka pooling dimensions validation ( #48057 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 12:57:36 -04:00
Nick Hill and GitHub
26587f9519
[BugFix][ModelRunner V2] Fix stale attn metadata in speculator prefill cudagraph capture ( #48261 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-13 09:39:15 -07:00
93e3bc8f30
[XPU][CI]Adjust timeout_in_minutes in Intel GPU CI ( #48418 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-13 23:11:16 +08:00
Yan Ma and GitHub
c2c9f7c5e2
remove force channels_last in Idefics3MultiModalProcessor ( #48467 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-07-13 14:18:57 +00:00
Omer Ullman Argov and GitHub
1be6e937b2
lower memory required for capturing cudagraphs for large cudagraph sizes ( #48483 )
...
Signed-off-by: Omer Ullman Argov <118735753+omera-nv@users.noreply.github.com >
2026-07-13 10:14:25 -04:00
Wentao Ye and GitHub
b3cfca996c
[Mypy Fix] Split mypy work ( #48490 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-13 12:42:42 +00:00
Bugen Zhao and GitHub
487dfb3418
[CI] Add SPDX license header to Rust/Protobuf sources ( #48472 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-13 10:22:47 +01:00
107a03ba63
[Core] Support fp32 lm_head for generation models via head_dtype (RFC #48305 §3.6) ( #48390 )
...
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-13 16:43:34 +08:00
56a357ed33
[Bugfix][KV Cache] Don't route uniform-page-size MLA+SWA models into DeepseekV4 packing ( #48256 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-13 08:16:24 +00:00
bea70c7cfc
[Attention] Make sliding-window support an explicit backend capability ( #48011 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-13 01:07:56 -07:00
Mohammad Miadh Angkad and GitHub
75fe92a316
[Distributed][Perf] Enable FlashInfer MNNVL allreduce RMS quant fusion ( #48064 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-07-13 15:02:59 +08:00
b7b58d1eba
[ROCm][CI] Cache Rust builds by source inputs ( #46527 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-07-13 01:14:08 -05:00
Canlin Guo and GitHub
36484e464a
[BugFix] Restore full tokens for Qwen MTP When MoE SP ( #48429 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-07-13 13:29:41 +08:00
9e57de7197
[CPU] Create Proper Numa topology for s390x ( #40714 )
...
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-13 12:58:43 +08:00
Yejing Lai and GitHub
8c5dafcd09
[Bugfix][UT]Fix EagleMiniCPMForCausalLM meet TypeError ( #48452 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-13 04:37:23 +00:00
05fa8183a6
[CPU][Spec Decode] Support DFlash speculative decoding for GDN models on CPU ( #46090 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-13 04:16:18 +00:00
d973cce3ca
Re-disable CUDA graph memory profiling on ROCm ( #48440 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-13 03:59:20 +00:00
775c1589ea
[Bugfix][ROCm] Keep TP all_gather on base-class collective ( #48446 )
...
Signed-off-by: fai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-13 03:53:53 +00:00
zzt and GitHub
2595d5cebc
[Model] Optimize Qwen3.5 on H20 ( #48350 )
...
Signed-off-by: zzt <zengzetang.zzt@antgroup.com >
2026-07-13 03:30:48 +00:00
ee5a89f4d7
[ROCm][MiniMax-M3] Add AITER sparse paged attention ( #47287 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-12 19:27:29 -07:00
e26264f3ef
[Kernel] Implement CUDA kernel for ReLUSquaredActivation (relu^2) ( #39058 )
...
Signed-off-by: Tanish Malekar <tanishmalekar32@gmail.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-12 19:18:03 -07:00
AlexHuang and GitHub
4c81772e8b
[Bugfix][KV Offloading] Fix stale transfer_jobs after reset_cache + harden job completion ( #48102 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
2026-07-12 20:00:04 +03:00
Bugen Zhao and GitHub
27c3e579f0
[CI][Rust Frontend] Pin cargo tool versions ( #48222 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-12 16:34:26 +01:00
8df14cfc8c
[EC Connector] Add EC Transfer Params ( #42433 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-12 14:35:33 +03:00
Jiangyun Zhu and GitHub
370b678a02
[CI][2/N] reduce CI time ( #48394 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-12 04:16:55 -07:00
5c0c987c03
Make tiering offload region DP-replica aware ( #47987 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-12 13:10:21 +03:00
Hugo Centeno and GitHub
5f8e73cb8b
[Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions ( #48330 )
...
Signed-off-by: hcenteno <hugo.centeno@estudiantat.upc.edu >
2026-07-12 09:39:27 +00:00
83762b77b0
[Frontend] Add /abort_requests to the RLHF dev API router ( #47173 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-12 14:21:02 +08:00
a02984ed47
[Perf][Qwen] Replace MOE all-reduce with reduce-scatter ( #47006 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: yewentao256 <zhyanwentao@126.com >
2026-07-12 06:14:49 +00:00
fc1c548093
Runtime Draft Weight Update for Speculative Decoding ( #46725 )
...
Signed-off-by: vx120 <893600387@qq.com >
Signed-off-by: vx120 <57470515+vx120@users.noreply.github.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: crp0128 <191679376@qq.com >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-11 22:51:53 -07:00
481e481be7
[2/N][Core] support partial prefix cache hit for hybrid model ( #46384 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-12 05:37:51 +00:00
zhao, zhenhui and GitHub
8e981630c9
[CI][CPU] Add Qwen2-VL multimodal tests for CPU backend and fix incompatibilities ( #48072 )
...
Signed-off-by: Zhenhui Zhao <zhenhui.zhao@intel.com >
2026-07-12 12:30:34 +08:00
Alejandro Paredes La Torre and GitHub
9a48eef89a
[Bugfix][LoRA] Support ark_linear base layer in _get_lora_device ( #47690 )
...
Signed-off-by: AlejandroParedesLT <alejandroparedeslatorre@gmail.com >
2026-07-12 00:13:50 +00:00
Jiangyun Zhu and GitHub
1ef1c7ebba
[CI] split tests to reduce CI time ( #48219 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-11 13:00:14 -07:00
54503ecec0
fix(processor): route MiMo-V2-Omni media fetch through MediaConnector ( #43117 )
...
Signed-off-by: Ievgen Bondarenko <ibondarenko@student.sierracollege.edu >
Signed-off-by: Ievgen (Jack) Bondarenko <ibondarenko@student.sierracollege.edu >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-11 15:52:52 +00:00
ErenAta16 and GitHub
0067311536
fix(entrypoints): stop resolve_items leaking in-flight media fetch tasks on partial failure ( #48333 )
...
Signed-off-by: ErenAta16 <erena6466@gmail.com >
2026-07-11 15:42:08 +00:00
51878e5b6e
[2/N][KV-Cache Layout Refactor] Pack K/V into the content dim across attention backends ( #44455 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-11 11:11:16 -04:00
Yejing Lai and GitHub
76fedaa2a5
[XPU][UT]Fix InternS1ProForConditionalGeneration AssertionError ( #48232 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-11 13:56:55 +00:00
19069bcbd5
FP32 router GEMV optimization ( #48335 )
...
Signed-off-by: peiyuanz <peiyuanz@inferact.ai >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: peiyuanz <peiyuanz@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: zhouzhou <zhouzhou@zhouzhoudeMacBook-Pro.local >
2026-07-11 13:07:48 +00:00
Harry Mellor and GitHub
1bd8f80a64
[CI] Point CI at Transformers release rather than release branch ( #48328 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-11 02:31:14 -07:00
0b6636cbcb
[XPU]remove is_xxx from moe class and bump up kernels ( #48079 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-11 09:27:13 +00:00
Harry Mellor and GitHub
4a6440acef
Bump Transformers version to 5.13.0 ( #47867 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-11 00:56:14 -07:00
Lucas Wilkinson and GitHub
bec0a4ede6
[Revert] [Build] Update vllm ...builds FA3 with torch stable API ( #48269 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-11 05:20:25 +00:00
3d99b0499a
[Logs] DP Supervisor Log Improvement ( #48278 )
...
Signed-off-by: Robert Shaw <robertgshaw2-redhat@ip-172-31-18-125.us-east-2.compute.internal >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@ip-172-31-18-125.us-east-2.compute.internal >
2026-07-11 12:07:00 +08:00
04d553f390
[Misc] Use meta tensor for KV cache stride calculation ( #47316 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-10 23:24:59 -04:00
9c18e90f6c
[BugFix] Fix packed HND KV cache reshape for FlashAttention ( #47314 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-10 23:22:39 -04:00
Jimmy Lee and GitHub
092387963c
[BugFix] weights processing peak memory reduction for nvfp4 MoE layers ( #46276 )
...
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com >
2026-07-11 02:05:35 +00:00
1bf3997eae
[Quantization] Bound peak memory when repacking FP4 MoE weights for Marlin ( #47851 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-10 19:46:13 -06:00
29fd688892
Add VLLM_FLASHINFER_AUTOTUNE_SKIP_OPS and skip CuTeDSL fp4_gemm autotuning by default ( #48268 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-10 18:13:08 -07:00
Ashwin Giridharan and GitHub
ed908cf0a0
[Bugfix] Fix thinking_token_budget not enforced after natural </think> re-entry ( #45984 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
2026-07-10 22:47:51 +00:00
26ff616bbf
[Bugfix][Test] Register Qwen/Qwen3.5-4B example model ( #48276 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-10 17:01:23 -04:00
gnovack and GitHub
f378f79b7c
handle topk_ids padding in align sum kernel ( #47785 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-10 13:33:28 -07:00
735def4fcf
[Bugfix] Fix FlashMLA dense fp8 metadata crash (num_sm_parts clamp) ( #48045 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-10 12:24:52 -07:00
c227aaa3f8
[ROCm] Enable DeepSeek-V4 DSpark speculative decoding on AMD (MI350X / MI355X, gfx950) ( #47419 )
...
Signed-off-by: larryli2-amd <larryli2@amd.com >
Signed-off-by: larryli2-amd <Larry.Li@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-10 23:22:19 +08:00
Michael Goin and GitHub
08dfd68610
[Model] Add LongCat-Flash-Lite (n-gram embedding) ( #47857 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-10 07:17:50 -07:00
Tyler Michael Smith and GitHub
978a6dfa3f
[Build/CI] Build arm64 PR and postmerge image builds for Blackwell SM10x and SM110 ( #48041 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-07-10 10:12:50 -04:00
85c09e9885
fix: correct load_weights track logic and enable weight integrity for… ( #41811 )
...
Signed-off-by: Yipeng Hu <i26268@metax-tech.com >
Signed-off-by: HuYiPeng <144002351+MynameFelix@users.noreply.github.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Yipeng Hu <i26268@metax-tech.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-10 14:08:20 +00:00
b12cca6a23
[Bugfix] Fix turboquant FP8 cast failure for BF16 models on Ampere GPUs ( #39988 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Hoseung Kim <ghyutjik123@gmail.com >
2026-07-10 06:55:28 -07:00
Wentao Ye and GitHub
e257faf87d
[Refactor] Remove unused rocm kernel combine_topk_swa_indices_ragged ( #48158 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-10 09:33:07 -04:00
FAN YUCHEN and GitHub
fabec87f63
[Model] Migrate MistralLarge3ForCausalLM to AutoWeightsLoader ( #48153 )
...
Signed-off-by: Yuchen Fan <functionhx@gmail.com >
2026-07-10 12:27:58 +00:00
7614b88ebd
[Bugfix][Spec Decode] Fix DFlash draft/target layer-count mismatch ( #48113 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-10 04:42:53 -07:00
Isotr0py and GitHub
68ea76e780
[Misc] Remove dead code in ViT functionality test ( #48220 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-10 11:17:42 +00:00
c241c7a2b0
[Rust Frontend] Add roundtrip fixtures for more chat parsers ( #47883 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-10 10:03:58 +00:00
e23b19309b
Deepstream video backend ( #42424 )
...
Signed-off-by: Viranjan Pagar <vpagar@nvidia.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-10 02:23:30 -07:00
f36284a8d2
[CI] Add TORCH_NIGHTLY=1 build mode (run full suite on torch nightly) ( #47180 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
Co-authored-by: Kevin H. Luu <khluu000@gmail.com >
2026-07-10 01:38:35 -07:00
Mingfei Guo and GitHub
424df4f65d
[Model][CI/Build] Cosmos3: enable registry tests and register Cosmos3-Super ( #48211 )
...
Signed-off-by: Mingfei Guo <1800012773@pku.edu.cn >
2026-07-10 16:35:15 +08:00
Bugen Zhao and GitHub
074bdd0d99
[Rust Frontend] Integrate MM video support ( #47959 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-10 08:15:33 +00:00
216ee58780
Add XPU nightly and release image publishing to DockerHub ( #48126 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: jun,du <jun.du@intel.com >
Co-authored-by: jun,du <jun.du@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-10 00:56:00 -07:00
433f291195
[CI] Right-size test-area timeouts from nightly durations ( #48186 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-10 00:53:16 -07:00
Chaojun Zhang and GitHub
28eaf05d56
[XPU] Enable v1/sample tests on XPU CI ( #44472 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-10 15:40:51 +08:00
Jiangyun Zhu and GitHub
300e33797f
[Perf] fuse more rmsnorm and all-reduce in qwen3.5 ( #46998 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-10 15:37:51 +08:00
5715fde12c
[Feature][Parser] Support include_reasoning param for non-Harmony models ( #44301 )
...
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-07-10 15:34:02 +08:00
e5588e49bc
[Core][KV events] Report prefix-cache-reused blocks in full report mode ( #45261 )
...
Signed-off-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 22:46:54 -07:00
95ed0feaa5
DCP supports hybrid attention ( #40996 )
...
Signed-off-by: YanXu <yancey.yx@alibaba-inc.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-09 21:34:45 -07:00
2d814a0082
[kv_offload] Emit tier-owned BlockStored events from FS/OBJ secondary tiers ( #47923 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-10 06:17:23 +03:00
88e5e2c57b
[CI/Build][AMD] Fix ROCm OOM in eagle_correctness_heavy by reserving CUDA graph memory ( #47366 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-10 02:14:38 +00:00
Augusto Yao and GitHub
feb384ada2
[bugfix] bge-m3-sparse-plugin mismatch requests ( #48112 )
...
Signed-off-by: augusto.yjh <augusto.yjh@antgroup.com >
2026-07-10 10:03:00 +08:00
a0f6d767e4
[ROCm][CI] Move remaining engine/samplers AMD steps to mi325_1 ( #48169 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-10 00:20:15 +00:00
gnovack and GitHub
f1a5adddb8
update marlin M size for EP ( #48144 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-09 23:52:52 +00:00
ap9272 and GitHub
cac3e70cd4
Correct model layer aliasing for Bert style models ( #43896 )
2026-07-09 19:46:22 -04:00
Lucas Wilkinson and GitHub
e12b91b032
[CI] Fix cargo-deny config flag ordering ( #48170 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-09 21:43:58 +00:00
Micah Williamson and GitHub
766469a4c4
[ROCm] Revert Part of [ROCm] Fix pooling startup workspace lock #47912 ( #48154 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-09 20:34:24 +00:00
Lucas Wilkinson and GitHub
ea0fa34f49
[CI] Increase extract hidden states TP2 timeout ( #48161 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-09 16:19:03 -04:00
ZihaoMu and GitHub
bbb0f945ff
[ROCm] Synchronize sparse MLA metadata before graph replay ( #47404 )
...
Signed-off-by: zihaomu <zmu@amd.com >
2026-07-09 14:59:56 -05:00
2ded1b24e7
[KV Connector][Mooncake] Apply SWA lookup mask before hashing/key build ( #47317 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 19:51:23 +00:00
b0dec2a11b
[ROCM][DSV32][Perf][MTP] Enable UNIFORM_BATCH CG mode in rocm_aiter_mla_sparse ( #45149 )
...
Signed-off-by: Teemu Virolainen <teemu.virolainen@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-09 14:35:10 -05:00
ff8d3488f2
[Bugfix][MRV2] Reset num_accepted_tokens on add_request in all modes ( #48132 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-09 18:17:11 +00:00
weishu and GitHub
2285cfca46
[KVConnector] MultiConnector: give every sub-connector the request's real blocks in update_state_after_alloc ( #46865 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-07-09 11:10:39 -07:00
e08a915146
[Bugfix] Preserve tensor causal metadata for grouped attention ( #48135 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-09 17:57:53 +00:00
Charlie Fu and GitHub
67e7ea8977
[ROCm][CI] Set all timeout_in_minutes to 180 ( #48146 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-07-09 17:52:26 +00:00
429f405748
[Bugfix] Guard CUDA-only rms_norm_per_block_quant in FUSED_OPS for non-CUDA builds ( #47296 )
...
Signed-off-by: Tsvika Shapira <tsvika@moonmath.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-09 10:10:53 -04:00
Brandon Pelfrey and GitHub
753c5039f0
Pin PyNvVideoCodec to tested 2.0.4 wheel ( #48056 )
2026-07-09 07:07:50 -07:00
299d2b5655
[CI] Annotate built Docker image tags on the Buildkite build page ( #48101 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-09 22:02:40 +08:00
85b3a7264b
[Bugfix][Model Runner V2] Order uniform decodes first so spec decodes aren't misclassified as prefills ( #47381 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-09 14:26:27 +01:00
Harry Mellor and GitHub
b83be00cdd
Migrate Olmo and Olmo2 to the Transformers modeling backend ( #48100 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 05:00:23 -07:00
412414d8e0
Remove PersimmonForCausalLM and FuyuForCausalLM model architectures ( #48096 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-09 04:59:08 -07:00
ae6170f874
[P/D][Bugfix] Fix PD async KV load lookahead handling for MTP spec decode ( #46694 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-09 10:01:22 +00:00
e87521626f
Sanitize server file paths from validation error responses ( #46415 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-09 17:46:29 +08:00
1cd75b3dd4
[Bugfix] Fix race condition in KVBlockZeroer ( #48085 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-09 09:18:19 +00:00
0206f10871
Add Intel XPU Docker release pipeline ( #47880 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: jun,du <jun.du@intel.com >
Co-authored-by: jun,du <jun.du@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-09 01:13:51 -07:00
ab7961a14a
Remove TeleChatForCausalLM ( #47989 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 00:34:54 -07:00
a07765c6bd
[Bugfix] Fix Qwen3-ASR transcription streaming postprocessing ( #42478 )
...
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
Signed-off-by: JooHo Lee <jooho414@gmail.com >
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
2026-07-09 00:33:27 -07:00
Li, Jiang and GitHub
1171467e91
[CPU] Fix Qwen-Next SSM type for AMX GDN ( #48073 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-09 15:09:31 +08:00
Chauncey and GitHub
529af88842
[KV Offloading] Add free block iterator for CPU offload scheduling ( #47849 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-09 06:59:54 +00:00
Chaojun Zhang and GitHub
b8c7c86533
[XPU][LoRA] Fix torch.compile DEVICE_LOST by avoiding view-mutation in LoRA shrink ( #47944 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-09 06:08:30 +00:00
2c17d33f42
[Bugfix][ROCm] Change AttentionCGSuppoort in TritonMLA to UNIFORM_SINGLE_TOKEN_DECODE ( #47144 )
...
Signed-off-by: Dino Music <Dino.Music@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-08 21:09:42 -05:00
bc44f9feb7
[ROCm][CI][MoE] Fix double-transpose of fused w3 expert weights ( #47874 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 16:59:28 -07:00
7802c20c4e
[KVConnector][NIXL] Support pipeline-parallel prefill in push mode ( #45880 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 16:49:23 -07:00
95d6d6f4bb
[Bugfix] Use int8 workspace for FlashInfer MLA decode ( #48046 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-08 23:39:40 +00:00
Harry Mellor and GitHub
56da398dac
Fix embed scaling + CUDA graphs in Transformers modelling backend ( #48010 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 00:14:33 +01:00
Andreas Karatzas and GitHub
26831949b4
[ROCm] Fix pooling startup workspace lock ( #47912 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-08 17:59:50 -05:00
6cf7b26bd4
[docs] Fix the docs build ( #48008 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-08 15:47:22 -07:00
Roberto L. Castro and GitHub
5f85975624
[Feat] Add runtime monitor for post-warmup TileLang compilation ( #46718 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
2026-07-08 22:11:28 +00:00
dcdd756d75
[CI] GSM8K eval integration test for KV offloading ( #46893 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 17:59:48 -04:00
Thien Tran and GitHub
0d2f4e7c9c
Allow FlashInfer A2A backends for TRTLLM FP8 MoE Modular ( #46661 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-08 14:58:39 -07:00
djramic and GitHub
49abadaedb
[ROCm][Bugfix] Fix empty-tensor .max() crash in AITER FA ( #47894 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-07-08 16:58:36 -05:00
Kaihang Jiang and GitHub
089e412878
[Perf] Integrate TRTLLM BF16 MoE Modular Kernel ( #45182 )
...
Signed-off-by: Kaihang Jiang <kaihangj@login-lyris02.lyris.clusters.nvidia.com >
2026-07-09 01:36:14 +04:00
Nick Hill and GitHub
a5d19cbb95
[Core] Move MRV1 late_interaction_runner.py out of MRV2 subtree ( #48014 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 18:30:11 +00:00
Chris Leonard and GitHub
8347c6e6e1
updated flash_attn GIT_TAG to point to torch Stable ABI FA3 commit ( #47995 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-08 10:56:29 -07:00
Shengqi Chen and GitHub
319db65b68
Merge branch 'main' into cuda-arch-fixup
2026-07-09 01:55:33 +08:00
b2cf70ea3a
[CI] BugFix Eval Small Models Distributed test for DiffusionGemma ( #47980 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-07-08 17:00:56 +00:00
almayne and GitHub
d1f1d86797
[Bugfix] Re-enable benchmarking of librispeech dataset. ( #47033 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
2026-07-08 16:19:26 +00:00
shawn and GitHub
f05603fa28
[Bugfix][DCP] Cast LSE to fp32 in a2a combine to fix bf16 bitcast crash ( #47801 )
...
Signed-off-by: Shawn Tsai <shawnyht@gmail.com >
2026-07-08 11:41:26 -04:00
c2ecd0f888
Fix FlashAttention MLA prefill V unpadding ( #42642 )
...
Signed-off-by: Martin Vit <martin@voipmonitor.org >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-08 15:22:20 +00:00
0d12618e98
[Spec Decode] Support hybrid (SWA + full attention) DFlash drafters ( #47914 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-08 11:12:45 -04:00
Tyler Michael Smith and GitHub
68b4a1d582
Fix NVML capability lookup for visible devices ( #47892 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-07-08 09:07:44 -04:00
572b25b03e
[Bug] Fix Batched DeepGEMM ( #47884 )
...
Signed-off-by: Robert Shaw <robertgshaw2-redhat@dgx-b200-02.mgmt.accl-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@dgx-b200-02.mgmt.accl-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:05:03 -04:00
9f2b3b093c
Improvement of Docker image build for IBM Power using prebuilt wheels from IBM published devpi index ( #46017 )
...
Signed-off-by: vivek sharma <vivsharm@redhat.com >
Signed-off-by: puneetsharma21 <puneet.sharma21@ibm.com >
Signed-off-by: Puneet Sharma <puneet.sharma21@ibm.com >
Co-authored-by: vivek sharma <vivsharm@redhat.com >
Co-authored-by: Puneet Sharma <puneet.sharma21@ibm.com >
Co-authored-by: depthfirst-app[bot] <184448029+depthfirst-app[bot]@users.noreply.github.com>
2026-07-08 13:01:14 +00:00
cd0de48d08
[Bugfix][V1] Free out-of-window blocks on the processed-token basis under async scheduling ( #47728 )
...
Signed-off-by: Saddss <28726669061@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Saddss <28726669061@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 13:34:19 +01:00
rasmith and GitHub
934eeaecfb
[CI/Build][BugFix][The Rock] Fix get_ssm_device_name to return sanitized, usable filename ( #47781 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-07-08 12:12:54 +00:00
Bugen Zhao and GitHub
2cae98dfa5
[Rust Frontend] Handle continue_final_message with renderer sentinel ( #47844 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-08 12:57:04 +01:00
db39d60010
Add tuned selective_state_update float32 config for AMD Instinct MI355 ( #47943 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:41:26 +00:00
a1ab51afb6
[Bugfix] Allocate HY V3 expert_bias in float32 to prevent silent downcasting ( #47797 )
...
Signed-off-by: 辰言 <oncwnuIWp30GguOyJ615Fqj8H-yc@git.weixin.qq.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: 辰言 <oncwnuIWp30GguOyJ615Fqj8H-yc@git.weixin.qq.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-07-08 11:25:53 +00:00
Thien Tran and GitHub
e7b3853bac
Remove router weight upcast for DSv2-related models ( #47970 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-08 11:19:10 +00:00
eeaf23107f
[ROCm] Add tuned selective_state_update float32 config for AMD Instinct MI300X ( #47947 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:09:02 +00:00
Canlin Guo and GitHub
285c08c036
[Model] Support MOSS-Transcribe-Diarize ( #47729 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-07-08 04:05:45 -07:00
1f4ad059d1
[ROCm] Add tuned selective_state_update float16 config for AMD Instinct MI300X ( #47945 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:03:58 +00:00
04a703e397
[Frontend] Support bad_words in the /v1/completions endpoint ( #46793 )
...
Signed-off-by: sungbin1015 <sbin@solbox.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:51:17 +00:00
Nicolò Lucchesi and GitHub
bd3bb4eb26
[Misc][Docs] Add human-readable integer support for more cli-args ( #47608 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 09:43:34 +00:00
Chaojun Zhang and GitHub
440002552e
[XPU] [Fusion passes] Disable fuse_rope_kvcache_cat_mla & qk_norm_rope_ fusion on XPU ( #47962 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-08 09:23:06 +00:00
99a85617bf
[Test] Skip DeepEP MoE layer tests without P2P access ( #47946 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:46:03 +01:00
Nicolò Lucchesi and GitHub
7c67da967f
Remove unused _get_kv_cache_config_deepseek_v4 alias ( #47969 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 01:18:03 -07:00
Nicolò Lucchesi and GitHub
d79855eaac
[Docs] kv_sharing_fast_prefill correction ( #47044 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 01:17:46 -07:00
51e5372f3d
[Model][HunyuanVL] Use native transformers processor and adapt to transformers 5.13 ( #47872 )
...
Co-authored-by: manayang <manayang@tencent.com >
2026-07-08 07:58:23 +00:00
Ace Eldeib and GitHub
7cc2e8e74f
fix: hash speculative draft model config ( #47911 )
...
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com >
Signed-off-by: Ace Eldeib <alexeldeib@gmail.com >
2026-07-08 08:36:30 +01:00
Hongxia Yang and GitHub
2c64b4c1cc
[ROCm] fixed aiter master flag and expert parallelism compatibility on minimax-m3-mxfp8 ( #47158 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-07-08 15:26:17 +08:00
d35eba302f
[Bugfix] Avoid leaking Pydantic repr in tool_choice error message ( #47028 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-08 15:00:59 +08:00
Nicklas Frahm and GitHub
c0e8e1f12a
[Bugfix] Register VLLM_BUILD_* and VLLM_IMAGE_TAG provenance env vars ( #45313 )
...
Signed-off-by: Nicklas Frahm <nicklas.frahm@gmail.com >
2026-07-08 06:21:12 +00:00
Zach Zhu and GitHub
5d5fab0061
[Bugfix][Frontend] Fix http_requests_total metric recording some 4xx errors as 5xx ( #44303 )
...
Signed-off-by: Zach Zhu <zzqshu@126.com >
2026-07-08 05:33:21 +00:00
2afa3f7e95
[Perf] Minimax M3 - Support cross-layer allreduce-norm fusion ( #47631 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-07-07 21:16:32 -07:00
80eb01e93d
[Bugfix] DSV4 TP16 garbage output ( #47493 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-07-07 21:04:33 -07:00
d9e57ea82e
[ROCm][Perf] MXFP8 dense-linear + grouped-MoE GEMM optimizations for MiniMax-M3 ( #46117 )
...
Signed-off-by: amd-ethany <amd-ethany@users.noreply.github.com >
Co-authored-by: amd-ethany <amd-ethany@users.noreply.github.com >
2026-07-08 04:03:34 +00:00
9021589498
[Minimax-M3] Using tok_sparse_select from MSA instead of triton kernels ( #47502 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 21:01:12 -07:00
Ting SUN and GitHub
0303f37a54
[Bugfix][Pooling] Align CrossEncoder token type ids after truncation ( #47772 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-08 03:59:22 +00:00
Walter Beller-Morales and GitHub
dd127d82ed
[Core][Engine] only materialize tokens when thinking budget is in req ( #47053 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-07-07 21:02:38 -06:00
0ca6eee743
[Core] Pass request context to CPU offload cache policy touch ( #47744 )
...
Signed-off-by: jacklin78911-collab <jacklin78911@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-08 05:56:25 +03:00
Isotr0py and GitHub
5e975eae1a
[Bugfix] Avoid blocking model launching when no system ffmpeg available for TorchCodec ( #47888 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-08 10:52:25 +08:00
Martin Hickey and GitHub
f7fc0ca993
[Frontend] Add endpoint plugins framework ( #47454 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-07-08 10:00:41 +08:00
Rahul Vishwakarma and GitHub
f7efab58ec
[CPU][Bugfix] Fix flaky ShortConv prefill test on ARM (uninitialized weights) ( #47848 )
...
Signed-off-by: Rahul Vishwakarma <Rahul.Vishwakarma2@ibm.com >
2026-07-07 18:20:09 -07:00
e97c3cb303
[Core] Persist and reuse the memory-profiling result across boots (opt-in) ( #47388 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 00:53:02 +00:00
4aceabf8c1
[ROCm][Bugfix] Key sparse-MLA persistent metadata on per-request context lengths ( #47766 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 19:22:34 -05:00
stefankoncarevic and GitHub
6e35c5e5af
[ROCm][CI] Minimize comment in RocmAttention q_scale check ( #47731 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
2026-07-07 19:16:08 -05:00
aad0fb741b
[CI/Build] Accept ready-run-all-tests label in pre-commit gate ( #47897 )
...
Signed-off-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-07 23:18:59 +00:00
yzong-rh and GitHub
7d2ce5750e
[Bugfix] Patch Hopper MXFP4 OOB scales reads leading to NaN ( #47910 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-07 22:51:22 +00:00
Juan Pérez de Algaba and GitHub
675f4295cd
fix(security): bound completion prompt list to prevent unbounded engine fan-out ( #47845 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-07 22:48:20 +00:00
Jason Li and GitHub
d99adcebdc
[BugFix] Fix ModelOpt quantization inference for fused siblings ( #47445 )
...
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
2026-07-08 03:19:42 +05:00
c8c2f838e7
Add tuned selective_state_update config for AMD Instinct MI355 ( #47767 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-07 22:19:08 +00:00
dd0d74cd92
[Doc] Surface the --kv-cache-memory suggestion at INFO and document fast-startup knobs ( #47374 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 15:05:07 -07:00
55da232db6
[Bugfix] Pad Mamba page size instead of scaling block_size in unify_kv_cache_spec_page_size ( #45207 )
...
Signed-off-by: Sahil170595 <147995121+Sahil170595@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 22:01:34 +00:00
Wentao Ye and GitHub
3f99883d97
[CI Bug Fix] Temp fix for v3.2 accuracy ( #47902 )
2026-07-07 16:36:03 -04:00
Nick Cao and GitHub
47c40bfe8a
[Doc] Fix manylinux tag in installation guide ( #47913 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
2026-07-07 20:34:05 +00:00
3dd910da42
[Bugfix] Allow non-contiguous query in FlashInfer FP8 query quantization ( #47908 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 20:11:34 +00:00
Benjamin Chislett and GitHub
7bd154375d
[Bugfix] Fix mamba+dflash for MRV2 ( #47698 )
2026-07-07 15:59:13 -04:00
Rishabh Saini and GitHub
2f3f441f84
fix: include topic frame in KV events replay response ( #45177 )
...
Signed-off-by: RishabhSaini <rishabhsaini01@gmail.com >
2026-07-07 14:48:23 -04:00
d6875196ad
[Bugfix] Exclude kv_cache_memory_bytes from CacheConfig.compute_hash ( #47356 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
2026-07-07 10:46:51 -07:00
Sting Lin and GitHub
abe41f28de
Upgrade tpu-inference to v0.24.0 ( #47835 )
...
Signed-off-by: StingLin <sting.lin@cienet.com >
2026-07-07 17:15:32 +00:00
Roberto L. Castro and GitHub
c3284c31f5
[Perf][3/N] Expand Triton kernel warmup coverage, Qwen ( #47546 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-07-07 17:06:59 +00:00
Robin and GitHub
c74e751824
[Doc] Fix grammatically incorrect error message in gpu_worker and xpu_worker ( #36715 )
...
Signed-off-by: Hongbin10 <jdmjdm1998@163.com >
2026-07-07 17:03:06 +00:00
liuzhenwei and GitHub
b93cbd7416
[XPU] Fix topk_sigmoid arg mismatch on XPU ( #47858 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-07 16:53:18 +00:00
bdc6f3bfa1
[Bug] Fix tmp directory for lm_eval ( #47755 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 16:40:12 +00:00
392d1b4d2e
[BugFix][LoRA] Refresh punica metadata when LoRA slots are reassigned under an unchanged mapping ( #47725 )
...
Signed-off-by: AmeenP <ameenp360@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 08:53:55 -07:00
21b396abe1
AGENTS MD: Add suggestion on how to incorporate tests ( #47784 )
...
Signed-off-by: Simon Mo <simon.mo@hey.com >
Co-authored-by: Cursor Agent <cursoragent@cursor.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-07-07 08:16:08 -07:00
liuzhenwei and GitHub
bdaf27519f
[XPU] Fix Event init failure w/ blocking ( #47868 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-07 22:54:03 +08:00
Eldar Kurtić and GitHub
beb4327c46
Enable causal masking for SWA in vllm-project/speculators models ( #47745 )
...
Signed-off-by: Eldar Kurtic <8884008+eldarkurtic@users.noreply.github.com >
2026-07-07 10:24:14 -04:00
c46ced1ee3
[kv_offload] Establish tier-owned KV event handling ( #46544 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Signed-off-by: Chang Guo <changg@nvidia.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:55:20 +03:00
65dcde1695
[Bugfix] Fix PD disagg + MTP correctness for Qwen3.5(GDN) ( #47466 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 13:51:22 +00:00
65a7b46284
[KV-Offloading] Support workload identity for objectstore secondary tier ( #47063 )
...
Signed-off-by: Pierangelo Di Pilato <pierdipi@redhat.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:30:16 +03:00
93e2ab7111
Disable dynamic speculative decoding when DP is enabled ( #45963 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 12:52:41 +00:00
Lanze Liu and GitHub
8b745527cd
[Bugfix] Fix UBatchWrapper CUDA graph key to sum all ubatches, not just first two ( #43161 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-07-07 12:42:31 +00:00
920469974a
[UX] Log worker exit code when process dies unexpectedly ( #38641 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-07 12:36:29 +00:00
8b91cd5b20
[Bugfix][Core] Close underlying iterator in merge_async_iterators single-iterator fast path ( #44726 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 05:13:41 -07:00
Harry Mellor and GitHub
dd94484577
Bump Transformers version to 5.10.4 ( #41359 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 05:13:28 -07:00
Shaun Kotek and GitHub
7ff656cc8b
fix: ensure no double load of lm head in nemotron mtp ( #47440 )
...
Signed-off-by: Shaun Kotek - Nvidia <skotek@nvidia.com >
2026-07-07 12:01:45 +00:00
danielafrimi and GitHub
0a2965b1b3
[BugFix] Fix ModelOpt mixed-precision quantization for sparse quantized_layers configs. ( #47318 )
...
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Signed-off-by: <dafrimi@nvidia.com >
2026-07-07 11:45:13 +00:00
Harry Mellor and GitHub
0ed05b6f82
[CI] Fix Transformers modeling backend LoRA test ( #47832 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 11:40:00 +00:00
Guan-Ming Chiu and GitHub
ed051fab54
[Bugfix] Reject sampling params unsupported by diffusion models ( #45418 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-07-07 11:25:36 +00:00
48fcfc926c
[KV Offload] Add ParentManager ABC for secondary tier callbacks ( #47274 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 13:51:18 +03:00
3354dba381
[Bugfix][KV offload] Store interior chunk-boundary blocks under MTP/Eagle ( #46972 )
...
Signed-off-by: Mikhail Kostryukov <mike@triptrack.net >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 13:16:52 +03:00
cbb5f045be
[ROCm][CI] Refresh ROCm base images when docker rocm_base changes ( #46904 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Codex <codex@example.invalid >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Codex <codex@example.invalid >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-07-07 03:10:50 -07:00
b3e85be663
fix: use configured max_logprobs instead of hardcoded 20 in derender validation ( #47834 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-07 09:42:47 +00:00
Summer Yang and GitHub
d3e69fd671
[Perf] Use blocking CUDA events to avoid busy polling cuda driver lock ( #47081 )
...
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-07 09:36:10 +00:00
c85d72076a
[HARDWARE][POWER] optimize math functions of VSX power ( #47321 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
2026-07-07 09:35:47 +00:00
c5b66233b2
[Bugfix][Spec Decode] Skip uniform spec-decode padding for diffusion models ( #47464 )
...
Signed-off-by: kl527 <kl527@cornell.edu >
Signed-off-by: Kyung Sub Lee (Daniel) <66861800+kl527@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 09:25:12 +00:00
066f02ae94
[MoE] FI autotuning: max bucket = max token count [e.g. DP_size*MNBT] ( #47427 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 12:08:36 +03:00
Jee Jee Li and GitHub
5d23ca47ab
[Kernel] Applies routed_scaling_factor internally ( #47408 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-07 02:00:54 -07:00
e55cc59e52
[Rust Frontend][CI] Unblock more end-to-end test cases ( #47735 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-07 08:27:20 +00:00
ba50b9763f
[Bugfix] Match the mapped filename in find_loaded_library ( #47586 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-07 08:06:29 +00:00
b4cfbc24d3
[Bugfix][Core] Fix host memory leak from undrained new_block_ids ( #44490 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-07 07:32:55 +00:00
Aritra Roy Gosthipaty and GitHub
1e823dc01d
[docs update] Update usage of hf cli for cache list and removal ( #47830 )
...
Signed-off-by: Aritra Roy Gosthipaty <aritra.born2fly@gmail.com >
2026-07-07 07:09:18 +00:00
8e61b646e2
fix(security): add resource bounds validation to derender endpoints ( #47260 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 14:58:26 +08:00
e040899a00
[KV Offloading] Add basic offloading metrics ( #45958 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivas Krovvidi <194645829+Srinivasoo7@users.noreply.github.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 09:26:28 +03:00
dd5c299fbe
[ROCm][Bugfix] Convert ModelOpt FP8 per-channel weights to e4m3fnuz on MI300/MI325 ( #47201 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 23:24:58 -07:00
xiangdong and GitHub
6db31c8e76
[XPU][CI]Adjust memory request for tests in Intel GPU CI ( #47758 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-07 05:56:40 +00:00
cbe9c40f99
[Bugfix] Forward callable hf_overrides to the draft model config ( #45352 )
...
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-06 21:12:38 -07:00
Andreas Karatzas and GitHub
2f71b2bd9f
[ROCm] Align mixed encoder-decoder KV cache views in V2 runner ( #47685 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-07 12:09:22 +08:00
32ab064621
[UX] Add model_class_overrides for development and debugging ( #47148 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 11:43:08 +08:00
Guan-Ming Chiu and GitHub
c64c356990
[Perf] Bound DiffusionGemma sampler transient via request-tiled logits ( #45672 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-07-07 03:42:01 +00:00
Tahsin Tunan and GitHub
34e6dfced8
[Rust Frontend] Stamp arrival_time at the frontend entry ( #47787 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-07 03:27:10 +00:00
Reid and GitHub
39a1d32b59
[Rust Frontend] Avoid extra copies for multimodal tensors ( #47581 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-07 03:09:55 +00:00
700e882eab
Add TorchCodec as a video decoding backend ( #46609 )
...
Signed-off-by: Nicolas Hug <contact@nicolas-hug.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-07-06 19:58:51 -07:00
a4f019fa25
fix(distributed): propagate distributed_timeout_seconds to NCCL device groups ( #45159 )
...
Signed-off-by: jialoop-git <joane8913456@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 02:52:51 +00:00
Rahul Vishwakarma and GitHub
9dd2465896
feat(cpu): add CPU support for Mamba ShortConv ( #35059 )
...
Signed-off-by: Rahul Vishwakarma <Rahul.Vishwakarma2@ibm.com >
2026-07-07 10:47:08 +08:00
Reid and GitHub
a46c9329e5
[Rust Frontend] Add DeepSeek V3.2 roundtrip fixture ( #47619 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-07 10:47:02 +08:00
Kyle Sayers and GitHub
445321fab4
[Bugfix] [Quantization] Fix loading for CT DSV2 ( #47780 )
2026-07-07 02:28:00 +00:00
69f3150981
[XPU] Fix PP accuracy on XPU device ( #47253 )
...
Signed-off-by: yisheng <yi.sheng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-07 09:17:27 +08:00
86db6c3070
[Frontend] add per-request timing metrics field to response body of Chat/Completions APIs ( #46768 )
...
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com >
Signed-off-by: Simon Mo <simon@inferact.ai >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: GPT-5.5 <noreply@cursor.com >
2026-07-06 17:48:29 -07:00
5769a7382c
[ROCm][CI][Bugfix] Fix flaky parallel tool-call streaming (test assertion + Mistral/Granite parsers) ( #47550 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 19:06:19 -04:00
Andreas Karatzas and GitHub
8484ca5d45
[ROCm][CI] Adding Rust parity ( #47478 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:05:39 -07:00
482e5524fe
[Bugfix][ROCm] Fix memory access fault in AITER MLA backend for DPA+FP8 KV ( #47276 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: nnyrhila <niko.nyrhila@amd.com >
2026-07-06 21:30:02 +00:00
567a78432d
[Bugfix] Fix dp mtp hang ( #40589 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: sherryC41 <sherry.c.c41@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-06 21:08:17 +00:00
d891b9bd51
[Quantization] add humming moe backend to all dense/moe oracles ( #41652 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-06 13:36:07 -07:00
04adc8843b
[Bugfix]Fix DeepSeek-V4 fp8_ds_mla KV cache reshape ( #47716 )
...
Co-authored-by: yy-fighting <23518844576@qq.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-07-06 12:56:44 -07:00
Harry Mellor and GitHub
ae098abe3f
[CI] Fix some errors on main ( #47726 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 19:40:23 +00:00
b1384f5ec6
Enable B12x backend for non-gated MoEs (like Nemotron) ( #43328 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-07-06 12:40:07 -07:00
b136cc2c2c
[Bugfix][Model] Add stability window to DiffusionGemma to match HF stability_threshold semantics ( #45965 )
...
Signed-off-by: Nathaniel McVicar <namcvica@microsoft.com >
Signed-off-by: Nathaniel McVicar <Nathaniel.McVicar@microsoft.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 19:39:12 +00:00
9fde043f54
[Kernel][Helion][1/N] Add Helion kernel for silu_and_mul_per_block_quant ( #43994 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 00:19:01 +08:00
24dd2aec81
[Bugfix] Preserve FP8 indexer WK pairs across incremental load_weights ( #46168 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-06 09:16:46 -07:00
Shengqi Chen and GitHub
83a7669827
Merge branch 'main' into cuda-arch-fixup
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-07 00:13:53 +08:00
Ranran and GitHub
3ee9eea928
[macOS][CPU][Installation] Fix the broken installation of vllm 0.24.0 in macos + cpu ( #47457 )
...
Signed-off-by: Ranran Haoran Zhang <ranranhaoranzhang@gmail.com >
2026-07-06 08:59:16 -07:00
5bce653e09
Make the Transformers modeling backend as fast as native vLLM ( #47187 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 16:59:14 +01:00
5ad11172b7
[perf]Add fused Kimi image preprocessing ( #47416 )
...
Signed-off-by: Kevin-XiongC <kevin_xiong1997@outlook.com >
Signed-off-by: Kevin_Xiong <kevin_xiong1997@outlook.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-06 08:46:32 -07:00
Wentao Ye and GitHub
f70caef48b
[Perf] Cache token_to_req_indices for dsv4, 5x~6x kernel performance improvement ( #47474 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-06 11:17:46 -04:00
8d8ec38361
[Bugfix][Spec Decode] Add missing draft_id_to_target_id to DSparkDeepseekV4ForCausalLM ( #47429 )
...
Signed-off-by: Laurent-Zhang <zhangdongsheng80@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 10:55:47 -04:00
Wentao Ye and GitHub
b1c6dba558
[Refactor] Remove multiple dead code ( #47329 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-06 07:54:08 -07:00
598d51153a
[Bugfix][Distributed] Delegate MNNVL allreduce one-shot selection ( #47589 )
...
Signed-off-by: jesco-absolut <team@srswti.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 07:47:06 -07:00
Yifan Qiao and GitHub
095adf1fdc
[Bugfix] Fix int32 overflow in triton_decode_attention page offsets ( #47671 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-06 10:36:15 -04:00
Harry Mellor and GitHub
51ee564e56
[CI] Skip test for checkpoint that was deleted ( #47748 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 07:24:09 -07:00
373eb314af
[Bugfix][Core] Fix num_output_placeholders underflow with async scheduling + spec decode ( #46066 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 13:50:38 +00:00
641cb59592
[Doc] Clarify fastokens availability ( #45813 )
...
Signed-off-by: LjjJzd <3542531707@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-06 13:33:05 +00:00
07f9baf756
Revert "[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )" ( #47668 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 14:18:33 +01:00
7a90eb98ab
[Bugfix] [Gemma4] Fix Gemma4 MTP draft model layers ignoring quant_config ( #47091 )
...
Signed-off-by: Ayushman Singh <40520701+ayush1399@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-07-06 14:04:00 +01:00
8f4c69b222
[Rust Frontend] Cache metric handles for scheduler & request stats ( #47444 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-06 13:02:59 +00:00
8b79971bb9
attention: pass None for unused args in unified attention TD path ( #43597 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: quinnlp <quinnlp@users.noreply.github.com >
2026-07-06 21:01:21 +08:00
Nick Hill and GitHub
f676808ba0
[CI] Use TTY for AMD CI tests for colored buildkite logs ( #47730 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-06 20:50:29 +08:00
Qiming Zhang and GitHub
98e4726a14
[fix][run_batch]: respect proxy env vars when downloading media URLs ( #47697 )
...
Signed-off-by: mauyuyuace <qiming1.zhang@intel.com >
2026-07-06 12:45:48 +00:00
BadrBasowid and GitHub
740f379fae
[ROCm][AITER] Directly Implement AITER Custom All-reduce in CudaCommunicator ( #46065 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-07-06 12:16:32 +00:00
Alexis K. and GitHub
40cc2e8327
[Bugfix] Return HTTP 422 for unprocessable image URLs instead of 500 ( #47165 )
...
Signed-off-by: Alexis Kinsella <alexis.kinsella@gmail.com >
2026-07-06 11:56:23 +00:00
ba22152096
fix(security): block request-level GPU video backend selection withou… ( #47259 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 02:36:49 -07:00
Yan Ma and GitHub
90ce3a09be
[bugfix] fix MOSS-Audio deepstack_input_embeds initialization in PP ( #47607 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-07-06 17:15:50 +08:00
26c754d847
[XPU][Bugfix] Do not transpose weight_scale_inv at load time ( #47116 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 17:15:26 +08:00
Sungjae Lee and GitHub
3d7f357ebf
[Doc] docs: fix note formatting for pooling models ( #47701 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
2026-07-06 09:01:10 +00:00
liuzhenwei and GitHub
736f1a5907
[XPU] Route mm_prefix models to Triton attention backend ( #47688 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-06 16:52:44 +08:00
Li, Jiang and GitHub
344609ab17
[CI/Build] Fix pre-commit check ( #47695 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-06 08:24:24 +00:00
xiaozhoupy and GitHub
d039c17114
[Bugfix] Recycle post-final-norm hidden in GLM MTP (single norm) ( #47448 )
2026-07-06 01:07:56 -07:00
xiangdong and GitHub
cdab28319f
[XPU][CI]Add agent tags for Basic Models Tests (Initialization) in Intel GPU CI ( #47675 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-06 15:15:45 +08:00
Qiming Zhang and GitHub
2fa10566e3
[Core][DP] Rotate load-balancer tie-break to avoid systematic engine bias ( #47420 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-07-06 07:09:16 +00:00
Andreas Karatzas and GitHub
fb265fc8fb
[ROCm][CI] Increasing parallelism in Basic Models Tests (Extra Initialization) ( #47591 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:06:16 +08:00
Andreas Karatzas and GitHub
8f0e75e16b
[ROCm][CI] Adding nixl multiconn ( #47481 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:04:58 +08:00
98ba9b9583
[Frontend] Support OpenAI Responses API namespace tools ( #47024 )
...
Signed-off-by: zhongjing123 <jimzhong5193@gmail.com >
Co-authored-by: zhongjing123 <jimzhong5193@gmail.com >
2026-07-06 06:21:27 +00:00
velonica0 and GitHub
990c2a0187
[RISC-V] Enable BF16 on VLEN=256 hardware ( #45243 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-07-06 06:05:16 +00:00
e433634c78
[Performance][Hardware][RISC-V] Reduce LMUL pressure in INT4 LUT dequant ( #47538 )
...
Signed-off-by: liutong <liutong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-06 05:58:56 +00:00
16f8110935
[Bugfix][CPU][RISC-V] Fix VLEN detection for RVV attention path ( #47532 )
...
Signed-off-by: liutong <liutong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-06 05:58:03 +00:00
d9c1767cd4
[INC][ARK] Direct Register Custom Op for ARK ( #46361 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 13:45:50 +08:00
Li, Jiang and GitHub
e9cc1fd093
[CI/Build][CPU] Remove global extra index ( #47687 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-06 13:42:01 +08:00
Fadi Arafeh and GitHub
f1073c050c
[CPU][BugFix] Multiple fixes to w4a8_int8 CPU MoE path ( #46739 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-07-06 05:39:20 +00:00
Qiming Zhang and GitHub
394edc8108
[XPU] limit max-num-seqs in test_lmeval.py for XPU ( #47682 )
...
Signed-off-by: mauyuyuace <qiming1.zhang@intel.com >
2026-07-06 05:34:16 +00:00
69715823df
[Test][XPU] Skip fork in kv_sharing_fast_prefill test on XPU ( #47406 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 11:32:26 +08:00
Chaojun Zhang and GitHub
6569df6a3e
[Test][LoRA] Use lightweight CPU reference and skip heavy cleanup in punica ops tests ( #47534 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-06 11:29:59 +08:00
f2aaf59151
[Feature] Support MTP speculative decoding for Bailing hybrid models ( #44880 )
...
Signed-off-by: zc02384840 <zc02384840@antgroup.com >
Co-authored-by: zc02384840 <zc02384840@antgroup.com >
2026-07-06 10:38:50 +08:00
95a248faed
[Attention Backend] HPC_ATTN backend support mtp and dynamic scheduled attention ( #47433 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
2026-07-05 18:18:25 -07:00
d2ec433e37
[XPU] Fix Eagle3 initialization on XPU ( #43957 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 08:46:05 +08:00
78a04c208d
[XPU] Fix CUDA API shims breaking Torch Dynamo during AOT compile ( #43092 )
...
Signed-off-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 08:29:20 +08:00
Spandan Tiwari and GitHub
b71218107f
[ROCm][Test] Fix test_per_token_group_quant_fp8 tolerance for 1-ULP FP8 rounding on gfx950 ( #46944 )
...
Signed-off-by: Spandan Tiwari <sptiwari@amd.com >
2026-07-05 18:02:30 -05:00
cc1d020d01
[MRV2] Enable mm prefix bidi attention support on MRV2 ( #46942 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-05 14:45:29 +00:00
Ting SUN and GitHub
8974ed89cd
[Bugfix][Voxtral Realtime] Fix token feedback timeout silent hang ( #44461 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 05:42:36 -07:00
fb2faceacd
[Bugfix][Model] Fix crash loading Mamba/Mamba2 checkpoints without an architectures field ( #46037 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Ting SUN <suntcrick@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-05 05:42:32 -07:00
b6cc46ec3b
[Feature] Support sequence parallel without the need for DP, 1.9%~5.0% E2E Throughput Improvement ( #47070 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com >
2026-07-05 05:41:30 -07:00
Lucas Wilkinson and GitHub
fa4321de3d
[Bugfix][TurboQuant] Preserve KV cache dtype in backend shape ( #47609 )
2026-07-05 08:20:48 +00:00
Ting SUN and GitHub
9226613043
[Bugfix][Pooling] Forward instruction to Jina reranker scoring prompts ( #47590 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 05:39:13 +00:00
34b560b725
[Bugfix][Gemma4] Fix FA4 mm_prefix mask: add sliding window and absolute q_idx ( #47332 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-07-04 17:46:40 -07:00
Ting SUN and GitHub
91b5647300
[Bugfix][Model] Allow Run:ai memory_limit sentinel values ( #47337 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 00:08:34 +00:00
Carl Persson and GitHub
4a6bf3c77f
[ROCm][CI] Fix Kernels and Kernels attention test failures ( #47519 )
...
Signed-off-by: Carl Persson <carl.persson@amd.com >
2026-07-04 15:59:51 -05:00
Ting SUN and GitHub
d2afe39647
[Bugfix][Frontend] Preserve default sampling params in batch chat ( #47597 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-04 19:06:39 +00:00
Wentao Ye and GitHub
2a9113f998
[Perf] Remove redundant op for GLM 5.2 ( #47198 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-04 13:25:02 -04:00
yzong-rh and GitHub
0cd6f767e3
[Bugfix][Frontend][gpt-oss] Recover raw tail when Harmony parser ends non-terminal ( #47379 )
2026-07-04 10:46:24 -04:00
Harry Mellor and GitHub
f1445f6dbd
[CI] Bump huggingface-hub from v1.10.2 to v1.22.0 ( #47551 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-04 07:45:45 -07:00
1d354c694e
[Misc] Validate Pooling cache_salt Values ( #46966 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-04 10:19:28 -04:00
Taneem Ibrahim and GitHub
2f21224527
[Misc] Update request-extras parity for batch chat completion ( #47333 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-07-04 10:19:04 -04:00
fa1fa968c4
[Misc] Forward request-level prompt extras for cross-encoder scoring ( #46939 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-04 10:18:36 -04:00
6eac8e0070
[Misc] Preserve cross-encoder pooling extra kwargs ( #47082 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-04 08:14:13 -04:00
1a308c449c
[XPU] Add W8A8 FP8 linear kernel with multi-granularity quant support ( #43645 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-04 18:10:01 +08:00
e7c9df9449
[Bugfix][Structured Output][Spec Decode] Constrain bitmask and trim grammar advance at the reasoning boundary ( #44297 )
...
Signed-off-by: Allen.Yu <yuyue0225sc@163.com >
Signed-off-by: yue.yu <yuyue0225sc@163.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-07-04 09:08:45 +00:00
gausah01 and GitHub
26eb87204d
[Bugfix] Fix CPU split-KV scratchpad sizing ( #45844 )
...
Signed-off-by: Gauri Sahnan <gauri.sahnan@arm.com >
2026-07-04 06:47:23 +00:00
4c3c17d43b
[ROCm] Disable persistent sparse-MLA kernel for chunked-prefill continuations ( #47567 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-04 01:21:43 -05:00
f329ce405b
[ROCm][CI][Bugfix] Use VllmRunner for voxtral_realtime tests to avoid OOM on AMD GPU ( #47536 )
...
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-04 12:26:10 +08:00
07516fda67
[MRV2][SD] Make Dynamic SD comatible with Full Cuda Graphs ( #45953 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-03 23:58:27 -04:00
67ff0ae30f
Support nvfp4 kv with kv-cache-dtype-skip-layers sliding_window ( #42890 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-04 02:29:13 +00:00
Bugen Zhao and GitHub
ab3b6d97aa
[Frontend] Limit SO_REUSEPORT to multi-worker serving ( #47529 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-04 01:26:24 +00:00
Ben Browning and GitHub
fb5291b35b
[Frontend] [Parser] Port DeepSeek V4 to streaming parser engine framework ( #45877 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-07-03 20:55:23 -04:00
labAxiaoming and GitHub
d6d39c111e
[GLM4V] Avoid GLM4V processor init during startup metadata reads ( #47155 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
2026-07-03 15:03:16 -07:00
379950191f
[Bugfix][Multimodal] Normalize direct PIL image inputs ( #47566 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-03 14:27:14 -07:00
576bf75d0e
[AMD][EPLB] Enable EPLB for Quark OCP MXFP4 MoE ( #47220 )
...
Signed-off-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 14:41:52 -05:00
Tres and GitHub
f006e5a24c
[CI][AMD] Allow git operations on previously created work trees ( #47554 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
2026-07-03 14:41:01 -05:00
f63dca6838
[ROCm] Fix encoder-decoder cross-attention KV layout aliasing ( #47035 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 13:53:29 -05:00
Bugen Zhao and GitHub
8651f043b8
[Rust Frontend] Speed up chat roundtrip tests ( #47523 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 19:25:06 +01:00
Andreas Karatzas and GitHub
3775d5fcab
[ROCm][CI] Adding test groups for parity with upstream ( #47479 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 19:15:01 +04:00
d7192cfccf
[CI Bugfix] Lazily import Qwen warmup dependencies ( #47539 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-03 23:10:49 +08:00
AgenticSpark and GitHub
978de83353
[Bugfix][CPU] Ship examples/ in the CPU release image ( #47447 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
2026-07-03 11:46:24 +00:00
wang.yuqi and GitHub
a14f57a3ac
[Frontend] Refine the entrypoint class's inheritance hierarchy. ( #47498 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-03 10:50:06 +00:00
18f658bb31
[Bugfix][Frontend] Fix batch chat endpoint corrupting logprobs when return_token_ids is set ( #47384 )
...
Signed-off-by: David Feng <fenghourun@meta.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-03 03:01:34 -07:00
Isotr0py and GitHub
400a9c386d
[Rust Frontend] Bump llm-multimodal version ( #47530 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-03 09:48:36 +00:00
Max de Bayser and GitHub
bbdcbe4686
Move Roberta remaining nn.Embedding to VocabParallelEmbedding ( #47452 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
2026-07-03 09:47:50 +00:00
Kalyanam Dewri and GitHub
4875b4456b
[Doc] Fix VLM2Vec benchmark chat template path ( #47517 )
...
Signed-off-by: kalyanamdewri <kalyanampriyam@gmail.com >
2026-07-03 08:24:45 +00:00
Shengqi Chen and GitHub
401bed48ad
Merge branch 'main' into cuda-arch-fixup
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-03 15:54:56 +08:00
Dakai An and GitHub
1f486d96a1
Add Triton Backend for Unlimited-OCR R-SWA ( #47102 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-07-03 00:11:50 -07:00
Bugen Zhao and GitHub
b790c84cde
[CI] Enable sccache for Rust build under CUDA/ROCm ( #45246 )
2026-07-02 23:45:41 -07:00
6429d5f527
[Rust Frontend] add repetition_detection support to sampling params ( #46684 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 14:06:23 +08:00
Chris Leonard and GitHub
fbc9ba6d30
New stable abi cleanup ( #46656 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-03 14:02:26 +08:00
xiangdong and GitHub
2dfaae752b
[XPU][CI]Fix dependency typo in Intel GPU CI ( #47510 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 04:11:47 +00:00
Evgeny Parshutin and GitHub
bd8d9021ce
[CPU][Build] Enable oneDNN ITT task collection by default for CPU primitive-level profiling ( #47467 )
...
Signed-off-by: Evgeny Parshutin <eugeny.parshutin@intel.com >
2026-07-03 04:00:19 +00:00
xiangdong and GitHub
3f0b773b30
[XPU][CI]Mv huggingface cache to larger disk in Intel GPU CI ( #47405 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 11:56:17 +08:00
Reid and GitHub
9b8e76589d
[Rust Frontend] Recover buffered text from incomplete tool calls at EOS ( #47289 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-03 03:45:03 +00:00
1aeabec355
[Bugfix][Rust Frontend] Tolerate out-of-vocab prompt ids in detokenizer ( #44682 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 03:41:53 +00:00
979f5511d7
[Bugfix][Gemma4] Keep image bidirectional attention within the sliding window ( #47217 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-07-02 19:57:41 -07:00
41de1380c2
[BugFix] Derive FlashInfer Q dtype from resolved per-group builder state ( #47485 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 19:33:28 -07:00
Nick Hill and GitHub
d85601c20f
[CI] Pin modelscope version to fix test breakage ( #47465 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:33:08 -07:00
Nick Hill and GitHub
276b837dc4
[ModelRunner V2][BugFix] Free all model refs on shutdown ( #47483 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:32:48 -07:00
34bf7b45a0
[CI] intel CI: add quantization and awq case for xpu ( #46456 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 09:51:56 +08:00
adamkbaranowski and GitHub
4c3c64fcf7
Add Laguna XS.2.1 DFlash drafter support ( #46853 )
...
Signed-off-by: Adam Baranowski <adam.baranowski@poolside.ai >
2026-07-02 18:09:27 -07:00
Andreas Karatzas and GitHub
442ccc6098
[ROCm][CI] Adding extract hs 2gpu ( #47482 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:59:38 -07:00
Andreas Karatzas and GitHub
6768fbc76f
[ROCm][CI] Adding qwen3 dp4 eplb ( #47480 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:58:56 -07:00
Andreas Karatzas and GitHub
407f406300
[ROCm][CI] Adding metadata ( #47477 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:45:23 -07:00
Harry Mellor and GitHub
e24d1b24fe
Fix Transformers modeling backend usage stats ( #47472 )
2026-07-02 12:51:23 -07:00
d29125c085
Xqa decode kernels ( #43232 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-02 12:32:05 -07:00
Michael Goin and GitHub
d715b3aa1e
Delete PagedAttention ( #47361 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 12:31:26 -07:00
Joe Rowell and GitHub
258f8de91f
[Bugfix][Tool Parser] poolside_v1: accept tool calls without newline after function name ( #47311 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-07-02 12:08:37 -07:00
Nick Hill and GitHub
e392bf7a68
[BugFix][MRV2] Ensure all req slots are accounted for when scheduling ( #46974 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:24 -07:00
Nick Hill and GitHub
443e68cfa6
[Bugfix] Fix pooled Whisper encoder sliding-window kernel size ( #47437 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:11 -07:00
Chauncey and GitHub
320ee285c9
[Model Runner V2][Perf] Warm up GLM-5.2 DSA indexer prefill metadata kernel ( #47285 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-02 16:31:31 +00:00
Bugen Zhao and GitHub
ec0ffaacc8
[Rust Frontend] Improve scheduler stats logging parity ( #47435 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 15:55:25 +01:00
Yuxuan Zhang and GitHub
178fd56094
support GLM-5.2 gate use FP32 ( #47410 )
...
Signed-off-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk >
2026-07-02 22:45:39 +08:00
a47f38f825
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in block verification kernels ( #47383 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 07:32:38 -07:00
Nick Hill and GitHub
3e158ae62d
[ModelRunner V2] Fix Mamba2 crash on non-spec-decode ( #47428 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 07:05:16 -07:00
a2f713002d
[ModelRunner V2] Enable by default for all dense models ( #44443 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 18:48:57 +08:00
TJian and GitHub
de2a8fc042
[ROCm] [PyTorch] Move to stable abi since ROCm upgraded to torch 2.11 ( #47128 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-02 18:34:07 +08:00
Michael Goin and GitHub
84b9c2762f
Update DeepGEMM tag to point to latest nv-dev branch for sm120 support ( #47304 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 18:33:44 +08:00
Bugen Zhao and GitHub
25fcb65d51
[Rust Frontend] Use enum-backed domain types for engine outputs and structured outputs ( #47283 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 10:41:46 +01:00
08a8a4af3f
feat(rust): expose profiler control routes in Rust frontend ( #46306 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 08:46:07 +00:00
b0b8a286dd
[Model] Add LLaVA-OneVision-2 (LlavaOnevision2ForConditionalGeneration) ( #44785 )
...
Signed-off-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-02 16:41:49 +08:00
3af8789559
[Feature] Universal speculative decoding for heterogeneous vocabularies (TLI) ( #38174 )
...
Signed-off-by: wan-danfeng <wandanfeng0802@gmail.com >
Signed-off-by: Wonderful <wandanfeng0802@gmail.com >
Co-authored-by: Wan_DF <wonderful199082@126.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-07-02 01:34:20 -07:00
Chaojun Zhang and GitHub
8357226f4f
[XPU][CI] Split test_punica_ops into separate pytest invocations for stability ( #47376 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-02 07:50:55 +00:00
Hiki and GitHub
2665ed704b
[Bugfix][Kernel] Correct FlashInfer CUTLASS MoE tuning token bound ( #46838 )
...
Signed-off-by: Haobin Guo <haobing@nvidia.com >
2026-07-02 05:11:00 +00:00
xaguilar-amd and GitHub
09663abde0
[ROCm][MLA] Fuse MLA q/kv RMSNorm + FP8 per-token quant in the FP8 attention path ( #44977 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
2026-07-02 13:00:41 +08:00
Giancarlo Delfin and GitHub
d63c8e9444
[BugFix][Spec Decode] Compact shared topk indices buffer after first MTP draft step ( #47238 )
2026-07-01 21:38:51 -07:00
Jee Jee Li and GitHub
1360c42fe6
[UX] Include NVTX in cuda.txt ( #47319 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 19:38:50 -07:00
d0a2584773
[Misc] Use functions instead of PTX for the PDL instruction ( #46984 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-01 19:38:35 -07:00
7fe7fa9cda
[CI][Bugfix] Rerun test_engine_log_metrics_ray on Ray GCS startup timeout ( #47208 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 21:32:09 -05:00
Michael Goin and GitHub
2b753ad200
[Spec Decode] DSpark speculators checkpoint support ( #47093 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-01 17:32:27 -07:00
e196268bad
[Docker] Remove unused Dockerfile.nightly_torch ( #47338 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:42 -07:00
e91f5f8439
[CI] Remove torch_nightly mirror tags (superseded by TORCH_NIGHTLY full-nightly build) ( #47342 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:06 -07:00
fa248139a0
[MoE] Plumb gemm1_alpha/beta/clamp_limit into TRT-LLM FP8 MoE ( #45723 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 14:34:05 -07:00
Yongye Zhu and GitHub
d3229431f9
[DSV4] Better MXFP8 quantization kernel ( #47229 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-07-01 14:33:51 -07:00
Nick Hill and GitHub
4787f2dd1b
[Bugfix] Don't read KV cache past seq_len in triton paged attn kernels ( #47305 )
2026-07-01 12:43:00 -07:00
Nick Hill and GitHub
8cfeb84dba
[ModelRunner V2] Warmup cross-attn properly in encoder-decoder case ( #47308 )
2026-07-01 12:36:48 -07:00
Chaitanya Sri Krishna Lolla and GitHub
5fd442187c
[ROCm][P/D] MoRIIO toy proxy: support JSON Content-Type for OpenAI clients. ( #46482 )
...
Signed-off-by: lcskrishna <lollachaitanya@gmail.com >
2026-07-01 19:17:05 +00:00
00eb7cefa3
[Bugfix] Prevent padding placeholders from reaching embeddings ( #47029 )
...
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 09:26:03 -07:00
Michał Ganczarenko and GitHub
c8bdcc0116
[Bench][BugFix] Fix empty decoder prompt for Cohere ASR in throughput benchmark ( #47135 )
...
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com >
2026-07-01 15:42:27 +00:00
f5a8d73377
[Spec Decode] DSpark ( #46995 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-01 08:30:24 -07:00
63fcce4de1
[Bugfix] Fix GraniteMoeShared weight loading broken by #41184 ( #47031 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 22:39:12 +08:00
Bugen Zhao and GitHub
c638f9216a
[Rust Frontend] Split engine core DTOs into separate modules ( #47265 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-01 15:28:21 +01:00
Chaojun Zhang and GitHub
13c49f9845
[xpu][lora]: Align LoRA implementation with Punica GPU: fix _apply_expand rank mismatch, add_inputs hardcode, and MoE EP ( #45368 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-01 22:14:04 +08:00
Nick Hill and GitHub
f1cf6b0086
[CI] Fix segfault in tracing test ( #47299 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 14:00:37 +00:00
Harry Mellor and GitHub
a78c15616f
Migrate GPTBigCode and Starcoder2 to the Transformers modeling backend ( #30966 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 13:41:36 +00:00
5c4db60f01
docs(security): document gRPC interface as insecure for private use only ( #45903 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <rbryant@redhat.com >
2026-07-01 12:39:57 +00:00
4e5ca89cfe
[ROCm][MiniMax-M3] Cross-layer lightning-indexer top-k sharing ( #47269 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 10:50:09 +00:00
Harry Mellor and GitHub
a22e0dfc69
[Model] Remove AyaVision, MusicFlamingo ( #47263 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:39:33 +00:00
stevenkuang and GitHub
cc56379e28
[Model] Support Hy3 token suffix and JSON Schema array types ( #47192 )
...
Signed-off-by: stevenkuang-tencent <stevenkuang@tencent.com >
2026-07-01 10:16:07 +00:00
024b06b0dc
[Bugfix] Expose usage field in GenerateResponse for disaggregated serving ( #42748 )
...
Signed-off-by: AIvashov <ivashov.aleksey@proton.me >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-01 10:00:19 +00:00
Harry Mellor and GitHub
e7d0fcbc09
[CI] Fix various failures on main ( #47197 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:35:34 +01:00
akii96 and GitHub
aa8bb5562e
[ROCm][Perf][Bugfix] DSv4 indexer: use platform FP8 dtype (fnuz) for Q-quant on gfx942 ( #46730 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-07-01 17:33:55 +08:00
Andy Lo and GitHub
fa4bec9056
[Bugfix] Fix pooled Whisper sliding-window KV sizing ( #47071 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-07-01 11:33:19 +02:00
dee5da1dec
[Test] Run SageMaker handler-override tests in-process via TestClient ( #47250 )
...
Signed-off-by: Jyothirmai Kottu <jkottu@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 09:14:00 +00:00
ed41aa270a
[ROCm][DSV4] Use aiter mHC pre/post as the default ROCm path ( #43950 )
...
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com >
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 16:27:42 +08:00
77a9c5ae28
Weight sync refactor + move sparse nccl engine ( #44353 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Signed-off-by: haoaaron <ahao@anyscale.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-01 01:25:19 -07:00
f651a8a9a4
[XPU][UT]Enable ut qk_norm_rope_fusion ( #42486 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 07:38:03 +00:00
Jee Jee Li and GitHub
8f82be5705
[CI/Build] Fix LoRA testing ( #47242 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 15:36:13 +08:00
Nils Matteson and GitHub
a461070d1c
[Core] Make sleep-mode backend capability flags communicator-agnostic ( #47243 )
2026-07-01 07:17:44 +00:00
4470ae84de
Remove mantis ( #46806 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 07:13:58 +00:00
Chauncey and GitHub
697c34b97b
[Bugfix] Fix beam search candidate indexing when logprobs count varies ( #47126 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-01 07:07:06 +00:00
Blas Rodriguez Irizar and GitHub
5b431b905c
[Rust Frontend] Coerce completion max_tokens: null to default ( #47166 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-07-01 06:41:33 +00:00
89e99202f2
[CPU][Perf]Added tanh AOR for faster gelu activations. ( #44639 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Signed-off-by: almayne <anna.mayne@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-30 23:24:40 -07:00
Micah Williamson and GitHub
b446792306
[ROCm][Bugfix] Fix Triton "out of resource: shared memory" Error In One-Shot LoRA MoE ( #47209 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:36 -07:00
Micah Williamson and GitHub
c3b1f9e827
[ROCm][CI] Enable LoRA TP Distributed Test Group In AMD CI ( #47193 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:32 -07:00
Jonathan Mamou and GitHub
df802a87b7
[CPU] Remove speculative decoding stream overrides from CPUModelRunner ( #47162 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
2026-07-01 06:12:49 +00:00
Nils Matteson and GitHub
93d8f834dd
[Core] Pluggable sleep-mode backend abstraction (RFC #34303 ) ( #44074 )
2026-06-30 22:00:53 -07:00
Maria Guevara and GitHub
aeb35b90f0
[Rust Frontend] Add error context in tool parser failures ( #46512 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-07-01 12:48:55 +08:00
Gabriel Wu and GitHub
9a08a5118e
fix: skip cooperative top-K on SM120 ( #47164 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-06-30 21:32:54 -07:00
c5200d3565
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE ( #46076 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: GirasoleY <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 00:32:20 -04:00
Matt and GitHub
3c1396bab6
[Hardware][AMD][CI] Toggle test coredumps on ROCm debug agent ( #47222 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 23:30:10 -05:00
Benjamin Chislett and GitHub
9969466a59
[Spec Decode] Support SWA + DFlash for MiMo ( #46104 )
2026-06-30 20:34:47 -07:00
achyuthan.s and GitHub
3406e8f83d
[Bugfix][Frontend][gpt-oss] Return raw output when Harmony parser ends non-terminal ( #47062 )
...
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
2026-07-01 01:46:01 +00:00
a264e41975
[Distributed] Default FlashInfer allreduce to mnnvl on single node ( #47219 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-30 18:35:56 -07:00
Woosuk Kwon and GitHub
f098ee70c7
[GLM5] Support FlashMLA FP8 KV cache (Hopper & Blackwell) ( #47090 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-30 18:13:21 -07:00
9294dd27eb
fix(reasoning): guard rfind in ernie45 streaming </response> branch ( #46255 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-07-01 01:01:14 +00:00
yzong-rh and GitHub
b1190d03cc
[Refactor][GPT-OSS] Harmony Responses API Refactor to use HarmonyParser ( #47185 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-30 19:23:20 -04:00
92c7fac640
[Perf] Restore zero-init of swizzled NVFP4 scale buffer to recover Blackwell decode throughput ( #45739 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-30 22:56:56 +00:00
Ting SUN and GitHub
ac521f6237
[Bugfix][Structured Outputs] Reject degenerate structured_outputs that crash EngineCore ( #45346 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-30 22:41:33 +00:00
28242824e0
[Bugfix][Frontend] Normalize constrained Harmony recipients ( #45657 )
...
Signed-off-by: shaojunjie <626650687@qq.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-30 17:33:10 -04:00
VectorPeak and GitHub
68294739d1
[Bugfix] Align OpenCV video metadata timeline ( #47099 )
...
Signed-off-by: VectorPeak <73048950+VectorPeak@users.noreply.github.com >
2026-06-30 20:43:42 +00:00
c8d2f3cb14
[Bugfix] compressed-tensors: allow int8 grouped WNA16 MoE on Marlin ( #47154 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 12:50:46 -07:00
Matt and GitHub
345b28ff2f
[Hardware][AMD][CI] Bump timeouts of various test groups on AMD CI ( #47195 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 14:30:53 -05:00
248d1fbb71
[Feat][1/N] CuTeDSL warmup infrastructure, FA4 MLA ( #46182 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-30 12:17:34 -07:00
11b26c5528
[Bugfix][Tool Parser] PoolsideV1: fix logprobs AttributeError on Responses API ( #47138 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 19:14:09 +00:00
Roberto L. Castro and GitHub
20434c472e
[Feat] Improve Triton JIT diagnostics ( #46621 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-30 18:50:15 +00:00
Andreas Karatzas and GitHub
c8f9c156a5
[ROCm][V1][MLA] Clone prefill backend state per metadata builder ( #46993 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 11:43:54 -07:00
953bba488d
[PERF] Extend NCCL symmetric memory to AllGather and ReduceScatter ( #46703 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: snordmann <snordmann@nvidia.com >
2026-06-30 11:38:18 -07:00
Wentao Ye and GitHub
3a9784b82c
[Feature] DP supervisor using rust frontend ( #47076 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 14:34:05 -04:00
Giancarlo Delfin and GitHub
3cecee40f3
[Model Runner V2][Spec Decode] Fix stale values in idx_mapping from CG num reqs padding ( #47066 )
2026-06-30 11:25:32 -07:00
a7732537f4
[Bugfix] Restore part of bugfix #42650 after accidental deletion in #43241 ( #47039 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
Co-authored-by: Zhanda Zhu <49645678+zhandaz@users.noreply.github.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-30 11:07:59 -07:00
727971f1c1
Add Medusa speculative decoding e2e test ( #41396 )
...
Signed-off-by: Anshika Ojha <anshikao@nvidia.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Anshika Ojha <215760622+ojhaanshika@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-06-30 18:02:22 +00:00
25671cb520
[Parser][Bugfix] Ensure tool call or other special tokens don't leak in non-streaming tool parsing ( #46875 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-30 13:46:53 -04:00
27d5f78b63
[CI] Move distributed small LM eval to B200 ( #47048 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 13:34:25 -04:00
liuzhenwei and GitHub
7a341fa109
[XPU] Support ZE_AFFINITY_MASK passthrough in xpu_disagg_acc_test ( #47105 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-30 17:06:12 +00:00
Charlie Fu and GitHub
f41e8ddc97
[ROCm][CI] Move PyTorch Compilation Unit Tests to MI300(gfx942) ( #47065 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-30 11:32:58 -05:00
245888ff77
[Feature] Detect all2all peer fault with fault tolerance backend and prevent corrupted output ( #43637 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 09:00:25 -07:00
e840f0d3f5
[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 08:39:59 -07:00
fcaa84efa7
[BugFix] Gate MRV2 mixed sparse-MLA warmup on max_num_seqs > 1 ( #47050 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: ziminghuang <ziminghuang@inferact.ai >
2026-06-30 16:31:27 +01:00
Wentao Ye and GitHub
9e84ec8648
[Refactor] Remove dead minimax allreduce rms kernel ( #46842 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 08:29:21 -07:00
d8f483dc30
[Spec Decode] Fix hidden-state extraction block size for hybrid verifiers ( #46301 )
...
Signed-off-by: Igor Margulis <igor.margulis@intel.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-06-30 08:19:51 -07:00
Nicolò Lucchesi and GitHub
dc148dc4d7
[CI][Bugfix] Fix Hybrid SSM NixlConnector PD prefix cache test (2 GPUs) ( #47157 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 23:14:13 +08:00
tc-mb and GitHub
7cf7cbcd95
[Bugfix] MiniCPM-V 4.6: fix grid rows/cols swap in placeholder generation ( #45918 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-30 08:12:44 -07:00
c231d1f290
fix(security): bound tokenizer work when explicit truncation_side is set ( #47007 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 23:08:51 +08:00
Giancarlo Delfin and GitHub
db808b3961
[Model Runner V2][Spec Decode] Implement block verification for rejection sampling ( #46781 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-30 08:07:24 -07:00
Arsalan Shakil and GitHub
00ebf19cca
[Bugfix][Quant] Raise actionable error instead of bare assert for group-size/TP mismatch ( #46230 ) ( #46236 )
...
Signed-off-by: Arsalan Shakil <shakil.arsalan@yahoo.com >
2026-06-30 14:57:14 +00:00
ded6676458
[Bugfix] Seed RayExecutorV2 TCPStore port by DP rank to avoid collisions ( #45960 )
...
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:37:34 -07:00
Bugen Zhao and GitHub
7a327f0b4f
[Rust Frontend] Simplify unit tests with shared TestTokenizer ( #47125 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 15:34:43 +01:00
Harry Mellor and GitHub
1ab9522935
Remove more unnecessary load_weights methods ( #47058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 15:22:16 +01:00
0fc2512094
[KV Offload] Pass ScheduleEndContext to on_schedule_end hook ( #46450 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 17:07:12 +03:00
Harry Mellor and GitHub
62c7d8009f
Forward fix nightly errors from #44589 ( #47151 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 14:02:34 +00:00
Isotr0py and GitHub
ab80b3dff4
[CI/Build] Bump PyNvVideoCodec version ( #47139 )
2026-06-30 06:38:46 -07:00
Qiming Zhang and GitHub
91055efd36
[XPU] C++ implementation for get_memory_info ( #47134 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-30 21:34:47 +08:00
Bugen Zhao and GitHub
3675bcff67
[Rust Frontend] Refactor TLS serve path with unified MaybeTlsListener ( #47101 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:31:58 +01:00
Shengqi Chen and GitHub
d492d1e697
Merge branch 'main' into cuda-arch-fixup
2026-06-30 21:30:29 +08:00
Bugen Zhao and GitHub
bdbd7278b6
[Rust Frontend] Extend renderer/parser roundtrip tests to support token ids ( #47110 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:27:45 +01:00
Harry Mellor and GitHub
5dc36a4fa5
[Model] Remove Tarsier, Tarsier2 ( #47143 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:20:33 +00:00
Shengqi Chen
f37e113590
[Build] Apply ruff format to CUDA arch regex
...
Keep the compiled CUDA arch regex on one line to match ruff-format output.
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 20:37:59 +08:00
aab7af0bcb
[Bugfix][ROCm][MLA] Pass q/kv dtypes to get_mla_metadata_v1 in FP8 decode ( #46997 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-30 05:31:16 -07:00
536047755e
Bump actions/checkout from 6.0.1 to 7.0.0 ( #33057 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 13:16:20 +01:00
Shengqi Chen
a10e369f06
[Build] Address CUDA arch review comments
...
Fix CUDA arch warning tests so they do not depend on importing the stable libtorch extension, which keeps the warning coverage active in lightweight CI environments and avoids mypy treating a fixture value as a base class.
Use regex for the compiled-arch parser, apply formatter output, and make the CUTLASS grouped GEMM Python support query fall back to false when the op is unavailable or unimplemented in the current build.
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 20:03:28 +08:00
1907d3854a
[Bugfix] Reject negative values for max_logprobs and long_prefill_token_threshold ( #44002 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:01:03 +01:00
Shengqi Chen
aa3f2efe42
[Temp] Cherry-pick #47139 to fix build
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 19:38:28 +08:00
Shengqi Chen and Codex
a32d1bff95
[Doc] Document CUDA wheel architecture coverage
...
Explain that pre-built CUDA wheels use the architecture lists selected by the release and build pipelines, which may be narrower than the full set vLLM can build from source.
Call out CUDA 12.9 architecture-specific wheel coverage, CUDA 13 family-specific targets, and the no-kernel-image error users may see when a wheel does not cover their GPU.
Co-authored-by: Codex <codex@openai.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 19:27:42 +08:00
Shengqi Chen and Codex
ee47a21fcd
[Build] Warn on uncovered CUDA device architectures
...
Expose the compiled CUDA arch list from the stable extension and check visible CUDA devices against it during CUDA platform startup. The runtime check distinguishes exact architecture targets from CUDA 13 family targets so users get an early warning before hitting missing kernel images.
The startup warning is limited to the NVML-backed CUDA platform path to preserve the existing no-CUDA-init import behavior for non-NVML environments.
Co-authored-by: Codex <codex@openai.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 19:27:42 +08:00
Shengqi Chen and Codex
0f471a3088
[Build] Scope stable CUDA kernel feature macros
...
Keep optional stable CUDA kernel feature macros on the source files that consume them instead of adding them to VLLM_GPU_FLAGS. This avoids perturbing unrelated compile commands and invalidating more cache entries when optional kernel families change.
Also align CUTLASS grouped MoE support with the SM10x/SM11x family so Thor works under both CUDA 12 SM101 and CUDA 13 SM110 reporting, and remove the stale ENABLE_CUTLASS_MLA definition left after the old CUTLASS MLA path was deleted.
Co-authored-by: Codex <codex@openai.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-30 19:27:42 +08:00
Chaojun Zhang and GitHub
ea9ddf59fc
[XPU][CI] Enable shared loader test ( #45977 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 11:20:33 +00:00
8cf7c4d8ad
[Attention Backend] add HPC-Ops Attention backend ( #46020 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 18:17:43 +08:00
8e9d70fdd5
[Kernel][XPU] Adjust kernel unit tests for XPU ( #45140 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-30 09:57:27 +00:00
Juan Pérez de Algaba and GitHub
364ee36af1
fix(security): prevent image decompression bomb OOM denial of service ( #47010 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-30 09:39:22 +00:00
Nicolò Lucchesi and GitHub
06fae69114
[Misc] Mistral label alert ( #47132 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 09:02:07 +00:00
14f8660a18
[CI/Build] Add CPU test dependency pre-commit hooks ( #47032 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 07:59:13 +00:00
aed541def4
[Bugfix][Responses] Set completed status for Harmony function calls ( #46945 )
...
Signed-off-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-30 07:55:14 +00:00
2bc20e8aba
[Frontend] Add Streaming Parser Engine and new Kimi k2.5/k2.6/k2.7 Parser ( #46610 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:53:17 +00:00
Chaojun Zhang and GitHub
8cc242335d
[XPU] Optimize XPU worker shutdown logic to prevent resource leak ( #46433 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 15:27:21 +08:00
Andreas Karatzas and GitHub
ba22cb6765
[ROCm][Ray][CI] Keep assigned GPU visible for weight transfer ( #47000 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:59:18 +08:00
Uros Markovic and GitHub
81bcced482
[Bugfix][ROCm] Preserve MoE weight padding for unquantized Triton path ( #46381 )
...
Signed-off-by: Uros Markovic <umarkovi@amd.com >
2026-06-30 14:47:57 +08:00
Kunshang Ji and GitHub
fb42e5219e
[Platform] Replace torch.cuda.mem_get_info with torch.accelerator.get_memory_info ( #44825 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-30 14:39:52 +08:00
Dakai An and GitHub
0feca7ffa8
PD disagg with Mooncake Connector: GDN support (Qwen3.5) and MLA support (Deepseek-V4-Flash) ( #46807 )
2026-06-29 23:29:04 -07:00
97b5ce5c39
[Bugfix] Raise VLLMValidationError for non-integer logit_bias keys ( #46612 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-30 06:18:59 +00:00
Andreas Karatzas and GitHub
4236514098
[ROCm][CI][Multimodal] Use ROCm-aware FA availability check for Unlimited-OCR ( #47004 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:03:13 +08:00
Blas Rodriguez Irizar and GitHub
e45c8a9f4b
[Rust Frontend] Start current wave for a stale DP FirstRequest ( #46833 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-30 05:13:09 +00:00
Wei Zhao and GitHub
b153dd3f28
[Bugfix] Use larger workspace size for Flashinfer MLA LSE ( #47074 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-29 22:11:03 -07:00
Reid and GitHub
930f8dc0a1
[Bugfix][Rust Frontend] Reject prompt_logprobs for streaming generate ( #46839 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 05:10:07 +00:00
Reid and GitHub
a16dbd5b85
[Rust Frontend] Avoid LoRA registry scans without active LoRA requests ( #47040 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 04:58:19 +00:00
bec232a914
Secondary tier implementation for PD disaggregation ( #42285 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 07:51:44 +03:00
b5c9e1ac33
[LoRA] Add language-backbone LoRA support for MiniCPM-V 4.6 ( #46740 )
...
Signed-off-by: linitra24 <Joy25810@foxmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-30 04:19:31 +00:00
ae2c4f3db7
[XPU][UT]Fix xpu pass_config.fuse_norm_quant assert issue ( #46804 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 21:13:44 -07:00
ganesh and GitHub
fca432e60a
[Bugfix] Propagate default stop_token_ids to per-request SamplingParams ( #35076 )
...
Signed-off-by: sriganesh123 <arjulasriganesh@gmail.com >
2026-06-30 12:10:09 +08:00
af1ee8c475
fix(config): reject negative max_logprobs (except -1) and long_prefill_token_threshold ( #44070 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 04:02:36 +00:00
5b4cb69523
[Bugfix][MLA] Fix LSE log-base mismatch in DCP + FlashInfer MLA decode ( #47079 )
...
Signed-off-by: girasoley <girasoleyang@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 19:15:02 -07:00
9fc0c08026
[ROCm][CI] Make tests/v1/shutdown an importable package ( #47085 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 21:01:27 -05:00
f2b5fabb23
[ROCm][CI] Move LM Eval Large Models (8 GPUs) to mi300 pool ( #47094 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 20:59:08 -05:00
b8cb75b149
[Rust Frontend] Add static HTTPS and mTLS support for HTTP and gRPC ( #45890 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 01:45:59 +00:00
Thien Tran and GitHub
43916891b2
[GDN] Improve kkt kernel of CuteDSL prefill backend ( #46346 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-29 18:34:18 -07:00
cda05ee8c4
[Bugfix][Reasoning] Fix thinking_token_budget not enforced on re-entry after forced end ( #43757 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-06-30 01:04:25 +00:00
weishu and GitHub
77654d080c
[KVTransfer] MultiConnector: merge kv_transfer_params dicts across connectors ( #46777 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-06-30 00:25:05 +00:00
Wentao Ye and GitHub
75698e60b3
[Bug] Fix sparse attention issue for GLM5.2 non-torch compile path ( #47083 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-29 15:45:53 -07:00
Andreas Karatzas and GitHub
8632c884dc
[ROCm][CI] Use spawn around the threaded OTLP test ( #47003 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 16:34:05 -05:00
c3734e8334
[CI][Bugfix] Add cohere_melody to ROCm test requirements ( #47072 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 16:29:47 -05:00
53f7553f09
[ROCm][DeepEP] Stabilize high-throughput DBO for DP+EP ( #46990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-29 14:28:02 -07:00
4eb227992a
[ROCm][CI] Make memory sampling less racy in tests and sleep mode ( #45490 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-29 14:26:41 -07:00
Micah Williamson and GitHub
ebcf511ec3
[ROCm][CI] Soft Fail Spec Decode Ngram + Suffix and Entrypoints Integration (LLM) AMD Mirrors ( #47067 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-29 16:24:08 -05:00
Matthew Bonanni and GitHub
8fc1b2d046
Fix FA4 dynamic_causal for full attention layers ( #46659 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-29 14:23:34 -07:00
Harry Mellor and GitHub
5316638a5e
Fix transient dependency issues caused by requirements/common.txt ( #47015 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 14:20:33 -07:00
zhrrr and GitHub
61ab70ec3b
[Model Runner V2] support mamba hybrid models align prefix cache ( #42406 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-29 14:09:16 -07:00
Woosuk Kwon and GitHub
a309d4fe60
Support DCP with FlashInfer MLA ( #43729 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-29 13:24:29 -07:00
72f639927f
[XPU] [RMSNorm] revert weightless change on xpu ( #46987 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 19:03:06 +00:00
Nick Hill and GitHub
8ad4a01825
[ModelRunner V2] Simplify recent UnlimitedOCR-related changes ( #46975 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-29 09:56:17 -07:00
Jee Jee Li and GitHub
7be582697b
[Bugfix] Fix DeepseekV2Model hidden_size ( #46986 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-29 16:44:05 +00:00
030c9523bd
[Perf][1/N] Expand Triton kernel warmup coverage, DSv4 ( #46634 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-29 16:40:34 +00:00
4708292d48
Bump flashinfer version to 0.6.13 ( #46683 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-29 09:30:57 -07:00
debec6440b
Add MiniMax-M3 modelopt nvfp4 support ( #46756 )
...
Signed-off-by: Xin Li <xinli@nvidia.com >
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
Co-authored-by: Xin Li <xinli@nvidia.com >
2026-06-29 09:29:39 -07:00
c8fb2963bd
[FS-Offloading] Batch Lookup in C ( #46713 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 09:28:32 -07:00
HDCharles and GitHub
379acd4e4f
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression ( #46860 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-06-29 15:55:42 +00:00
Martin Hickey and GitHub
07d33e575b
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager ( #44657 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 16:42:35 +01:00
36bbecd643
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" ( #46958 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 07:54:34 -07:00
Nicolò Lucchesi and GitHub
6149187a4c
[Kernel] Triton MLA logits workspace ( #46819 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-29 07:54:29 -07:00
Xiaohong (Sean) Chen and GitHub
49e28e8e91
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope ( #44010 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-06-29 22:54:15 +08:00
0ca39c4f1f
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone ( #46973 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-29 10:00:31 -04:00
Blas Rodriguez Irizar and GitHub
6185d73882
[Rust Frontend] Keep literal "null" string for string-typed tool params ( #46827 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-29 13:46:33 +00:00
bc8481af09
[MoE Refactor] Standardize Humming MoE experts + utilities ( #43373 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-29 06:19:29 -07:00
59575da46d
[XPU] exclude unsupported models for test_tensor_sechma.py ( #47008 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 12:30:28 +00:00
wang.yuqi and GitHub
3483240b7e
[Frontend] Consolidate scale out entrypoints ( #44512 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 03:18:53 -07:00
Roberto L. Castro and GitHub
eddfd4cf21
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen ( #46750 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-29 10:10:07 +00:00
Martin Hickey and GitHub
a4e3cb40d0
[mypy] Enable mypy for tests directory ( #47018 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 09:29:09 +00:00
soaringk and GitHub
ab132ee98b
Fix model info cache for package models ( #46567 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-29 09:17:54 +00:00
e186107870
[Bugfix] Use native SiLU activation in CPU fused MoE ( #45961 )
...
Signed-off-by: Alden Lobo <alden.lobo@arm.com >
Co-authored-by: Alden Lobo <alden.lobo@arm.com >
2026-06-29 09:12:20 +00:00
0e207dac78
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models ( #46835 )
...
Signed-off-by: John Langford <jl@hunch.net >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:59:15 +00:00
wang.yuqi and GitHub
9e86352c60
[CI Failure] Add transformers version check for openai/privacy-filter ( #47011 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 08:57:26 +00:00
Harry Mellor and GitHub
5051698e41
Remove unnecessary load_weights methods ( #44589 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 01:52:23 -07:00
Andreas Karatzas and GitHub
db28ae2d07
[ROCm][CI] Explicitly tear down multimodal offline LLMs ( #46999 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 07:59:24 +00:00
Harry Mellor and GitHub
f6bb8682ee
Fix docs on main ( #47009 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 15:50:57 +08:00
4559c43a95
[MM][CG] Gemma3 Encoder CUDA Graph ( #43591 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-29 04:52:00 +00:00
Bugen Zhao and GitHub
5274c1181d
[Rust Frontend] Add Harmony Renderer for GPT-OSS ( #46800 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-29 03:39:04 +00:00
Yuwen Zhou and GitHub
58d6a6e60a
[CPU] Support cpu compressed-tensor w8a8 int8 moe ( #42920 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-06-29 03:04:05 +00:00
a2abce646f
[EPLB] Mask padding in EPLB load recording ( #38128 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-28 19:43:58 -07:00
Harry Mellor and GitHub
311ad689ad
Remove boilerplate missed by #46820 ( #46956 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:11:17 +08:00
Woosuk Kwon and GitHub
0472436541
[Spec Decode] Avoid redundant hidden-states gather in draft prefill ( #46968 )
2026-06-28 17:04:01 -07:00
4dfbf1503b
[Model] Add support for openai/privacy-filter ( #41026 )
...
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-28 16:18:22 -07:00
Wei Zhao and GitHub
95528527ea
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 ( #46855 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-28 14:36:23 -07:00
c2127a25c7
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm ( #46895 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:50:30 -05:00
03c6d01c30
[OCP MX ] Add back emulation to available OCP MX backends list ( #46629 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:43:19 -05:00
Woosuk Kwon and GitHub
4b643c463e
[GLM5] Fix minor typo ( #46961 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-28 08:37:00 -07:00
7544286b04
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE ( #46552 )
...
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-28 15:19:28 +00:00
Woosuk Kwon and GitHub
89876b0c54
[GLM5] Implement op fusion for GLM5/DSV3.2 ( #46876 )
2026-06-28 08:17:39 -07:00
Wentao Ye and GitHub
5c91039c41
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement ( #46635 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-28 14:55:54 +00:00
5ecae3266c
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) ( #45033 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-28 07:52:00 -07:00
6eb63a1da6
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers ( #46600 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 01:37:44 -07:00
09841ae705
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation ( #46846 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-28 00:07:33 -07:00
Matt and GitHub
a2a92cbbaa
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection ( #46930 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-28 00:07:14 -07:00
35e6c86caa
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL ( #46034 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-28 00:06:43 -07:00
c7ca0bccae
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 ( #44313 )
...
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com >
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com >
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
2026-06-28 00:04:08 -07:00
c6741b2ad4
[Model] Support Unlimited OCR ( #46564 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-27 23:09:18 -07:00
a65f93fb2e
[ROCm][CI] Add ci_base metadata for external cache orchestration ( #46886 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-28 12:51:19 +08:00
Chauncey and GitHub
11a12305c0
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models ( #46786 )
2026-06-27 18:38:07 -07:00
798185d438
[KV-Offloading] Fix tensors_per_block stride ( #46888 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-27 21:01:45 -04:00
Matt and GitHub
9036c89ee4
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now ( #46928 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-27 17:30:49 -05:00
Giancarlo Delfin and GitHub
b6caeb5a09
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance ( #46878 )
2026-06-27 14:09:25 -07:00
Taneem Ibrahim and GitHub
8bf064f8d3
Fixed chunked embedding aggregation with request-id metadata ( #46782 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-27 20:57:47 +00:00
ea2ead1db3
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod ( #46818 )
...
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
2026-06-27 20:23:59 +00:00
Wentao Ye and GitHub
56aa067bf0
[CI Bug] Fix h100 AssertionError: Cold-start child failed ( #46927 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-27 20:17:33 +00:00
xiaolinchen and GitHub
35e3850fa9
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) ( #46915 )
...
Signed-off-by: wentian-byte <2990624738@qq.com >
2026-06-27 14:30:10 -04:00
51a99565c3
[ROCm][Perf] Fused shared expert for Minimax M3 ( #46474 )
...
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-27 12:34:17 +00:00
867fd5e8ed
[ROCm][Perf] Use flydsl moe with Minimax-M3 mxfp8 weights on gfx950 and implemented moe-backend selection ( #46184 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
2026-06-27 10:22:57 +00:00
9fd00ee006
[ROCm][CI] Move remaining mi250_2 tests out of the MI250 queue ( #46905 )
...
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-27 17:08:54 +08:00
091d13976c
[ROCm][CI] Add TRITON_ATTN score absolute tolerance floor ( #46891 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-27 06:35:50 +00:00
Wentao Ye and GitHub
b588f66dc2
[GLM5.2 Perf] fused_indexer_q_rope_quant triton kernel, 1.9% ~ 3.3% E2E Throughput improvement. ( #46862 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-26 22:16:20 -07:00
Benjamin Chislett and GitHub
455f25aa13
[CLI] Add flag to print TTFT and TPS in vllm chat ( #46775 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-26 22:15:10 -07:00
d706dec904
fix: Correct reasoning-end detection for prompt history ( #44551 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Signed-off-by: Jason Ozuzu <jasonozuzu@cohere.com >
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: Walter Beller-Morales <walterbm@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-26 22:15:06 -07:00
Divakar Verma and GitHub
68ee8300a0
[ROCm][CI]Fix test_concat_and_cache_mla_rope_fused on ROCm ( #46409 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 12:38:13 +08:00
ddd3855a28
[MoE Backend] add HPC-Ops MoE backend ( #45924 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: youkaichao <youkaichao@gmail.com >
2026-06-27 11:18:07 +08:00
Divakar Verma and GitHub
00e045b7c7
[ROCm][CI TG] refactor and fix deepep_moe test group ( #46758 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:45:23 +08:00
Divakar Verma and GitHub
17a71d8702
[ROCm][CI] Relax fused layernorm quant test tolerances for one-ULP outliers ( #46658 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:44:29 +08:00
weizhoublue and GitHub
2e058851d3
fix(docker): eliminate race conditions in shared buildkit cache mounts ( #44984 )
2026-06-26 19:43:17 -07:00
Dāvis and GitHub
1a92dfcce4
[Build] Show error message when using ROCm with LTO and different compilers ( #35232 )
2026-06-26 19:43:00 -07:00
Chris Leonard and GitHub
d0f800811b
[Build] Update vllm to point to vllm-project/flash-attention commit that builds FA3 with torch stable API. ( #46644 )
2026-06-26 19:42:46 -07:00
Nick Hill and GitHub
c6dd32a810
[ModelRunner V2] Support realtime embeddings ( #46762 )
2026-06-26 19:42:27 -07:00
af16446bf3
Vram semaphore infra ( #44465 )
...
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-26 17:32:51 -07:00
Harry Mellor and GitHub
3f67477497
[CI] Don't try and download files that we already know don't exist ( #46854 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-26 23:56:39 +00:00
Nick Hill and GitHub
1d41009e81
[ModelRunner V2] Fix cross-attention block table sizing ( #46753 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:34:21 -07:00
Nick Hill and GitHub
b94f212e37
[ModelRunner V2] Deduplicate ModelState init logic ( #46776 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:32:45 -07:00
Harry Mellor and GitHub
d8eb734d94
Fix Transformers backend FP8 MoE and remove some boilerplate ( #46820 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-27 00:16:05 +01:00
2ff76a5e85
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 ( #46760 )
...
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-26 21:58:40 +00:00
Yifan Qiao and GitHub
75fdcc82a5
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas ( #46873 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-26 21:48:53 +00:00
yzong-rh and GitHub
77f8796d16
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. ( #46437 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-26 17:18:47 -04:00
c40d307731
[Core] Remove FlashAttention block size restriction for hybrid models ( #36701 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-26 21:16:39 +00:00
Woosuk Kwon and GitHub
65e655d295
[GLM-5] Add DSV3.2/GLM5 to vllm/models/ ( #46808 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-26 14:09:05 -07:00
Charlie Fu and GitHub
6e2fb02fe5
[ROCm][CI] Fix rlhf_nccl.py on ROCm ( #46851 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-26 15:41:49 -05:00
Micah Williamson and GitHub
274325dd43
[ROCm][CI] Remove V1 Sample + Logits from mi250 Queue ( #46867 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-26 15:38:38 -05:00
Matt and GitHub
95e6442a6b
[Hardware][AMD][CI] Fix Kernels Quantization test timeout ( #46859 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-26 15:19:16 -05:00
701a23d99f
[Bugfix][Model] Support tensor parallelism for DiffusionGemma ( #45719 ) ( #46177 )
...
Signed-off-by: Carlos Alvarado <carlos-alvarado@outlook.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-26 20:05:04 +00:00
Ben Browning and GitHub
dccb412e2c
[Bugfix][Parser] Pass token IDs to parser.parse() in Responses API and batch serving ( #46843 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-26 19:29:52 +00:00
c6554f321c
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build ( #46769 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 14:32:21 -04:00
Julien Denize and GitHub
3d3b96488f
Migrate Voxtral to mistral-common 1.11.5 audio API ( #46705 )
...
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-06-26 11:06:31 -07:00
Nick Hill and GitHub
658b54efe4
[ModelRunner V2] Update scheduler tests to cover MRV2 paths ( #46771 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 09:36:31 -07:00
Li, Jiang and GitHub
abc71548ef
[CI/Build][CPU] Add test image cache clean-up ( #46831 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-26 23:28:49 +08:00
Nick Hill and GitHub
4e07ca2c92
[Core] Add VLLM_GPU_SYNC_CHECK env var ( #44800 )
2026-06-26 08:24:33 -07:00
Bugen Zhao and GitHub
e71bc6da85
[Rust Frontend] Use oss-harmony for Harmony output processing ( #46799 )
2026-06-26 08:24:13 -07:00
fxmarty-amd and GitHub
37ce34922f
[CI] Fix failing CUDA graph capture in Triton MOE ( #46735 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-26 07:21:20 -07:00
c2507fb293
[ROCm] [MoE] [Perf] Shared-expert fusion for bias-routed MoE; enable on MiniMax-M3 mxfp8 model ( #46545 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 07:05:20 -07:00
TJian and GitHub
8921c4be88
[ROCm] [Performance] Optimize aiter moe for DeepSeekV4 ( #46122 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 06:43:27 -07:00
8e394244a5
[ROCm]Enable AITER MoE backend for MiniMax-M3-MXFP4 ( #46419 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-26 06:35:35 -07:00
TJian and GitHub
302954e5f6
[ROCm] [CI] fix transcription flakiness AMD: Entrypoints Integration (API Server OpenAI - Part 1) (mi325_1) ( #46823 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 21:33:35 +08:00
Hyunkyun Moon and GitHub
950ee4c2e4
[API] Add token offsets to render endpoints (/v1/.../render) ( #44226 )
...
Signed-off-by: HyunKyun Moon <mhg5303@gmail.com >
2026-06-26 05:02:52 -07:00
d980a3cc6e
[ROCm] Fix AITER_UNIFIED_ATTN Dispatching After AITER Bump ( #46780 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-26 02:09:56 -07:00
bf292b5f6b
[Docs] Remove BambaForCausalLM from supported hybrid models list ( #46071 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-26 08:02:50 +00:00
wang.yuqi and GitHub
5e3dad04b1
[Misc] Move the legacy api_server.py to the examples directory. ( #46783 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-26 07:43:29 +00:00
Joe Rowell and GitHub
63e161f296
[Bugfix][Tool Parser] PoolsideV1: fix string whitespace and required named tool choice ( #46486 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-06-26 06:05:16 +00:00
Tiezhen WANG and GitHub
c7645bce04
Remove grok model arch from vllm ( #46706 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
2026-06-25 23:02:10 -07:00
35a49fcfc2
[CI][Bugfix] Spawn engine in mm cache sleep test to fix ROCm HIP error ( #46749 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-26 00:38:26 -05:00
peizhang56 and GitHub
915e99ec67
[ROCm][Bugfix] Fix HIP fork re-init in multimodal offline examples ( #46741 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
2026-06-26 00:37:47 -05:00
Nick Hill and GitHub
5b33041746
[ModelRunner V2] Fix whisper test ( #46773 )
2026-06-25 22:10:36 -07:00
Matt and GitHub
1a4984520e
[Hardware][AMD][CI] Fix AMD CI image build ( #46792 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 22:05:12 -07:00
Reid and GitHub
e312c5cb25
[Rust Frontend] Make Granite4 string argument scanning incremental ( #46507 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-26 03:54:03 +00:00
Matti4 and GitHub
1502cf6274
Fix relative allowed local media paths ( #45263 )
2026-06-25 20:45:20 -07:00
d350fa8ddd
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens ( #46733 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-26 03:41:33 +00:00
dbc49b6b99
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 ( #45166 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
Signed-off-by: ovidiusm <ovidium@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-25 19:33:42 -07:00
fxmarty-amd and GitHub
552a9dbe59
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears ( #44667 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-25 19:33:00 -07:00
02a1f23711
[DFlash] Fuse precompute kv per-layer rmsnorms ( #46761 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:32:07 -07:00
652d962bc9
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation ( #46448 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:30:07 -07:00
Giancarlo Delfin and GitHub
5314665bad
[Model Runner V2][DFlash] Enable dflash attention backend selection ( #46770 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 19:29:25 -07:00
Michael Goin and GitHub
3daea7ceb9
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models ( #46759 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 19:03:09 -07:00
Wentao Ye and GitHub
cc7981599e
[Refactor] Remove dead kernel code ( #46405 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 18:09:56 -07:00
Nick Hill and GitHub
32bb3195f0
[ModelRunner V2] Bound memory for large logprobs requests ( #46746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-25 18:04:06 -07:00
ad28d605e6
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) ( #45544 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-25 17:46:28 -07:00
Bugen Zhao and GitHub
ae7c8ec223
[Rust Frontend] Switch rustls to native-tls/OpenSSL ( #46696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:19:44 -07:00
Bugen Zhao and GitHub
1d3f4cb3a4
[Rust Frontend] Extract renderer fixture test utilities ( #46719 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:12:38 -07:00
Bugen Zhao and GitHub
f9e684499f
[Rust Frontend] Migrate gemma4 to unified parser ( #46602 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 16:59:57 -07:00
Giancarlo Delfin and GitHub
c53994e134
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling ( #46665 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 23:46:10 +00:00
Matt and GitHub
27da2a2ac4
[Hardware][AMD][CI] Use Triton-based AITER MHA for LM Eval Qwen-3.5 Models Tests ( #46691 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 17:08:04 -05:00
Michael Goin and GitHub
a2e8ec3d52
[CI] Depend GPQA Eval DGX Spark job on arm64 image build ( #46736 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 17:07:04 -04:00
e8c24a7695
[Kernel] Vectorized fp32 moe_sum reduction and support any topk ( #46643 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 14:02:28 -07:00
Andreas Karatzas and GitHub
2a6f8f0c05
[ROCm][CI] Fine-tuning queues and test names ( #39238 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-25 13:24:09 -07:00
Robert Shaw and GitHub
c5e3c40877
Fix P/D with DP Supervisor ( #46628 )
...
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-25 13:13:08 -07:00
Wentao Ye and GitHub
8b4d93ba2b
[Perf] Remove redundant clone for GLM, Deepseek etc ( #46651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 13:09:00 -07:00
Michael Goin and GitHub
e8e7b592d1
[Kernel][MoE] Tune block-FP8 fused MoE for low-batch decode ( #46642 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 12:38:28 -07:00
Rohan Potdar and GitHub
e53a17232c
[ROCm]: Bump aiter to 0.1.16.post2 ( #46692 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-25 11:53:26 -07:00
Flora Feng and GitHub
96eb8ddc41
[CI] Re-enable skipped glm and seedoss parser tests ( #46671 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-25 13:11:41 -04:00
Gabriel Wu and GitHub
8fa36fbbeb
[Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 ( #46506 )
2026-06-25 09:12:00 -07:00
Ranran and GitHub
e45b279928
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3Next ( #46316 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
2026-06-25 09:05:04 -07:00
d490b98162
[Core] Avoid mixed length specdec batches via padding ( #45237 )
...
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
2026-06-25 08:34:44 -07:00
haoyangli0109 and GitHub
1744adc256
[ROCM] [Communication] Add INT3 quantization method for quickreduce ( #45666 )
...
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com >
2026-06-25 15:14:15 +00:00
Divakar Verma and GitHub
cdfa2fd7e9
[ROCm][CI] rm duplicate Distributed Torchrun ci test ( #46729 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-25 09:58:17 -05:00
6f3da461d1
[Pooling] Fix Cohere embed billed image token accounting for mixed-content inputs ( #46093 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 10:44:29 -04:00
Russell Bryant and GitHub
d3130d878c
[CI] Pin GitHub Actions to commit hashes in macos-smoke-test.yml ( #38290 )
2026-06-25 13:48:44 +00:00
9bfd878a48
[MoE] [MoE Refactor] Add moe kernel oracle abc 37753 ( #43461 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Signed-off-by: qyYue1389 <yueqiuyang1389@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 09:34:03 -04:00
Matt and GitHub
2365b7a8e7
[Hardware][AMD][CI] Mirror Basic Models (Others) and Weight Loading Multiple GPU test groups ( #46668 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 08:25:09 -05:00
15be78732b
[NIXL][Mamba] Add Mamba1 support to NIXL P/D disaggregation ( #45019 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 05:50:41 -07:00
92221485aa
[CPU][CI/Build] Allow more CPU CI agents ( #46702 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:33:39 +08:00
xiangdong and GitHub
a6f41ab678
[XPU][CI]Refine .buildkite/ci_config_intel.yaml for Intel GPU CI ( #46674 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-25 08:58:26 +00:00
c63cd4906c
[ROCm][ [Perf] sparse attention optimization on minimax-m3 ( #46546 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: yueliu14 <yue.liu4@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-25 16:56:00 +08:00
638b1a99cc
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM ( #45269 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-25 08:18:10 +00:00
72adb20a6a
[Model] Remove AquilaForCausalLM, AquilaModel ( #46605 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 08:08:26 +00:00
2396d91e93
[CPU][Spec Decode] Enable DFlash SD for CPU ( #44029 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 15:32:48 +08:00
9b215ae60b
[Rust Frontend] Forward VLLM_ENGINE_READY_TIMEOUT_S via --args-json ( #44610 )
...
Signed-off-by: kai <kai@example.com >
Co-authored-by: 图灵 <tuling.wk@alibaba-inc.com >
2026-06-25 07:25:08 +00:00
Bugen Zhao and GitHub
4d3b4b9b01
[Rust Frontend] Make ToolParserOutput a seq of ToolParserEvent to preserve order ( #46584 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 06:27:07 +00:00
Matthias Gehre and GitHub
77c1d9fe9b
[ROCm][Perf] Tune wvSplitK on gfx1151 ( #40784 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-06-25 14:17:46 +08:00
Jeff (Junze) Ma and GitHub
36fd7e8b86
[SimpleCPUOffloadConnector] Fix remaining global→block conversions under PCP/DCP ( #46394 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
2026-06-24 23:05:24 -07:00
fc61c6fc26
[Perf] Enable + tune FlashInfer fused allreduce at world_size=16 on SM 10.3 (GB300) ( #46392 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:04:17 -07:00
Matt and GitHub
e2af449c39
[Hardware][AMD][CI] Move Metrics, Tracing (2 GPUs) & make optional ( #46686 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 05:49:33 +00:00
3f5a1e1733
[ROCm][CI] Expand basic correctness target suites ( #46573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matt <156021403+mawong-amd@users.noreply.github.com >
2026-06-25 12:18:57 +08:00
710ebaa189
[ROCm][Bugfix] Fix chunk alignment when using context parallelism with TRITON_MLA ( #46114 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:07:28 -04:00
1aad125815
[CPU] Enable chunked prefill and prefix caching for qwen3.5 ( #46202 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-25 03:49:21 +00:00
dc55936f64
[AMD][CI] Fix Pipeline + Context Parallelism test group ( #46650 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 22:23:42 -05:00
Bugen Zhao and GitHub
76c3c4ff63
[Rust Frontend] Introduce unified parser interface & combined parser ( #46583 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 03:17:31 +00:00
efb5acffd5
[Bugfix] fix: stream Mimimax m2 tool call string arguments ( #46382 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-25 03:12:45 +00:00
6e3a983cf3
[ROCm] Remove erroneous inclusion of gptq_marlin as supported quant scheme on ROCm ( #46655 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 21:27:19 -05:00
Xin Yang and GitHub
1273a8f05a
[Kernel] Add swap AB optimization to fused_moe_kernel ( #36559 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-06-25 01:44:30 +00:00
9e88e969c0
[Perf][KVConnector][Mooncake] Parallelize KV load with a receive-thread pool ( #45971 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 18:25:12 -07:00
dda3aca47f
[Speculative Decoding] Propagate norm_output and fc_norm config for Eagle3 speculators ( #46488 )
...
Signed-off-by: Orestis Zambounis <orestis.zambounis@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:51:33 +00:00
Jee Jee Li and GitHub
23aed9b0ee
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel ( #46508 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-25 08:42:51 +08:00
Maxwill Lin and GitHub
cd347298e8
[Frontend] Port seed_oss to the streaming parser engine as a Qwen3 subclass ( #46314 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 20:08:42 -04:00
Yifan Qiao and GitHub
b69816043a
[Bugfix][MooncakeStore] track resumed requests via scheduler's resumed_req_ids ( #46595 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 23:50:56 +00:00
Kaihang Jiang and GitHub
fc7fc421e9
[Kernel][MoE] Allow FlashInfer MXINT4 MoE for gated SiLU ( #46518 )
...
Signed-off-by: Kaihang Jiang <kaihangj@nvidia.com >
2026-06-24 18:32:50 -05:00
cyq and GitHub
e06a83445c
[Bugfix] Normalize slashes in Helion GPU names ( #46101 )
...
Signed-off-by: cyq <15000851237@163.com >
2026-06-24 18:22:49 -05:00
d7ab9be775
[Bugfix] Support -1 (invalid/non-local) slots in topk_ids for Triton MoE ( #46408 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:59:42 -07:00
6a1570711c
[Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing ( #46406 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:52:09 -07:00
Micah Williamson and GitHub
d6696e2385
[ROCm] Begin Deprecation Window for CUDA_VISIBLE_DEVICES on ROCm ( #46636 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-24 20:40:28 +00:00
Chauncey and GitHub
84c2f9f0fb
[Frontend] Fix Kimi K2 tool call IDs for required tool choice ( #46344 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-24 19:59:40 +00:00
49f2104c53
[Feature] Support DCP with FP8 KV cache in MLA decode path ( #44044 )
...
Signed-off-by: shivampr <shivampr.dev@gmail.com >
Signed-off-by: Shivam <shivamprasad91@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 19:28:17 +00:00
d511b5bae9
Chore: Fix minor doc sentence, grammar, quote errors ( #40469 )
...
Signed-off-by: Ashwin Phadke <23502062+ashwin-phadke@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-24 18:58:24 +00:00
3c43237233
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels ( #46560 )
...
Signed-off-by: xiaojun.wei <jessiewei747@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-24 11:00:57 -07:00
56ca5997ea
Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference ( #46389 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-24 13:53:54 -04:00
Aarushi Jain and GitHub
cf57311187
Run DeepSeek-V2-Lite prefetch-offload eval eager on ROCm ( #46386 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-24 12:25:57 -05:00
Lucas Wilkinson and GitHub
e7df232288
[KV Offload] Gate packed HMA KV cache on cross-layer config ( #46252 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-06-24 11:55:30 -04:00
b3a688cb9e
[ROCm] Fix OOB During Model Warmup With ROCM_ATTN and MRV2 ( #46548 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-24 10:53:21 -05:00
Wentao Ye and GitHub
1cd3e0e945
[Bug] Fix IndentationError: expected an indented block after 'with' statement ( #46627 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-24 23:14:17 +08:00
Yiwei Hu and GitHub
f889325c51
[KV Offload] Use background thread for mmap / cpu_tensors pinning ( #45850 )
...
Signed-off-by: Sorryhorizon <arikara6666@gmail.com >
2026-06-24 18:13:27 +03:00
bb61177e49
[KV Offloading] Replace bool|None lookup return with LookupResult enum ( #46363 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 18:06:08 +03:00
7f99e80c3b
[Perf][ThinkingBudget] reduce search space for thinking tokens ( #46425 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:02:25 +08:00
2801b11156
[Test] Pin block_size in auto-fit max_model_len test ( #45914 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:56:21 +08:00
007b5a52ed
[Log] Update to log once ( #46511 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-24 14:45:16 +00:00
Cyrus Leung and GitHub
24d5186138
[Bugfix] Re-enable FP8 MoE on NVIDIA Thor ( #46339 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-24 07:35:46 -07:00
Nemani Harsha Vardhan and GitHub
7dc036058b
[Doc] Document Qwen3.6 (dense + MoE) ViT CUDA graph support ( #44720 )
...
Signed-off-by: harsha20032020 <nhvardhan2020@gmail.com >
2026-06-24 14:35:08 +00:00
61ee183d28
[ROCm] Fix AITER FP8 quantization schema tests ( #46414 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:29:19 +08:00
84c62e1cbd
[Model Runner V2][MM] Support EVS ( #46535 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:18:56 -04:00
Fadi Arafeh and GitHub
061043eaca
[CPU][Perf] Accelerate unquantized MoE for AArch64 ( #46353 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-24 14:14:35 +00:00
93ec645878
[Bugfix] Fix illegal memory access from a forward during a partial wake_up ( #44483 )
...
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 22:12:23 +08:00
Kunshang Ji and GitHub
563c628968
[XPU] bump up vllm_xpu_kernels to v0.1.10.1 ( #46607 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 10:05:31 -04:00
0bc479e6eb
[Perf][LoRA] Replace O(n) list.index() with a dict in convert_mapping ( #46542 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 21:41:46 +08:00
Tae Jeong and GitHub
62890e204c
Fix duplicated logging when loading a corrupt or partial video ( #46467 )
...
Signed-off-by: hhhhhhhhhhhhhhhhho <man2719@naver.com >
2026-06-24 06:14:13 -07:00
Nicolò Lucchesi and GitHub
a2cb08b3d5
[Misc][PD] Disable bidirectional xfer mode for NixlPushConnector ( #46473 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-24 21:14:05 +08:00
cf9fd6457e
Fix KV offload request-finished lifecycle contract ( #46284 )
...
Signed-off-by: test test <2260891073@qq.com >
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 15:42:40 +03:00
Kunshang Ji and GitHub
d4448b511d
[XPU][Docker] switch to ubuntu 24.04 as base image ( #45973 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 20:39:20 +08:00
f1a6703edd
[Bugfix][Config] Keep pydantic validation for fields with a TYPE_CHECKING Literal alias ( #46220 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 12:25:50 +00:00
Roy Wang and GitHub
160c80a34c
[Rust Frontend] Raise frontend JSON body limit ( #46582 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-06-24 12:15:31 +00:00
Martin Hickey and GitHub
f237e16b41
[KV Offload] Replace OffloadingHandler with OffloadingWorker ( #45053 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-24 14:44:24 +03:00
70749fdcca
[Feature] Triton INT4 per-token-head KV cache quantization ( #40835 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:21:25 +00:00
d20dbf921b
[Mooncake] Only check and store new KV cache range ( #46412 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 03:10:50 -07:00
ede54b926e
set AttentionCGSupport.UNIFORM_BATCH for fa2 on xpu ( #46555 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 18:05:02 +08:00
52fbe12283
[Perf][Multimodal] Avoid building a full timestamps list in video frame sampling ( #46543 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 09:38:27 +00:00
Dakai An and GitHub
dc0d318177
[Attention] Add FLASH_ATTN_MLA_SPARSE backend for Hopper sparse MLA ( #46189 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-06-24 09:33:10 +00:00
soaringk and GitHub
d7c1821b5a
[Model][MiniMax-M3] Add pipeline parallelism support ( #45810 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-24 08:23:03 +00:00
4cd1a84c88
[Model] Remove BaiChuanForCausalLM and BaichuanForCausalLM ( #46362 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 16:13:57 +08:00
Mohammad Miadh Angkad and GitHub
191826ec61
[CI/Build] Fix topk histogram build on SM75 ( #46550 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-24 00:51:11 -07:00
Andreas Karatzas and GitHub
549c7074cd
[ROCm][CI] Skip the MoE Marlin tile-padding helper assertion ( #46580 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 07:31:33 +00:00
489abadfb8
feat: support to OpenMOSS-Team ( #44124 )
...
Signed-off-by: nagisa-kun <1434936049@qq.com >
Signed-off-by: nagisa19 <1434936049@qq.com >
Signed-off-by: nagisa <1434936049@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 00:08:13 -07:00
Woosuk Kwon and GitHub
96de8bb389
[MoE] Free unused MXFP4 scales in OAI Triton Backend ( #46549 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-24 00:06:41 -07:00
Jee Jee Li and GitHub
9d6fdc2901
[Kernel] GLM5 Router GEMM ( #46385 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:54:50 -07:00
Benjamin Chislett and GitHub
4c5bc41ba6
[Bugfix][Spec Decode] Fix probabilistic sampling for parallel drafting ( #45956 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-24 05:36:23 +00:00
Michał Ganczarenko and GitHub
ac1fa74616
[Bugfix] Fix NemotronLayerNorm1P hardcoded cuda device type ( #46495 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
2026-06-24 13:21:02 +08:00
Sting Lin and GitHub
556bc4e3a0
Upgrade tpu-inference to v0.23.0 ( #46568 )
2026-06-23 21:15:14 -07:00
Wei Zhao and GitHub
05a0caba91
[Mooncake] Optimize lookup pool key string construction ( #46188 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-24 11:51:53 +08:00
Nick Hill and GitHub
7ee4d22009
[Spec Decode] Reject placeholder (-1) draft tokens in rejection sampler ( #46533 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 03:32:32 +00:00
ce9f64020b
[Rust Frontend] Pass effective reasoning_parser_kwargs for structured output ( #46360 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 03:13:44 +00:00
4ed8eaafb0
[Rust Frontend] Integrate xgrammar-structural-tag for strict and required tool calling ( #46057 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 10:46:49 +08:00
6af0559ddb
[Core][DP] Throttle prefills based on local prefill work ( #46532 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 02:27:12 +00:00
e2bdc24612
[ROCm][Bugfix] Fix use_v2_model_runner inside Ray driver thread ( #45998 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 08:41:56 +08:00
Andreas Karatzas and GitHub
bcbeaac786
[ROCm][CI] Stage C-II of gating additional test groups ( #46537 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 17:36:40 -07:00
Maxwill Lin and GitHub
e48f2aa4ca
[Bugfix][Frontend] Emit a content block for empty Anthropic completions ( #46525 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 00:04:26 +00:00
Roberto L. Castro and GitHub
d86c66c981
[Feat] Add runtime monitor for post-warmup CuTeDSL compilation ( #46167 )
2026-06-23 23:33:17 +00:00
Nico Holmberg and GitHub
80e511772f
[ROCm][Bugfix][Perf] enable shared expert fusion for Qwen3.5 ( #44434 )
...
Signed-off-by: Nico Holmberg <nico.holmberg@amd.com >
2026-06-23 23:19:51 +00:00
Roberto L. Castro and GitHub
855cd4d787
[Perf][DSv4/DSv3.2] Add cluster-cooperative topK kernel for low-latency scenarios ( #43008 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-23 16:11:00 -07:00
3cc871aaf1
[Perf] Skip detokenization in online beam search ( #46422 )
...
Signed-off-by: Guy Stone <guys@spotify.com >
Signed-off-by: Guy Stone <guystone3@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 15:46:09 -07:00
0a3e2dbc09
[Optimization] Skip DP padding tokens in MoE ( #46428 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 14:54:46 -07:00
84f13374b3
[CI] Fix test_auto_gptq on ROCm CI ( #46164 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 16:38:06 -05:00
Micah Williamson and GitHub
b28103e1ca
[ROCm][CI] Shard LM Eval Qwen3-5 Models (B200-MI355) in AMD CI ( #46520 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 16:32:05 -05:00
Wentao Ye and GitHub
abc33134fa
[CI Test] Mark batch invariance test flaky ( #46530 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 21:01:34 +00:00
6617db1bfb
[Bugfix][Frontend] Emit non-ASCII tool-call arguments without \uXXXX escapes ( #46308 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-23 20:43:11 +00:00
899d72a58c
[Bugfix][ToolParser] Handle braces in required tool streaming strings ( #45389 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-23 20:29:34 +00:00
Yongye Zhu and GitHub
11b56b2ff2
[Kernel] Add FlashInferCutedslMxfp8LinearKernel (cute-dsl mm_mxfp8) ( #46393 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-23 12:45:49 -07:00
0d4d164488
[Bugfix] Allow flashinfer_cutlass as a clamped NVFP4 MoE backend ( #46492 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-23 12:43:36 -07:00
Mike G and GitHub
0775b882ba
[NVFP4 MoE/Deepseek V4] Marlin: wire SwiGLU clamp + allow it for clamped models on non-Blackwell ( #45836 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-23 12:21:19 -07:00
7c2e08451a
[Docker] Remove redundant flashinfer download-cubin step ( #46517 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 12:16:51 -07:00
Giancarlo Delfin and GitHub
ef361de916
[Model Runer V2][DFlash] Fix lm head sharing for dflash ( #46435 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-23 19:09:06 +00:00
Yan Ma and GitHub
acce57d8dd
Deprecate old FP8 online MoE quantization class ( #44514 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:53:38 -07:00
68afd78897
[Bugfix][ROCm] Fix cumem sleep and teardown ( #46203 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 02:45:31 +08:00
37a682d392
[Kernel] Extend Marlin thread-tile padding to MoE (WNA16 + FP8/MXFP8) ( #45703 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 11:45:10 -07:00
Rui Yin and GitHub
d8e422ccda
[Bugfix] Parse MiniMax M3 streaming reasoning by text markers ( #45718 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-23 14:43:58 -04:00
fxmarty-amd and GitHub
e368415daa
[AMD][OCP MX][CI] Fix tests to not dispatch on UNFUSED_TRITON backend on MI300, improve w_mxfp4_a_fp8 emulation support ( #46142 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-23 14:25:27 -04:00
Andreas Karatzas and GitHub
ceae5bcbda
[ROCm][CI] Fix nixl tests ( #45219 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 13:11:40 -05:00
6691f087a6
[Minimax-M3] BF16/FP8 Indexer using MSA ( #45892 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-23 10:28:49 -07:00
f4d5f73ffa
[Bugfix]: Fix unquantized gpt-oss weight loading broken by FusedMoE r… ( #45818 )
...
Signed-off-by: priyansh jain <priyansh.jain2@amd.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-23 16:56:17 +00:00
fd50a66015
[CI][ROCm] Skip unsupported test cases on ROCm ( #46160 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 11:35:49 -05:00
84586c9acc
[ROCm][CI] fix fp8 range in vit_fp8_quant ( #46410 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 11:34:21 -05:00
Taneem Ibrahim and GitHub
40e5522121
[Docs] Add Qwen3 forced alignment online example ( #46197 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-23 11:59:45 -04:00
Willow Lopez and GitHub
f3410b3bb1
fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility ( #45404 )
...
Signed-off-by: Oxygen <1391083091@qq.com >
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-23 11:46:23 -04:00
568874fec2
[ROCm][CI] pass merge-base to container for python-only wheel metadata ( #45869 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 15:44:43 +00:00
275b43183c
[MyPy] Fix mypy for vllm/benchmarks ( #39896 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-23 15:22:29 +00:00
Yan Ma and GitHub
547d2c40d7
Add weights padding for fp8 per-block online quantization ( #44763 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:08:17 -04:00
2aaaf3febd
[ROCm][Test] Fix stale test_gfx950_moe MXFP4 oracle tests ( #46260 )
...
Signed-off-by: Spandan Tiwari <23646532+spandantiwari@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 23:07:46 +08:00
Micah Williamson and GitHub
156b12667c
[ROCm][CI] Skip Quark mxfp4 tests unless Quark version is compatible with Torch version ( #46431 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 22:26:39 +08:00
Jee Jee Li and GitHub
9f6f296428
[CI/Build] Remove BaiChuanForCausalLM from the LoRA test ( #46494 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:09:48 +08:00
e51e700470
[LoRA] Gate all_gather on fully_sharded_loras inside _mcp_apply; rewrite regression test ( #45715 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-23 07:08:33 -07:00
f59db63732
[Bugfix] GPT-OSS Autodrop reasoning in Response API and cleanup ( #45048 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 09:36:33 -04:00
Rukhaiya2004 and GitHub
9f5117820f
[HARDWARE][POWER] Enable fp16 support for PowerPC ( #46135 )
...
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
2026-06-23 13:24:49 +00:00
1bf149f334
Filter Pydantic-internal markers from validation error param ( #46457 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Rehmann <114791389+mahadrehmann@users.noreply.github.com >
2026-06-23 13:20:50 +00:00
2a675a7b9f
[Bugfix] Responses API assistant EasyInputMessageParam input ( #44361 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 08:54:46 -04:00
7d47cff933
[Bugfix][KV Offload] Fix swap_blocks_batch on the default stream ( #46379 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-23 05:45:27 -07:00
091bc1026e
[KV Offloading] Add tiering metric plumbing ( #45959 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-23 15:10:36 +03:00
3554ada5d8
[CPU][Bugfix][Speculative Decoding] Accept USE_FP64_GUMBEL in CPU recovered-tokens sampler ( #46069 )
...
Signed-off-by: hillel.darshan <hillel.darshan@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-23 11:54:07 +00:00
wang.yuqi and GitHub
31ca9504b1
[Frontend] Split ServingRender into renderer and entrypoint. ( #44285 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-23 11:19:09 +00:00
d32575a2d2
[ROCm][P/D] Support MoRIIO heterogeneous TP fan-in ( #46332 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-23 10:33:23 +00:00
Juan Pérez de Algaba and GitHub
83fa302ca4
fix(security): prevent infinite loop in split_audio with NaN audio sa… ( #46463 )
2026-06-23 10:24:51 +00:00
frida-andersson and GitHub
20b5af55c1
[ROCm][Perf] DSv3.2: fuse MLA Q concat+fp8-quant in forward_mqa ( #43673 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-06-23 18:12:04 +08:00
Qiming Zhang and GitHub
901a3b091c
fix gpt_oss pp>1 with ep ( #46441 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-23 16:59:11 +08:00
2d721ab5d8
[Rust Frontend] Align Rust allowed_token_ids validation with Python ( #46348 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 08:32:33 +00:00
accaa434f3
[Rust Frontend] Support echo for token-ID completion prompts ( #46219 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-23 08:04:41 +00:00
Sunny Yuan and GitHub
a04654da23
Doc: fix missing GLM-5.x in supported models ( #46452 )
...
Signed-off-by: Sunny Yuan <y.zichen@wustl.edu >
2026-06-23 07:42:27 +00:00
Bugen Zhao and GitHub
25bc3be49c
[Rust Frontend] Correct --reasoning-parser semantics ( #46359 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 15:38:39 +08:00
Ting SUN and GitHub
a46f3eb232
[Bugfix][Model Runner V2] Preserve all allowed_token_ids in the logit bias kernel ( #46245 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-23 07:01:13 +00:00