Bugen Zhao and OpenAI Codex
e1a763558c
[CI] Discover Rust coverage artifacts from build metadata
...
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-23 06:44:16 +00:00
Bugen Zhao and OpenAI Codex
84aeec9f22
[CI] Simplify Rust coverage reporting
...
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-22 08:18:14 +00:00
Bugen Zhao and OpenAI Codex
82a770ddbd
[CI] Simplify Rust coverage aggregation
...
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-22 02:58:45 +00:00
Bugen Zhao
a09a9bace1
[CI] Disable redundant Codecov file fixes
2026-07-21 13:57:38 +00:00
Bugen Zhao
6c20d467a2
[CI] Run Codecov from repository root
2026-07-21 13:34:22 +00:00
Bugen Zhao and OpenAI Codex
cb59d0a351
[CI] Collect Rust coverage in Buildkite
...
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-21 13:01:56 +00:00
Bugen Zhao and OpenAI Codex
0ab1bded36
[CI] Instrument Rust artifacts for coverage
...
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-21 12:22:25 +00:00
Umut Polat and GitHub
040cbf95cc
[Misc] Use VLLMValidationError in chat completion tool and batch validators ( #49214 )
...
Signed-off-by: Umut Polat <52835619+umut-polat@users.noreply.github.com >
2026-07-21 11:38:20 +00:00
5b3762a7f0
[Bugfix][CPU] Fix Clang OpenMP build on macOS ( #49021 )
...
Signed-off-by: markyangcc <mmdou3@163.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-21 09:58:52 +00:00
bastefaniak and GitHub
4d30c510ce
[bugfix] Fix Cosmos3 Edge checkpoint weights filtering, video loading, prompt expansion ( #49190 )
...
Signed-off-by: Bartosz Stefaniak <bstefaniak@nvidia.com >
2026-07-21 17:18:36 +08:00
6700813f86
[3/N][KV-Cache Layout Refactor] Standardize Mamba cache; drop get_transfer_cache_regions ( #44456 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-21 09:16:15 +00:00
Bugen Zhao and GitHub
eb44b3aaa4
[Rust][Benchmark] Use async HTTP clients ( #49295 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-21 16:53:57 +08:00
Nicolò Lucchesi and GitHub
7a98c7a392
[Misc] Remove old now unsupported max_num_partial_prefills and max_long_partial_prefills ( #49244 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-21 08:52:52 +00:00
Lena Onyshchenko and GitHub
0d9e60619b
[Misc][Docs] Fix XPU compute-runtime driver link version mismatch ( #49299 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
2026-07-21 08:45:41 +00:00
1134545b6f
Revert "[Sampler] Stop upcasting logits to fp32 in apply_sampling_params" ( #48641 ) ( #49033 )
...
Co-authored-by: vllm-agent <vllm-agent@users.noreply.github.com >
2026-07-21 09:36:45 +01:00
3e0c887511
[Bugfix] Fix Ovis2_5 special tokens for transformers v5 ( #47298 )
...
Signed-off-by: mgrunwal <milosz.grunwald@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-21 08:09:20 +00:00
Stefan Kaestle and GitHub
adfbbc1005
Propagate Flash Attention cache configuration to Ray workers ( #49177 )
...
Signed-off-by: Stefan Kaestle <skaestle@nvidia.com >
2026-07-21 07:47:53 +00:00
Roy Wang and GitHub
adc98f04d0
[Misc] Add @esmeetu to codeowners for rust/src/bench ( #49298 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-07-21 07:44:24 +00:00
8def3cdde2
[Bugfix] Propagate quant_config to LFM2 ShortConv projections ( #48917 )
...
Signed-off-by: Alex Yuan <alex.yuan@liquid.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-21 06:51:55 +00:00
616c9bd0f4
[Frontend] Support additional sampling parameters for translation API ( #45839 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-21 05:53:56 +00:00
Bugen Zhao and GitHub
8688a06d67
[Rust][Benchmark] Use tracing for logs ( #48937 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-21 05:18:13 +00:00
f25953cc59
[Bugfix][Rust Frontend] Handle zero-column logprobs payloads without panicking ( #49113 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Feathbow <feathbow@gmail.com >
2026-07-21 04:30:29 +00:00
d9aa35161d
Update BGE-M3 token expectations for leading spaces ( #49269 )
...
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com >
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com >
Co-authored-by: Codex <noreply@openai.com >
2026-07-21 03:49:17 +00:00
6bcda970fd
[CI][NIXL] Isolate concurrent engine internal ports ( #49129 )
...
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-20 22:28:11 -05:00
Isotr0py and GitHub
ea0e9c8f2e
[MRV2] Add encoder cache profiling implementation ( #47985 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-20 20:18:14 -07:00
Chauncey and GitHub
94ed0bf4e0
[Bugfix][KV Offloading] Handle queued request aborts without allocated KV blocks ( #49146 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-21 11:16:26 +08:00
1940c8441e
[Rust Frontend][gRPC] Add engine-aware health reporting ( #48992 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Connor Carpenter <connorc@nvidia.com >
2026-07-21 10:54:39 +08:00
Simon Mo and GitHub
72d16aee15
[CI] Exercise FA3 FP8 attention on SM90 ( #49231 )
...
Signed-off-by: Simon Mo <simon@inferact.ai >
2026-07-21 10:26:55 +08:00
Kunshang Ji and GitHub
e78a0c8e59
[XPU][Doc] Update XPU docker image documents ( #49148 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-21 10:11:27 +08:00
Chris Leonard and GitHub
97a98006b0
Update qutlass cmake for stable abi ( #47879 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-20 18:31:10 -07:00
0a684ab0c0
[Bugfix] Fix WSL circular import from pin_memory warning_once ( #48444 )
...
Signed-off-by: AlejandroParedesLT <alejandroparedeslatorre@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-20 18:30:55 -07:00
0d9210a502
Fixes non-coalesced HBM access in marlin_int4_fp8_preprocess_kernel_awq ( #47268 )
...
Signed-off-by: xjx <493337577@qq.com >
Signed-off-by: flutist-alibaba <30485581+flutist@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-20 18:30:38 -07:00
1d874867ea
[Misc][Docs] Fix broken protocol link in speech_to_text doc ( #47212 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-21 01:05:05 +00:00
2e2e626b40
[Bugfix] Count per-group blocks in get_max_concurrency_for_kv_cache_config ( #48317 )
...
Signed-off-by: David Orman <ormandj@corenode.com >
Co-authored-by: Luke Alonso <lalonso@gmail.com >
Co-authored-by: Martin Vit <martin@voipmonitor.org >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-21 00:29:10 +00:00
Nick Hill and GitHub
af91f4b3e4
[Cleanup] Remove unused StructuredOutputRequest.status field ( #49235 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-21 00:06:16 +00:00
2396a61108
[Attention][MLA][DCP] Query replication for MLA decode (DeepSeek-V2/R1 + Kimi-K2.5) ( #45964 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-20 23:51:27 +00:00
97a668152b
[RL Infra][FlashInfer] Enable router replay output from FlashInfer monolithic MoE kernel ( #44214 )
...
Signed-off-by: Xuanyu Zhang <xuanyu.zhang@mistral.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-20 16:45:10 -07:00
58b2012aa2
[copy of #45208 ] CuMem slept-L1 fragmentation accounting ( #49208 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: Justin Wood <justin.m.wood@me.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: haosdent <haosdent@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Justin Wood <jwood@me.com >
2026-07-20 23:08:11 +00:00
Ning Xie and GitHub
b7c20d0cfa
[chore] adjust logo be more friendly to white background terminal ( #48938 )
...
Signed-off-by: Andy Xie <andy.xning@gmail.com >
2026-07-20 15:15:28 -07:00
TJian and GitHub
a2b1f9fc3b
[ROCm] [Release] [Bugfix] Fix the per commit wheel release pipeline. ( #49245 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-20 22:12:38 +00:00
642076d26c
Support loading sample_from_anchor flag from speculators config ( #48639 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-07-20 14:52:43 -07:00
Charlie Fu and GitHub
5feb3950e5
[ROCm][CI] fix test_rocm_quick_reduce.py ( #49234 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-07-20 16:39:33 -05:00
4ec199b66a
[Bugfix][Spec-Decode] Populate draft seq_lens_cpu_upper_bound for spec-decode attention metadata ( #44492 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Signed-off-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-20 20:50:58 +00:00
7ca017778f
[Feat][Perf] Add new warmup infrastructure for JITs ( #47451 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 13:21:55 -07:00
fbfe58133d
[Bugfix][KV Offload] Preserve reachable tails for hybrid SWA groups ( #48911 )
...
Signed-off-by: Colton Ottley <colton@ottleyengineering.com >
Co-authored-by: Colton Ottley <colton@ottleyengineering.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-20 22:12:36 +03:00
9dd62d80ab
Cosmos3 FP8 ModelOpt/Diffusers remapping ( #48952 )
...
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com >
Signed-off-by: wkutak <wkutak@nvidia.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-20 11:31:48 -07:00
f878367898
[Revert][Bugfix] Restore MiniCPM-V 4.6 ViT QKV weight loader ( #49193 )
...
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: wjinxu <1299461899@qq.com >
2026-07-20 18:17:46 +00:00
bd091079cb
[Attention] FlashAttention 4 SM100 FP8 kv cache support ( #42569 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 10:53:27 -07:00
b23bd73f54
[XPU]add sycl path for Mhc ( #47245 )
...
Signed-off-by: root <xiaolong.guo@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:32:54 +00:00
Bugen Zhao and GitHub
e2d7adeb64
[Rust Frontend] Bump xgrammar-structural-tag and enable local extension ( #49161 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-20 16:22:24 +01:00
Isotr0py and GitHub
15cb8e140d
[Multimodal] Allow keeping original image mode for ImageIO ( #49159 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-20 13:42:45 +00:00
f007cceb42
[KV Offload] Support self-describing KV events with TieringOffloadingSpec ( #48679 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 16:41:58 +03:00
0a5069e4e3
[Bugfix][Gemma4] Fix ModelOpt mixed-precision MoE config mapping ( #48563 )
...
Signed-off-by: wangqian <601731555@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 06:39:28 -07:00
8ce53a616e
[Bugfix] Zero new KV blocks for quantized + sliding-window hybrid caches ( #47574 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-20 13:18:17 +00:00
Lena Onyshchenko and GitHub
ae10e855ab
[Misc][Docs] Remove duplicate CodeGeex4 row in XPU model table ( #47210 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
2026-07-20 10:05:36 +00:00
hcl and GitHub
530ee36a0d
fix(openai): reject non-numeric logprobs with 400 instead of 500 ( #49144 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
2026-07-20 10:04:50 +00:00
Salt Sato and GitHub
d835ad572c
[Bugfix][Rust Frontend] Map missing prompt logprobs for single-token prompts in chat and raw generate ( #49111 )
...
Signed-off-by: Feathbow <feathbow@gmail.com >
2026-07-20 10:00:06 +00:00
47d0597ca2
[Misc][Docs] Fix broken csrc kernel links in fusions doc ( #47211 )
...
Signed-off-by: oonyshch <xonyshch@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-20 09:44:25 +00:00
Reid and GitHub
818cf61e91
[Rust Frontend] Fix macro-based content format detection ( #49042 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-20 09:39:13 +00:00
c01618fdc8
[Rust][Benchmark] Integrate vllm-bench to vllm-rs & vllm CLI ( #48930 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-20 09:31:25 +00:00
823eaf667d
[XPU] FP8 o_proj with fp8_bmm and load-time scale transpose ( #48334 )
...
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 16:32:03 +08:00
f1f1259692
[Rust Frontend] Use zero-copy slicing for multimodal tensors ( #48781 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-07-20 16:28:25 +08:00
df13b5aef5
[XPU] [MoE] add quant input when prepare for fusedmoe ( #47122 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:47:26 +08:00
4938d44a3b
[CPU] fixes heterogeneous NIXL KV transfer into CPU_ATTN decode workers ( #47871 )
...
Signed-off-by: Spycsh <sihan.chen@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-20 07:33:13 +00:00
37bf988c2f
[XPU][Bugfix] Fix GroupCoordinator device_index ( #47295 )
...
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-20 15:25:56 +08:00
aoshen02 and GitHub
9459fc6471
[Bugfix][RL] Set vLLM config during weight reload ( #45989 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-20 15:02:56 +08:00
5245c80564
[Doc] Document blocks_per_chunk in the KV offloading guide ( #49100 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-07-20 09:48:43 +03:00
9bc266d923
[Bugfix][KV Offload] Propagate EAGLE mode to SimpleCPU coordinator ( #49071 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-20 06:39:11 +00:00
5c9f6557d7
[Hardware][CPU] Enable granite-4 model on cpu ( #47641 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akashkaothalkar@dhcp-9-123-5-76.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akashkaothalkar@Akashs-MBP.lan >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@dhcp-9-123-5-76.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@Akashs-MBP.lan >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-20 06:15:16 +00:00
dcfebf93f4
[Bugfix] Fix logprobs token-string collision from SentencePiece space… ( #48674 )
...
Signed-off-by: Allen Shen <aoshen@inferact.ai >
Co-authored-by: mvanhorn <mvanhorn@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-20 12:17:18 +08:00
752bd10647
[ROCm][CI] Fix sparse MLA metadata sync fixture ( #49128 )
...
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-19 23:02:03 -05:00
Thien Tran and GitHub
2730b657c4
[Bugfix] Fix broken NVVM caused by CuteDSL 4.6.0 ( #49108 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-19 19:45:56 -07:00
1dcbbd9cac
[CI] Move compatible 1xL4 jobs to H200 35GB MIG ( #43024 )
...
Signed-off-by: Simon Mo <simon@inferact.ai >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-19 19:21:25 -07:00
ace9fda495
[CI/Build][BugFix][The Rock][AMD] Add spawn method in vision examples to avoid reinitialization ( #47932 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-19 13:41:52 -05:00
TJian and GitHub
ef0aa7ca2f
[ROCm] [Release] [Per-commit] Reenable per commit rocm wheel ( #49044 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-19 13:38:04 -05:00
Taneem Ibrahim and GitHub
e6d1310b2a
[Bugfix] Reject removed pooling parameters ( #48984 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-07-19 05:18:03 -07:00
yzong-rh and GitHub
ac5f38a0f7
[Refactor] Extract StructuredOutputsParams creation logic from Request.to_sampling_params ( #49003 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-19 05:18:00 -07:00
b6ff8a2f50
[Core] Add MRV2 virtual-batch PCP for MLA ( #46570 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Codex <noreply@openai.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-19 02:53:15 +00:00
9243e0124e
[Multimodal] Automatically fallback to ViT DP when TP is unavailable ( #49046 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-18 14:41:04 -07:00
Andreas Karatzas and GitHub
df362b2d6d
[ROCm][CI] Ensure sliding window tests release GPU memory ( #49055 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-18 20:44:05 +00:00
SYLAR and GitHub
7c2acd38b7
[Bugfix] Qwen3-VL/Qwen-Omni: honor max_pixels/min_pixels for video prompts ( #49015 )
2026-07-18 10:29:11 -07:00
yzong-rh and GitHub
a287eb163f
[Front-end] [Messages] Populate num_cache_creation_tokens ( #48535 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-18 13:04:35 -04:00
frida-andersson and GitHub
e94243893d
[ROCm][DSv3.2][Perf] Cap sparse MLA decode KV-splits with a work-per-split heuristic ( #46832 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-07-18 09:39:37 -07:00
29c0ec4d63
[ci] Move 3 entrypoints tests to h200_35gb queue ( #43164 )
...
Signed-off-by: Simon Mo <simon@inferact.ai >
Signed-off-by: Simon Mo <simon@simon-mac-mini-9.local >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-18 08:43:49 -07:00
Michael Goin and GitHub
c7ce03bcbd
[Bugfix] Bump tml-fa4 for cutlass-dsl 4.6 API compatibility ( #48988 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-18 05:59:33 -07:00
Harry Mellor and GitHub
c233d90aa8
Remove even more unnecessary load_weights methods ( #48496 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-18 08:40:27 +00:00
d96aee0951
[Bugfix] Re-sync parameter tp_rank after process_weights_after_loading (fix replicated / disable_tp weight reload) ( #48025 )
...
Signed-off-by: Alex Xu <alexxu@roblox.com >
Co-authored-by: YQ-Wang <yiqingwang@roblox.com >
Co-authored-by: alexhxu <alex.xu1015@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-18 08:40:06 +00:00
Francesco Fusco and GitHub
c71a583aa9
[Perf][Hybrid] Vectorize _copy_mamba_state_block to uint64 for temporal ( #48110 )
2026-07-18 04:43:09 +00:00
xuebwang-amd and GitHub
f12b80c6ef
[ROCm][Bugfix] Fix GPT-OSS Quark MXFP4 MoE loading - emulation buffer not block-aligned ( #43979 )
...
Signed-off-by: xuebwang-amd <xuebwang@amd.com >
2026-07-18 03:49:41 +00:00
Jee Jee Li and GitHub
da64db78b9
[LoRA] Optimize TrtLlmLoRAExperts ( #48759 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-18 10:26:14 +08:00
425c4eafb0
[Sampler] Stop upcasting logits to fp32 in apply_sampling_params ( #48641 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-17 18:15:20 -07:00
Michael Goin and GitHub
02c01f442b
[Model] Use standard ModelOpt config for Inkling NVFP4 ( #48990 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-17 18:13:14 -07:00
fae543015c
[Frontend]Flatten beam-search beams with itertools.chain instead of sum ( #48829 )
...
Signed-off-by: Wang Xingda <wangxingda1993@126.com >
Co-authored-by: 王兴达 <wangxingda@360itdeMacBook-Pro.local >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-17 23:09:14 +01:00
c9be3a8aa1
[Kernel][Helion] Disable warp specialization in rms_norm_per_block_quant B200 configs ( #48797 )
...
Signed-off-by: Shangdi Yu <shangdiy@meta.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-17 21:39:09 +00:00
41ea2dd44a
[Bugfix][V1/V2] Fix prompt_logprobs to respect logprobs_mode ( #47680 )
...
Signed-off-by: Wojciech Wais <wojciech.wais@gmail.com >
Signed-off-by: Federico Kamelhar <209537060+fede-kamel@users.noreply.github.com >
Signed-off-by: Allen Shen <aoshen@inferact.ai >
Co-authored-by: Wojciech Wais <wojciech.wais@gmail.com >
Co-authored-by: Federico Kamelhar <209537060+fede-kamel@users.noreply.github.com >
2026-07-17 21:58:59 +01:00
088c0be268
[CI] Fix macOS wheel release annotation context ( #48771 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-17 13:44:48 -07:00
fcd2255d16
[Hardware][GPU] Profiler config additional to increase it scope and annotation details ( #37524 )
...
Signed-off-by: devalshahamd <deval.shah@amd.com >
Signed-off-by: Deval Shah <devashah@amd.com >
Signed-off-by: Deval Shah <deval.shah@amd.com >
Co-authored-by: Deval Shah <devashah@amd.com >
2026-07-17 13:38:59 -07:00
Wentao Ye and GitHub
b5433b6f50
[Perf] Optimize dsv4 routing using specialized kernel, 2.94% E2E TPOT improvement ( #48660 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-17 13:35:06 -07:00
cc25f028b7
[Loader] Improve InstantTensor loading ( #46868 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-17 16:30:02 -04:00
c4cd2bd544
[Bugfix] MoRIIO toy P/D proxy: fix DP-rank index aliasing + harden for high-concurrency bursts ( #46115 )
...
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: QinPR <1905873179@qq.com >
Co-authored-by: Peiran Qin <66068739+QinPR@users.noreply.github.com >
2026-07-17 12:35:04 -07:00
5784507da4
[Attention] Allow selecting a different attention backend per KV-cache group ( #48012 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-17 15:19:02 -04:00
labAxiaoming and GitHub
bf578e1abd
[Bugfix][GLM4V] Fix video dummy profiling and memory usage ( #48729 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
2026-07-18 01:44:45 +08:00
Fangzhou Ai and GitHub
efed8a1e83
[ROCm][Perf][DSV4] Improve sparse decode reduction occupancy on gfx950 ( #48788 )
...
Signed-off-by: fai <fangzhouai@gmail.com >
2026-07-17 10:24:59 -07:00
11d291511a
[Bugfix][Tool Parser] Preserve whitespace in parameter values (MiniMax M2, Qwen3, MiniCPM5 XML) ( #48846 )
...
Signed-off-by: mosya415 <263250241+mosya415@users.noreply.github.com >
Signed-off-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
Co-authored-by: mosya415 <263250241+mosya415@users.noreply.github.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-17 16:45:41 +00:00
877dae9c68
[Refactor] Remove deepseek dead code ( #48780 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-17 14:57:13 +00:00
passtoor-agi and GitHub
c4dd6d78fd
Fix: Restore data_parallel_size > 1 for use_sequence_parallel_moe ( #48849 )
...
Signed-off-by: passtoor-agi <305788622+passtoor-agi@users.noreply.github.com >
2026-07-17 10:27:50 -04:00
JooHo Lee and GitHub
ce2aecc4dc
[Performance] Use CuTe-DSL for FlashInfer MXFP4 quantization ( #48417 )
...
Signed-off-by: BWAAEEEK <jooho414@gmail.com >
2026-07-17 06:53:48 -07:00
f38f3d11fb
[Bugfix][KV Offloading] Offload last block at request finish and prevent reuse race ( #48596 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-17 16:50:49 +03:00
d4b4562917
[XPU] Bump vllm_xpu_kernels to v0.1.11.1 ( #48942 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-17 20:43:23 +08:00
7b3192523e
[Bugfix]Fix transformer backend failed: AttributeError: 'Parameter' object has no attribute 'weight_loader' ( #48699 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-17 12:43:44 +01:00
Yejing Lai and GitHub
4c6e2e4b30
[XPU][UT]fix _POSSIBLE_KERNELS error on XPU ( #47516 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-17 11:19:05 +00:00
liuzhenwei and GitHub
8502958810
[XPU] support HND layout ( #47975 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-17 10:54:34 +00:00
ce4bdcbda4
[Bugfix] Enable FlashAttention MLA prefill for Mistral Small 4 head dims ( #48855 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-17 18:07:00 +08:00
liuzhenwei and GitHub
d5b1ec2684
[XPU] allow forcing flash attn for mm_prefix ( #48828 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-17 09:44:18 +00:00
867ff69733
[CI] Gate non-default release wheel builds ( #48772 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-17 02:16:44 -07:00
Sage and GitHub
109b736b86
[docs] preserve page path in stable-docs announcement link ( #48839 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-07-17 08:56:45 +00:00
69d4f5ef63
[Bugfix][Multimodal] Fix Qwen3-Omni use_audio_in_video with mixed image/video inputs ( #46213 )
...
Signed-off-by: wendadawen <wendadawen@qq.com >
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
2026-07-17 08:31:16 +00:00
426d48bfa1
[KV Offload] Add optional tier locality to FS/OBJ KV events ( #48281 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-17 10:31:52 +03:00
26c909ed74
[Model] Support TranslateGemma-12b-it ( #41599 )
...
Signed-off-by: Zhang Jian <jianmusings@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-17 07:17:59 +00:00
fb1d8ccaf5
[rl] Stateful Trainer Send: New Abstractions [1/N] ( #48042 )
...
Signed-off-by: haoaaron <ahao@anyscale.com >
Signed-off-by: Aaron Hao <ahao@anyscale.com >
Co-authored-by: Sumanth R Hegde <39546518+SumanthRH@users.noreply.github.com >
2026-07-17 15:11:06 +08:00
9354f22204
[Rust][Benchmark] Port in vllm-bench ( #48107 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-17 14:25:29 +08:00
aoshen02 and GitHub
17fdd42100
[Bugfix][Attention] Preserve post-load tensors across weight reloads ( #48251 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-17 14:15:26 +08:00
472d330c21
Add blocks_per_chunk configuration for KV offloading to support heterogeneous KV cache groups ( #48878 )
...
Signed-off-by: Debasish-87 <22btics06@suiit.ac.in >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-17 09:00:12 +03:00
3b6c96a101
[Bugfix][Pooling] Fix wrong scores for chunked prefill under torch.compile ( #48901 )
...
Signed-off-by: seewoo <seewoo@ucsc.edu >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-17 05:19:12 +00:00
Martin Hickey and GitHub
4d4e04f452
[Render] Add round trip parity test and docs for derender ( #48617 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-07-17 05:02:46 +00:00
Micah Williamson and GitHub
67fe73b2b4
[CI] Extend max-model-len for test_parsable_context to allow reasoning to finish ( #48873 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-17 11:36:44 +08:00
+1
ee8f36d0b3
[Warmup] Show CuTeDSL compilation progress ( #48881 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 20:17:51 -07:00
+1
f3e9497e92
[Model] Add Inkling LoRA support [4/N] ( #48884 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-17 09:52:15 +08:00
Thien Tran and GitHub
fe784ff22e
[M3] Improve indexer for long-context decode (sm100) ( #48582 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-16 18:12:15 -07:00
Daoyuan Li and GitHub
b88abb5036
[Misc] Remove orphaned env vars and stale env-var references ( #44749 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-07-17 00:00:47 +00:00
67f9046e4a
[Bugfix] Sparse MLA: enable fp8_ds_mla dense prefill ( #48642 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 22:44:04 +00:00
f17be06fbe
[Perf] Optimize clamp to clamp_ ( #48143 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 18:41:07 -04:00
2cab53ddee
[Model][Hardware][AMD]: Part 1/2 -> Enable e2e QK Norm + RoPE + KV Cache runtime fusion for Qwen3-30B-A3B on ROCM_AITER_FA, and ROCM_AITER_UNIFIED_ATTN ( #42749 )
...
Signed-off-by: Jack Hu <Jack.Hu@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-07-16 17:39:04 -05:00
ab0a20d151
[Docs] Add Phi-3.5-mini-instruct to batch invariance tested models ( #46396 )
...
Signed-off-by: Yuval Luria <yuvalluria@users.noreply.github.com >
Co-authored-by: Yuval Luria <yuvalluria@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-07-16 18:05:58 -04:00
4a394bfcda
[Spec Decode][DSpark] Add Gemma4-12B DSpark draft model ( #47216 )
...
Signed-off-by: DiegoCao <DiegoCao@users.noreply.github.com >
Co-authored-by: DiegoCao <DiegoCao@users.noreply.github.com >
2026-07-16 21:51:47 +00:00
Michael Goin and GitHub
c95c663049
[Quant] Add nvfp4_per_token online MoE quantization ( #48538 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-16 14:25:27 -07:00
HDCharles and GitHub
ab3c1aedf3
[Bugfix] Fix activation quantization dispatch for WNA4Int/WNA8Int ( #48785 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-07-16 17:13:02 -04:00
+1
fb5ec0dc9e
[Model] Add Inkling MTP=1 support [3/N] ( #48869 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 13:27:21 -07:00
971dac2caa
[Bugfix][KV-transfer] MoRIIO: retry RDMA send-queue-full backpressure instead of failing the read ( #47495 )
...
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io >
Signed-off-by: Edwin Lim <edwinlim0919@gmail.com >
Signed-off-by: harishk-mangoboost <harish.kambhampaty@mangoboost.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: harishk-mangoboost <harish.kambhampaty@mangoboost.io >
2026-07-16 20:02:27 +00:00
Shangdi Yu and GitHub
efa2e424f6
[Helion] Fix degenerate scale_ub in kernel input generators ( #48868 )
...
Signed-off-by: Shangdi Yu <shangdiy@meta.com >
2026-07-16 19:57:45 +00:00
02bf9c7907
Fix Quark mxfp4 quantized model loading issue under mtp ( #46757 )
...
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-16 14:58:47 -04:00
Woosuk Kwon and GitHub
f61163e6c7
[Model] Add Hopper FA4 relative attention for Inkling ( #48858 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-07-16 11:26:52 -07:00
Wentao Ye and GitHub
626c90b2d5
[Refactor] Move fla to third party ( #48500 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-16 19:22:36 +01:00
+1
251f7e478e
[Model] Add PW CUDA graph support for Inkling [2/N] ( #48822 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-16 10:28:56 -07:00
ce65385618
[KV Offload] Split tiering_lookup_delay into sync/async histograms ( #47679 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-07-16 20:03:45 +03:00
music-dino and GitHub
7d56fe2adc
[ROCm][CI] Avoid HIP init at config time via lazy aiter import in Quark OCP-MX ( #48015 )
...
Signed-off-by: Dino Music <Dino.Music@amd.com >
2026-07-16 16:47:08 +00:00
Zhongdongming Dai and GitHub
75bdad40b5
[Bug][Quantization] Fix humming is_layer_skipped for compressed-tensors "re:" ignore entries ( #48507 )
...
Signed-off-by: Zhongdongming Dai <zhongdongmin@nvidia.com >
2026-07-16 07:47:42 -07:00
d08eebad16
[Perf][MoE] Write FlashInfer combine into final output ( #47156 )
...
Signed-off-by: snordmann <snordmann@nvidia.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-16 17:21:03 +03:00
wang.yuqi and GitHub
3e90d015ba
[Frontend] Overlap preprocessing and computation for pooling models offline inference ( #47699 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-16 14:00:20 +00:00
7cd1d57b74
[CI/Build][Docker] Bump nvidia-cutlass-dsl to 4.6.0 and drop packaging workarounds ( #47442 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-16 13:51:41 +00:00
b8168e33e0
[ROCm][Perf][DSV4] Enable split sparse decode on gfx942 ( #46275 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-16 12:28:20 +00:00
ovidiusm and GitHub
d803b44dbe
[NIXL] Bump nixl to 1.3.1 ( #47559 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
2026-07-16 14:24:33 +02:00
530852f959
[KV Connector] Fix PD async scheduling race condition for hybrid attn models ( #48481 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: llx-08 <2596671364@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-16 11:42:37 +01:00
Nicolò Lucchesi and GitHub
a317bc5739
[Misc][Nixl] Unify _logical_to_remote_kernel_block_ids ( #48717 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-16 18:36:25 +08:00
a9531edfa6
[KV Offload] Define clean backend configuration boundary ( #48150 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Signed-off-by: Chang Guo <cguo51@asu.edu >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-16 13:27:05 +03:00
Reid and GitHub
8c3393f373
[Bugfix][Rust Frontend] Limit chat top_logprobs in responses ( #48134 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-16 09:41:44 +00:00
Ilia Yastrebov and GitHub
9f8cbfd8eb
Vectorize prep xfer list creation ( #48209 )
...
Signed-off-by: Ilia Yastrebov <iyastrebov@nvidia.com >
2026-07-16 11:39:05 +02:00
ea1d65fe6d
[Rust Frontend] Add Seed-OSS tool parser ( #47741 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-07-16 17:28:02 +08:00
f44f3d6f79
[Rust Frontend] Wait for mock engine endpoints before ZMQ connect ( #47965 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-16 09:20:31 +00:00
cc706b05a5
[Bugfix][Rust Frontend] Detokenizer: avoid leaking prompt on zero-generated-token completions ( #47707 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: xiaguan <751080330@qq.com >
2026-07-16 09:11:21 +00:00
Thien Tran and GitHub
85e296950c
BF16x3 router GEMM ( #47973 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-16 17:04:20 +08:00
Reid and GitHub
dc9f845ddc
[Rust Frontend] Fix mock engine test shutdown race ( #48738 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-16 07:27:43 +00:00
Elvir Crnčević and GitHub
12f2c515a7
[Bugfix] Fix offloading set_ overflow for packed non-uniform KV caches ( #48530 )
...
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
2026-07-16 10:04:39 +03:00
+1
6570c9800c
[Model] Add Inkling model support [1/N] ( #48799 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <32987265+TheEpicDolphin@users.noreply.github.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-15 23:40:07 -07:00
8bfd683901
[Spec Decode] Add kv_cache_dtype to speculative_config to control separately from target ( #48787 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-15 23:49:23 -06:00
Micah Williamson and GitHub
7dc2698632
[ROCm][CI] Set "highest" matmul precision for reference hf_runner in test_bert_for_masked_lm ( #48784 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-16 05:49:15 +00:00
ErenAta16 and GitHub
59b964f37d
fix(lora): validate LoRA rank is positive in PEFTHelper ( #48437 )
...
Signed-off-by: ErenAta16 <erena6466@gmail.com >
2026-07-16 05:22:52 +00:00
6a9f24aa8c
[ROCm][CI] Fix cuda graph mem profile issue ( #48764 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-16 04:23:18 +00:00
ba47bb5be1
Bump flashinfer version to 0.6.14 ( #47669 )
...
Signed-off-by: AmeenP <ameenp360@gmail.com >
Signed-off-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Pavani Majety <pmajety@nvidia.com >
2026-07-15 21:00:40 -07:00
df8a0900df
[BugFix] Don't apply weight in batch-invariant RMSNorm when has_weight=False ( #48741 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: Michael Gokhman <michael.gokhman@yahoo.com >
2026-07-16 11:59:03 +08:00
2db39c7049
[Bugfix][Spec Decode] Fix eagle3 first-layer qkv_proj prefix for quantized drafts ( #48068 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-16 11:34:41 +08:00
rongfu.leng and GitHub
3935829f89
[Docs] fix error key name ( #48802 )
...
Signed-off-by: rongfu.leng <lenronfu@gmail.com >
2026-07-16 03:16:16 +00:00
qli88 and GitHub
7746961277
[CI] Fix flaky lora test ( #47375 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Signed-off-by: qli88 <qiang.li2@amd.com >
2026-07-16 02:23:03 +00:00
qli88 and GitHub
5de1add806
[feature]Add int4 quantization support for emulation moe backend ( #48451 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-07-16 02:00:12 +00:00
Mike G and GitHub
915dffaa5f
[Attention] Mirror Triton KV dtype checks in MLA ( #47060 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-07-16 01:54:52 +00:00
nemanjaudovic and GitHub
81e13a0591
[Compilation] Skip x.size(dim) in _decompose_size_nodes ( #42543 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-07-15 17:59:14 -07:00
BRIJ RAJ KISHORE and GitHub
f95e3f0edb
[Tests] Gate Step3VL under Transformers v5 ( #44349 )
...
Signed-off-by: brijrajk <22271048+brijrajk@users.noreply.github.com >
2026-07-15 17:59:10 -07:00
5a65ba5f17
[Refactor] Move iteration logging to the frontend ( #46647 )
...
Signed-off-by: maxyanghu <hyoung2991@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Shang Wang <shangw@nvidia.com >
2026-07-15 17:59:05 -07:00
9d1c695be5
[XPU] Add DSpark speculative decoding support for DeepSeek-V4 ( #47677 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 17:59:02 -07:00
3c1bc1fc0d
[ROCm][Perf] Optimize sparse attention prefill kernel for DeepSeek-V4 ( #48519 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-07-15 17:58:59 -07:00
Michael Goin and GitHub
3a5e88e629
[Bugfix] Fix local speculators with dots in the name from classifying as custom_class ( #48754 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-15 17:58:06 -07:00
0becb7486b
[BugFix][MLA] Support kv_cache_dtype_skip_layers for MLA attention ( #47309 )
...
Signed-off-by: liuruikang <liuruikang.cs@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-16 00:06:11 +00:00
Wentao Ye and GitHub
2dab187f75
[Perf] Optimize fused_topk_bias for DSv4, 1.5~2x kernel performance improvement ( #47463 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-15 23:40:55 +00:00
Giuseppe Grossi and GitHub
015b0320de
Add giuseppegrossi to rocm label auto cc action ( #48643 )
...
Signed-off-by: giuseppegrossi <ggrossi@amd.com >
2026-07-15 16:38:56 -07:00
4238b011a7
[Feature] Migrate moe sp support to non-torch compiled path for GLM5.2 ( #47881 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 23:33:15 +00:00
kliuae and GitHub
eb33ff34dd
[ROCm][Perf] DSv4 two-stage compressor kernel for HCA prefill ( #47718 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
2026-07-15 23:31:51 +00:00
Mike G and GitHub
2bd8957627
[Bugfix][NVFP4 MoE] Pad gated intermediate to 64 for FlashInfer TRT-LLM shuffle (M%128) ( #46880 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-07-15 17:37:15 -04:00
Nicolò Lucchesi and GitHub
3034c8d389
[CI][PD] Add optional/nightly DSv4 Disaggregated eval ( #42310 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-15 21:04:54 +00:00
ecf4aa5ce2
[Bugfix] Fix FlashInfer non-causal draft attention (DFlash/DSpark) on Blackwell ( #48167 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-15 12:44:01 -07:00
49e777cf08
[CI][ROCm] Retry failed Docker build steps once ( #48773 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 14:31:23 -05:00
b7950e798f
[Bugfix] Initialize draft CUDA-graph keys for the native draft_model proposer ( #47460 )
...
Signed-off-by: Alagappan Valliappan <avalliappan@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 15:09:00 -04:00
de100ffb62
[Docs] Document pooling config resolution ( #48497 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 14:24:16 -04:00
Sage and GitHub
43cd340247
[Fix] Align OpenAI vllm_xargs value types across request schemas ( #48252 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
Signed-off-by: Sage <80211083+sagearc@users.noreply.github.com >
2026-07-15 17:48:24 +00:00
1d99f0f421
[ROCm][BugFix] Triton W4A16 handling for GPTQ/AutoGPTQ qzeros layout ( #47770 )
...
Signed-off-by: giuseppegrossi <ggrossi@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 11:55:47 -05:00
Andreas Karatzas and GitHub
0885b51981
[CI][ROCm] Stabilize ci_base hash calculation and image handoff ( #48746 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 10:56:56 -05:00
Xiaohong (Sean) Chen and GitHub
6036bf110a
[Kernel][Helion] Add Helion kernel benchmark script ( #48512 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-07-15 15:43:06 +00:00
Xiaohong (Sean) Chen and GitHub
2fa63e0fff
[Kernel][Helion] Helion kernel lazy registration ( #48264 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-07-15 15:42:46 +00:00
61141ed265
[Hardware][XPU] Register batch-invariant kernels for XPU ( #41934 )
...
Signed-off-by: tzielinski-habana <tomasz.zielinski@intel.com >
Signed-off-by: Tomasz Zielinski <85164140+tzielinski-habana@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Chendi.Xue <chendi.xue@intel.com >
2026-07-15 11:19:44 -04:00
05eed72aec
[ROCm] Re-enable cudagraph memory profiling, captured on the current stream ( #48526 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 10:03:48 -05:00
Gopala-Krishna Char and GitHub
5810e884f1
[Model] Add RobertaForTokenClassification / XLMRobertaForTokenClassification ( #47991 )
...
Signed-off-by: krishy91 <crgkc.r@gmail.com >
2026-07-15 14:30:33 +00:00
615834ee58
[KVOffload][P2P] Well-known default host/port env vars and per-DP-rank control port ( #47636 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 15:22:56 +03:00
Chaojun Zhang and GitHub
5811ed6a05
[Test][kv_offload] Fix flaky drain() helper in test_fs_tier.py ( #48545 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-15 14:53:47 +03:00
Tahsin Tunan and GitHub
1b30ae4ca4
[Rust Frontend] Fix flaky tls_handshake_timeout_drops_silent_client test ( #47873 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-15 11:05:38 +00:00
Tahsin Tunan and GitHub
4e04bcbce6
[Rust Frontend] Tolerate whitespace before the outer brace in JSON tool-call parsers ( #48034 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-15 11:03:37 +00:00
Nicolò Lucchesi and GitHub
66b6c684ab
[PD][Bugfix] Fix validation of cache shape for attn backends enforcing different kernel_block_size ( #48125 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-15 18:26:02 +08:00
c0302d9497
[Bugfix] Fix parallel_tool_calls=null crash in Responses API from_request() ( #48098 )
...
Signed-off-by: mahadrehmann <mahadrehman04@gmail.com >
Signed-off-by: Mahad Rehman <114791389+mahadrehmann@users.noreply.github.com >
Co-authored-by: muhammadfawaz1 <135441198+professorsab@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-07-15 18:01:17 +08:00
Jee Jee Li and GitHub
313fae3e89
[Bugfix] Fix GLM5 config ( #48711 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-15 09:55:39 +00:00
7aab6e2684
[ROCm][Bugfix] Enable the fp32 head_dtype torch.mm fast path on ROCm ( #48688 )
...
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 08:18:22 +00:00
9dd2e72828
fix flaky multi example connector consistency ( #48206 )
...
Signed-off-by: aarushjain29 <aarushi.jain2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 09:20:34 +02:00
Giuseppe Grossi and GitHub
d119beb1b9
[ROCm] Add tuned selective_state_update config for AMD MI350 ( #48159 )
...
Signed-off-by: Giuseppe Grossi <ggrossi@amd.com >
2026-07-15 10:18:09 +03:00
12a8057bfe
[CI/Build] Split release artifact annotations by type ( #48600 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-07-15 00:00:52 -07:00
e281ac663a
[Rust Frontend] Integrate MM audio support ( #48554 )
...
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-15 15:00:17 +08:00
adce068118
[ROCm][CI] fix test_common.py ( #48676 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 06:42:06 +00:00
b6770d7b54
[ROCm] Run init test engine in-process to avoid KV-cache OOM ( #48527 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 06:39:31 +00:00
3b39fd284a
[Bugfix][Spec Decode] Support heterogeneous QK fusion geometry ( #48671 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-14 22:37:10 -07:00
6472131298
[Bugfix] Set kv_quant_mode on the generic MLA KV-cache spec ( #48379 )
...
Signed-off-by: Mikhail Kostryukov <mike@triptrack.net >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-15 03:36:52 +00:00
37aa52821d
Build with ABI stable FlashMLA ( #48174 )
...
Signed-off-by: Jane Xu <janeyx@meta.com >
Signed-off-by: Shengqi Chen <i@harrychen.xyz >
Co-authored-by: Shengqi Chen <i@harrychen.xyz >
2026-07-14 20:29:28 -07:00
96d2ceda4b
[Security] Replace diskcache to eliminate pickle deserialization ( #44549 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-14 20:29:24 -07:00
Jee Jee Li and GitHub
fdf2cf66d3
[LoRA][1/N] Integrate flashinfer MoE LoRA for BF16 model ( #48632 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-15 10:54:00 +08:00
HDCharles and GitHub
9b2be4e9a5
[Quant] Enable humming w[2-7]a[4,8] inference with compressed-tensors ( #46390 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-07-14 20:22:31 -06:00
Andreas Karatzas and GitHub
3ad85e0de4
[CI][AMD] Configure MI300 tests for native execution without DinD ( #48387 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 02:14:16 +00:00
4f7fffb92f
[Core][LoRA] Support fp32 lm_head (head_dtype) on the LoRA path ( #48525 )
...
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-15 09:51:09 +08:00
6e073440b1
[ROCm][CI] Remove mxfp4 test skips after amd-quark 0.12 release ( #47330 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-15 01:25:05 +00:00
gnovack and GitHub
f7aadae5e5
add pad-aware reduce path ( #48385 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-14 18:05:50 -07:00
442c421e79
[Perf] Remove redundant repeat and copy for dsv4, 1.8% E2E TPOT improvement. ( #48137 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-15 00:48:10 +00:00
0bd6b85a1f
[Bugfix] Preserve unloaded non-persistent buffers during layerwise reload ( #44371 )
...
Signed-off-by: Joan Velja <joan.velja22@gmail.com >
Co-authored-by: Dakai An <77474977+andakai@users.noreply.github.com >
2026-07-14 17:46:29 -07:00
aoshen02 and GitHub
3ca242d1b6
[Bugfix][R3] Exclude draft routers from expert capture ( #48622 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-07-14 17:45:35 -07:00
Joe Rowell and GitHub
7e950521b3
fix: size FlashInfer prefill workspace to batch head footprint ( #48428 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-07-14 17:18:41 -07:00
Micah Williamson and GitHub
0f0f28b537
[Bugfix][CI] Fix test_head_dtype quant_method test on ROCm ( #48654 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-14 18:31:36 -05:00
520a20ba4e
[Bugfix] MoRIIO toy P/D proxy: add /health ( #45222 )
...
Signed-off-by: Chaemin Lim <chaemin.lim@mangoboost.io >
Signed-off-by: Edwin Lim <edwinlim0919@gmail.com >
Co-authored-by: Edwin Lim <edwin.lim@mangoboost.io >
Co-authored-by: Jaeyoun Kim <jaeyoun.kim@mangoboost.io >
Co-authored-by: Edwin Lim <edwinlim0919@gmail.com >
2026-07-14 22:56:46 +00:00
9182e86971
Log fully resolved pooling config at startup ( #48030 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-14 22:00:46 +00:00
Matthew Bonanni and GitHub
313d01f507
[CI][Bugfix] Fix FlashAttention reported MLA dimension support ( #48631 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-14 21:33:02 +00:00
Divakar Verma and GitHub
05d4f8bba3
[ROCm][CI] fix flashinfer import check ( #48647 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-07-14 20:54:19 +00:00
Michael Goin and GitHub
0b54201a04
[CI] Build macOS arm64 CPU wheel natively on the macmini queue ( #48289 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-14 19:40:26 +00:00
32e632dfeb
[Reasoning] Optimize TPOT for thinking budget when used with speculative decoding ( #46662 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-14 18:55:40 +00:00
7ffb98e248
[ROCm] Retune MI355 selective_state_update float32 config on the unified effective_batch grid ( #48373 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-14 18:26:35 +00:00
cdaa40d2a8
[KV Offload] Split cpu_cache_usage_perc into write/read usage gauges ( #47666 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-14 20:13:41 +03:00
ca3618bc69
[Doc] Sync four function docstrings with their signatures ( #45437 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-14 13:10:13 -04:00
Michael Goin and GitHub
b2f7d2560a
[Bugfix] Make MLA+SWA check the layer's backend, not the model config ( #48520 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-14 09:53:34 -07:00
Wentao Ye and GitHub
1ff9429655
[CI Bug] Fully solve accuracy issue for DSv3.2 + MTP + Sequence Parallel ( #48036 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-14 10:00:24 -04:00
af453e5647
[Bugfix] Gemma4 parser: classify channel-less output consistently in streaming and non-streaming ( #48262 )
...
Signed-off-by: Adhithya Balakrishnan <adhithya.b2004@gmail.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-14 09:30:16 -04:00
32aef44388
[Bugfix] Include inline per-token-head scales in offloaded page transfer width ( #48411 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-07-14 16:07:26 +03:00
7a74a9662b
[NIXL] Avoid reading expired blocks in bidirectional turn-2 read ( #47021 )
...
Signed-off-by: Tomer Gilad <tgilad@nvidia.com >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-14 13:03:41 +00:00
karthik and GitHub
b6754f536e
[Model] Enable LoRA support for tower and connector in LlavaNextVideo ( #48594 )
...
Signed-off-by: gangula-karthik <gkarthik923@gmail.com >
2026-07-14 20:09:38 +08:00
Juan Pérez de Algaba and GitHub
793cf79c89
[Bugfix][Security] Fix concurrent sparse invariant race bypassing CVE remediation ( #48583 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-14 11:08:24 +00:00
50ac1c7bab
[Misc] Rename VLLM_TRITON_ATTN_USE_TD to VLLM_TRITON_USE_TD ( #45781 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-14 10:32:57 +00:00
f04d3f640e
[Test] Enable KV cache events for HMA models in CPU offloading test ( #47754 )
...
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-14 12:22:27 +03:00
xiangdong and GitHub
0a9396a25e
[XPU][CI] Add tests/v1/e2e/general/test_correctness_sliding_window.py in Intel GPU CI ( #47231 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Signed-off-by: xiangdong <40376367+zxd1997066@users.noreply.github.com >
2026-07-14 08:50:16 +00:00
038ec293b1
[Bugfix] Return 400 instead of 500 when multimodal data is sent to a text-only model ( #48473 )
...
Signed-off-by: Hoang Nguyen Tien <hoang.nguyentien.2601@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-14 08:15:43 +00:00
894ebb27f5
Add Cosmos3 Edge Reasoner model ( #48291 )
...
Signed-off-by: Bartosz Stefaniak <bstefaniak@nvidia.com >
Co-authored-by: Bartosz Stefaniak <bstefaniak@nvidia.com >
2026-07-14 08:14:50 +00:00
Juan Pérez de Algaba and GitHub
c9a788eedc
fix(security): guard lm-format-enforcer regex compile with timeout ( #47595 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-14 07:18:11 +00:00
0762f2afeb
[Perf][Feat] Add generic cuteDSL LL BF16 router (GEMM) ( #42562 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-13 23:01:21 -07:00
31be872f55
[ROCm] Retune MI355 selective_state_update float16 config on the unified effective_batch grid ( #48372 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-14 05:16:29 +00:00
wangxiyuan and GitHub
94c0ef3001
[Misc] Clean up "swap_space" ( #48549 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-07-14 04:43:45 +00:00
Matt Woodson and GitHub
af1f036a70
[Bugfix] Skip minimax_m3 tool parser tests when Rust extension is absent ( #48523 )
...
Signed-off-by: Matt Woodson <mwoodson@redhat.com >
2026-07-14 04:43:22 +00:00
95aab66e95
[ROCm][MiniMax-M3][Spec Decode] Support speculative decode with AITER sparse PA ( #47984 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-07-14 04:12:53 +00:00
nemanjaudovic and GitHub
dcf4072da9
[Perf][ROCm] Fix GDN KKT warmup regression on RDNA by avoiding fp32 tl.dot ( #45000 )
...
Signed-off-by: Saeid Rostami <srostami@amd.com >
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-07-13 20:48:54 -07:00
382bbd5144
[ROCm][Kernel] Add HybridW4A16LinearKernel: Triton prefill + HIP skinny decode ( #40977 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-13 20:22:00 -07:00
b50ef9c6ed
[ROCm][MiniMax-M2] Dispatch fused QK-norm + AllReduce via AITER ( #44849 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: Pawel Kowalski <pawel.kowalski@amd.com >
2026-07-14 03:10:16 +00:00
Dan Blanaru and GitHub
9e289c553c
up FI fp8 moe topk to 32 ( #44462 )
2026-07-14 02:58:16 +00:00
c4f5cd60da
[1/N] Add dense MHA path for sparse MLA short sequences ( #47327 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-14 00:29:56 +00:00
0b0ef8d7eb
[Quantization][INC][ARK] Support INT2 XPU WOQ Linear ( #47521 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-14 08:29:45 +08:00
21472f32ea
add pad-aware swiglu limit kernel ( #48287 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 16:48:19 -07:00
fec64fea75
[BugFix] Correct OTEL span start time for Dynamo compilation ( #40698 )
...
Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-07-13 16:25:03 -07:00
8b8af2caf7
[Frontend] Expose logprob_token_ids on Python OpenAI endpoints ( #43463 )
...
Signed-off-by: Lang Zhao <lang.zhao@galileo.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-13 14:40:21 -07:00
Snehlata and GitHub
7738ef35b8
[Feat] Add Support for BertForMaskedLM to vLLM ( #48463 )
...
Signed-off-by: atalhens <sneh.lata@nutanix.com >
2026-07-13 20:56:25 +00:00
9a21f0d1a3
[BugFix] Initialize model_config for Qwen3-VL MoE ( #44863 )
...
Signed-off-by: wenpengw-nv <wenpengw@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-13 13:43:53 -07:00
Nick Hill and GitHub
8ac8375270
[Core] Preserve Marconi caching with selective hybrid cache retention ( #47782 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-13 21:24:20 +01:00
shanjiaz and GitHub
7dc447dda7
Added sliding window attention support for qwen-eagle3 architecture ( #47568 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-07-13 20:20:44 +00:00
7fc97042c3
Add DCP + Eagle support for Tokenspeed MLA backends ( #48180 )
...
Signed-off-by: Pavani Majety <pmajety@nvidia.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-13 11:46:02 -07:00
Micah Williamson and GitHub
18c4067a54
[ROCm][CI] Unblock AMD: Language Models Test (Extended Pooling) ( #48513 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-13 18:38:10 +00:00
550218b136
[Bugfix][Frontend] Flush engine reasoning parser at engine-reasoning → tool streaming boundary ( #47606 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com >
2026-07-13 14:06:10 -04:00
Gavin Morris and GitHub
5c342876a6
[Doc] Add DeepseekV32ForCausalLM to supported_models.md ( #48293 )
...
Signed-off-by: Gavin Morris <gmorriscs@gmail.com >
2026-07-13 17:43:59 +00:00
9427c45386
[ROCm][CI] Transformers: pass only one of input_ids/inputs_embeds ( #48258 )
...
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-13 17:28:50 +00:00
43c8cbf79b
[EC Connector] CPU Offloading EC Connector ( #47423 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-13 20:09:41 +03:00
62286308c9
[Misc] Improve Matryoshka pooling dimensions validation ( #48057 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 12:57:36 -04:00
Nick Hill and GitHub
26587f9519
[BugFix][ModelRunner V2] Fix stale attn metadata in speculator prefill cudagraph capture ( #48261 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-13 09:39:15 -07:00
93e3bc8f30
[XPU][CI]Adjust timeout_in_minutes in Intel GPU CI ( #48418 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-13 23:11:16 +08:00
Yan Ma and GitHub
c2c9f7c5e2
remove force channels_last in Idefics3MultiModalProcessor ( #48467 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-07-13 14:18:57 +00:00
Omer Ullman Argov and GitHub
1be6e937b2
lower memory required for capturing cudagraphs for large cudagraph sizes ( #48483 )
...
Signed-off-by: Omer Ullman Argov <118735753+omera-nv@users.noreply.github.com >
2026-07-13 10:14:25 -04:00
Wentao Ye and GitHub
b3cfca996c
[Mypy Fix] Split mypy work ( #48490 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-13 12:42:42 +00:00
Bugen Zhao and GitHub
487dfb3418
[CI] Add SPDX license header to Rust/Protobuf sources ( #48472 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-13 10:22:47 +01:00
107a03ba63
[Core] Support fp32 lm_head for generation models via head_dtype (RFC #48305 §3.6) ( #48390 )
...
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-13 16:43:34 +08:00
56a357ed33
[Bugfix][KV Cache] Don't route uniform-page-size MLA+SWA models into DeepseekV4 packing ( #48256 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-13 08:16:24 +00:00
bea70c7cfc
[Attention] Make sliding-window support an explicit backend capability ( #48011 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-13 01:07:56 -07:00
Mohammad Miadh Angkad and GitHub
75fe92a316
[Distributed][Perf] Enable FlashInfer MNNVL allreduce RMS quant fusion ( #48064 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-07-13 15:02:59 +08:00
b7b58d1eba
[ROCm][CI] Cache Rust builds by source inputs ( #46527 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-07-13 01:14:08 -05:00
Canlin Guo and GitHub
36484e464a
[BugFix] Restore full tokens for Qwen MTP When MoE SP ( #48429 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-07-13 13:29:41 +08:00
9e57de7197
[CPU] Create Proper Numa topology for s390x ( #40714 )
...
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-13 12:58:43 +08:00
Yejing Lai and GitHub
8c5dafcd09
[Bugfix][UT]Fix EagleMiniCPMForCausalLM meet TypeError ( #48452 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-13 04:37:23 +00:00
05fa8183a6
[CPU][Spec Decode] Support DFlash speculative decoding for GDN models on CPU ( #46090 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-07-13 04:16:18 +00:00
d973cce3ca
Re-disable CUDA graph memory profiling on ROCm ( #48440 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-13 03:59:20 +00:00
775c1589ea
[Bugfix][ROCm] Keep TP all_gather on base-class collective ( #48446 )
...
Signed-off-by: fai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-13 03:53:53 +00:00
zzt and GitHub
2595d5cebc
[Model] Optimize Qwen3.5 on H20 ( #48350 )
...
Signed-off-by: zzt <zengzetang.zzt@antgroup.com >
2026-07-13 03:30:48 +00:00
ee5a89f4d7
[ROCm][MiniMax-M3] Add AITER sparse paged attention ( #47287 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-12 19:27:29 -07:00
e26264f3ef
[Kernel] Implement CUDA kernel for ReLUSquaredActivation (relu^2) ( #39058 )
...
Signed-off-by: Tanish Malekar <tanishmalekar32@gmail.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-12 19:18:03 -07:00
AlexHuang and GitHub
4c81772e8b
[Bugfix][KV Offloading] Fix stale transfer_jobs after reset_cache + harden job completion ( #48102 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
2026-07-12 20:00:04 +03:00
Bugen Zhao and GitHub
27c3e579f0
[CI][Rust Frontend] Pin cargo tool versions ( #48222 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-12 16:34:26 +01:00
8df14cfc8c
[EC Connector] Add EC Transfer Params ( #42433 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-12 14:35:33 +03:00
Jiangyun Zhu and GitHub
370b678a02
[CI][2/N] reduce CI time ( #48394 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-12 04:16:55 -07:00
5c0c987c03
Make tiering offload region DP-replica aware ( #47987 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-12 13:10:21 +03:00
Hugo Centeno and GitHub
5f8e73cb8b
[Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions ( #48330 )
...
Signed-off-by: hcenteno <hugo.centeno@estudiantat.upc.edu >
2026-07-12 09:39:27 +00:00
83762b77b0
[Frontend] Add /abort_requests to the RLHF dev API router ( #47173 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-12 14:21:02 +08:00
a02984ed47
[Perf][Qwen] Replace MOE all-reduce with reduce-scatter ( #47006 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: yewentao256 <zhyanwentao@126.com >
2026-07-12 06:14:49 +00:00
fc1c548093
Runtime Draft Weight Update for Speculative Decoding ( #46725 )
...
Signed-off-by: vx120 <893600387@qq.com >
Signed-off-by: vx120 <57470515+vx120@users.noreply.github.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: crp0128 <191679376@qq.com >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-11 22:51:53 -07:00
481e481be7
[2/N][Core] support partial prefix cache hit for hybrid model ( #46384 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-12 05:37:51 +00:00
zhao, zhenhui and GitHub
8e981630c9
[CI][CPU] Add Qwen2-VL multimodal tests for CPU backend and fix incompatibilities ( #48072 )
...
Signed-off-by: Zhenhui Zhao <zhenhui.zhao@intel.com >
2026-07-12 12:30:34 +08:00
Alejandro Paredes La Torre and GitHub
9a48eef89a
[Bugfix][LoRA] Support ark_linear base layer in _get_lora_device ( #47690 )
...
Signed-off-by: AlejandroParedesLT <alejandroparedeslatorre@gmail.com >
2026-07-12 00:13:50 +00:00
Jiangyun Zhu and GitHub
1ef1c7ebba
[CI] split tests to reduce CI time ( #48219 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-11 13:00:14 -07:00
54503ecec0
fix(processor): route MiMo-V2-Omni media fetch through MediaConnector ( #43117 )
...
Signed-off-by: Ievgen Bondarenko <ibondarenko@student.sierracollege.edu >
Signed-off-by: Ievgen (Jack) Bondarenko <ibondarenko@student.sierracollege.edu >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-11 15:52:52 +00:00
ErenAta16 and GitHub
0067311536
fix(entrypoints): stop resolve_items leaking in-flight media fetch tasks on partial failure ( #48333 )
...
Signed-off-by: ErenAta16 <erena6466@gmail.com >
2026-07-11 15:42:08 +00:00
51878e5b6e
[2/N][KV-Cache Layout Refactor] Pack K/V into the content dim across attention backends ( #44455 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-07-11 11:11:16 -04:00
Yejing Lai and GitHub
76fedaa2a5
[XPU][UT]Fix InternS1ProForConditionalGeneration AssertionError ( #48232 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-07-11 13:56:55 +00:00
19069bcbd5
FP32 router GEMV optimization ( #48335 )
...
Signed-off-by: peiyuanz <peiyuanz@inferact.ai >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: peiyuanz <peiyuanz@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: zhouzhou <zhouzhou@zhouzhoudeMacBook-Pro.local >
2026-07-11 13:07:48 +00:00
Harry Mellor and GitHub
1bd8f80a64
[CI] Point CI at Transformers release rather than release branch ( #48328 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-11 02:31:14 -07:00
0b6636cbcb
[XPU]remove is_xxx from moe class and bump up kernels ( #48079 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-11 09:27:13 +00:00
Harry Mellor and GitHub
4a6440acef
Bump Transformers version to 5.13.0 ( #47867 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-11 00:56:14 -07:00
Lucas Wilkinson and GitHub
bec0a4ede6
[Revert] [Build] Update vllm ...builds FA3 with torch stable API ( #48269 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-11 05:20:25 +00:00
3d99b0499a
[Logs] DP Supervisor Log Improvement ( #48278 )
...
Signed-off-by: Robert Shaw <robertgshaw2-redhat@ip-172-31-18-125.us-east-2.compute.internal >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@ip-172-31-18-125.us-east-2.compute.internal >
2026-07-11 12:07:00 +08:00
04d553f390
[Misc] Use meta tensor for KV cache stride calculation ( #47316 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-10 23:24:59 -04:00
9c18e90f6c
[BugFix] Fix packed HND KV cache reshape for FlashAttention ( #47314 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-10 23:22:39 -04:00
Jimmy Lee and GitHub
092387963c
[BugFix] weights processing peak memory reduction for nvfp4 MoE layers ( #46276 )
...
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com >
2026-07-11 02:05:35 +00:00
1bf3997eae
[Quantization] Bound peak memory when repacking FP4 MoE weights for Marlin ( #47851 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-10 19:46:13 -06:00
29fd688892
Add VLLM_FLASHINFER_AUTOTUNE_SKIP_OPS and skip CuTeDSL fp4_gemm autotuning by default ( #48268 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-10 18:13:08 -07:00
Ashwin Giridharan and GitHub
ed908cf0a0
[Bugfix] Fix thinking_token_budget not enforced after natural </think> re-entry ( #45984 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
2026-07-10 22:47:51 +00:00
26ff616bbf
[Bugfix][Test] Register Qwen/Qwen3.5-4B example model ( #48276 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-10 17:01:23 -04:00
gnovack and GitHub
f378f79b7c
handle topk_ids padding in align sum kernel ( #47785 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-10 13:33:28 -07:00
735def4fcf
[Bugfix] Fix FlashMLA dense fp8 metadata crash (num_sm_parts clamp) ( #48045 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-10 12:24:52 -07:00
c227aaa3f8
[ROCm] Enable DeepSeek-V4 DSpark speculative decoding on AMD (MI350X / MI355X, gfx950) ( #47419 )
...
Signed-off-by: larryli2-amd <larryli2@amd.com >
Signed-off-by: larryli2-amd <Larry.Li@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-10 23:22:19 +08:00
Michael Goin and GitHub
08dfd68610
[Model] Add LongCat-Flash-Lite (n-gram embedding) ( #47857 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-10 07:17:50 -07:00
Tyler Michael Smith and GitHub
978a6dfa3f
[Build/CI] Build arm64 PR and postmerge image builds for Blackwell SM10x and SM110 ( #48041 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-07-10 10:12:50 -04:00
85c09e9885
fix: correct load_weights track logic and enable weight integrity for… ( #41811 )
...
Signed-off-by: Yipeng Hu <i26268@metax-tech.com >
Signed-off-by: HuYiPeng <144002351+MynameFelix@users.noreply.github.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Yipeng Hu <i26268@metax-tech.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-10 14:08:20 +00:00
b12cca6a23
[Bugfix] Fix turboquant FP8 cast failure for BF16 models on Ampere GPUs ( #39988 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Hoseung Kim <ghyutjik123@gmail.com >
2026-07-10 06:55:28 -07:00
Wentao Ye and GitHub
e257faf87d
[Refactor] Remove unused rocm kernel combine_topk_swa_indices_ragged ( #48158 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-10 09:33:07 -04:00
FAN YUCHEN and GitHub
fabec87f63
[Model] Migrate MistralLarge3ForCausalLM to AutoWeightsLoader ( #48153 )
...
Signed-off-by: Yuchen Fan <functionhx@gmail.com >
2026-07-10 12:27:58 +00:00
7614b88ebd
[Bugfix][Spec Decode] Fix DFlash draft/target layer-count mismatch ( #48113 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-10 04:42:53 -07:00
Isotr0py and GitHub
68ea76e780
[Misc] Remove dead code in ViT functionality test ( #48220 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-10 11:17:42 +00:00
c241c7a2b0
[Rust Frontend] Add roundtrip fixtures for more chat parsers ( #47883 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-10 10:03:58 +00:00
e23b19309b
Deepstream video backend ( #42424 )
...
Signed-off-by: Viranjan Pagar <vpagar@nvidia.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-07-10 02:23:30 -07:00
f36284a8d2
[CI] Add TORCH_NIGHTLY=1 build mode (run full suite on torch nightly) ( #47180 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
Co-authored-by: Kevin H. Luu <khluu000@gmail.com >
2026-07-10 01:38:35 -07:00
Mingfei Guo and GitHub
424df4f65d
[Model][CI/Build] Cosmos3: enable registry tests and register Cosmos3-Super ( #48211 )
...
Signed-off-by: Mingfei Guo <1800012773@pku.edu.cn >
2026-07-10 16:35:15 +08:00
Bugen Zhao and GitHub
074bdd0d99
[Rust Frontend] Integrate MM video support ( #47959 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-10 08:15:33 +00:00
216ee58780
Add XPU nightly and release image publishing to DockerHub ( #48126 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: jun,du <jun.du@intel.com >
Co-authored-by: jun,du <jun.du@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-10 00:56:00 -07:00
433f291195
[CI] Right-size test-area timeouts from nightly durations ( #48186 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-10 00:53:16 -07:00
Chaojun Zhang and GitHub
28eaf05d56
[XPU] Enable v1/sample tests on XPU CI ( #44472 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-10 15:40:51 +08:00
Jiangyun Zhu and GitHub
300e33797f
[Perf] fuse more rmsnorm and all-reduce in qwen3.5 ( #46998 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-07-10 15:37:51 +08:00
5715fde12c
[Feature][Parser] Support include_reasoning param for non-Harmony models ( #44301 )
...
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-07-10 15:34:02 +08:00
e5588e49bc
[Core][KV events] Report prefix-cache-reused blocks in full report mode ( #45261 )
...
Signed-off-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 22:46:54 -07:00
95ed0feaa5
DCP supports hybrid attention ( #40996 )
...
Signed-off-by: YanXu <yancey.yx@alibaba-inc.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-09 21:34:45 -07:00
2d814a0082
[kv_offload] Emit tier-owned BlockStored events from FS/OBJ secondary tiers ( #47923 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-10 06:17:23 +03:00
88e5e2c57b
[CI/Build][AMD] Fix ROCm OOM in eagle_correctness_heavy by reserving CUDA graph memory ( #47366 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-10 02:14:38 +00:00
Augusto Yao and GitHub
feb384ada2
[bugfix] bge-m3-sparse-plugin mismatch requests ( #48112 )
...
Signed-off-by: augusto.yjh <augusto.yjh@antgroup.com >
2026-07-10 10:03:00 +08:00
a0f6d767e4
[ROCm][CI] Move remaining engine/samplers AMD steps to mi325_1 ( #48169 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-10 00:20:15 +00:00
gnovack and GitHub
f1a5adddb8
update marlin M size for EP ( #48144 )
...
Signed-off-by: gnovack <novackgm@gmail.com >
2026-07-09 23:52:52 +00:00
ap9272 and GitHub
cac3e70cd4
Correct model layer aliasing for Bert style models ( #43896 )
2026-07-09 19:46:22 -04:00
Lucas Wilkinson and GitHub
e12b91b032
[CI] Fix cargo-deny config flag ordering ( #48170 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-09 21:43:58 +00:00
Micah Williamson and GitHub
766469a4c4
[ROCm] Revert Part of [ROCm] Fix pooling startup workspace lock #47912 ( #48154 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-07-09 20:34:24 +00:00
Lucas Wilkinson and GitHub
ea0fa34f49
[CI] Increase extract hidden states TP2 timeout ( #48161 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-07-09 16:19:03 -04:00
ZihaoMu and GitHub
bbb0f945ff
[ROCm] Synchronize sparse MLA metadata before graph replay ( #47404 )
...
Signed-off-by: zihaomu <zmu@amd.com >
2026-07-09 14:59:56 -05:00
2ded1b24e7
[KV Connector][Mooncake] Apply SWA lookup mask before hashing/key build ( #47317 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 19:51:23 +00:00
b0dec2a11b
[ROCM][DSV32][Perf][MTP] Enable UNIFORM_BATCH CG mode in rocm_aiter_mla_sparse ( #45149 )
...
Signed-off-by: Teemu Virolainen <teemu.virolainen@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-09 14:35:10 -05:00
ff8d3488f2
[Bugfix][MRV2] Reset num_accepted_tokens on add_request in all modes ( #48132 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-09 18:17:11 +00:00
weishu and GitHub
2285cfca46
[KVConnector] MultiConnector: give every sub-connector the request's real blocks in update_state_after_alloc ( #46865 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-07-09 11:10:39 -07:00
e08a915146
[Bugfix] Preserve tensor causal metadata for grouped attention ( #48135 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Codex <codex@openai.com >
2026-07-09 17:57:53 +00:00
Charlie Fu and GitHub
67e7ea8977
[ROCm][CI] Set all timeout_in_minutes to 180 ( #48146 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-07-09 17:52:26 +00:00
429f405748
[Bugfix] Guard CUDA-only rms_norm_per_block_quant in FUSED_OPS for non-CUDA builds ( #47296 )
...
Signed-off-by: Tsvika Shapira <tsvika@moonmath.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-09 10:10:53 -04:00
Brandon Pelfrey and GitHub
753c5039f0
Pin PyNvVideoCodec to tested 2.0.4 wheel ( #48056 )
2026-07-09 07:07:50 -07:00
299d2b5655
[CI] Annotate built Docker image tags on the Buildkite build page ( #48101 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-07-09 22:02:40 +08:00
85b3a7264b
[Bugfix][Model Runner V2] Order uniform decodes first so spec decodes aren't misclassified as prefills ( #47381 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-09 14:26:27 +01:00
Harry Mellor and GitHub
b83be00cdd
Migrate Olmo and Olmo2 to the Transformers modeling backend ( #48100 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 05:00:23 -07:00
412414d8e0
Remove PersimmonForCausalLM and FuyuForCausalLM model architectures ( #48096 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-09 04:59:08 -07:00
ae6170f874
[P/D][Bugfix] Fix PD async KV load lookahead handling for MTP spec decode ( #46694 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-09 10:01:22 +00:00
e87521626f
Sanitize server file paths from validation error responses ( #46415 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-09 17:46:29 +08:00
1cd75b3dd4
[Bugfix] Fix race condition in KVBlockZeroer ( #48085 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-09 09:18:19 +00:00
0206f10871
Add Intel XPU Docker release pipeline ( #47880 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: jun,du <jun.du@intel.com >
Co-authored-by: jun,du <jun.du@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-09 01:13:51 -07:00
ab7961a14a
Remove TeleChatForCausalLM ( #47989 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 00:34:54 -07:00
a07765c6bd
[Bugfix] Fix Qwen3-ASR transcription streaming postprocessing ( #42478 )
...
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
Signed-off-by: JooHo Lee <jooho414@gmail.com >
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
2026-07-09 00:33:27 -07:00
Li, Jiang and GitHub
1171467e91
[CPU] Fix Qwen-Next SSM type for AMX GDN ( #48073 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-09 15:09:31 +08:00
Chauncey and GitHub
529af88842
[KV Offloading] Add free block iterator for CPU offload scheduling ( #47849 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-09 06:59:54 +00:00
Chaojun Zhang and GitHub
b8c7c86533
[XPU][LoRA] Fix torch.compile DEVICE_LOST by avoiding view-mutation in LoRA shrink ( #47944 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-09 06:08:30 +00:00
2c17d33f42
[Bugfix][ROCm] Change AttentionCGSuppoort in TritonMLA to UNIFORM_SINGLE_TOKEN_DECODE ( #47144 )
...
Signed-off-by: Dino Music <Dino.Music@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-08 21:09:42 -05:00
bc44f9feb7
[ROCm][CI][MoE] Fix double-transpose of fused w3 expert weights ( #47874 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 16:59:28 -07:00
7802c20c4e
[KVConnector][NIXL] Support pipeline-parallel prefill in push mode ( #45880 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 16:49:23 -07:00
95d6d6f4bb
[Bugfix] Use int8 workspace for FlashInfer MLA decode ( #48046 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-08 23:39:40 +00:00
Harry Mellor and GitHub
56da398dac
Fix embed scaling + CUDA graphs in Transformers modelling backend ( #48010 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 00:14:33 +01:00
Andreas Karatzas and GitHub
26831949b4
[ROCm] Fix pooling startup workspace lock ( #47912 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-08 17:59:50 -05:00
6cf7b26bd4
[docs] Fix the docs build ( #48008 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-08 15:47:22 -07:00
Roberto L. Castro and GitHub
5f85975624
[Feat] Add runtime monitor for post-warmup TileLang compilation ( #46718 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
2026-07-08 22:11:28 +00:00
dcdd756d75
[CI] GSM8K eval integration test for KV offloading ( #46893 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 17:59:48 -04:00
Thien Tran and GitHub
0d2f4e7c9c
Allow FlashInfer A2A backends for TRTLLM FP8 MoE Modular ( #46661 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-08 14:58:39 -07:00
djramic and GitHub
49abadaedb
[ROCm][Bugfix] Fix empty-tensor .max() crash in AITER FA ( #47894 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-07-08 16:58:36 -05:00
Kaihang Jiang and GitHub
089e412878
[Perf] Integrate TRTLLM BF16 MoE Modular Kernel ( #45182 )
...
Signed-off-by: Kaihang Jiang <kaihangj@login-lyris02.lyris.clusters.nvidia.com >
2026-07-09 01:36:14 +04:00
Nick Hill and GitHub
a5d19cbb95
[Core] Move MRV1 late_interaction_runner.py out of MRV2 subtree ( #48014 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 18:30:11 +00:00
Chris Leonard and GitHub
8347c6e6e1
updated flash_attn GIT_TAG to point to torch Stable ABI FA3 commit ( #47995 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-08 10:56:29 -07:00
b2cf70ea3a
[CI] BugFix Eval Small Models Distributed test for DiffusionGemma ( #47980 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-07-08 17:00:56 +00:00
almayne and GitHub
d1f1d86797
[Bugfix] Re-enable benchmarking of librispeech dataset. ( #47033 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
2026-07-08 16:19:26 +00:00
shawn and GitHub
f05603fa28
[Bugfix][DCP] Cast LSE to fp32 in a2a combine to fix bf16 bitcast crash ( #47801 )
...
Signed-off-by: Shawn Tsai <shawnyht@gmail.com >
2026-07-08 11:41:26 -04:00
c2ecd0f888
Fix FlashAttention MLA prefill V unpadding ( #42642 )
...
Signed-off-by: Martin Vit <martin@voipmonitor.org >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-08 15:22:20 +00:00
0d12618e98
[Spec Decode] Support hybrid (SWA + full attention) DFlash drafters ( #47914 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-08 11:12:45 -04:00
Tyler Michael Smith and GitHub
68b4a1d582
Fix NVML capability lookup for visible devices ( #47892 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-07-08 09:07:44 -04:00
572b25b03e
[Bug] Fix Batched DeepGEMM ( #47884 )
...
Signed-off-by: Robert Shaw <robertgshaw2-redhat@dgx-b200-02.mgmt.accl-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@dgx-b200-02.mgmt.accl-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:05:03 -04:00
9f2b3b093c
Improvement of Docker image build for IBM Power using prebuilt wheels from IBM published devpi index ( #46017 )
...
Signed-off-by: vivek sharma <vivsharm@redhat.com >
Signed-off-by: puneetsharma21 <puneet.sharma21@ibm.com >
Signed-off-by: Puneet Sharma <puneet.sharma21@ibm.com >
Co-authored-by: vivek sharma <vivsharm@redhat.com >
Co-authored-by: Puneet Sharma <puneet.sharma21@ibm.com >
Co-authored-by: depthfirst-app[bot] <184448029+depthfirst-app[bot]@users.noreply.github.com>
2026-07-08 13:01:14 +00:00
cd0de48d08
[Bugfix][V1] Free out-of-window blocks on the processed-token basis under async scheduling ( #47728 )
...
Signed-off-by: Saddss <28726669061@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Saddss <28726669061@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 13:34:19 +01:00
rasmith and GitHub
934eeaecfb
[CI/Build][BugFix][The Rock] Fix get_ssm_device_name to return sanitized, usable filename ( #47781 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-07-08 12:12:54 +00:00
Bugen Zhao and GitHub
2cae98dfa5
[Rust Frontend] Handle continue_final_message with renderer sentinel ( #47844 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-08 12:57:04 +01:00
db39d60010
Add tuned selective_state_update float32 config for AMD Instinct MI355 ( #47943 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:41:26 +00:00
a1ab51afb6
[Bugfix] Allocate HY V3 expert_bias in float32 to prevent silent downcasting ( #47797 )
...
Signed-off-by: 辰言 <oncwnuIWp30GguOyJ615Fqj8H-yc@git.weixin.qq.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: 辰言 <oncwnuIWp30GguOyJ615Fqj8H-yc@git.weixin.qq.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-07-08 11:25:53 +00:00
Thien Tran and GitHub
e7b3853bac
Remove router weight upcast for DSv2-related models ( #47970 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-07-08 11:19:10 +00:00
eeaf23107f
[ROCm] Add tuned selective_state_update float32 config for AMD Instinct MI300X ( #47947 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:09:02 +00:00
Canlin Guo and GitHub
285c08c036
[Model] Support MOSS-Transcribe-Diarize ( #47729 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-07-08 04:05:45 -07:00
1f4ad059d1
[ROCm] Add tuned selective_state_update float16 config for AMD Instinct MI300X ( #47945 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-08 11:03:58 +00:00
04a703e397
[Frontend] Support bad_words in the /v1/completions endpoint ( #46793 )
...
Signed-off-by: sungbin1015 <sbin@solbox.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:51:17 +00:00
Nicolò Lucchesi and GitHub
bd3bb4eb26
[Misc][Docs] Add human-readable integer support for more cli-args ( #47608 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 09:43:34 +00:00
Chaojun Zhang and GitHub
440002552e
[XPU] [Fusion passes] Disable fuse_rope_kvcache_cat_mla & qk_norm_rope_ fusion on XPU ( #47962 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-08 09:23:06 +00:00
99a85617bf
[Test] Skip DeepEP MoE layer tests without P2P access ( #47946 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:46:03 +01:00
Nicolò Lucchesi and GitHub
7c67da967f
Remove unused _get_kv_cache_config_deepseek_v4 alias ( #47969 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 01:18:03 -07:00
Nicolò Lucchesi and GitHub
d79855eaac
[Docs] kv_sharing_fast_prefill correction ( #47044 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 01:17:46 -07:00
51e5372f3d
[Model][HunyuanVL] Use native transformers processor and adapt to transformers 5.13 ( #47872 )
...
Co-authored-by: manayang <manayang@tencent.com >
2026-07-08 07:58:23 +00:00
Ace Eldeib and GitHub
7cc2e8e74f
fix: hash speculative draft model config ( #47911 )
...
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com >
Signed-off-by: Ace Eldeib <alexeldeib@gmail.com >
2026-07-08 08:36:30 +01:00
Hongxia Yang and GitHub
2c64b4c1cc
[ROCm] fixed aiter master flag and expert parallelism compatibility on minimax-m3-mxfp8 ( #47158 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-07-08 15:26:17 +08:00
d35eba302f
[Bugfix] Avoid leaking Pydantic repr in tool_choice error message ( #47028 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-08 15:00:59 +08:00
Nicklas Frahm and GitHub
c0e8e1f12a
[Bugfix] Register VLLM_BUILD_* and VLLM_IMAGE_TAG provenance env vars ( #45313 )
...
Signed-off-by: Nicklas Frahm <nicklas.frahm@gmail.com >
2026-07-08 06:21:12 +00:00
Zach Zhu and GitHub
5d5fab0061
[Bugfix][Frontend] Fix http_requests_total metric recording some 4xx errors as 5xx ( #44303 )
...
Signed-off-by: Zach Zhu <zzqshu@126.com >
2026-07-08 05:33:21 +00:00
2afa3f7e95
[Perf] Minimax M3 - Support cross-layer allreduce-norm fusion ( #47631 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-07-07 21:16:32 -07:00
80eb01e93d
[Bugfix] DSV4 TP16 garbage output ( #47493 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-07-07 21:04:33 -07:00
d9e57ea82e
[ROCm][Perf] MXFP8 dense-linear + grouped-MoE GEMM optimizations for MiniMax-M3 ( #46117 )
...
Signed-off-by: amd-ethany <amd-ethany@users.noreply.github.com >
Co-authored-by: amd-ethany <amd-ethany@users.noreply.github.com >
2026-07-08 04:03:34 +00:00
9021589498
[Minimax-M3] Using tok_sparse_select from MSA instead of triton kernels ( #47502 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 21:01:12 -07:00
Ting SUN and GitHub
0303f37a54
[Bugfix][Pooling] Align CrossEncoder token type ids after truncation ( #47772 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-08 03:59:22 +00:00
Walter Beller-Morales and GitHub
dd127d82ed
[Core][Engine] only materialize tokens when thinking budget is in req ( #47053 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-07-07 21:02:38 -06:00
0ca6eee743
[Core] Pass request context to CPU offload cache policy touch ( #47744 )
...
Signed-off-by: jacklin78911-collab <jacklin78911@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-08 05:56:25 +03:00
Isotr0py and GitHub
5e975eae1a
[Bugfix] Avoid blocking model launching when no system ffmpeg available for TorchCodec ( #47888 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-08 10:52:25 +08:00
Martin Hickey and GitHub
f7fc0ca993
[Frontend] Add endpoint plugins framework ( #47454 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-07-08 10:00:41 +08:00
Rahul Vishwakarma and GitHub
f7efab58ec
[CPU][Bugfix] Fix flaky ShortConv prefill test on ARM (uninitialized weights) ( #47848 )
...
Signed-off-by: Rahul Vishwakarma <Rahul.Vishwakarma2@ibm.com >
2026-07-07 18:20:09 -07:00
e97c3cb303
[Core] Persist and reuse the memory-profiling result across boots (opt-in) ( #47388 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 00:53:02 +00:00
4aceabf8c1
[ROCm][Bugfix] Key sparse-MLA persistent metadata on per-request context lengths ( #47766 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 19:22:34 -05:00
stefankoncarevic and GitHub
6e35c5e5af
[ROCm][CI] Minimize comment in RocmAttention q_scale check ( #47731 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
2026-07-07 19:16:08 -05:00
aad0fb741b
[CI/Build] Accept ready-run-all-tests label in pre-commit gate ( #47897 )
...
Signed-off-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: AmeenP <ameen@primeintellect.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-07 23:18:59 +00:00
yzong-rh and GitHub
7d2ce5750e
[Bugfix] Patch Hopper MXFP4 OOB scales reads leading to NaN ( #47910 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-07 22:51:22 +00:00
Juan Pérez de Algaba and GitHub
675f4295cd
fix(security): bound completion prompt list to prevent unbounded engine fan-out ( #47845 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-07 22:48:20 +00:00
Jason Li and GitHub
d99adcebdc
[BugFix] Fix ModelOpt quantization inference for fused siblings ( #47445 )
...
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
2026-07-08 03:19:42 +05:00
c8c2f838e7
Add tuned selective_state_update config for AMD Instinct MI355 ( #47767 )
...
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com >
2026-07-07 22:19:08 +00:00
dd0d74cd92
[Doc] Surface the --kv-cache-memory suggestion at INFO and document fast-startup knobs ( #47374 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 15:05:07 -07:00
55da232db6
[Bugfix] Pad Mamba page size instead of scaling block_size in unify_kv_cache_spec_page_size ( #45207 )
...
Signed-off-by: Sahil170595 <147995121+Sahil170595@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 22:01:34 +00:00
Wentao Ye and GitHub
3f99883d97
[CI Bug Fix] Temp fix for v3.2 accuracy ( #47902 )
2026-07-07 16:36:03 -04:00
Nick Cao and GitHub
47c40bfe8a
[Doc] Fix manylinux tag in installation guide ( #47913 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
2026-07-07 20:34:05 +00:00
3dd910da42
[Bugfix] Allow non-contiguous query in FlashInfer FP8 query quantization ( #47908 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 20:11:34 +00:00
Benjamin Chislett and GitHub
7bd154375d
[Bugfix] Fix mamba+dflash for MRV2 ( #47698 )
2026-07-07 15:59:13 -04:00
Rishabh Saini and GitHub
2f3f441f84
fix: include topic frame in KV events replay response ( #45177 )
...
Signed-off-by: RishabhSaini <rishabhsaini01@gmail.com >
2026-07-07 14:48:23 -04:00
d6875196ad
[Bugfix] Exclude kv_cache_memory_bytes from CacheConfig.compute_hash ( #47356 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
2026-07-07 10:46:51 -07:00
Sting Lin and GitHub
abe41f28de
Upgrade tpu-inference to v0.24.0 ( #47835 )
...
Signed-off-by: StingLin <sting.lin@cienet.com >
2026-07-07 17:15:32 +00:00
Roberto L. Castro and GitHub
c3284c31f5
[Perf][3/N] Expand Triton kernel warmup coverage, Qwen ( #47546 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-07-07 17:06:59 +00:00
Robin and GitHub
c74e751824
[Doc] Fix grammatically incorrect error message in gpu_worker and xpu_worker ( #36715 )
...
Signed-off-by: Hongbin10 <jdmjdm1998@163.com >
2026-07-07 17:03:06 +00:00
liuzhenwei and GitHub
b93cbd7416
[XPU] Fix topk_sigmoid arg mismatch on XPU ( #47858 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-07 16:53:18 +00:00
bdc6f3bfa1
[Bug] Fix tmp directory for lm_eval ( #47755 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 16:40:12 +00:00
392d1b4d2e
[BugFix][LoRA] Refresh punica metadata when LoRA slots are reassigned under an unchanged mapping ( #47725 )
...
Signed-off-by: AmeenP <ameenp360@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 08:53:55 -07:00
21b396abe1
AGENTS MD: Add suggestion on how to incorporate tests ( #47784 )
...
Signed-off-by: Simon Mo <simon.mo@hey.com >
Co-authored-by: Cursor Agent <cursoragent@cursor.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-07-07 08:16:08 -07:00
liuzhenwei and GitHub
bdaf27519f
[XPU] Fix Event init failure w/ blocking ( #47868 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-07 22:54:03 +08:00
Eldar Kurtić and GitHub
beb4327c46
Enable causal masking for SWA in vllm-project/speculators models ( #47745 )
...
Signed-off-by: Eldar Kurtic <8884008+eldarkurtic@users.noreply.github.com >
2026-07-07 10:24:14 -04:00
c46ced1ee3
[kv_offload] Establish tier-owned KV event handling ( #46544 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Signed-off-by: Chang Guo <changg@nvidia.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:55:20 +03:00
65dcde1695
[Bugfix] Fix PD disagg + MTP correctness for Qwen3.5(GDN) ( #47466 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 13:51:22 +00:00
65a7b46284
[KV-Offloading] Support workload identity for objectstore secondary tier ( #47063 )
...
Signed-off-by: Pierangelo Di Pilato <pierdipi@redhat.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:30:16 +03:00
93e2ab7111
Disable dynamic speculative decoding when DP is enabled ( #45963 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 12:52:41 +00:00
Lanze Liu and GitHub
8b745527cd
[Bugfix] Fix UBatchWrapper CUDA graph key to sum all ubatches, not just first two ( #43161 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-07-07 12:42:31 +00:00
920469974a
[UX] Log worker exit code when process dies unexpectedly ( #38641 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-07 12:36:29 +00:00
8b91cd5b20
[Bugfix][Core] Close underlying iterator in merge_async_iterators single-iterator fast path ( #44726 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 05:13:41 -07:00
Harry Mellor and GitHub
dd94484577
Bump Transformers version to 5.10.4 ( #41359 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 05:13:28 -07:00
Shaun Kotek and GitHub
7ff656cc8b
fix: ensure no double load of lm head in nemotron mtp ( #47440 )
...
Signed-off-by: Shaun Kotek - Nvidia <skotek@nvidia.com >
2026-07-07 12:01:45 +00:00
danielafrimi and GitHub
0a2965b1b3
[BugFix] Fix ModelOpt mixed-precision quantization for sparse quantized_layers configs. ( #47318 )
...
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Signed-off-by: <dafrimi@nvidia.com >
2026-07-07 11:45:13 +00:00
Harry Mellor and GitHub
0ed05b6f82
[CI] Fix Transformers modeling backend LoRA test ( #47832 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 11:40:00 +00:00
Guan-Ming Chiu and GitHub
ed051fab54
[Bugfix] Reject sampling params unsupported by diffusion models ( #45418 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-07-07 11:25:36 +00:00
48fcfc926c
[KV Offload] Add ParentManager ABC for secondary tier callbacks ( #47274 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 13:51:18 +03:00
3354dba381
[Bugfix][KV offload] Store interior chunk-boundary blocks under MTP/Eagle ( #46972 )
...
Signed-off-by: Mikhail Kostryukov <mike@triptrack.net >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 13:16:52 +03:00
cbb5f045be
[ROCm][CI] Refresh ROCm base images when docker rocm_base changes ( #46904 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Codex <codex@example.invalid >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Codex <codex@example.invalid >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-07-07 03:10:50 -07:00
b3e85be663
fix: use configured max_logprobs instead of hardcoded 20 in derender validation ( #47834 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-07 09:42:47 +00:00
Summer Yang and GitHub
d3e69fd671
[Perf] Use blocking CUDA events to avoid busy polling cuda driver lock ( #47081 )
...
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
2026-07-07 09:36:10 +00:00
c85d72076a
[HARDWARE][POWER] optimize math functions of VSX power ( #47321 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
2026-07-07 09:35:47 +00:00
c5b66233b2
[Bugfix][Spec Decode] Skip uniform spec-decode padding for diffusion models ( #47464 )
...
Signed-off-by: kl527 <kl527@cornell.edu >
Signed-off-by: Kyung Sub Lee (Daniel) <66861800+kl527@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 09:25:12 +00:00
066f02ae94
[MoE] FI autotuning: max bucket = max token count [e.g. DP_size*MNBT] ( #47427 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 12:08:36 +03:00
Jee Jee Li and GitHub
5d23ca47ab
[Kernel] Applies routed_scaling_factor internally ( #47408 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-07 02:00:54 -07:00
e55cc59e52
[Rust Frontend][CI] Unblock more end-to-end test cases ( #47735 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-07 08:27:20 +00:00
ba50b9763f
[Bugfix] Match the mapped filename in find_loaded_library ( #47586 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-07-07 08:06:29 +00:00
b4cfbc24d3
[Bugfix][Core] Fix host memory leak from undrained new_block_ids ( #44490 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-07 07:32:55 +00:00
Aritra Roy Gosthipaty and GitHub
1e823dc01d
[docs update] Update usage of hf cli for cache list and removal ( #47830 )
...
Signed-off-by: Aritra Roy Gosthipaty <aritra.born2fly@gmail.com >
2026-07-07 07:09:18 +00:00
8e61b646e2
fix(security): add resource bounds validation to derender endpoints ( #47260 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 14:58:26 +08:00
e040899a00
[KV Offloading] Add basic offloading metrics ( #45958 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivas Krovvidi <194645829+Srinivasoo7@users.noreply.github.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 09:26:28 +03:00
dd5c299fbe
[ROCm][Bugfix] Convert ModelOpt FP8 per-channel weights to e4m3fnuz on MI300/MI325 ( #47201 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 23:24:58 -07:00
xiangdong and GitHub
6db31c8e76
[XPU][CI]Adjust memory request for tests in Intel GPU CI ( #47758 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-07 05:56:40 +00:00
cbe9c40f99
[Bugfix] Forward callable hf_overrides to the draft model config ( #45352 )
...
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-06 21:12:38 -07:00
Andreas Karatzas and GitHub
2f71b2bd9f
[ROCm] Align mixed encoder-decoder KV cache views in V2 runner ( #47685 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-07 12:09:22 +08:00
32ab064621
[UX] Add model_class_overrides for development and debugging ( #47148 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 11:43:08 +08:00
Guan-Ming Chiu and GitHub
c64c356990
[Perf] Bound DiffusionGemma sampler transient via request-tiled logits ( #45672 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-07-07 03:42:01 +00:00
Tahsin Tunan and GitHub
34e6dfced8
[Rust Frontend] Stamp arrival_time at the frontend entry ( #47787 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-07-07 03:27:10 +00:00
Reid and GitHub
39a1d32b59
[Rust Frontend] Avoid extra copies for multimodal tensors ( #47581 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-07 03:09:55 +00:00
700e882eab
Add TorchCodec as a video decoding backend ( #46609 )
...
Signed-off-by: Nicolas Hug <contact@nicolas-hug.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-07-06 19:58:51 -07:00
a4f019fa25
fix(distributed): propagate distributed_timeout_seconds to NCCL device groups ( #45159 )
...
Signed-off-by: jialoop-git <joane8913456@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 02:52:51 +00:00
Rahul Vishwakarma and GitHub
9dd2465896
feat(cpu): add CPU support for Mamba ShortConv ( #35059 )
...
Signed-off-by: Rahul Vishwakarma <Rahul.Vishwakarma2@ibm.com >
2026-07-07 10:47:08 +08:00
Reid and GitHub
a46c9329e5
[Rust Frontend] Add DeepSeek V3.2 roundtrip fixture ( #47619 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-07 10:47:02 +08:00
Kyle Sayers and GitHub
445321fab4
[Bugfix] [Quantization] Fix loading for CT DSV2 ( #47780 )
2026-07-07 02:28:00 +00:00
69f3150981
[XPU] Fix PP accuracy on XPU device ( #47253 )
...
Signed-off-by: yisheng <yi.sheng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-07 09:17:27 +08:00
86db6c3070
[Frontend] add per-request timing metrics field to response body of Chat/Completions APIs ( #46768 )
...
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com >
Signed-off-by: Simon Mo <simon@inferact.ai >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: GPT-5.5 <noreply@cursor.com >
2026-07-06 17:48:29 -07:00
5769a7382c
[ROCm][CI][Bugfix] Fix flaky parallel tool-call streaming (test assertion + Mistral/Granite parsers) ( #47550 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 19:06:19 -04:00
Andreas Karatzas and GitHub
8484ca5d45
[ROCm][CI] Adding Rust parity ( #47478 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:05:39 -07:00
482e5524fe
[Bugfix][ROCm] Fix memory access fault in AITER MLA backend for DPA+FP8 KV ( #47276 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: nnyrhila <niko.nyrhila@amd.com >
2026-07-06 21:30:02 +00:00
567a78432d
[Bugfix] Fix dp mtp hang ( #40589 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: sherryC41 <sherry.c.c41@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-06 21:08:17 +00:00
d891b9bd51
[Quantization] add humming moe backend to all dense/moe oracles ( #41652 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-06 13:36:07 -07:00
04adc8843b
[Bugfix]Fix DeepSeek-V4 fp8_ds_mla KV cache reshape ( #47716 )
...
Co-authored-by: yy-fighting <23518844576@qq.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-07-06 12:56:44 -07:00
Harry Mellor and GitHub
ae098abe3f
[CI] Fix some errors on main ( #47726 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 19:40:23 +00:00
b1384f5ec6
Enable B12x backend for non-gated MoEs (like Nemotron) ( #43328 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-07-06 12:40:07 -07:00
b136cc2c2c
[Bugfix][Model] Add stability window to DiffusionGemma to match HF stability_threshold semantics ( #45965 )
...
Signed-off-by: Nathaniel McVicar <namcvica@microsoft.com >
Signed-off-by: Nathaniel McVicar <Nathaniel.McVicar@microsoft.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 19:39:12 +00:00
9fde043f54
[Kernel][Helion][1/N] Add Helion kernel for silu_and_mul_per_block_quant ( #43994 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 00:19:01 +08:00
24dd2aec81
[Bugfix] Preserve FP8 indexer WK pairs across incremental load_weights ( #46168 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-06 09:16:46 -07:00
Ranran and GitHub
3ee9eea928
[macOS][CPU][Installation] Fix the broken installation of vllm 0.24.0 in macos + cpu ( #47457 )
...
Signed-off-by: Ranran Haoran Zhang <ranranhaoranzhang@gmail.com >
2026-07-06 08:59:16 -07:00
5bce653e09
Make the Transformers modeling backend as fast as native vLLM ( #47187 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 16:59:14 +01:00
5ad11172b7
[perf]Add fused Kimi image preprocessing ( #47416 )
...
Signed-off-by: Kevin-XiongC <kevin_xiong1997@outlook.com >
Signed-off-by: Kevin_Xiong <kevin_xiong1997@outlook.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-07-06 08:46:32 -07:00
Wentao Ye and GitHub
f70caef48b
[Perf] Cache token_to_req_indices for dsv4, 5x~6x kernel performance improvement ( #47474 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-06 11:17:46 -04:00
8d8ec38361
[Bugfix][Spec Decode] Add missing draft_id_to_target_id to DSparkDeepseekV4ForCausalLM ( #47429 )
...
Signed-off-by: Laurent-Zhang <zhangdongsheng80@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 10:55:47 -04:00
Wentao Ye and GitHub
b1c6dba558
[Refactor] Remove multiple dead code ( #47329 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-06 07:54:08 -07:00
598d51153a
[Bugfix][Distributed] Delegate MNNVL allreduce one-shot selection ( #47589 )
...
Signed-off-by: jesco-absolut <team@srswti.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 07:47:06 -07:00
Yifan Qiao and GitHub
095adf1fdc
[Bugfix] Fix int32 overflow in triton_decode_attention page offsets ( #47671 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-06 10:36:15 -04:00
Harry Mellor and GitHub
51ee564e56
[CI] Skip test for checkpoint that was deleted ( #47748 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 07:24:09 -07:00
373eb314af
[Bugfix][Core] Fix num_output_placeholders underflow with async scheduling + spec decode ( #46066 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 13:50:38 +00:00
641cb59592
[Doc] Clarify fastokens availability ( #45813 )
...
Signed-off-by: LjjJzd <3542531707@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-07-06 13:33:05 +00:00
07f9baf756
Revert "[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )" ( #47668 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-06 14:18:33 +01:00
7a90eb98ab
[Bugfix] [Gemma4] Fix Gemma4 MTP draft model layers ignoring quant_config ( #47091 )
...
Signed-off-by: Ayushman Singh <40520701+ayush1399@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-07-06 14:04:00 +01:00
8f4c69b222
[Rust Frontend] Cache metric handles for scheduler & request stats ( #47444 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-06 13:02:59 +00:00
8b79971bb9
attention: pass None for unused args in unified attention TD path ( #43597 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: quinnlp <quinnlp@users.noreply.github.com >
2026-07-06 21:01:21 +08:00
Nick Hill and GitHub
f676808ba0
[CI] Use TTY for AMD CI tests for colored buildkite logs ( #47730 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-06 20:50:29 +08:00
Qiming Zhang and GitHub
98e4726a14
[fix][run_batch]: respect proxy env vars when downloading media URLs ( #47697 )
...
Signed-off-by: mauyuyuace <qiming1.zhang@intel.com >
2026-07-06 12:45:48 +00:00
BadrBasowid and GitHub
740f379fae
[ROCm][AITER] Directly Implement AITER Custom All-reduce in CudaCommunicator ( #46065 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-07-06 12:16:32 +00:00
Alexis K. and GitHub
40cc2e8327
[Bugfix] Return HTTP 422 for unprocessable image URLs instead of 500 ( #47165 )
...
Signed-off-by: Alexis Kinsella <alexis.kinsella@gmail.com >
2026-07-06 11:56:23 +00:00
ba22152096
fix(security): block request-level GPU video backend selection withou… ( #47259 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 02:36:49 -07:00
Yan Ma and GitHub
90ce3a09be
[bugfix] fix MOSS-Audio deepstack_input_embeds initialization in PP ( #47607 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-07-06 17:15:50 +08:00
26c754d847
[XPU][Bugfix] Do not transpose weight_scale_inv at load time ( #47116 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 17:15:26 +08:00
Sungjae Lee and GitHub
3d7f357ebf
[Doc] docs: fix note formatting for pooling models ( #47701 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
2026-07-06 09:01:10 +00:00
liuzhenwei and GitHub
736f1a5907
[XPU] Route mm_prefix models to Triton attention backend ( #47688 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-07-06 16:52:44 +08:00
Li, Jiang and GitHub
344609ab17
[CI/Build] Fix pre-commit check ( #47695 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-06 08:24:24 +00:00
xiaozhoupy and GitHub
d039c17114
[Bugfix] Recycle post-final-norm hidden in GLM MTP (single norm) ( #47448 )
2026-07-06 01:07:56 -07:00
xiangdong and GitHub
cdab28319f
[XPU][CI]Add agent tags for Basic Models Tests (Initialization) in Intel GPU CI ( #47675 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-06 15:15:45 +08:00
Qiming Zhang and GitHub
2fa10566e3
[Core][DP] Rotate load-balancer tie-break to avoid systematic engine bias ( #47420 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-07-06 07:09:16 +00:00
Andreas Karatzas and GitHub
fb265fc8fb
[ROCm][CI] Increasing parallelism in Basic Models Tests (Extra Initialization) ( #47591 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:06:16 +08:00
Andreas Karatzas and GitHub
8f0e75e16b
[ROCm][CI] Adding nixl multiconn ( #47481 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-06 15:04:58 +08:00
98ba9b9583
[Frontend] Support OpenAI Responses API namespace tools ( #47024 )
...
Signed-off-by: zhongjing123 <jimzhong5193@gmail.com >
Co-authored-by: zhongjing123 <jimzhong5193@gmail.com >
2026-07-06 06:21:27 +00:00
velonica0 and GitHub
990c2a0187
[RISC-V] Enable BF16 on VLEN=256 hardware ( #45243 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-07-06 06:05:16 +00:00
e433634c78
[Performance][Hardware][RISC-V] Reduce LMUL pressure in INT4 LUT dequant ( #47538 )
...
Signed-off-by: liutong <liutong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-06 05:58:56 +00:00
16f8110935
[Bugfix][CPU][RISC-V] Fix VLEN detection for RVV attention path ( #47532 )
...
Signed-off-by: liutong <liutong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-06 05:58:03 +00:00
d9c1767cd4
[INC][ARK] Direct Register Custom Op for ARK ( #46361 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 13:45:50 +08:00
Li, Jiang and GitHub
e9cc1fd093
[CI/Build][CPU] Remove global extra index ( #47687 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-07-06 13:42:01 +08:00
Fadi Arafeh and GitHub
f1073c050c
[CPU][BugFix] Multiple fixes to w4a8_int8 CPU MoE path ( #46739 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-07-06 05:39:20 +00:00
Qiming Zhang and GitHub
394edc8108
[XPU] limit max-num-seqs in test_lmeval.py for XPU ( #47682 )
...
Signed-off-by: mauyuyuace <qiming1.zhang@intel.com >
2026-07-06 05:34:16 +00:00
69715823df
[Test][XPU] Skip fork in kv_sharing_fast_prefill test on XPU ( #47406 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 11:32:26 +08:00
Chaojun Zhang and GitHub
6569df6a3e
[Test][LoRA] Use lightweight CPU reference and skip heavy cleanup in punica ops tests ( #47534 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-06 11:29:59 +08:00
f2aaf59151
[Feature] Support MTP speculative decoding for Bailing hybrid models ( #44880 )
...
Signed-off-by: zc02384840 <zc02384840@antgroup.com >
Co-authored-by: zc02384840 <zc02384840@antgroup.com >
2026-07-06 10:38:50 +08:00
95a248faed
[Attention Backend] HPC_ATTN backend support mtp and dynamic scheduled attention ( #47433 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
2026-07-05 18:18:25 -07:00
d2ec433e37
[XPU] Fix Eagle3 initialization on XPU ( #43957 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 08:46:05 +08:00
78a04c208d
[XPU] Fix CUDA API shims breaking Torch Dynamo during AOT compile ( #43092 )
...
Signed-off-by: Łukasz Ślusarczyk <lukasz.slusarczyk@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-06 08:29:20 +08:00
Spandan Tiwari and GitHub
b71218107f
[ROCm][Test] Fix test_per_token_group_quant_fp8 tolerance for 1-ULP FP8 rounding on gfx950 ( #46944 )
...
Signed-off-by: Spandan Tiwari <sptiwari@amd.com >
2026-07-05 18:02:30 -05:00
cc1d020d01
[MRV2] Enable mm prefix bidi attention support on MRV2 ( #46942 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-05 14:45:29 +00:00
Ting SUN and GitHub
8974ed89cd
[Bugfix][Voxtral Realtime] Fix token feedback timeout silent hang ( #44461 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 05:42:36 -07:00
fb2faceacd
[Bugfix][Model] Fix crash loading Mamba/Mamba2 checkpoints without an architectures field ( #46037 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Ting SUN <suntcrick@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-05 05:42:32 -07:00
b6cc46ec3b
[Feature] Support sequence parallel without the need for DP, 1.9%~5.0% E2E Throughput Improvement ( #47070 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com >
2026-07-05 05:41:30 -07:00
Lucas Wilkinson and GitHub
fa4321de3d
[Bugfix][TurboQuant] Preserve KV cache dtype in backend shape ( #47609 )
2026-07-05 08:20:48 +00:00
Ting SUN and GitHub
9226613043
[Bugfix][Pooling] Forward instruction to Jina reranker scoring prompts ( #47590 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 05:39:13 +00:00
34b560b725
[Bugfix][Gemma4] Fix FA4 mm_prefix mask: add sliding window and absolute q_idx ( #47332 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-07-04 17:46:40 -07:00
Ting SUN and GitHub
91b5647300
[Bugfix][Model] Allow Run:ai memory_limit sentinel values ( #47337 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-05 00:08:34 +00:00
Carl Persson and GitHub
4a6bf3c77f
[ROCm][CI] Fix Kernels and Kernels attention test failures ( #47519 )
...
Signed-off-by: Carl Persson <carl.persson@amd.com >
2026-07-04 15:59:51 -05:00
Ting SUN and GitHub
d2afe39647
[Bugfix][Frontend] Preserve default sampling params in batch chat ( #47597 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-04 19:06:39 +00:00
Wentao Ye and GitHub
2a9113f998
[Perf] Remove redundant op for GLM 5.2 ( #47198 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-07-04 13:25:02 -04:00
yzong-rh and GitHub
0cd6f767e3
[Bugfix][Frontend][gpt-oss] Recover raw tail when Harmony parser ends non-terminal ( #47379 )
2026-07-04 10:46:24 -04:00
Harry Mellor and GitHub
f1445f6dbd
[CI] Bump huggingface-hub from v1.10.2 to v1.22.0 ( #47551 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-04 07:45:45 -07:00
1d354c694e
[Misc] Validate Pooling cache_salt Values ( #46966 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-04 10:19:28 -04:00
Taneem Ibrahim and GitHub
2f21224527
[Misc] Update request-extras parity for batch chat completion ( #47333 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-07-04 10:19:04 -04:00
fa1fa968c4
[Misc] Forward request-level prompt extras for cross-encoder scoring ( #46939 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-04 10:18:36 -04:00
6eac8e0070
[Misc] Preserve cross-encoder pooling extra kwargs ( #47082 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-04 08:14:13 -04:00
1a308c449c
[XPU] Add W8A8 FP8 linear kernel with multi-granularity quant support ( #43645 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-07-04 18:10:01 +08:00
e7c9df9449
[Bugfix][Structured Output][Spec Decode] Constrain bitmask and trim grammar advance at the reasoning boundary ( #44297 )
...
Signed-off-by: Allen.Yu <yuyue0225sc@163.com >
Signed-off-by: yue.yu <yuyue0225sc@163.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-07-04 09:08:45 +00:00
gausah01 and GitHub
26eb87204d
[Bugfix] Fix CPU split-KV scratchpad sizing ( #45844 )
...
Signed-off-by: Gauri Sahnan <gauri.sahnan@arm.com >
2026-07-04 06:47:23 +00:00
4c3c17d43b
[ROCm] Disable persistent sparse-MLA kernel for chunked-prefill continuations ( #47567 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-04 01:21:43 -05:00
f329ce405b
[ROCm][CI][Bugfix] Use VllmRunner for voxtral_realtime tests to avoid OOM on AMD GPU ( #47536 )
...
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-04 12:26:10 +08:00
07516fda67
[MRV2][SD] Make Dynamic SD comatible with Full Cuda Graphs ( #45953 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-03 23:58:27 -04:00
67ff0ae30f
Support nvfp4 kv with kv-cache-dtype-skip-layers sliding_window ( #42890 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-04 02:29:13 +00:00
Bugen Zhao and GitHub
ab3b6d97aa
[Frontend] Limit SO_REUSEPORT to multi-worker serving ( #47529 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-04 01:26:24 +00:00
Ben Browning and GitHub
fb5291b35b
[Frontend] [Parser] Port DeepSeek V4 to streaming parser engine framework ( #45877 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-07-03 20:55:23 -04:00
labAxiaoming and GitHub
d6d39c111e
[GLM4V] Avoid GLM4V processor init during startup metadata reads ( #47155 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
2026-07-03 15:03:16 -07:00
379950191f
[Bugfix][Multimodal] Normalize direct PIL image inputs ( #47566 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-03 14:27:14 -07:00
576bf75d0e
[AMD][EPLB] Enable EPLB for Quark OCP MXFP4 MoE ( #47220 )
...
Signed-off-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: okorzh-amd <okorzh-amd@users.noreply.github.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 14:41:52 -05:00
Tres and GitHub
f006e5a24c
[CI][AMD] Allow git operations on previously created work trees ( #47554 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
2026-07-03 14:41:01 -05:00
f63dca6838
[ROCm] Fix encoder-decoder cross-attention KV layout aliasing ( #47035 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 13:53:29 -05:00
Bugen Zhao and GitHub
8651f043b8
[Rust Frontend] Speed up chat roundtrip tests ( #47523 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 19:25:06 +01:00
Andreas Karatzas and GitHub
3775d5fcab
[ROCm][CI] Adding test groups for parity with upstream ( #47479 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-03 19:15:01 +04:00
d7192cfccf
[CI Bugfix] Lazily import Qwen warmup dependencies ( #47539 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-03 23:10:49 +08:00
AgenticSpark and GitHub
978de83353
[Bugfix][CPU] Ship examples/ in the CPU release image ( #47447 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
2026-07-03 11:46:24 +00:00
wang.yuqi and GitHub
a14f57a3ac
[Frontend] Refine the entrypoint class's inheritance hierarchy. ( #47498 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-03 10:50:06 +00:00
18f658bb31
[Bugfix][Frontend] Fix batch chat endpoint corrupting logprobs when return_token_ids is set ( #47384 )
...
Signed-off-by: David Feng <fenghourun@meta.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-03 03:01:34 -07:00
Isotr0py and GitHub
400a9c386d
[Rust Frontend] Bump llm-multimodal version ( #47530 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-03 09:48:36 +00:00
Max de Bayser and GitHub
bbdcbe4686
Move Roberta remaining nn.Embedding to VocabParallelEmbedding ( #47452 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
2026-07-03 09:47:50 +00:00
Kalyanam Dewri and GitHub
4875b4456b
[Doc] Fix VLM2Vec benchmark chat template path ( #47517 )
...
Signed-off-by: kalyanamdewri <kalyanampriyam@gmail.com >
2026-07-03 08:24:45 +00:00
Dakai An and GitHub
1f486d96a1
Add Triton Backend for Unlimited-OCR R-SWA ( #47102 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-07-03 00:11:50 -07:00
Bugen Zhao and GitHub
b790c84cde
[CI] Enable sccache for Rust build under CUDA/ROCm ( #45246 )
2026-07-02 23:45:41 -07:00
6429d5f527
[Rust Frontend] add repetition_detection support to sampling params ( #46684 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 14:06:23 +08:00
Chris Leonard and GitHub
fbc9ba6d30
New stable abi cleanup ( #46656 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-03 14:02:26 +08:00
xiangdong and GitHub
2dfaae752b
[XPU][CI]Fix dependency typo in Intel GPU CI ( #47510 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 04:11:47 +00:00
Evgeny Parshutin and GitHub
bd8d9021ce
[CPU][Build] Enable oneDNN ITT task collection by default for CPU primitive-level profiling ( #47467 )
...
Signed-off-by: Evgeny Parshutin <eugeny.parshutin@intel.com >
2026-07-03 04:00:19 +00:00
xiangdong and GitHub
3f0b773b30
[XPU][CI]Mv huggingface cache to larger disk in Intel GPU CI ( #47405 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 11:56:17 +08:00
Reid and GitHub
9b8e76589d
[Rust Frontend] Recover buffered text from incomplete tool calls at EOS ( #47289 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-03 03:45:03 +00:00
1aeabec355
[Bugfix][Rust Frontend] Tolerate out-of-vocab prompt ids in detokenizer ( #44682 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 03:41:53 +00:00
979f5511d7
[Bugfix][Gemma4] Keep image bidirectional attention within the sliding window ( #47217 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-07-02 19:57:41 -07:00
41de1380c2
[BugFix] Derive FlashInfer Q dtype from resolved per-group builder state ( #47485 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 19:33:28 -07:00
Nick Hill and GitHub
d85601c20f
[CI] Pin modelscope version to fix test breakage ( #47465 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:33:08 -07:00
Nick Hill and GitHub
276b837dc4
[ModelRunner V2][BugFix] Free all model refs on shutdown ( #47483 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:32:48 -07:00
34bf7b45a0
[CI] intel CI: add quantization and awq case for xpu ( #46456 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 09:51:56 +08:00
adamkbaranowski and GitHub
4c3c64fcf7
Add Laguna XS.2.1 DFlash drafter support ( #46853 )
...
Signed-off-by: Adam Baranowski <adam.baranowski@poolside.ai >
2026-07-02 18:09:27 -07:00
Andreas Karatzas and GitHub
442ccc6098
[ROCm][CI] Adding extract hs 2gpu ( #47482 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:59:38 -07:00
Andreas Karatzas and GitHub
6768fbc76f
[ROCm][CI] Adding qwen3 dp4 eplb ( #47480 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:58:56 -07:00
Andreas Karatzas and GitHub
407f406300
[ROCm][CI] Adding metadata ( #47477 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:45:23 -07:00
Harry Mellor and GitHub
e24d1b24fe
Fix Transformers modeling backend usage stats ( #47472 )
2026-07-02 12:51:23 -07:00
d29125c085
Xqa decode kernels ( #43232 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-02 12:32:05 -07:00
Michael Goin and GitHub
d715b3aa1e
Delete PagedAttention ( #47361 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 12:31:26 -07:00
Joe Rowell and GitHub
258f8de91f
[Bugfix][Tool Parser] poolside_v1: accept tool calls without newline after function name ( #47311 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-07-02 12:08:37 -07:00
Nick Hill and GitHub
e392bf7a68
[BugFix][MRV2] Ensure all req slots are accounted for when scheduling ( #46974 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:24 -07:00
Nick Hill and GitHub
443e68cfa6
[Bugfix] Fix pooled Whisper encoder sliding-window kernel size ( #47437 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:11 -07:00
Chauncey and GitHub
320ee285c9
[Model Runner V2][Perf] Warm up GLM-5.2 DSA indexer prefill metadata kernel ( #47285 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-02 16:31:31 +00:00
Bugen Zhao and GitHub
ec0ffaacc8
[Rust Frontend] Improve scheduler stats logging parity ( #47435 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 15:55:25 +01:00
Yuxuan Zhang and GitHub
178fd56094
support GLM-5.2 gate use FP32 ( #47410 )
...
Signed-off-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk >
2026-07-02 22:45:39 +08:00
a47f38f825
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in block verification kernels ( #47383 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 07:32:38 -07:00
Nick Hill and GitHub
3e158ae62d
[ModelRunner V2] Fix Mamba2 crash on non-spec-decode ( #47428 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 07:05:16 -07:00
a2f713002d
[ModelRunner V2] Enable by default for all dense models ( #44443 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 18:48:57 +08:00
TJian and GitHub
de2a8fc042
[ROCm] [PyTorch] Move to stable abi since ROCm upgraded to torch 2.11 ( #47128 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-02 18:34:07 +08:00
Michael Goin and GitHub
84b9c2762f
Update DeepGEMM tag to point to latest nv-dev branch for sm120 support ( #47304 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 18:33:44 +08:00
Bugen Zhao and GitHub
25fcb65d51
[Rust Frontend] Use enum-backed domain types for engine outputs and structured outputs ( #47283 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 10:41:46 +01:00
08a8a4af3f
feat(rust): expose profiler control routes in Rust frontend ( #46306 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 08:46:07 +00:00
b0b8a286dd
[Model] Add LLaVA-OneVision-2 (LlavaOnevision2ForConditionalGeneration) ( #44785 )
...
Signed-off-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-02 16:41:49 +08:00
3af8789559
[Feature] Universal speculative decoding for heterogeneous vocabularies (TLI) ( #38174 )
...
Signed-off-by: wan-danfeng <wandanfeng0802@gmail.com >
Signed-off-by: Wonderful <wandanfeng0802@gmail.com >
Co-authored-by: Wan_DF <wonderful199082@126.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-07-02 01:34:20 -07:00
Chaojun Zhang and GitHub
8357226f4f
[XPU][CI] Split test_punica_ops into separate pytest invocations for stability ( #47376 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-02 07:50:55 +00:00
Hiki and GitHub
2665ed704b
[Bugfix][Kernel] Correct FlashInfer CUTLASS MoE tuning token bound ( #46838 )
...
Signed-off-by: Haobin Guo <haobing@nvidia.com >
2026-07-02 05:11:00 +00:00
xaguilar-amd and GitHub
09663abde0
[ROCm][MLA] Fuse MLA q/kv RMSNorm + FP8 per-token quant in the FP8 attention path ( #44977 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
2026-07-02 13:00:41 +08:00
Giancarlo Delfin and GitHub
d63c8e9444
[BugFix][Spec Decode] Compact shared topk indices buffer after first MTP draft step ( #47238 )
2026-07-01 21:38:51 -07:00
Jee Jee Li and GitHub
1360c42fe6
[UX] Include NVTX in cuda.txt ( #47319 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 19:38:50 -07:00
d0a2584773
[Misc] Use functions instead of PTX for the PDL instruction ( #46984 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-01 19:38:35 -07:00
7fe7fa9cda
[CI][Bugfix] Rerun test_engine_log_metrics_ray on Ray GCS startup timeout ( #47208 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 21:32:09 -05:00
Michael Goin and GitHub
2b753ad200
[Spec Decode] DSpark speculators checkpoint support ( #47093 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-01 17:32:27 -07:00
e196268bad
[Docker] Remove unused Dockerfile.nightly_torch ( #47338 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:42 -07:00
e91f5f8439
[CI] Remove torch_nightly mirror tags (superseded by TORCH_NIGHTLY full-nightly build) ( #47342 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:06 -07:00
fa248139a0
[MoE] Plumb gemm1_alpha/beta/clamp_limit into TRT-LLM FP8 MoE ( #45723 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 14:34:05 -07:00
Yongye Zhu and GitHub
d3229431f9
[DSV4] Better MXFP8 quantization kernel ( #47229 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-07-01 14:33:51 -07:00
Nick Hill and GitHub
4787f2dd1b
[Bugfix] Don't read KV cache past seq_len in triton paged attn kernels ( #47305 )
2026-07-01 12:43:00 -07:00
Nick Hill and GitHub
8cfeb84dba
[ModelRunner V2] Warmup cross-attn properly in encoder-decoder case ( #47308 )
2026-07-01 12:36:48 -07:00
Chaitanya Sri Krishna Lolla and GitHub
5fd442187c
[ROCm][P/D] MoRIIO toy proxy: support JSON Content-Type for OpenAI clients. ( #46482 )
...
Signed-off-by: lcskrishna <lollachaitanya@gmail.com >
2026-07-01 19:17:05 +00:00
00eb7cefa3
[Bugfix] Prevent padding placeholders from reaching embeddings ( #47029 )
...
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 09:26:03 -07:00
Michał Ganczarenko and GitHub
c8bdcc0116
[Bench][BugFix] Fix empty decoder prompt for Cohere ASR in throughput benchmark ( #47135 )
...
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com >
2026-07-01 15:42:27 +00:00
f5a8d73377
[Spec Decode] DSpark ( #46995 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-01 08:30:24 -07:00
63fcce4de1
[Bugfix] Fix GraniteMoeShared weight loading broken by #41184 ( #47031 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 22:39:12 +08:00
Bugen Zhao and GitHub
c638f9216a
[Rust Frontend] Split engine core DTOs into separate modules ( #47265 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-01 15:28:21 +01:00
Chaojun Zhang and GitHub
13c49f9845
[xpu][lora]: Align LoRA implementation with Punica GPU: fix _apply_expand rank mismatch, add_inputs hardcode, and MoE EP ( #45368 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-01 22:14:04 +08:00
Nick Hill and GitHub
f1cf6b0086
[CI] Fix segfault in tracing test ( #47299 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 14:00:37 +00:00
Harry Mellor and GitHub
a78c15616f
Migrate GPTBigCode and Starcoder2 to the Transformers modeling backend ( #30966 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 13:41:36 +00:00
5c4db60f01
docs(security): document gRPC interface as insecure for private use only ( #45903 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <rbryant@redhat.com >
2026-07-01 12:39:57 +00:00
4e5ca89cfe
[ROCm][MiniMax-M3] Cross-layer lightning-indexer top-k sharing ( #47269 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 10:50:09 +00:00
Harry Mellor and GitHub
a22e0dfc69
[Model] Remove AyaVision, MusicFlamingo ( #47263 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:39:33 +00:00
stevenkuang and GitHub
cc56379e28
[Model] Support Hy3 token suffix and JSON Schema array types ( #47192 )
...
Signed-off-by: stevenkuang-tencent <stevenkuang@tencent.com >
2026-07-01 10:16:07 +00:00
024b06b0dc
[Bugfix] Expose usage field in GenerateResponse for disaggregated serving ( #42748 )
...
Signed-off-by: AIvashov <ivashov.aleksey@proton.me >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-01 10:00:19 +00:00
Harry Mellor and GitHub
e7d0fcbc09
[CI] Fix various failures on main ( #47197 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:35:34 +01:00
akii96 and GitHub
aa8bb5562e
[ROCm][Perf][Bugfix] DSv4 indexer: use platform FP8 dtype (fnuz) for Q-quant on gfx942 ( #46730 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-07-01 17:33:55 +08:00
Andy Lo and GitHub
fa4bec9056
[Bugfix] Fix pooled Whisper sliding-window KV sizing ( #47071 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-07-01 11:33:19 +02:00
dee5da1dec
[Test] Run SageMaker handler-override tests in-process via TestClient ( #47250 )
...
Signed-off-by: Jyothirmai Kottu <jkottu@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 09:14:00 +00:00
ed41aa270a
[ROCm][DSV4] Use aiter mHC pre/post as the default ROCm path ( #43950 )
...
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com >
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 16:27:42 +08:00
77a9c5ae28
Weight sync refactor + move sparse nccl engine ( #44353 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Signed-off-by: haoaaron <ahao@anyscale.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-01 01:25:19 -07:00
f651a8a9a4
[XPU][UT]Enable ut qk_norm_rope_fusion ( #42486 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 07:38:03 +00:00
Jee Jee Li and GitHub
8f82be5705
[CI/Build] Fix LoRA testing ( #47242 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 15:36:13 +08:00
Nils Matteson and GitHub
a461070d1c
[Core] Make sleep-mode backend capability flags communicator-agnostic ( #47243 )
2026-07-01 07:17:44 +00:00
4470ae84de
Remove mantis ( #46806 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 07:13:58 +00:00
Chauncey and GitHub
697c34b97b
[Bugfix] Fix beam search candidate indexing when logprobs count varies ( #47126 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-01 07:07:06 +00:00
Blas Rodriguez Irizar and GitHub
5b431b905c
[Rust Frontend] Coerce completion max_tokens: null to default ( #47166 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-07-01 06:41:33 +00:00
89e99202f2
[CPU][Perf]Added tanh AOR for faster gelu activations. ( #44639 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Signed-off-by: almayne <anna.mayne@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-30 23:24:40 -07:00
Micah Williamson and GitHub
b446792306
[ROCm][Bugfix] Fix Triton "out of resource: shared memory" Error In One-Shot LoRA MoE ( #47209 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:36 -07:00
Micah Williamson and GitHub
c3b1f9e827
[ROCm][CI] Enable LoRA TP Distributed Test Group In AMD CI ( #47193 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:32 -07:00
Jonathan Mamou and GitHub
df802a87b7
[CPU] Remove speculative decoding stream overrides from CPUModelRunner ( #47162 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
2026-07-01 06:12:49 +00:00
Nils Matteson and GitHub
93d8f834dd
[Core] Pluggable sleep-mode backend abstraction (RFC #34303 ) ( #44074 )
2026-06-30 22:00:53 -07:00
Maria Guevara and GitHub
aeb35b90f0
[Rust Frontend] Add error context in tool parser failures ( #46512 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-07-01 12:48:55 +08:00
Gabriel Wu and GitHub
9a08a5118e
fix: skip cooperative top-K on SM120 ( #47164 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-06-30 21:32:54 -07:00
c5200d3565
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE ( #46076 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: GirasoleY <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 00:32:20 -04:00
Matt and GitHub
3c1396bab6
[Hardware][AMD][CI] Toggle test coredumps on ROCm debug agent ( #47222 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 23:30:10 -05:00
Benjamin Chislett and GitHub
9969466a59
[Spec Decode] Support SWA + DFlash for MiMo ( #46104 )
2026-06-30 20:34:47 -07:00
achyuthan.s and GitHub
3406e8f83d
[Bugfix][Frontend][gpt-oss] Return raw output when Harmony parser ends non-terminal ( #47062 )
...
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
2026-07-01 01:46:01 +00:00
a264e41975
[Distributed] Default FlashInfer allreduce to mnnvl on single node ( #47219 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-30 18:35:56 -07:00
Woosuk Kwon and GitHub
f098ee70c7
[GLM5] Support FlashMLA FP8 KV cache (Hopper & Blackwell) ( #47090 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-30 18:13:21 -07:00
9294dd27eb
fix(reasoning): guard rfind in ernie45 streaming </response> branch ( #46255 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-07-01 01:01:14 +00:00
yzong-rh and GitHub
b1190d03cc
[Refactor][GPT-OSS] Harmony Responses API Refactor to use HarmonyParser ( #47185 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-30 19:23:20 -04:00
92c7fac640
[Perf] Restore zero-init of swizzled NVFP4 scale buffer to recover Blackwell decode throughput ( #45739 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-30 22:56:56 +00:00
Ting SUN and GitHub
ac521f6237
[Bugfix][Structured Outputs] Reject degenerate structured_outputs that crash EngineCore ( #45346 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-30 22:41:33 +00:00
28242824e0
[Bugfix][Frontend] Normalize constrained Harmony recipients ( #45657 )
...
Signed-off-by: shaojunjie <626650687@qq.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-30 17:33:10 -04:00
VectorPeak and GitHub
68294739d1
[Bugfix] Align OpenCV video metadata timeline ( #47099 )
...
Signed-off-by: VectorPeak <73048950+VectorPeak@users.noreply.github.com >
2026-06-30 20:43:42 +00:00
c8d2f3cb14
[Bugfix] compressed-tensors: allow int8 grouped WNA16 MoE on Marlin ( #47154 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 12:50:46 -07:00
Matt and GitHub
345b28ff2f
[Hardware][AMD][CI] Bump timeouts of various test groups on AMD CI ( #47195 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 14:30:53 -05:00
248d1fbb71
[Feat][1/N] CuTeDSL warmup infrastructure, FA4 MLA ( #46182 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-30 12:17:34 -07:00
11b26c5528
[Bugfix][Tool Parser] PoolsideV1: fix logprobs AttributeError on Responses API ( #47138 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 19:14:09 +00:00
Roberto L. Castro and GitHub
20434c472e
[Feat] Improve Triton JIT diagnostics ( #46621 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-30 18:50:15 +00:00
Andreas Karatzas and GitHub
c8f9c156a5
[ROCm][V1][MLA] Clone prefill backend state per metadata builder ( #46993 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 11:43:54 -07:00
953bba488d
[PERF] Extend NCCL symmetric memory to AllGather and ReduceScatter ( #46703 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: snordmann <snordmann@nvidia.com >
2026-06-30 11:38:18 -07:00
Wentao Ye and GitHub
3a9784b82c
[Feature] DP supervisor using rust frontend ( #47076 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 14:34:05 -04:00
Giancarlo Delfin and GitHub
3cecee40f3
[Model Runner V2][Spec Decode] Fix stale values in idx_mapping from CG num reqs padding ( #47066 )
2026-06-30 11:25:32 -07:00
a7732537f4
[Bugfix] Restore part of bugfix #42650 after accidental deletion in #43241 ( #47039 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
Co-authored-by: Zhanda Zhu <49645678+zhandaz@users.noreply.github.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-30 11:07:59 -07:00
727971f1c1
Add Medusa speculative decoding e2e test ( #41396 )
...
Signed-off-by: Anshika Ojha <anshikao@nvidia.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Anshika Ojha <215760622+ojhaanshika@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-06-30 18:02:22 +00:00
25671cb520
[Parser][Bugfix] Ensure tool call or other special tokens don't leak in non-streaming tool parsing ( #46875 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-30 13:46:53 -04:00
27d5f78b63
[CI] Move distributed small LM eval to B200 ( #47048 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 13:34:25 -04:00
liuzhenwei and GitHub
7a341fa109
[XPU] Support ZE_AFFINITY_MASK passthrough in xpu_disagg_acc_test ( #47105 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-30 17:06:12 +00:00
Charlie Fu and GitHub
f41e8ddc97
[ROCm][CI] Move PyTorch Compilation Unit Tests to MI300(gfx942) ( #47065 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-30 11:32:58 -05:00
245888ff77
[Feature] Detect all2all peer fault with fault tolerance backend and prevent corrupted output ( #43637 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 09:00:25 -07:00
e840f0d3f5
[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 08:39:59 -07:00
fcaa84efa7
[BugFix] Gate MRV2 mixed sparse-MLA warmup on max_num_seqs > 1 ( #47050 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: ziminghuang <ziminghuang@inferact.ai >
2026-06-30 16:31:27 +01:00
Wentao Ye and GitHub
9e84ec8648
[Refactor] Remove dead minimax allreduce rms kernel ( #46842 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 08:29:21 -07:00
d8f483dc30
[Spec Decode] Fix hidden-state extraction block size for hybrid verifiers ( #46301 )
...
Signed-off-by: Igor Margulis <igor.margulis@intel.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-06-30 08:19:51 -07:00
Nicolò Lucchesi and GitHub
dc148dc4d7
[CI][Bugfix] Fix Hybrid SSM NixlConnector PD prefix cache test (2 GPUs) ( #47157 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 23:14:13 +08:00
tc-mb and GitHub
7cf7cbcd95
[Bugfix] MiniCPM-V 4.6: fix grid rows/cols swap in placeholder generation ( #45918 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-30 08:12:44 -07:00
c231d1f290
fix(security): bound tokenizer work when explicit truncation_side is set ( #47007 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 23:08:51 +08:00
Giancarlo Delfin and GitHub
db808b3961
[Model Runner V2][Spec Decode] Implement block verification for rejection sampling ( #46781 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-30 08:07:24 -07:00
Arsalan Shakil and GitHub
00ebf19cca
[Bugfix][Quant] Raise actionable error instead of bare assert for group-size/TP mismatch ( #46230 ) ( #46236 )
...
Signed-off-by: Arsalan Shakil <shakil.arsalan@yahoo.com >
2026-06-30 14:57:14 +00:00
ded6676458
[Bugfix] Seed RayExecutorV2 TCPStore port by DP rank to avoid collisions ( #45960 )
...
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:37:34 -07:00
Bugen Zhao and GitHub
7a327f0b4f
[Rust Frontend] Simplify unit tests with shared TestTokenizer ( #47125 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 15:34:43 +01:00
Harry Mellor and GitHub
1ab9522935
Remove more unnecessary load_weights methods ( #47058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 15:22:16 +01:00
0fc2512094
[KV Offload] Pass ScheduleEndContext to on_schedule_end hook ( #46450 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 17:07:12 +03:00
Harry Mellor and GitHub
62c7d8009f
Forward fix nightly errors from #44589 ( #47151 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 14:02:34 +00:00
Isotr0py and GitHub
ab80b3dff4
[CI/Build] Bump PyNvVideoCodec version ( #47139 )
2026-06-30 06:38:46 -07:00
Qiming Zhang and GitHub
91055efd36
[XPU] C++ implementation for get_memory_info ( #47134 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-30 21:34:47 +08:00
Bugen Zhao and GitHub
3675bcff67
[Rust Frontend] Refactor TLS serve path with unified MaybeTlsListener ( #47101 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:31:58 +01:00
Bugen Zhao and GitHub
bdbd7278b6
[Rust Frontend] Extend renderer/parser roundtrip tests to support token ids ( #47110 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:27:45 +01:00
Harry Mellor and GitHub
5dc36a4fa5
[Model] Remove Tarsier, Tarsier2 ( #47143 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:20:33 +00:00
aab7af0bcb
[Bugfix][ROCm][MLA] Pass q/kv dtypes to get_mla_metadata_v1 in FP8 decode ( #46997 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-30 05:31:16 -07:00
536047755e
Bump actions/checkout from 6.0.1 to 7.0.0 ( #33057 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 13:16:20 +01:00
1907d3854a
[Bugfix] Reject negative values for max_logprobs and long_prefill_token_threshold ( #44002 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:01:03 +01:00
Chaojun Zhang and GitHub
ea9ddf59fc
[XPU][CI] Enable shared loader test ( #45977 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 11:20:33 +00:00
8cf7c4d8ad
[Attention Backend] add HPC-Ops Attention backend ( #46020 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 18:17:43 +08:00
8e9d70fdd5
[Kernel][XPU] Adjust kernel unit tests for XPU ( #45140 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-30 09:57:27 +00:00
Juan Pérez de Algaba and GitHub
364ee36af1
fix(security): prevent image decompression bomb OOM denial of service ( #47010 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-30 09:39:22 +00:00
Nicolò Lucchesi and GitHub
06fae69114
[Misc] Mistral label alert ( #47132 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 09:02:07 +00:00
14f8660a18
[CI/Build] Add CPU test dependency pre-commit hooks ( #47032 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 07:59:13 +00:00
aed541def4
[Bugfix][Responses] Set completed status for Harmony function calls ( #46945 )
...
Signed-off-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-30 07:55:14 +00:00
2bc20e8aba
[Frontend] Add Streaming Parser Engine and new Kimi k2.5/k2.6/k2.7 Parser ( #46610 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:53:17 +00:00
Chaojun Zhang and GitHub
8cc242335d
[XPU] Optimize XPU worker shutdown logic to prevent resource leak ( #46433 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 15:27:21 +08:00
Andreas Karatzas and GitHub
ba22cb6765
[ROCm][Ray][CI] Keep assigned GPU visible for weight transfer ( #47000 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:59:18 +08:00
Uros Markovic and GitHub
81bcced482
[Bugfix][ROCm] Preserve MoE weight padding for unquantized Triton path ( #46381 )
...
Signed-off-by: Uros Markovic <umarkovi@amd.com >
2026-06-30 14:47:57 +08:00
Kunshang Ji and GitHub
fb42e5219e
[Platform] Replace torch.cuda.mem_get_info with torch.accelerator.get_memory_info ( #44825 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-30 14:39:52 +08:00
Dakai An and GitHub
0feca7ffa8
PD disagg with Mooncake Connector: GDN support (Qwen3.5) and MLA support (Deepseek-V4-Flash) ( #46807 )
2026-06-29 23:29:04 -07:00
97b5ce5c39
[Bugfix] Raise VLLMValidationError for non-integer logit_bias keys ( #46612 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-30 06:18:59 +00:00
Andreas Karatzas and GitHub
4236514098
[ROCm][CI][Multimodal] Use ROCm-aware FA availability check for Unlimited-OCR ( #47004 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:03:13 +08:00
Blas Rodriguez Irizar and GitHub
e45c8a9f4b
[Rust Frontend] Start current wave for a stale DP FirstRequest ( #46833 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-30 05:13:09 +00:00
Wei Zhao and GitHub
b153dd3f28
[Bugfix] Use larger workspace size for Flashinfer MLA LSE ( #47074 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-29 22:11:03 -07:00
Reid and GitHub
930f8dc0a1
[Bugfix][Rust Frontend] Reject prompt_logprobs for streaming generate ( #46839 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 05:10:07 +00:00
Reid and GitHub
a16dbd5b85
[Rust Frontend] Avoid LoRA registry scans without active LoRA requests ( #47040 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 04:58:19 +00:00
bec232a914
Secondary tier implementation for PD disaggregation ( #42285 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 07:51:44 +03:00
b5c9e1ac33
[LoRA] Add language-backbone LoRA support for MiniCPM-V 4.6 ( #46740 )
...
Signed-off-by: linitra24 <Joy25810@foxmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-30 04:19:31 +00:00
ae2c4f3db7
[XPU][UT]Fix xpu pass_config.fuse_norm_quant assert issue ( #46804 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 21:13:44 -07:00
ganesh and GitHub
fca432e60a
[Bugfix] Propagate default stop_token_ids to per-request SamplingParams ( #35076 )
...
Signed-off-by: sriganesh123 <arjulasriganesh@gmail.com >
2026-06-30 12:10:09 +08:00
af1ee8c475
fix(config): reject negative max_logprobs (except -1) and long_prefill_token_threshold ( #44070 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 04:02:36 +00:00
5b4cb69523
[Bugfix][MLA] Fix LSE log-base mismatch in DCP + FlashInfer MLA decode ( #47079 )
...
Signed-off-by: girasoley <girasoleyang@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 19:15:02 -07:00
9fc0c08026
[ROCm][CI] Make tests/v1/shutdown an importable package ( #47085 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 21:01:27 -05:00
f2b5fabb23
[ROCm][CI] Move LM Eval Large Models (8 GPUs) to mi300 pool ( #47094 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 20:59:08 -05:00
b8cb75b149
[Rust Frontend] Add static HTTPS and mTLS support for HTTP and gRPC ( #45890 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 01:45:59 +00:00
Thien Tran and GitHub
43916891b2
[GDN] Improve kkt kernel of CuteDSL prefill backend ( #46346 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-29 18:34:18 -07:00
cda05ee8c4
[Bugfix][Reasoning] Fix thinking_token_budget not enforced on re-entry after forced end ( #43757 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-06-30 01:04:25 +00:00
weishu and GitHub
77654d080c
[KVTransfer] MultiConnector: merge kv_transfer_params dicts across connectors ( #46777 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-06-30 00:25:05 +00:00
Wentao Ye and GitHub
75698e60b3
[Bug] Fix sparse attention issue for GLM5.2 non-torch compile path ( #47083 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-29 15:45:53 -07:00
Andreas Karatzas and GitHub
8632c884dc
[ROCm][CI] Use spawn around the threaded OTLP test ( #47003 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 16:34:05 -05:00
c3734e8334
[CI][Bugfix] Add cohere_melody to ROCm test requirements ( #47072 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 16:29:47 -05:00
53f7553f09
[ROCm][DeepEP] Stabilize high-throughput DBO for DP+EP ( #46990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-29 14:28:02 -07:00
4eb227992a
[ROCm][CI] Make memory sampling less racy in tests and sleep mode ( #45490 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-29 14:26:41 -07:00
Micah Williamson and GitHub
ebcf511ec3
[ROCm][CI] Soft Fail Spec Decode Ngram + Suffix and Entrypoints Integration (LLM) AMD Mirrors ( #47067 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-29 16:24:08 -05:00
Matthew Bonanni and GitHub
8fc1b2d046
Fix FA4 dynamic_causal for full attention layers ( #46659 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-29 14:23:34 -07:00
Harry Mellor and GitHub
5316638a5e
Fix transient dependency issues caused by requirements/common.txt ( #47015 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 14:20:33 -07:00
zhrrr and GitHub
61ab70ec3b
[Model Runner V2] support mamba hybrid models align prefix cache ( #42406 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-29 14:09:16 -07:00
Woosuk Kwon and GitHub
a309d4fe60
Support DCP with FlashInfer MLA ( #43729 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-29 13:24:29 -07:00
72f639927f
[XPU] [RMSNorm] revert weightless change on xpu ( #46987 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 19:03:06 +00:00
Nick Hill and GitHub
8ad4a01825
[ModelRunner V2] Simplify recent UnlimitedOCR-related changes ( #46975 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-29 09:56:17 -07:00
Jee Jee Li and GitHub
7be582697b
[Bugfix] Fix DeepseekV2Model hidden_size ( #46986 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-29 16:44:05 +00:00
030c9523bd
[Perf][1/N] Expand Triton kernel warmup coverage, DSv4 ( #46634 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-29 16:40:34 +00:00
4708292d48
Bump flashinfer version to 0.6.13 ( #46683 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-29 09:30:57 -07:00
debec6440b
Add MiniMax-M3 modelopt nvfp4 support ( #46756 )
...
Signed-off-by: Xin Li <xinli@nvidia.com >
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
Co-authored-by: Xin Li <xinli@nvidia.com >
2026-06-29 09:29:39 -07:00
c8fb2963bd
[FS-Offloading] Batch Lookup in C ( #46713 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 09:28:32 -07:00
HDCharles and GitHub
379acd4e4f
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression ( #46860 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-06-29 15:55:42 +00:00
Martin Hickey and GitHub
07d33e575b
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager ( #44657 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 16:42:35 +01:00
36bbecd643
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" ( #46958 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 07:54:34 -07:00
Nicolò Lucchesi and GitHub
6149187a4c
[Kernel] Triton MLA logits workspace ( #46819 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-29 07:54:29 -07:00
Xiaohong (Sean) Chen and GitHub
49e28e8e91
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope ( #44010 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-06-29 22:54:15 +08:00
0ca39c4f1f
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone ( #46973 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-29 10:00:31 -04:00
Blas Rodriguez Irizar and GitHub
6185d73882
[Rust Frontend] Keep literal "null" string for string-typed tool params ( #46827 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-29 13:46:33 +00:00
bc8481af09
[MoE Refactor] Standardize Humming MoE experts + utilities ( #43373 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-29 06:19:29 -07:00
59575da46d
[XPU] exclude unsupported models for test_tensor_sechma.py ( #47008 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 12:30:28 +00:00
wang.yuqi and GitHub
3483240b7e
[Frontend] Consolidate scale out entrypoints ( #44512 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 03:18:53 -07:00
Roberto L. Castro and GitHub
eddfd4cf21
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen ( #46750 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-29 10:10:07 +00:00
Martin Hickey and GitHub
a4e3cb40d0
[mypy] Enable mypy for tests directory ( #47018 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 09:29:09 +00:00
soaringk and GitHub
ab132ee98b
Fix model info cache for package models ( #46567 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-29 09:17:54 +00:00
e186107870
[Bugfix] Use native SiLU activation in CPU fused MoE ( #45961 )
...
Signed-off-by: Alden Lobo <alden.lobo@arm.com >
Co-authored-by: Alden Lobo <alden.lobo@arm.com >
2026-06-29 09:12:20 +00:00
0e207dac78
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models ( #46835 )
...
Signed-off-by: John Langford <jl@hunch.net >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:59:15 +00:00
wang.yuqi and GitHub
9e86352c60
[CI Failure] Add transformers version check for openai/privacy-filter ( #47011 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 08:57:26 +00:00
Harry Mellor and GitHub
5051698e41
Remove unnecessary load_weights methods ( #44589 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 01:52:23 -07:00
Andreas Karatzas and GitHub
db28ae2d07
[ROCm][CI] Explicitly tear down multimodal offline LLMs ( #46999 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 07:59:24 +00:00
Harry Mellor and GitHub
f6bb8682ee
Fix docs on main ( #47009 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 15:50:57 +08:00
4559c43a95
[MM][CG] Gemma3 Encoder CUDA Graph ( #43591 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-29 04:52:00 +00:00
Bugen Zhao and GitHub
5274c1181d
[Rust Frontend] Add Harmony Renderer for GPT-OSS ( #46800 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-29 03:39:04 +00:00
Yuwen Zhou and GitHub
58d6a6e60a
[CPU] Support cpu compressed-tensor w8a8 int8 moe ( #42920 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-06-29 03:04:05 +00:00
a2abce646f
[EPLB] Mask padding in EPLB load recording ( #38128 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-28 19:43:58 -07:00
Harry Mellor and GitHub
311ad689ad
Remove boilerplate missed by #46820 ( #46956 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:11:17 +08:00
Woosuk Kwon and GitHub
0472436541
[Spec Decode] Avoid redundant hidden-states gather in draft prefill ( #46968 )
2026-06-28 17:04:01 -07:00
4dfbf1503b
[Model] Add support for openai/privacy-filter ( #41026 )
...
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-28 16:18:22 -07:00
Wei Zhao and GitHub
95528527ea
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 ( #46855 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-28 14:36:23 -07:00
c2127a25c7
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm ( #46895 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:50:30 -05:00
03c6d01c30
[OCP MX ] Add back emulation to available OCP MX backends list ( #46629 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:43:19 -05:00
Woosuk Kwon and GitHub
4b643c463e
[GLM5] Fix minor typo ( #46961 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-28 08:37:00 -07:00
7544286b04
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE ( #46552 )
...
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-28 15:19:28 +00:00
Woosuk Kwon and GitHub
89876b0c54
[GLM5] Implement op fusion for GLM5/DSV3.2 ( #46876 )
2026-06-28 08:17:39 -07:00
Wentao Ye and GitHub
5c91039c41
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement ( #46635 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-28 14:55:54 +00:00
5ecae3266c
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) ( #45033 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-28 07:52:00 -07:00
6eb63a1da6
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers ( #46600 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 01:37:44 -07:00
09841ae705
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation ( #46846 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-28 00:07:33 -07:00
Matt and GitHub
a2a92cbbaa
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection ( #46930 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-28 00:07:14 -07:00
35e6c86caa
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL ( #46034 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-28 00:06:43 -07:00
c7ca0bccae
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 ( #44313 )
...
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com >
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com >
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
2026-06-28 00:04:08 -07:00
c6741b2ad4
[Model] Support Unlimited OCR ( #46564 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-27 23:09:18 -07:00
a65f93fb2e
[ROCm][CI] Add ci_base metadata for external cache orchestration ( #46886 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-28 12:51:19 +08:00
Chauncey and GitHub
11a12305c0
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models ( #46786 )
2026-06-27 18:38:07 -07:00
798185d438
[KV-Offloading] Fix tensors_per_block stride ( #46888 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-27 21:01:45 -04:00
Matt and GitHub
9036c89ee4
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now ( #46928 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-27 17:30:49 -05:00
Giancarlo Delfin and GitHub
b6caeb5a09
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance ( #46878 )
2026-06-27 14:09:25 -07:00
Taneem Ibrahim and GitHub
8bf064f8d3
Fixed chunked embedding aggregation with request-id metadata ( #46782 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-27 20:57:47 +00:00
ea2ead1db3
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod ( #46818 )
...
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
2026-06-27 20:23:59 +00:00
Wentao Ye and GitHub
56aa067bf0
[CI Bug] Fix h100 AssertionError: Cold-start child failed ( #46927 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-27 20:17:33 +00:00
xiaolinchen and GitHub
35e3850fa9
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) ( #46915 )
...
Signed-off-by: wentian-byte <2990624738@qq.com >
2026-06-27 14:30:10 -04:00
51a99565c3
[ROCm][Perf] Fused shared expert for Minimax M3 ( #46474 )
...
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-27 12:34:17 +00:00
867fd5e8ed
[ROCm][Perf] Use flydsl moe with Minimax-M3 mxfp8 weights on gfx950 and implemented moe-backend selection ( #46184 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
2026-06-27 10:22:57 +00:00
9fd00ee006
[ROCm][CI] Move remaining mi250_2 tests out of the MI250 queue ( #46905 )
...
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-27 17:08:54 +08:00
091d13976c
[ROCm][CI] Add TRITON_ATTN score absolute tolerance floor ( #46891 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-27 06:35:50 +00:00
Wentao Ye and GitHub
b588f66dc2
[GLM5.2 Perf] fused_indexer_q_rope_quant triton kernel, 1.9% ~ 3.3% E2E Throughput improvement. ( #46862 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-26 22:16:20 -07:00
Benjamin Chislett and GitHub
455f25aa13
[CLI] Add flag to print TTFT and TPS in vllm chat ( #46775 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-26 22:15:10 -07:00
d706dec904
fix: Correct reasoning-end detection for prompt history ( #44551 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Signed-off-by: Jason Ozuzu <jasonozuzu@cohere.com >
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: Walter Beller-Morales <walterbm@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-26 22:15:06 -07:00
Divakar Verma and GitHub
68ee8300a0
[ROCm][CI]Fix test_concat_and_cache_mla_rope_fused on ROCm ( #46409 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 12:38:13 +08:00
ddd3855a28
[MoE Backend] add HPC-Ops MoE backend ( #45924 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: youkaichao <youkaichao@gmail.com >
2026-06-27 11:18:07 +08:00
Divakar Verma and GitHub
00e045b7c7
[ROCm][CI TG] refactor and fix deepep_moe test group ( #46758 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:45:23 +08:00
Divakar Verma and GitHub
17a71d8702
[ROCm][CI] Relax fused layernorm quant test tolerances for one-ULP outliers ( #46658 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:44:29 +08:00
weizhoublue and GitHub
2e058851d3
fix(docker): eliminate race conditions in shared buildkit cache mounts ( #44984 )
2026-06-26 19:43:17 -07:00
Dāvis and GitHub
1a92dfcce4
[Build] Show error message when using ROCm with LTO and different compilers ( #35232 )
2026-06-26 19:43:00 -07:00
Chris Leonard and GitHub
d0f800811b
[Build] Update vllm to point to vllm-project/flash-attention commit that builds FA3 with torch stable API. ( #46644 )
2026-06-26 19:42:46 -07:00
Nick Hill and GitHub
c6dd32a810
[ModelRunner V2] Support realtime embeddings ( #46762 )
2026-06-26 19:42:27 -07:00
af16446bf3
Vram semaphore infra ( #44465 )
...
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-26 17:32:51 -07:00
Harry Mellor and GitHub
3f67477497
[CI] Don't try and download files that we already know don't exist ( #46854 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-26 23:56:39 +00:00
Nick Hill and GitHub
1d41009e81
[ModelRunner V2] Fix cross-attention block table sizing ( #46753 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:34:21 -07:00
Nick Hill and GitHub
b94f212e37
[ModelRunner V2] Deduplicate ModelState init logic ( #46776 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:32:45 -07:00
Harry Mellor and GitHub
d8eb734d94
Fix Transformers backend FP8 MoE and remove some boilerplate ( #46820 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-27 00:16:05 +01:00
2ff76a5e85
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 ( #46760 )
...
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-26 21:58:40 +00:00
Yifan Qiao and GitHub
75fdcc82a5
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas ( #46873 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-26 21:48:53 +00:00
yzong-rh and GitHub
77f8796d16
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. ( #46437 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-26 17:18:47 -04:00
c40d307731
[Core] Remove FlashAttention block size restriction for hybrid models ( #36701 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-26 21:16:39 +00:00
Woosuk Kwon and GitHub
65e655d295
[GLM-5] Add DSV3.2/GLM5 to vllm/models/ ( #46808 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-26 14:09:05 -07:00
Charlie Fu and GitHub
6e2fb02fe5
[ROCm][CI] Fix rlhf_nccl.py on ROCm ( #46851 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-26 15:41:49 -05:00
Micah Williamson and GitHub
274325dd43
[ROCm][CI] Remove V1 Sample + Logits from mi250 Queue ( #46867 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-26 15:38:38 -05:00
Matt and GitHub
95e6442a6b
[Hardware][AMD][CI] Fix Kernels Quantization test timeout ( #46859 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-26 15:19:16 -05:00
701a23d99f
[Bugfix][Model] Support tensor parallelism for DiffusionGemma ( #45719 ) ( #46177 )
...
Signed-off-by: Carlos Alvarado <carlos-alvarado@outlook.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-26 20:05:04 +00:00
Ben Browning and GitHub
dccb412e2c
[Bugfix][Parser] Pass token IDs to parser.parse() in Responses API and batch serving ( #46843 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-26 19:29:52 +00:00
c6554f321c
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build ( #46769 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 14:32:21 -04:00
Julien Denize and GitHub
3d3b96488f
Migrate Voxtral to mistral-common 1.11.5 audio API ( #46705 )
...
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-06-26 11:06:31 -07:00
Nick Hill and GitHub
658b54efe4
[ModelRunner V2] Update scheduler tests to cover MRV2 paths ( #46771 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 09:36:31 -07:00
Li, Jiang and GitHub
abc71548ef
[CI/Build][CPU] Add test image cache clean-up ( #46831 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-26 23:28:49 +08:00
Nick Hill and GitHub
4e07ca2c92
[Core] Add VLLM_GPU_SYNC_CHECK env var ( #44800 )
2026-06-26 08:24:33 -07:00
Bugen Zhao and GitHub
e71bc6da85
[Rust Frontend] Use oss-harmony for Harmony output processing ( #46799 )
2026-06-26 08:24:13 -07:00
fxmarty-amd and GitHub
37ce34922f
[CI] Fix failing CUDA graph capture in Triton MOE ( #46735 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-26 07:21:20 -07:00
c2507fb293
[ROCm] [MoE] [Perf] Shared-expert fusion for bias-routed MoE; enable on MiniMax-M3 mxfp8 model ( #46545 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 07:05:20 -07:00
TJian and GitHub
8921c4be88
[ROCm] [Performance] Optimize aiter moe for DeepSeekV4 ( #46122 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 06:43:27 -07:00
8e394244a5
[ROCm]Enable AITER MoE backend for MiniMax-M3-MXFP4 ( #46419 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-26 06:35:35 -07:00
TJian and GitHub
302954e5f6
[ROCm] [CI] fix transcription flakiness AMD: Entrypoints Integration (API Server OpenAI - Part 1) (mi325_1) ( #46823 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 21:33:35 +08:00
Hyunkyun Moon and GitHub
950ee4c2e4
[API] Add token offsets to render endpoints (/v1/.../render) ( #44226 )
...
Signed-off-by: HyunKyun Moon <mhg5303@gmail.com >
2026-06-26 05:02:52 -07:00
d980a3cc6e
[ROCm] Fix AITER_UNIFIED_ATTN Dispatching After AITER Bump ( #46780 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-26 02:09:56 -07:00
bf292b5f6b
[Docs] Remove BambaForCausalLM from supported hybrid models list ( #46071 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-26 08:02:50 +00:00
wang.yuqi and GitHub
5e3dad04b1
[Misc] Move the legacy api_server.py to the examples directory. ( #46783 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-26 07:43:29 +00:00
Joe Rowell and GitHub
63e161f296
[Bugfix][Tool Parser] PoolsideV1: fix string whitespace and required named tool choice ( #46486 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-06-26 06:05:16 +00:00
Tiezhen WANG and GitHub
c7645bce04
Remove grok model arch from vllm ( #46706 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
2026-06-25 23:02:10 -07:00
35a49fcfc2
[CI][Bugfix] Spawn engine in mm cache sleep test to fix ROCm HIP error ( #46749 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-26 00:38:26 -05:00
peizhang56 and GitHub
915e99ec67
[ROCm][Bugfix] Fix HIP fork re-init in multimodal offline examples ( #46741 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
2026-06-26 00:37:47 -05:00
Nick Hill and GitHub
5b33041746
[ModelRunner V2] Fix whisper test ( #46773 )
2026-06-25 22:10:36 -07:00
Matt and GitHub
1a4984520e
[Hardware][AMD][CI] Fix AMD CI image build ( #46792 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 22:05:12 -07:00
Reid and GitHub
e312c5cb25
[Rust Frontend] Make Granite4 string argument scanning incremental ( #46507 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-26 03:54:03 +00:00
Matti4 and GitHub
1502cf6274
Fix relative allowed local media paths ( #45263 )
2026-06-25 20:45:20 -07:00
d350fa8ddd
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens ( #46733 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-26 03:41:33 +00:00
dbc49b6b99
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 ( #45166 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
Signed-off-by: ovidiusm <ovidium@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-25 19:33:42 -07:00
fxmarty-amd and GitHub
552a9dbe59
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears ( #44667 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-25 19:33:00 -07:00
02a1f23711
[DFlash] Fuse precompute kv per-layer rmsnorms ( #46761 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:32:07 -07:00
652d962bc9
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation ( #46448 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:30:07 -07:00
Giancarlo Delfin and GitHub
5314665bad
[Model Runner V2][DFlash] Enable dflash attention backend selection ( #46770 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 19:29:25 -07:00
Michael Goin and GitHub
3daea7ceb9
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models ( #46759 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 19:03:09 -07:00
Wentao Ye and GitHub
cc7981599e
[Refactor] Remove dead kernel code ( #46405 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 18:09:56 -07:00
Nick Hill and GitHub
32bb3195f0
[ModelRunner V2] Bound memory for large logprobs requests ( #46746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-25 18:04:06 -07:00
ad28d605e6
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) ( #45544 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-25 17:46:28 -07:00
Bugen Zhao and GitHub
ae7c8ec223
[Rust Frontend] Switch rustls to native-tls/OpenSSL ( #46696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:19:44 -07:00
Bugen Zhao and GitHub
1d3f4cb3a4
[Rust Frontend] Extract renderer fixture test utilities ( #46719 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:12:38 -07:00
Bugen Zhao and GitHub
f9e684499f
[Rust Frontend] Migrate gemma4 to unified parser ( #46602 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 16:59:57 -07:00
Giancarlo Delfin and GitHub
c53994e134
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling ( #46665 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 23:46:10 +00:00
Matt and GitHub
27da2a2ac4
[Hardware][AMD][CI] Use Triton-based AITER MHA for LM Eval Qwen-3.5 Models Tests ( #46691 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 17:08:04 -05:00
Michael Goin and GitHub
a2e8ec3d52
[CI] Depend GPQA Eval DGX Spark job on arm64 image build ( #46736 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 17:07:04 -04:00
e8c24a7695
[Kernel] Vectorized fp32 moe_sum reduction and support any topk ( #46643 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 14:02:28 -07:00
Andreas Karatzas and GitHub
2a6f8f0c05
[ROCm][CI] Fine-tuning queues and test names ( #39238 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-25 13:24:09 -07:00
Robert Shaw and GitHub
c5e3c40877
Fix P/D with DP Supervisor ( #46628 )
...
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-25 13:13:08 -07:00
Wentao Ye and GitHub
8b4d93ba2b
[Perf] Remove redundant clone for GLM, Deepseek etc ( #46651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 13:09:00 -07:00
Michael Goin and GitHub
e8e7b592d1
[Kernel][MoE] Tune block-FP8 fused MoE for low-batch decode ( #46642 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 12:38:28 -07:00
Rohan Potdar and GitHub
e53a17232c
[ROCm]: Bump aiter to 0.1.16.post2 ( #46692 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-25 11:53:26 -07:00
Flora Feng and GitHub
96eb8ddc41
[CI] Re-enable skipped glm and seedoss parser tests ( #46671 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-25 13:11:41 -04:00
Gabriel Wu and GitHub
8fa36fbbeb
[Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 ( #46506 )
2026-06-25 09:12:00 -07:00
Ranran and GitHub
e45b279928
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3Next ( #46316 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
2026-06-25 09:05:04 -07:00
d490b98162
[Core] Avoid mixed length specdec batches via padding ( #45237 )
...
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
2026-06-25 08:34:44 -07:00
haoyangli0109 and GitHub
1744adc256
[ROCM] [Communication] Add INT3 quantization method for quickreduce ( #45666 )
...
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com >
2026-06-25 15:14:15 +00:00
Divakar Verma and GitHub
cdfa2fd7e9
[ROCm][CI] rm duplicate Distributed Torchrun ci test ( #46729 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-25 09:58:17 -05:00
6f3da461d1
[Pooling] Fix Cohere embed billed image token accounting for mixed-content inputs ( #46093 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 10:44:29 -04:00
Russell Bryant and GitHub
d3130d878c
[CI] Pin GitHub Actions to commit hashes in macos-smoke-test.yml ( #38290 )
2026-06-25 13:48:44 +00:00
9bfd878a48
[MoE] [MoE Refactor] Add moe kernel oracle abc 37753 ( #43461 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Signed-off-by: qyYue1389 <yueqiuyang1389@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 09:34:03 -04:00
Matt and GitHub
2365b7a8e7
[Hardware][AMD][CI] Mirror Basic Models (Others) and Weight Loading Multiple GPU test groups ( #46668 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 08:25:09 -05:00
15be78732b
[NIXL][Mamba] Add Mamba1 support to NIXL P/D disaggregation ( #45019 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 05:50:41 -07:00
92221485aa
[CPU][CI/Build] Allow more CPU CI agents ( #46702 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:33:39 +08:00
xiangdong and GitHub
a6f41ab678
[XPU][CI]Refine .buildkite/ci_config_intel.yaml for Intel GPU CI ( #46674 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-25 08:58:26 +00:00
c63cd4906c
[ROCm][ [Perf] sparse attention optimization on minimax-m3 ( #46546 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: yueliu14 <yue.liu4@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-25 16:56:00 +08:00
638b1a99cc
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM ( #45269 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-25 08:18:10 +00:00
72adb20a6a
[Model] Remove AquilaForCausalLM, AquilaModel ( #46605 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 08:08:26 +00:00
2396d91e93
[CPU][Spec Decode] Enable DFlash SD for CPU ( #44029 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 15:32:48 +08:00
9b215ae60b
[Rust Frontend] Forward VLLM_ENGINE_READY_TIMEOUT_S via --args-json ( #44610 )
...
Signed-off-by: kai <kai@example.com >
Co-authored-by: 图灵 <tuling.wk@alibaba-inc.com >
2026-06-25 07:25:08 +00:00
Bugen Zhao and GitHub
4d3b4b9b01
[Rust Frontend] Make ToolParserOutput a seq of ToolParserEvent to preserve order ( #46584 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 06:27:07 +00:00
Matthias Gehre and GitHub
77c1d9fe9b
[ROCm][Perf] Tune wvSplitK on gfx1151 ( #40784 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-06-25 14:17:46 +08:00
Jeff (Junze) Ma and GitHub
36fd7e8b86
[SimpleCPUOffloadConnector] Fix remaining global→block conversions under PCP/DCP ( #46394 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
2026-06-24 23:05:24 -07:00
fc61c6fc26
[Perf] Enable + tune FlashInfer fused allreduce at world_size=16 on SM 10.3 (GB300) ( #46392 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:04:17 -07:00
Matt and GitHub
e2af449c39
[Hardware][AMD][CI] Move Metrics, Tracing (2 GPUs) & make optional ( #46686 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 05:49:33 +00:00
3f5a1e1733
[ROCm][CI] Expand basic correctness target suites ( #46573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matt <156021403+mawong-amd@users.noreply.github.com >
2026-06-25 12:18:57 +08:00
710ebaa189
[ROCm][Bugfix] Fix chunk alignment when using context parallelism with TRITON_MLA ( #46114 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:07:28 -04:00
1aad125815
[CPU] Enable chunked prefill and prefix caching for qwen3.5 ( #46202 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-25 03:49:21 +00:00
dc55936f64
[AMD][CI] Fix Pipeline + Context Parallelism test group ( #46650 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 22:23:42 -05:00
Bugen Zhao and GitHub
76c3c4ff63
[Rust Frontend] Introduce unified parser interface & combined parser ( #46583 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 03:17:31 +00:00
efb5acffd5
[Bugfix] fix: stream Mimimax m2 tool call string arguments ( #46382 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-25 03:12:45 +00:00
6e3a983cf3
[ROCm] Remove erroneous inclusion of gptq_marlin as supported quant scheme on ROCm ( #46655 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 21:27:19 -05:00
Xin Yang and GitHub
1273a8f05a
[Kernel] Add swap AB optimization to fused_moe_kernel ( #36559 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-06-25 01:44:30 +00:00
9e88e969c0
[Perf][KVConnector][Mooncake] Parallelize KV load with a receive-thread pool ( #45971 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 18:25:12 -07:00
dda3aca47f
[Speculative Decoding] Propagate norm_output and fc_norm config for Eagle3 speculators ( #46488 )
...
Signed-off-by: Orestis Zambounis <orestis.zambounis@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:51:33 +00:00
Jee Jee Li and GitHub
23aed9b0ee
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel ( #46508 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-25 08:42:51 +08:00
Maxwill Lin and GitHub
cd347298e8
[Frontend] Port seed_oss to the streaming parser engine as a Qwen3 subclass ( #46314 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 20:08:42 -04:00
Yifan Qiao and GitHub
b69816043a
[Bugfix][MooncakeStore] track resumed requests via scheduler's resumed_req_ids ( #46595 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 23:50:56 +00:00
Kaihang Jiang and GitHub
fc7fc421e9
[Kernel][MoE] Allow FlashInfer MXINT4 MoE for gated SiLU ( #46518 )
...
Signed-off-by: Kaihang Jiang <kaihangj@nvidia.com >
2026-06-24 18:32:50 -05:00
cyq and GitHub
e06a83445c
[Bugfix] Normalize slashes in Helion GPU names ( #46101 )
...
Signed-off-by: cyq <15000851237@163.com >
2026-06-24 18:22:49 -05:00
d7ab9be775
[Bugfix] Support -1 (invalid/non-local) slots in topk_ids for Triton MoE ( #46408 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:59:42 -07:00
6a1570711c
[Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing ( #46406 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:52:09 -07:00
Micah Williamson and GitHub
d6696e2385
[ROCm] Begin Deprecation Window for CUDA_VISIBLE_DEVICES on ROCm ( #46636 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-24 20:40:28 +00:00
Chauncey and GitHub
84c2f9f0fb
[Frontend] Fix Kimi K2 tool call IDs for required tool choice ( #46344 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-24 19:59:40 +00:00
49f2104c53
[Feature] Support DCP with FP8 KV cache in MLA decode path ( #44044 )
...
Signed-off-by: shivampr <shivampr.dev@gmail.com >
Signed-off-by: Shivam <shivamprasad91@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 19:28:17 +00:00
d511b5bae9
Chore: Fix minor doc sentence, grammar, quote errors ( #40469 )
...
Signed-off-by: Ashwin Phadke <23502062+ashwin-phadke@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-24 18:58:24 +00:00
3c43237233
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels ( #46560 )
...
Signed-off-by: xiaojun.wei <jessiewei747@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-24 11:00:57 -07:00
56ca5997ea
Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference ( #46389 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-24 13:53:54 -04:00
Aarushi Jain and GitHub
cf57311187
Run DeepSeek-V2-Lite prefetch-offload eval eager on ROCm ( #46386 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-24 12:25:57 -05:00
Lucas Wilkinson and GitHub
e7df232288
[KV Offload] Gate packed HMA KV cache on cross-layer config ( #46252 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-06-24 11:55:30 -04:00
b3a688cb9e
[ROCm] Fix OOB During Model Warmup With ROCM_ATTN and MRV2 ( #46548 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-24 10:53:21 -05:00
Wentao Ye and GitHub
1cd3e0e945
[Bug] Fix IndentationError: expected an indented block after 'with' statement ( #46627 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-24 23:14:17 +08:00
Yiwei Hu and GitHub
f889325c51
[KV Offload] Use background thread for mmap / cpu_tensors pinning ( #45850 )
...
Signed-off-by: Sorryhorizon <arikara6666@gmail.com >
2026-06-24 18:13:27 +03:00
bb61177e49
[KV Offloading] Replace bool|None lookup return with LookupResult enum ( #46363 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 18:06:08 +03:00
7f99e80c3b
[Perf][ThinkingBudget] reduce search space for thinking tokens ( #46425 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:02:25 +08:00
2801b11156
[Test] Pin block_size in auto-fit max_model_len test ( #45914 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:56:21 +08:00
007b5a52ed
[Log] Update to log once ( #46511 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-24 14:45:16 +00:00
Cyrus Leung and GitHub
24d5186138
[Bugfix] Re-enable FP8 MoE on NVIDIA Thor ( #46339 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-24 07:35:46 -07:00
Nemani Harsha Vardhan and GitHub
7dc036058b
[Doc] Document Qwen3.6 (dense + MoE) ViT CUDA graph support ( #44720 )
...
Signed-off-by: harsha20032020 <nhvardhan2020@gmail.com >
2026-06-24 14:35:08 +00:00
61ee183d28
[ROCm] Fix AITER FP8 quantization schema tests ( #46414 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:29:19 +08:00
84c62e1cbd
[Model Runner V2][MM] Support EVS ( #46535 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:18:56 -04:00
Fadi Arafeh and GitHub
061043eaca
[CPU][Perf] Accelerate unquantized MoE for AArch64 ( #46353 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-24 14:14:35 +00:00
93ec645878
[Bugfix] Fix illegal memory access from a forward during a partial wake_up ( #44483 )
...
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 22:12:23 +08:00
Kunshang Ji and GitHub
563c628968
[XPU] bump up vllm_xpu_kernels to v0.1.10.1 ( #46607 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 10:05:31 -04:00
0bc479e6eb
[Perf][LoRA] Replace O(n) list.index() with a dict in convert_mapping ( #46542 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 21:41:46 +08:00
Tae Jeong and GitHub
62890e204c
Fix duplicated logging when loading a corrupt or partial video ( #46467 )
...
Signed-off-by: hhhhhhhhhhhhhhhhho <man2719@naver.com >
2026-06-24 06:14:13 -07:00
Nicolò Lucchesi and GitHub
a2cb08b3d5
[Misc][PD] Disable bidirectional xfer mode for NixlPushConnector ( #46473 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-24 21:14:05 +08:00
cf9fd6457e
Fix KV offload request-finished lifecycle contract ( #46284 )
...
Signed-off-by: test test <2260891073@qq.com >
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 15:42:40 +03:00
Kunshang Ji and GitHub
d4448b511d
[XPU][Docker] switch to ubuntu 24.04 as base image ( #45973 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 20:39:20 +08:00
f1a6703edd
[Bugfix][Config] Keep pydantic validation for fields with a TYPE_CHECKING Literal alias ( #46220 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 12:25:50 +00:00
Roy Wang and GitHub
160c80a34c
[Rust Frontend] Raise frontend JSON body limit ( #46582 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-06-24 12:15:31 +00:00
Martin Hickey and GitHub
f237e16b41
[KV Offload] Replace OffloadingHandler with OffloadingWorker ( #45053 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-24 14:44:24 +03:00
70749fdcca
[Feature] Triton INT4 per-token-head KV cache quantization ( #40835 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:21:25 +00:00
d20dbf921b
[Mooncake] Only check and store new KV cache range ( #46412 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 03:10:50 -07:00
ede54b926e
set AttentionCGSupport.UNIFORM_BATCH for fa2 on xpu ( #46555 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 18:05:02 +08:00
52fbe12283
[Perf][Multimodal] Avoid building a full timestamps list in video frame sampling ( #46543 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 09:38:27 +00:00
Dakai An and GitHub
dc0d318177
[Attention] Add FLASH_ATTN_MLA_SPARSE backend for Hopper sparse MLA ( #46189 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-06-24 09:33:10 +00:00
soaringk and GitHub
d7c1821b5a
[Model][MiniMax-M3] Add pipeline parallelism support ( #45810 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-24 08:23:03 +00:00
4cd1a84c88
[Model] Remove BaiChuanForCausalLM and BaichuanForCausalLM ( #46362 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 16:13:57 +08:00
Mohammad Miadh Angkad and GitHub
191826ec61
[CI/Build] Fix topk histogram build on SM75 ( #46550 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-24 00:51:11 -07:00
Andreas Karatzas and GitHub
549c7074cd
[ROCm][CI] Skip the MoE Marlin tile-padding helper assertion ( #46580 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 07:31:33 +00:00
489abadfb8
feat: support to OpenMOSS-Team ( #44124 )
...
Signed-off-by: nagisa-kun <1434936049@qq.com >
Signed-off-by: nagisa19 <1434936049@qq.com >
Signed-off-by: nagisa <1434936049@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 00:08:13 -07:00
Woosuk Kwon and GitHub
96de8bb389
[MoE] Free unused MXFP4 scales in OAI Triton Backend ( #46549 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-24 00:06:41 -07:00
Jee Jee Li and GitHub
9d6fdc2901
[Kernel] GLM5 Router GEMM ( #46385 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:54:50 -07:00
Benjamin Chislett and GitHub
4c5bc41ba6
[Bugfix][Spec Decode] Fix probabilistic sampling for parallel drafting ( #45956 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-24 05:36:23 +00:00
Michał Ganczarenko and GitHub
ac1fa74616
[Bugfix] Fix NemotronLayerNorm1P hardcoded cuda device type ( #46495 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
2026-06-24 13:21:02 +08:00
Sting Lin and GitHub
556bc4e3a0
Upgrade tpu-inference to v0.23.0 ( #46568 )
2026-06-23 21:15:14 -07:00
Wei Zhao and GitHub
05a0caba91
[Mooncake] Optimize lookup pool key string construction ( #46188 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-24 11:51:53 +08:00
Nick Hill and GitHub
7ee4d22009
[Spec Decode] Reject placeholder (-1) draft tokens in rejection sampler ( #46533 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 03:32:32 +00:00
ce9f64020b
[Rust Frontend] Pass effective reasoning_parser_kwargs for structured output ( #46360 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 03:13:44 +00:00
4ed8eaafb0
[Rust Frontend] Integrate xgrammar-structural-tag for strict and required tool calling ( #46057 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 10:46:49 +08:00
6af0559ddb
[Core][DP] Throttle prefills based on local prefill work ( #46532 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 02:27:12 +00:00
e2bdc24612
[ROCm][Bugfix] Fix use_v2_model_runner inside Ray driver thread ( #45998 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 08:41:56 +08:00
Andreas Karatzas and GitHub
bcbeaac786
[ROCm][CI] Stage C-II of gating additional test groups ( #46537 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 17:36:40 -07:00
Maxwill Lin and GitHub
e48f2aa4ca
[Bugfix][Frontend] Emit a content block for empty Anthropic completions ( #46525 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 00:04:26 +00:00
Roberto L. Castro and GitHub
d86c66c981
[Feat] Add runtime monitor for post-warmup CuTeDSL compilation ( #46167 )
2026-06-23 23:33:17 +00:00
Nico Holmberg and GitHub
80e511772f
[ROCm][Bugfix][Perf] enable shared expert fusion for Qwen3.5 ( #44434 )
...
Signed-off-by: Nico Holmberg <nico.holmberg@amd.com >
2026-06-23 23:19:51 +00:00
Roberto L. Castro and GitHub
855cd4d787
[Perf][DSv4/DSv3.2] Add cluster-cooperative topK kernel for low-latency scenarios ( #43008 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-23 16:11:00 -07:00
3cc871aaf1
[Perf] Skip detokenization in online beam search ( #46422 )
...
Signed-off-by: Guy Stone <guys@spotify.com >
Signed-off-by: Guy Stone <guystone3@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 15:46:09 -07:00
0a3e2dbc09
[Optimization] Skip DP padding tokens in MoE ( #46428 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 14:54:46 -07:00
84f13374b3
[CI] Fix test_auto_gptq on ROCm CI ( #46164 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 16:38:06 -05:00
Micah Williamson and GitHub
b28103e1ca
[ROCm][CI] Shard LM Eval Qwen3-5 Models (B200-MI355) in AMD CI ( #46520 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 16:32:05 -05:00
Wentao Ye and GitHub
abc33134fa
[CI Test] Mark batch invariance test flaky ( #46530 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 21:01:34 +00:00
6617db1bfb
[Bugfix][Frontend] Emit non-ASCII tool-call arguments without \uXXXX escapes ( #46308 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-23 20:43:11 +00:00
899d72a58c
[Bugfix][ToolParser] Handle braces in required tool streaming strings ( #45389 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-23 20:29:34 +00:00
Yongye Zhu and GitHub
11b56b2ff2
[Kernel] Add FlashInferCutedslMxfp8LinearKernel (cute-dsl mm_mxfp8) ( #46393 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-23 12:45:49 -07:00
0d4d164488
[Bugfix] Allow flashinfer_cutlass as a clamped NVFP4 MoE backend ( #46492 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-23 12:43:36 -07:00
Mike G and GitHub
0775b882ba
[NVFP4 MoE/Deepseek V4] Marlin: wire SwiGLU clamp + allow it for clamped models on non-Blackwell ( #45836 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-23 12:21:19 -07:00
7c2e08451a
[Docker] Remove redundant flashinfer download-cubin step ( #46517 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 12:16:51 -07:00
Giancarlo Delfin and GitHub
ef361de916
[Model Runer V2][DFlash] Fix lm head sharing for dflash ( #46435 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-23 19:09:06 +00:00
Yan Ma and GitHub
acce57d8dd
Deprecate old FP8 online MoE quantization class ( #44514 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:53:38 -07:00
68afd78897
[Bugfix][ROCm] Fix cumem sleep and teardown ( #46203 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 02:45:31 +08:00
37a682d392
[Kernel] Extend Marlin thread-tile padding to MoE (WNA16 + FP8/MXFP8) ( #45703 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 11:45:10 -07:00
Rui Yin and GitHub
d8e422ccda
[Bugfix] Parse MiniMax M3 streaming reasoning by text markers ( #45718 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-23 14:43:58 -04:00
fxmarty-amd and GitHub
e368415daa
[AMD][OCP MX][CI] Fix tests to not dispatch on UNFUSED_TRITON backend on MI300, improve w_mxfp4_a_fp8 emulation support ( #46142 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-23 14:25:27 -04:00
Andreas Karatzas and GitHub
ceae5bcbda
[ROCm][CI] Fix nixl tests ( #45219 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 13:11:40 -05:00
6691f087a6
[Minimax-M3] BF16/FP8 Indexer using MSA ( #45892 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-23 10:28:49 -07:00
f4d5f73ffa
[Bugfix]: Fix unquantized gpt-oss weight loading broken by FusedMoE r… ( #45818 )
...
Signed-off-by: priyansh jain <priyansh.jain2@amd.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-23 16:56:17 +00:00
fd50a66015
[CI][ROCm] Skip unsupported test cases on ROCm ( #46160 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 11:35:49 -05:00
84586c9acc
[ROCm][CI] fix fp8 range in vit_fp8_quant ( #46410 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 11:34:21 -05:00
Taneem Ibrahim and GitHub
40e5522121
[Docs] Add Qwen3 forced alignment online example ( #46197 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-23 11:59:45 -04:00
Willow Lopez and GitHub
f3410b3bb1
fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility ( #45404 )
...
Signed-off-by: Oxygen <1391083091@qq.com >
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-23 11:46:23 -04:00
568874fec2
[ROCm][CI] pass merge-base to container for python-only wheel metadata ( #45869 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 15:44:43 +00:00
275b43183c
[MyPy] Fix mypy for vllm/benchmarks ( #39896 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-23 15:22:29 +00:00
Yan Ma and GitHub
547d2c40d7
Add weights padding for fp8 per-block online quantization ( #44763 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:08:17 -04:00
2aaaf3febd
[ROCm][Test] Fix stale test_gfx950_moe MXFP4 oracle tests ( #46260 )
...
Signed-off-by: Spandan Tiwari <23646532+spandantiwari@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 23:07:46 +08:00
Micah Williamson and GitHub
156b12667c
[ROCm][CI] Skip Quark mxfp4 tests unless Quark version is compatible with Torch version ( #46431 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 22:26:39 +08:00
Jee Jee Li and GitHub
9f6f296428
[CI/Build] Remove BaiChuanForCausalLM from the LoRA test ( #46494 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:09:48 +08:00
e51e700470
[LoRA] Gate all_gather on fully_sharded_loras inside _mcp_apply; rewrite regression test ( #45715 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-23 07:08:33 -07:00
f59db63732
[Bugfix] GPT-OSS Autodrop reasoning in Response API and cleanup ( #45048 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 09:36:33 -04:00
Rukhaiya2004 and GitHub
9f5117820f
[HARDWARE][POWER] Enable fp16 support for PowerPC ( #46135 )
...
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
2026-06-23 13:24:49 +00:00
1bf149f334
Filter Pydantic-internal markers from validation error param ( #46457 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Rehmann <114791389+mahadrehmann@users.noreply.github.com >
2026-06-23 13:20:50 +00:00
2a675a7b9f
[Bugfix] Responses API assistant EasyInputMessageParam input ( #44361 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 08:54:46 -04:00
7d47cff933
[Bugfix][KV Offload] Fix swap_blocks_batch on the default stream ( #46379 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-23 05:45:27 -07:00
091bc1026e
[KV Offloading] Add tiering metric plumbing ( #45959 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-23 15:10:36 +03:00
3554ada5d8
[CPU][Bugfix][Speculative Decoding] Accept USE_FP64_GUMBEL in CPU recovered-tokens sampler ( #46069 )
...
Signed-off-by: hillel.darshan <hillel.darshan@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-23 11:54:07 +00:00
wang.yuqi and GitHub
31ca9504b1
[Frontend] Split ServingRender into renderer and entrypoint. ( #44285 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-23 11:19:09 +00:00
d32575a2d2
[ROCm][P/D] Support MoRIIO heterogeneous TP fan-in ( #46332 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-23 10:33:23 +00:00
Juan Pérez de Algaba and GitHub
83fa302ca4
fix(security): prevent infinite loop in split_audio with NaN audio sa… ( #46463 )
2026-06-23 10:24:51 +00:00
frida-andersson and GitHub
20b5af55c1
[ROCm][Perf] DSv3.2: fuse MLA Q concat+fp8-quant in forward_mqa ( #43673 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-06-23 18:12:04 +08:00
Qiming Zhang and GitHub
901a3b091c
fix gpt_oss pp>1 with ep ( #46441 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-23 16:59:11 +08:00
2d721ab5d8
[Rust Frontend] Align Rust allowed_token_ids validation with Python ( #46348 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 08:32:33 +00:00
accaa434f3
[Rust Frontend] Support echo for token-ID completion prompts ( #46219 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-23 08:04:41 +00:00
Sunny Yuan and GitHub
a04654da23
Doc: fix missing GLM-5.x in supported models ( #46452 )
...
Signed-off-by: Sunny Yuan <y.zichen@wustl.edu >
2026-06-23 07:42:27 +00:00
Bugen Zhao and GitHub
25bc3be49c
[Rust Frontend] Correct --reasoning-parser semantics ( #46359 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 15:38:39 +08:00
Ting SUN and GitHub
a46f3eb232
[Bugfix][Model Runner V2] Preserve all allowed_token_ids in the logit bias kernel ( #46245 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-23 07:01:13 +00:00
6c427dd401
[BugFix] Omit empty tool_calls from OpenAI chat responses ( #44105 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-23 13:43:53 +08:00
3ce5823762
[Refactor] Responses API parser state into conversation context ( #46030 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 13:42:58 +08:00
Woosuk Kwon and GitHub
04c2a8deac
[DeepEP V2] Fill invalid recv_topk_idx with -1 ( #46432 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-22 21:45:49 -07:00
7e47fb72b5
[ROCm][P/D] Fix MoRIIO WRITE mode for mixed KV layouts ( #46290 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
2026-06-23 12:12:51 +08:00
a8481be7a9
[Rust Frontend][Perf] Use dedicated runtime for HTTP/request-processing/ZMQ ( #46051 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 04:03:20 +00:00
Kunshang Ji and GitHub
9d3317172c
[XPU][CI]fix xpu kv cache layout test ( #46429 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-23 03:43:29 +00:00
430a95ae3a
[v1][kvcache] Honor prefix-cache retention interval for Mamba/linear attention ( #45845 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 19:51:11 -07:00
Mike G and GitHub
56e5797511
[Quant] Enable modelopt_mixed on Turing (SM75) ( #45375 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-22 19:30:49 -07:00
8db12169a4
fix: stream Qwen3 tool call string arguments ( #46351 )
...
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-06-23 10:26:37 +08:00
33f50773cb
[Doc] Fix typos, grammar, and broken commands across docs ( #46398 )
...
Signed-off-by: MichaelCaoo <a992033227@163.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 02:01:22 +00:00
Micah Williamson and GitHub
fa36f86d77
[CI] Torch 2.11 flaky test_spec_decode_logprobs and gritlm tests ( #45772 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 01:26:54 +00:00
8207ce0850
[Bugfix] Fix humming lm_head crash and FusedMoE weight_shape coercion ( #46420 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 18:19:29 -07:00
e48592066e
[DeepEP V2] Bound num_max_tokens_per_rank in do_expand=False ( #46404 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Roy Wang <jasonailu87@gmail.com >
Co-authored-by: gnovack <novackgm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 18:14:53 -07:00
91ba720b75
[ROCm][CI] Only require q_scale==1.0 for fp8 query in RocmAttention ( #46148 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 18:25:43 -05:00
fxmarty-amd and GitHub
6ead164e52
[CI] Add TP=4 requirement to test_mixed_precision_model_accuracies ( #46161 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-22 18:19:43 -05:00
c97e8f99d6
[ROCm][Quantization][4/N] refactor quark_moe fp8 w/ oracle ( #43721 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:58:03 -07:00
183b5f27ea
[Bugfix][V1][TurboQuant] Reserve workspace before CUDA graph capture ( #44053 )
...
Signed-off-by: Guipeng Zhang <zhangguipeng23z@ict.ac.cn >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 15:47:48 -07:00
ca5b24695b
Fix static actorder handling for compressed-tensors WNA16 MoE ( #41161 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:46:46 -07:00
Charlie Fu and GitHub
6f6bd3b8fe
[ROCm][CI] Increase the max wait time for server startup ( #46417 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-22 17:46:31 -05:00
Andreas Karatzas and GitHub
70ef4d3009
[ROCm][CI] Purging away redundant test group definitions ( #46418 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:42:47 -07:00
e2fe837572
[CI] Fix CPU-Multi-Modal Model Tests timeout by adding a 4th shard ( #46388 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 22:08:00 +00:00
Aarushi Jain and GitHub
fbf9ff7cf4
[CI][ROCm] Restrict MLA cross-layer KV cache test to supported backends on ROCm ( #46401 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-22 17:05:26 -05:00
6cc2c9ba3a
[CI] Add DGX Spark GPQA smoke test ( #39541 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 14:52:38 -07:00
c0b2d8f471
[Bugfix] FusedMoE: coerce shape-(1,) per-tensor scales to 0-D scalar … ( #43362 )
...
Signed-off-by: Varshith <kvarshithgowda@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 13:26:53 -07:00
Mohammad Miadh Angkad and GitHub
d1a38c2762
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend ( #42235 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-22 16:17:18 -04:00
2b4a7491ec
[ROCm][CI] Query total device memory via amdsmi to avoid HIP init ( #46141 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:12:24 -05:00
Saddss and GitHub
82ede09a5a
[Bugfix][KVConnector] Fix SimpleCPUOffloadConnector GPU->CPU store race ( #46278 )
2026-06-22 13:08:47 -07:00
Nick Hill and GitHub
fbf520cf3a
[MRV2] Generalize use of WhisperModelState ( #46096 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-22 12:40:02 -07:00
44d95069e9
Enable DeepSeek V4 and GLM-5.1 on SM120 ( #43477 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-22 11:54:14 -07:00
3ce15fd574
[v1][kvconnector] DecodeBenchConnector: fill list/tuple (Mamba/KDA) KV caches ( #45080 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 11:54:00 -07:00
e4b3da3feb
[Quantization][CI] add humming lm-eval test ( #43752 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:23:55 -07:00
3e6529cc0e
[Bugfix][Spec Decode] Fix EAGLE drafter multimodal encoder cache misses ( #46315 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 18:14:02 +00:00
ac614587f5
[EPLB] Enable nixl eplb communicator for elastic ep ( #45013 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-22 10:54:08 -07:00
f2069b005b
[Pooling] Validate non-negative rerank top_n ( #46119 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:40:47 -04:00
Martin Hickey and GitHub
ccd49f6821
[MyPy] Fix mypy for vllm/lora ( #41722 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-22 10:57:09 -04:00
Li, Jiang and GitHub
1c7bc18318
[Bugfix][CPU] Fix CPU model runner v2 ( #46365 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-22 22:52:05 +08:00
AlexHuang and GitHub
9a938df64e
[Test][KV Offloading] Add unit tests for OffloadingSpecFactory and SecondaryTierFactory ( #46355 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
2026-06-22 17:45:04 +03:00
Liangliang Ma and GitHub
3da4a1b124
[XPU] add awq format for INCXPULinear ( #43404 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
2026-06-22 22:29:13 +08:00
6871738777
[Doc] Document pull request limit ( #46376 )
...
Signed-off-by: simon-mo <simon.mo@hey.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 14:04:56 +00:00
Yifan Qiao and GitHub
aa4990a9a2
[Attention] Re-enable cross-layer KV cache layout for MLA via stride-aware kernels ( #45111 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-22 06:57:02 -07:00
a4610da0c6
[docs] link security docs from AGENTS ( #46373 )
...
Add a security-review routing sentence to AGENTS.md that points agents to SECURITY.md, docs/usage/security.md, and docs/contributing/vulnerability_management.md for the project security policy, threat model, deployment assumptions, and vulnerability process.
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 06:28:25 -07:00
liuzhenwei and GitHub
09cdcf34aa
[XPU] update nixl to v1.2.0 ( #46327 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 20:55:06 +08:00
d2c671c29b
[CPU][RISC-V] Add RVV micro GEMM for WNA16 ( #44324 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-22 12:53:54 +00:00
xiangdong and GitHub
b5a2adec4b
[XPU][CI]Skip v1/spec_decode/test_speculators_correctness.py in intel GPU nightly ( #46356 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 19:30:41 +08:00
78739e3bda
[Bugfix] Reject matryoshka embedding dimensions above hidden size ( #46313 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-22 10:16:35 +00:00
Tuukka Sarvi and GitHub
89accad2cc
[ROCm][DSV4] Disable TileLang MHC dispatch on gfx942 ( #45931 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-06-22 09:26:54 +00:00
3c8e49596c
[Model] ColQwen3.5: fix retrieval correctness (bias + bidirectional) ( #46108 )
...
Signed-off-by: Athrael Soju <athrael.soju@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 17:25:54 +08:00
cec2ec1176
[Bugfix] Avoid racy accepted counts in async spec decode ( #45100 )
...
Signed-off-by: Weiwei Sun <68775773+sunnweiwei@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-22 08:53:16 +00:00
liuzhenwei and GitHub
435f82d61a
[Bugfix] Fix Llama4ForCausalLM initialization test failure ( #46341 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 08:40:43 +00:00
Roger Wang and GitHub
1c4b51b990
Temporarily skip M3 on CI ( #46352 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:35:31 -07:00
2e2c47928b
[Doc] Update MiniMax-M3 ( #45940 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:23:27 -07:00
80abe0de7d
[Rust Frontend] Support thinking_token_budget for chat and completions ( #46137 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-22 16:00:02 +08:00
a9f7b2d41c
[feature][kv_offload] Self-describing KV events for OffloadingConnector ( #43468 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 07:27:46 +00:00
d14e551a53
[Model] Remove MiniMaxText01, MiniMaxVL01, MiniMaxForCausalLM ( #45993 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:20:46 +08:00
68567ef2df
[CPUOffloadingManager] Maintain evictable list in LRUCachePolicy ( #46216 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-22 06:54:44 +00:00
6bc6f2d86d
[1/N][Core] add partial prefix cache primitives ( #45939 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 23:43:10 -07:00
wang.yuqi and GitHub
1eb2cc961e
[Frontend] Refactor ServingTokenization entrypoint. ( #46022 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-22 06:27:58 +00:00
31124749d1
[Bugfix] [Rust Frontend] Fix stop string truncation with repeated matches ( #46113 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-22 14:11:29 +08:00
Ma Jian and GitHub
9037498c22
[DSV4][XPU] Pass gemm1_clamp_limit to XpuFusedMoe ( #44517 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-22 12:57:10 +08:00
db32b53e30
[SpecDecode] Support DFlash with FlashInfer ( #43081 )
...
Signed-off-by: gss <2783977641@qq.com >
Co-authored-by: gss <2783977641@qq.com >
2026-06-22 04:55:30 +00:00
xiangdong and GitHub
b529bfd6c5
[XPU][CI] Add agent_tags for Intel GPU CI ( #45768 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 10:33:17 +08:00
Micah Williamson and GitHub
f3df7a7231
[ROCm][CI] Enable kv_connector unit tests on ROCm ( #45955 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-22 05:08:44 +03:00
485bbe1c6f
[CI] Fix missing tp_size attribute on RoutedExperts ( #46163 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 18:46:49 -06:00
Matt and GitHub
a19ff2218a
[Hardware][AMD][CI] Fix Spec Decode Eagle test group ( #46018 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:40:02 -05:00
Matt and GitHub
4f0d0049a0
[Hardware][AMD][CI] Fix Kernels Attention test groups ( #46080 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:10:51 -05:00
13b83d77ad
[ROCm][CI] skip test_double_aiter_rms_quant_fusion ( #45967 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 16:53:11 -05:00
Matt and GitHub
50241602fd
[Hardware][AMD][CI] Fix gfx942 Kernels MoE test group ( #46298 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 16:45:37 -05:00
Ting SUN and GitHub
12fe2a9aac
[Bugfix][Qwen3-VL] Fix multi-video crash with list-valued fps/num_frames ( #46305 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 14:31:23 -07:00
Benjamin Chislett and GitHub
89bd2c14d3
[Spec Decode] Add Qwen3 architecture support for EAGLE3 ( #43132 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-21 13:55:26 -07:00
ZedongLiu and GitHub
9c450b1027
[Kernel][Bugfix] Fix INT8 per-token-head KV cache rounding in Triton reshape-and-cache ( #45361 )
...
Signed-off-by: ZedongLiu <113341356+Zedong-Liu@users.noreply.github.com >
2026-06-21 15:59:40 -04:00
635c38338a
[Multimodal] Add Qwen2-VL/Qwen2.5-VL processor-mapped video loader ( #45555 )
...
Signed-off-by: Ranran <hzz5361@psu.edu >
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-21 18:56:50 +00:00
c441ad1c07
[KV Offloading] Add labeled metrics support ( #45957 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-21 18:04:01 +00:00
Jee Jee Li and GitHub
745bba5ea8
[Model]Fix MiniMaxM2ForCausalLM perf regression ( #45935 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-22 00:28:52 +08:00
2cac89f9da
[Spec Decode] Support mixed KV page sizes for DFlash ( #45181 )
...
Signed-off-by: Alex Steiner <asteiner@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 22:45:14 +08:00
3e6e33526d
[Disagg] return routed_experts on streaming generate responses ( #44638 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-21 07:37:10 -07:00
b91b7726e0
[ROCm][P/D] Support MiniMax-M3 mixed KV layouts in MoRIIO READ mode ( #46039 )
...
Signed-off-by: Jun Kang Chow <junkangchow@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-21 12:55:19 +00:00
Palaiologos1453 and GitHub
d3ad8e8bcd
[Bugfix] Defer offload reads while transfers are pending ( #46231 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-21 14:30:13 +03:00
b80ce9dd2f
[CI][test] Replace InternVL2-1B with InternVL3-1B in test_pipeline_parallel.py ( #46241 )
...
Signed-off-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-21 15:11:19 +08:00
b5495cc5f9
Fix memory pointer overflow in Mamba state buffers ( #44665 )
...
Signed-off-by: Shifani Rajabose <shifani.rajabose@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-21 14:00:50 +08:00
Ting SUN and GitHub
183a430c13
[Bugfix][Model Runner V2] Fix min_tokens off-by-one in the V2 GPU sampler ( #46243 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 05:06:49 +00:00
Matt and GitHub
a346d589f5
[Bugfix] Fix NVFP4/OCP MX MoE emulation ( #46254 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-20 23:13:10 -05:00
Nick Hill and GitHub
7df3d7dada
[Core] Ensure memory is pinned prior to async h2d copy ( #45424 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 20:02:24 -07:00
8dd1b702f2
[Misc] Fix stale doc URL and docstring module path ( #35530 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-20 23:57:01 +00:00
f57ac274b2
[Render] Add reasoning/tool parsing to /derender + fix byte-fallback FFFD ( #45919 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-20 19:43:32 -04:00
6e919960af
[Perf] Skip/shrink all_token_ids copy in scheduler for non-async and V2 runner ( #45840 )
...
Signed-off-by: amanchugh89 <amanchugh.89@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 22:36:57 +00:00
Jonathan Chen and GitHub
c88d3d4775
[SimpleCPUOffloadConnector] PCP + DCP support ( #39831 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
2026-06-20 15:01:06 -07:00
Yifan Qiao and GitHub
ab7fcbdd5d
[Perf][KVConnector][Mooncake] Compact chunk-hash keys and zero-copy lookup wire format ( #45969 )
2026-06-20 15:00:11 -07:00
3b4a76b63f
[KV-Offloading] : Expose CPU cache usage metric ( #45737 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-20 21:21:55 +00:00
cc22621b51
[KV Offload] Support packed HMA KV cache layout ( #46205 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:19:40 +00:00
77148992cf
[Bugfix] Move extract_layer_index back inside is_v32 guard ( #46199 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 21:19:10 +00:00
891cc4b9c5
[Frontend] Report cache usage in Anthropic /v1/messages API ( #40912 )
...
Signed-off-by: mistral0105 <zhangshuoming17@mails.ucas.ac.cn >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:12:48 +00:00
TJian and GitHub
1bdf9810aa
[ROCm] [Bugfix] Bugfix ROCm Sparse Indexer ( #46222 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-20 13:38:42 -07:00
ebfbcfe46a
Stop setting CUDA_VISIBLE_DEVICES internally in vLLM, add device_ids arg ( #45026 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: kourosh hakhamaneshi <kouroshHakha@users.noreply.github.com >
2026-06-20 13:38:10 -07:00
e9de72fe6c
[Bugfix] Guard model_config access in _log_compilation_config ( #46198 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 19:26:38 +00:00
d272418f45
[Perf] Optimize Qwen3-VL multi-video prompt processing ( #46026 )
...
Signed-off-by: Sirius29 <422058530@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 07:09:18 -07:00
Sumanth R Hegde and GitHub
7ff7f5c8eb
Revert "Fix Stale Encoder Cache After Weight Update" ( #46125 )
2026-06-20 07:09:09 -07:00
dced290769
[Hardware][AMD][CI] Fix e2e core test group ( #46024 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 02:04:35 -05:00
JasonLi314 and GitHub
93bad11912
[Bugfix] Fix gridDim.y overflow for large row counts ( #45255 )
...
Signed-off-by: Jason Li <li.jason.cs@gmail.com >
2026-06-19 23:27:45 -04:00
djramic and GitHub
0fbf42af84
[ROCm] Fix VRAM not freed in test_phi3v ( #46046 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 17:20:59 -05:00
Charlie Fu and GitHub
e6cd8913dd
[ROCm][CI] Skip Qwen3.5-35B-A3B-MXFP4-AITER-TP2 for non gfx950 ( #46109 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-19 17:20:10 -05:00
Ben Browning and GitHub
859e4d436b
[Bugfix][Parser] Fix U+FFFD leak at reasoning-to-content transition in engine parsers ( #46159 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 22:09:28 +00:00
Micah Williamson and GitHub
4a083cc858
[ROCm][CI] Pin test_rocm_compressed_tensors_w8a8 to TRITON_ATTN ( #46180 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-19 15:20:06 -05:00
Vadim Gimpelson and GitHub
ca7e1f2c43
Move CI failure diagnosis docs into ci-fails-buildkite skill ( #45975 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-06-19 20:12:40 +00:00
djramic and GitHub
dec860fb19
[ROCm] Use vLLM's fp8 quant max in AITER hipBLASLt accuracy test ( #46176 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 13:24:02 -05:00
Harry Mellor and GitHub
0a49fb2b13
Fix dead link in docs ( #46181 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-19 18:16:09 +00:00
Ben Browning and GitHub
4a8abf37c7
[Test] Migrate test_openai_schema.py to schemathesis 4.x ( #46173 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 18:05:18 +00:00
01192139bf
[DSv4] Pack KV caches into contiguous per-block allocations for DeepSeek V4 ( #44577 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-19 12:55:42 -04:00
Chris Leonard and GitHub
b9a7cd464c
[12/n] final _C library kernel migration ( #45415 )
2026-06-19 06:57:26 -07:00
69bdd34542
[Bugfix] Fall back to Pydantic loc for param in validation errors ( #46038 )
...
Signed-off-by: professorsab <135441198+professorsab@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-19 19:11:11 +08:00
Kunshang Ji and GitHub
ec67d7ae61
[xpu] bump up vllm-xpu-kernels v0.1.10 and upgrade 2618 umd ( #40367 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-19 15:37:20 +08:00
ecf9d83520
[AMD][CI] Fix Language Models Test (Extended Generation) failures ( #45509 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-19 12:06:56 +08:00
Samuel Shen and GitHub
c9135db27c
[Docs] Update stale LMCache examples ( #45762 )
...
Signed-off-by: Samuel Shen <slshen@tensormesh.ai >
2026-06-19 03:21:36 +00:00
2a6c6b9429
[DeepSeek-V4] Support TEP=16 for the block-FP8 shared expert ( #46001 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:10:12 -07:00
Jared Wen and GitHub
ab66606993
[bugfix]Indexer init skip and MTP TopK share for iteration ( #45895 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-19 09:57:51 +08:00
9ea3a4015b
[Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded ( #42120 )
...
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 18:26:09 -07:00
Flora Feng and GitHub
560fb8b867
[Cohere] Remove dead prepare_structured_tag override in Cohere parser ( #46099 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-19 01:02:11 +00:00
Wentao Ye and GitHub
675cd5d228
[Model Runner V2] Fix MRv2 memory leak test ( #46095 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-19 00:36:40 +00:00
7f616c327d
[Bugfix] [Parser] Fix empty tool block silently dropping subsequent content ( #46091 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-18 23:17:18 +00:00
Ivy Xu and GitHub
c3c6d723fd
[Perf] Remove unused loggers in reasoning/ ( #45988 )
...
Signed-off-by: Ivy <fakeshadow1337@gmail.com >
2026-06-18 22:24:29 +00:00
41dcf49ca5
[Bugfix][KV Connector] Disable Mooncake TP put-striding when DCP > 1 ( #45371 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:13:44 -07:00
35e4dd4a69
[KV Connector][Mooncake] Async lookup to reduce scheduler overhead ( #45659 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-18 21:44:02 +00:00
4ce2d01453
fix(anthropic): auto-detect template support for mid-conversation system messages ( #46025 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:19:11 -04:00
Woosuk Kwon and GitHub
16908e132e
[MRV2] Make FP32 Gumbel sampling more accurate ( #45996 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-18 19:42:09 +00:00
Wentao Ye and GitHub
225936a1dd
[CI Bug] Revert #42379 to fix CI Multi-Modal Models (Extended Generation 1) ( #46070 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:37:39 -07:00
f6ba720963
(security) Upgrade Starlette to >= 1.0.1 to fix CVE-2026-48710 ( #45675 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-18 12:35:13 -07:00
Wentao Ye and GitHub
b53b1c7ffe
[Model Runner V2] Migration to support quantized model by default [5/N] ( #44446 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:20:44 -07:00
79ca54d221
[Bugfix][Quantization] Don't reject fp8_e5m2 KV cache for non-fp8 quantized checkpoints ( #45040 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:18:25 -04:00
Ben Browning and GitHub
09f3cd5c10
[Bugfix] [Parser] Fix Qwen3 latent bug in partial params dropping values containing < ( #46047 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 18:04:06 +00:00
ea6078fe6a
[KV Connector][Offloading] Disable parallel-agnostic fs-tier cache on V2 model runner ( #46044 )
...
Signed-off-by: Itay Etelis <etelis2019@gmail.com >
Co-authored-by: Itay Etelis <etelis2019@gmail.com >
2026-06-18 20:43:35 +03:00
Palaiologos1453 and GitHub
a0df04e477
[Tests] Add Qwen3 streaming parser delta boundary cases ( #45708 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-18 17:37:39 +00:00
stefankoncarevic and GitHub
e2352c2974
[ROCm][Spec Decode] Fix probabilistic draft probs test attention backend ( #45706 )
...
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com >
2026-06-18 11:59:37 -05:00
qli88 and GitHub
25faa1f4cc
[CI]Enable mxfp4 lora test for ROCm platform ( #43802 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-18 16:59:09 +00:00
Humphrey and GitHub
4583630b56
[Bugfix][Kernel] Check output alignment in vectorize_with_alignment (fixes misaligned-address crash for non-multiple-of-8 head sizes) ( #45466 )
...
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com >
2026-06-18 16:58:22 +00:00
Divakar Verma and GitHub
21da47dabe
[ROCm][CI] move lora%N test to mi300 and gate ( #45970 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-19 00:50:32 +08:00
6c379b9e54
[Frontend] Add Streaming Parser Engine and new GLM4.7/GLM5.1/GLM5.2 Parser ( #45915 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-19 00:42:10 +08:00
Rohan Potdar and GitHub
5099474633
[Bugfix][ROCm] Fix rocm_aiter_per_tensor_quant custom op aliasing ( #45747 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-18 11:30:21 -05:00
Yuwen Zhou and GitHub
058cc0a8b6
[Bugfix] Restore is_sym guard for zp in GPTQ/CT MoE to fix symmetric quant regression ( #45656 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-18 16:20:29 +00:00
837db7605e
[Bugfix][Tool Parser] Handle non-finite numbers in coerce_to_schema_type ( #43984 )
...
Signed-off-by: ashishpatel26 <shriganesh.patel@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:00:20 +00:00
Mark McLoughlin and GitHub
bf2a393034
Temporarily remove @markmc from CODEOWNERS ( #46053 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-18 14:15:43 +00:00
d682968aa9
[Model] Remove BambaForCausalLM ( #45990 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-18 06:51:00 -07:00
021cdf72bc
Fix _riscv_supports_rvv_vlen128() to detect RVV on hardware without zvl flags ( #43179 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn >
2026-06-18 21:22:35 +08:00
Ashar and GitHub
4cb5e746b6
[Rust Frontend]: Add /get_world_size route with static parallel size ( #44801 )
2026-06-18 13:10:20 +00:00
Jee Jee Li and GitHub
22cc891108
[Kernel] Add PDL support for DeepGEMM kernel ( #46006 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-18 20:49:01 +08:00
afdcbd5d39
[ROCm][DSv4] Functional fixes for DeepSeek V4 on MI300X/MI325X ( #45681 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Jin Tao <jintao12@amd.com >
2026-06-18 12:21:14 +00:00
8d4f54966c
fix(quantization): Fix AWQ dequantize on Intel XPU and refactor AutoAWQ config ( #42727 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@tencent.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:12:28 +08:00
351c72d6e5
[CPU] Skip Triton kernel monkey-patches when Triton-CPU is available ( #44991 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-18 18:59:30 +08:00
Tahsin Tunan and GitHub
7299e6509e
[Rust Frontend] Return model metadata fields in /v1/models ( #45950 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-18 10:29:21 +00:00
littlecircle0730 and GitHub
08985351f3
Fix Stale Encoder Cache After Weight Update ( #45093 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
2026-06-18 09:32:10 +00:00
Wei Zhao and GitHub
5fd3b276f8
[Mooncake] Skip KV lookup for non-reachable SWA blocks ( #45444 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-18 02:23:20 -07:00
1e9f04da14
fix(anthropic): preserve inline system message position for prefix caching ( #44602 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
2026-06-18 15:58:11 +08:00
702214146c
[Bugfix][Frontend] Fix Anthropic count_tokens decorator order driving server load negative ( #44725 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 23:56:46 -07:00
a331589394
[XPU] Update nixl to v0.10.1 in Dockerfile ( #40287 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:01:26 +08:00
Micah Williamson and GitHub
e945169207
Revert "[Kernel] Add PDL support for DeepGEMM kernel" ( #45999 )
2026-06-17 22:59:48 -07:00
554352a311
[Test][KV Connector] Add request_finished fence population tests for offloading scheduler ( #45679 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@future.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:52 +03:00
421c1ec448
[KV Offloading] Remove dummy worker-side stats from OffloadingConnector ( #45905 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@alexai.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:28 +03:00
b4c80ec0fd
[Refactor] Remove dead cutlass mxfp8 code ( #44681 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-17 21:18:25 -07:00
Ronen Schaffer and GitHub
f428718ffe
[Fix][KV offload] Defer on_request_finished until in-flight transfers drain ( #45823 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-18 07:05:46 +03:00
Jee Jee Li and GitHub
4403af8fb5
[Kernel] Add PDL support for DeepGEMM kernel ( #42996 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-17 20:37:17 -07:00
d57888efa4
[SimpleCPUOffloadConnector]: Add support for reset_cache() ( #39726 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
Signed-off-by: Jonathan <chenleejonathan@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 19:47:12 -07:00
ed938ad7db
[CPUOffloading] Guard CPU eviction check ( #45757 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-18 05:34:59 +03:00
Reid and GitHub
731fb3323d
[Rust Frontend] Validate tokenized bad_words vocabulary range ( #45876 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-18 02:28:45 +00:00
8dd8b6ed78
[XPU] Fix FP8 block-scaled scheme selection on non-CUDA platforms ( #43958 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-18 10:16:20 +08:00
e1a5fc406b
[Rust Frontend][Perf] O(n) argument scan in tool parser ( #45826 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-18 01:42:35 +00:00
Ace Eldeib and GitHub
b4092176b9
[Bugfix] Complete one-shot fused all-reduce PDL at end to avoid NaN ( #45448 )
2026-06-18 00:54:39 +00:00
Jee Jee Li and GitHub
ebbb2d55ac
[CI/Build][Bugfix] Fix SD LoRA ( #45941 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 00:34:15 +00:00
liuzhenwei and GitHub
2959a9273a
[XPU][CI] add model runner v2 into CI ( #44650 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-18 00:28:34 +00:00
1797576237
Revert "[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture" ( #45309 ) ( #45972 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-17 17:20:59 -07:00
0d339cf135
[Bugfix] Fix NixlConnector handshake block_len validation for GQA-replicated KV heads ( #45879 )
...
Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com >
Co-authored-by: waynehacking8 <waynehacking8@gmail.com >
2026-06-17 15:11:29 -07:00
5fd21eb0b2
[BUG] fix hidden states nan for hybrid attention models ( #45849 )
...
Signed-off-by: shanjiaz <hezhao@redhat.com >
Co-authored-by: shanjiaz <hezhao@redhat.com >
2026-06-17 18:02:24 -04:00
Ting SUN and GitHub
9d4b87f4f0
[Bugfix][Model] Validate DefaultModelLoader / LoadConfig and fail with clear errors ( #45196 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-17 21:46:33 +00:00
58b2e89642
[Bugfix][Gemma4] Render reasoning on assistant turns without tool_calls ( #45867 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-17 20:44:15 +00:00
Wentao Ye and GitHub
2659f60a1a
[Refactor] Remove dead quantization code and tests ( #45454 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 16:12:01 -04:00
091386a99b
[Bugfix] MiniMax-M3 (AMD): add packed_modules_mapping and pass swiglu… ( #45794 )
...
Signed-off-by: wangjiaxin99 <jiaxwang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-06-17 19:15:46 +00:00
qli88 and GitHub
d112eb1ac7
[feature] MiniMax-M3-MXFP4 support added ( #45896 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-17 18:50:48 +00:00
Wentao Ye and GitHub
2a47a9ff0f
[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture, 26.8% ~ 27.9% E2E TTFT improvement ( #45309 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 09:34:53 -07:00
Wentao Ye and GitHub
9c7c74bf10
[Log] Update deepgemm log ( #45857 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 15:34:22 +00:00
danisereb and GitHub
5e27b2baf4
[Bugfix] Pass TP group to FlashInfer all-reduce fusion ( #45917 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-06-17 15:24:26 +00:00
zhanqiuhu and GitHub
eb0fdeb1e8
[Bugfix][PD] Fix DSV4 disaggregated serving ( #45831 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-06-17 15:17:14 +00:00
46f74e144b
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_dynamic_per_token_quant ( #34432 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-17 23:03:54 +08:00
Wentao Ye and GitHub
0a7bacdcac
[DSv4 Perf] DSv4 flashinfer sparse index cache for metadata, 2%~4% TTFT improvement ( #45863 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 10:55:48 -04:00
amirkl94 and GitHub
8b2b566ea7
Feature: Enable Flashinfer non-gated MoE bf16 ( #43853 )
...
Signed-off-by: Amir Klein <203507526+amirkl94@users.noreply.github.com >
2026-06-17 14:32:49 +00:00
xaguilar-amd and GitHub
0b131b16c9
[ROCm][AITER][Quark] Tag per-channel FP8 weights as PER_CHANNEL so AITER pre-shuffled GEMM is selected ( #44626 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
2026-06-17 14:05:34 +00:00
bcb518ad7a
[quant][autoround]Refactor INC quantization into package with INCScheme orchestrator ( #40601 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
Co-authored-by: n1ck-guo <heng.guo@intel.com >
Co-authored-by: Zhenzhong1 <zhenzhong.xu@intel.com >
2026-06-17 21:51:32 +08:00
Chaojun Zhang and GitHub
06e1e0885c
[XPU] Fix test_logprobs_e2e import error: pin lm-eval[api]>=0.4.12 ( #44469 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-17 12:26:47 +00:00
Isotr0py and GitHub
1a59078c87
[CI/Build] Avoid duplicate ViT CG test introduced by accident ( #45654 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-17 12:23:44 +00:00
Oğuzhan KIR and GitHub
fa85ead2f3
[MM][Perf][CG] Support ViT full CUDA graph for Kimi-VL ( #41992 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
2026-06-17 12:14:01 +00:00
e28e8c8782
[ROCm][Quant] Minimax-M3: Enable fp8_per_channel for bf16 weights on mi300x ( #45854 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-17 12:02:40 +00:00
Angelo Ruocco and GitHub
ee0fd6984a
docs, kv_offloading: add docs for selective offload ( #45279 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
2026-06-17 14:58:00 +03:00
vllmellm and GitHub
d537122398
[ROCm][Bugfix]: Fallback GFX942 sparse MLA ops to Triton ( #45782 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-17 11:41:29 +00:00
Juan Pérez de Algaba and GitHub
3d20275bb4
fix(security): enforce audio decode duration limit in chat completions path ( #45908 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-17 11:07:13 +00:00
f694d43b33
[Bugfix][test] Use Salesforce/wikitext for ppl tests ( #45913 )
...
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-17 10:37:19 +00:00
Nikhilesh Chhetri and GitHub
3c6084bb0d
[Bugfix][Gemma4] Pre-initialise streaming reasoning state when prompt ends inside an open <|channel> ( fixes #45834 ) ( #45852 )
...
Signed-off-by: nikhilesh-csa <nchhetri@csa1.com >
2026-06-17 06:16:02 -04:00
Joel Smith and GitHub
68ff30d40e
[Bugfix] Fixes MiniCPM-O resampler device placement to avoid tensor device mismatch ( #42332 )
...
Signed-off-by: j9smith <j.smith9103@outlook.com >
2026-06-17 08:35:27 +00:00
6d8fff5698
[KV Connector][Offloading] Avoid blocking the engine to flush offloads on idle ( #45595 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-17 11:35:07 +03:00
e2c58570ea
[Rust Frontend] Support hybrid/external DP LB in Python supervised bootstrap ( #45805 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-17 07:32:40 +00:00
Taneem Ibrahim and GitHub
43fa24e832
[Misc] Validate Cohere Embed Mixed Content Payloads ( #45873 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-17 06:57:12 +00:00
arghyadeep sarkar and GitHub
93bbe94d3a
[Kernel] Add weightless RMSNorm CUDA kernels for has_weight=False ( #41430 ) ( #44109 )
...
Signed-off-by: hello-args <args.sarkar@gmail.com >
2026-06-16 23:45:55 -07:00
Will Eaton and GitHub
17bc144556
[Rust Frontend] Add serde defaults for omit_defaults fields in EngineCoreSamplingParams ( #45848 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-17 06:40:49 +00:00
Sahil Singh and GitHub
295232a26a
[Rust Frontend] Add /abort_requests endpoint ( #44382 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-17 06:40:47 +00:00
Reid and GitHub
56e4345226
[Rust Frontend] Support prompt-only completions ( #44938 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-17 06:38:06 +00:00
Nick Hill and GitHub
e9993a52aa
[BugFix][CI] Fix scheduler plugin test ( #45897 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 06:30:49 +00:00
a46abb7ae6
[Bugfix][Quantization] Reject unsupported compressed tensors KV cache schemes ( #45312 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 05:08:41 +00:00
4c62663315
[M3] Enable FP8 sparse GQA ( #45744 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-16 21:38:03 -07:00
d78650cf97
[CI][NIXL] Pin NIXL to 1.2.0 ( #45843 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
Signed-off-by: Itay Alroy <75032521+itayalroy@users.noreply.github.com >
Co-authored-by: ovidiusm <ovidium@nvidia.com >
2026-06-16 21:29:34 -07:00
5bdc01bcc3
[M3] Tune Triton indexer score decode for spec-decode ( #45743 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 21:07:34 -07:00
liangel-02 and GitHub
20a5f8b43b
[FlexAttention] make custom mask mods fully cudagraphable ( #45232 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-06-17 11:53:12 +08:00
7b5d60cc37
[Bugfix][V1] Clean up compiled-model bytecode hooks on VllmRunner exit ( #45195 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 20:31:17 -07:00
Nick Hill and GitHub
14b438a98b
[ModelRunnerV2] Various model/config compatibility fixes ( #45868 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 03:23:01 +00:00
2785a5e0e6
[Bugfix][ROCm] Fix FP8 per-tensor scale rank mismatch causing Inductor assertion failure ( #44912 )
...
Signed-off-by: nehmathe2 <nehmathe2@gmail.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: nehmathe <nehmathe@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:17:42 -07:00
efd15e192a
[Bugfix][ROCm] Fix MiniMax-M3 FP8 KV cache dtype ( #45720 )
...
Signed-off-by: Cam Quilici <cjquilici@gmail.com >
Signed-off-by: Cameron Quilici <cjquilici@gmail.com >
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-17 03:14:45 +00:00
556b063e45
[XPU] Fix test_spec_decode_logprobs: use FLASH_ATTN for XPU in GPU_DETERMINISM_KWARGS ( #44468 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 11:07:04 +08:00
aa0ac8a661
[CI] Run pre-commit on self-hosted vllm-runners ( #45865 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 19:49:22 -07:00
Federico and GitHub
b831374cf1
[Bugfix][Gemma4] Fix parsing when thinking is disabled ( #45832 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-17 02:41:36 +00:00
71bc19dbdd
[Bugfix] Fix MoE model load OOM in FlashInfer_TRTLLM backend with sleep mode ( #45589 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 19:36:51 -07:00
Kunshang Ji and GitHub
ef2c40dc00
[XPU][CI] fix server test file path ( #45870 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 09:06:25 +08:00
4bf699d310
[Kernel] Support DS Mamba tail copy for MTP align mode ( #45473 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Co-authored-by: Thomas Parnell <tom.parnell@gmail.com >
2026-06-16 22:50:30 +00:00
Stan Wozniak and GitHub
520828789c
Apply LRU policy only to proper cache entries ( #42656 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-16 21:49:15 +00:00
9d4dc4ca2f
[Kernel] Support GLM-5 dimensions for TRT-LLM ragged MLA prefill ( #43525 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 20:49:47 +00:00
Federico and GitHub
b9684d99e9
[Bugfix] Gemma4: skip forced JSON for required/named tool choice ( #45795 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-16 20:38:29 +00:00
Divakar Verma and GitHub
4fadf9c92c
[ROCm][CI] fix multimodel run cmds ( #45858 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-16 15:31:52 -05:00
Nick Hill and GitHub
d8d95998dc
[Core] Add prefill step cadence for better non-PD DP balancing ( #44558 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 13:17:18 -07:00
Flora Feng and GitHub
475a6ad18a
[Misc] Update Mergify tool-calling label ( #45853 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-16 19:08:00 +00:00
Hongxia Yang and GitHub
f2beaa80c8
[ROCm][Quant] mxfp8 moe/linear gfx950 tuning for MiniMax-M3 ( #45725 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-06-16 18:50:40 +00:00
8e27a9c215
[PERF] Fuse multi-group block table staged writes ( #44944 )
...
Signed-off-by: jesse <szxfml@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 10:53:27 -07:00
7d567172fc
[Bugfix] Fix Qwen3 prompt tool-call reasoning false positive ( #45763 )
...
Signed-off-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
Co-authored-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
2026-06-16 17:48:01 +00:00
Chauncey and GitHub
f00e163f35
[Frontend] Add Streaming Parser Engine and new MinimaxM2 Parser ( #45701 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-16 13:38:17 -04:00
44b2512767
[KV Connector][Mooncake] Add cache_prefix to namespace store keys ( #45767 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 10:24:20 -07:00
188c68798e
[KVConnector][MoRIIO] Allow overriding the advertised host IP ( #45488 )
...
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 17:18:37 +00:00
c45f681932
[Bugfix][Core] Fall back when numactl --membind is blocked in constrained containers ( #45438 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-16 09:49:56 -07:00
89e8645a9e
[Model] Remove Dots1ForCausalLM ( #45637 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-17 00:32:18 +08:00
Wentao Ye and GitHub
88a9cdd439
[Model Runner V2] Enable GraniteMOE for MRv2 by default ( #45461 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-16 09:31:32 -07:00
Micah Williamson and GitHub
6f612fbedf
[ROCm][CI] Patch conftest to resolve occasional OOMs ( #45722 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-16 10:00:15 -05:00
Sting Lin and GitHub
506ec6d656
Upgrade tpu-inference to v0.22.1 ( #45793 )
2026-06-16 07:54:57 -07:00
a52205bccf
[Model] Add HrmTextForCausalLM (Hierarchical Reasoning Model — Text) ( #43098 )
...
Signed-off-by: Wuyifei <wuyifei@me.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-16 22:41:41 +08:00
3d34f8cbdc
[ROCm][Cleanup] Remove stale AITER FA hybrid KV-cache TODO ( #44178 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-16 07:28:06 -07:00
Carl Y and GitHub
eb04c769d3
feat: MLA prefill enable FA4 fp8 output ( #43050 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-06-16 07:10:59 -07:00
ce3ef17bec
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_per_block_quant ( #36895 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-16 22:09:52 +08:00
bf5149b516
[Bugfix] Fix FlashMLA sparse accuracy with topk_length and zero-init padding ( #36616 )
...
Signed-off-by: AjAnubolu <anuboluajay@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 07:09:00 -07:00
Tahsin Tunan and GitHub
cca3365b73
[Rust Frontend] Add CORS support ( #45753 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-16 13:47:11 +00:00
040df8f2ea
[CI] Fix attention benchmark smoke test ( #45728 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 13:43:37 +00:00
ced32bb474
[Perf] Add VLLM_TRITON_FORCE_FIRST_CONFIG to skip Triton autotuning ( #42425 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 15:16:45 +02:00
c5e5c33fcd
[Bugfix][MoE] Restore routed output unpadding before shared expert add ( #45707 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 16:06:28 +03:00
Mike G and GitHub
a8c86eeb16
[Quant] Support modelopt_mixed on Ampere (SM80/SM86) ( #45306 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-16 08:43:44 -04:00
Andreas Karatzas and GitHub
7e179e4bc0
[ROCm][CI] Gate incompatible HF references on Transformers v5 ( #41532 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:34:11 +08:00
405c7cf283
[ZenCPU] Add zencpu Platform Runtime Logging and Docs ( #42726 )
...
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-16 08:23:12 -04:00
3f53e2138f
[Refactor] Remove Fp8OnlineLinearMethod as scheduled ( #45463 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 04:35:58 -07:00
Hank Han and GitHub
d53f4593ce
[KV Connector][Mooncake] Pipeline-parallel support for PD-disaggregated serving with Mooncake connector ( #44528 )
...
Signed-off-by: hanhan.hank <hanhan.hank@bytedance.com >
Signed-off-by: Hank Han <hanhan7630@outlook.com >
2026-06-16 04:35:38 -07:00
ad32608e24
[MM][Perf][CG] Support dual-path ViT full CUDA graph for DeepSeek-OCR ( #43586 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 04:35:20 -07:00
Thien Tran and GitHub
b2cfae777d
Add Triton recompile detection ( #45631 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-16 18:25:28 +08:00
wangxiyuan and GitHub
3f1ff1ff14
[Misc]Clean up useless test ( #45792 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-06-16 09:53:08 +00:00
c69c73418a
[XPU][CI] add intel xpu cases for nightly CI ( #44372 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-16 16:35:08 +08:00
Thomas Parnell and GitHub
ebf3a6d705
[Bugfix] Fix trtllm fused allreduce+rms_norm for transformers backend ( #45307 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-16 08:34:27 +00:00
wang.yuqi and GitHub
c4fd9794e9
[Frontend] Remove AsyncMicrobatchTokenizer. ( #45759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-16 08:02:11 +00:00
7ad894c86a
[Bugfix] Prevent cuMemcpyBatchAsync segfault with MTP and KV offloading ( #44784 )
...
Signed-off-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-16 07:58:39 +00:00
Li, Jiang and GitHub
a7fdfeef72
[CPU] Support Gemma Diffusion ( #45690 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-16 14:39:56 +08:00
Jimmy Lee and GitHub
8bf374955f
[Bug Fix] Allow pinned memory for WSL2 ( #41496 )
...
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com >
2026-06-16 05:56:26 +00:00
Cyrus Leung and GitHub
9096659edb
[Cleanup] Remove dead env ( #45777 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-15 22:56:23 -07:00
Taneem Ibrahim and GitHub
81d8f4ebac
[Misc] Added validation for Cohere /v2/embed input field exclusivity ( #45640 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-16 05:42:43 +00:00
a9a8a32dcd
Register parsed config classes before tokenizer init ( #40299 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-16 05:33:08 +00:00
9d808e2309
[Core] Use fastsafetensors ParallelLoader for weight loading ( #40183 )
...
Signed-off-by: Git Bisector <gitbisector@gmail.com >
Signed-off-by: gitbisector <gitbisector@gmail.com >
Signed-off-by: git bisector <gitbisector@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-15 22:32:05 -07:00
Ben Browning and GitHub
f3858d5422
[Frontend] [Parser] Migrate Nemotron V3 to streaming parser engine ( #45755 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-16 05:31:21 +00:00
Bugen Zhao and GitHub
259ff891be
[Rust Frontend] Require ModelConfig.vocab_size to be present ( #45696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 05:30:25 +00:00
6607a80dab
[Bugfix][Gemma4] Fix offline parser truncation, adjust_request token leak, and chat template sync ( #45553 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-16 04:31:53 +00:00
liuzhenwei and GitHub
b8bd773fe4
[XPU] Fix Triton attn fp8/bf16 check failing ( #45758 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-16 12:31:20 +08:00
Ruinan Ma and GitHub
2addbb9cc9
[BugFix] Support async scheduling with prompt embeds for multimodal models ( #45673 )
...
Signed-off-by: Ruinan Ma <r7ma3088@gmail.com >
2026-06-16 04:12:54 +00:00
Isotr0py and GitHub
e3cfea2e1b
[Multimodal] Add Qwen3-VL video loader ( #44412 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 03:45:34 +00:00
Bugen Zhao and GitHub
f99260d2aa
[Rust Frontend] Lower out-of-vocab validation to text layer ( #45685 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 03:37:58 +00:00
Bugen Zhao and GitHub
3f65e21e32
[Rust Frontend] Support max_logprobs validation ( #45674 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 10:57:56 +08:00
xx-thomas and GitHub
b00e76ff72
[Misc][Model] add io processor for query/document embeddings from ColBERT (jinaai/jina-colbert-v2) ( #45210 )
...
Signed-off-by: thomas <thomas.varghese@columbia.edu >
2026-06-16 01:32:32 +00:00
Woosuk Kwon and GitHub
f4359a70f9
[DSV4][Minor] Fix supported KV cache dtypes ( #44892 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-16 00:14:51 +00:00
Itay Alroy and GitHub
3afe659b6b
[EP] Enable DBO with NIXL EP ( #45275 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 23:37:22 +00:00
Itay Alroy and GitHub
16e91176cf
[EP] Query NIXL EP top-k index dtype ( #45298 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:50:18 +00:00
Itay Alroy and GitHub
ab8b0fe338
nixl_ep: Skip post-receive quantization for NVFP4 ( #45606 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:42:05 +00:00
d467a2a7f2
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer ( #45357 )
...
Signed-off-by: llx-08 <2596671364@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-15 21:36:09 +00:00
76a373eff4
[Frontend] Replace legacy Gemma4 parsers with engine-based implementation ( #45588 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 21:34:07 +00:00
Zang Peiyu and GitHub
25ee659db0
Fix parallel_tool_calls: null treated as false instead of default true ( #44955 )
...
Signed-off-by: factnn <166481866+factnn@users.noreply.github.com >
2026-06-15 21:14:10 +00:00
eacff17c8d
[Model Runner V2][Bugfix] Fix MRV2 LoRA warmup ( #35536 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-15 13:17:23 -07:00
Flora Feng and GitHub
cd9078fe59
[Frontend] Skip structural tags for auto tool_choice without strict mode ( #45600 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 19:55:31 +00:00
Wentao Ye and GitHub
e18fe932ca
[Perf] Optimize DSv4 prefill chunk planning, 4.0% E2E Throughput Improvement ( #45061 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-15 19:50:21 +00:00
51ec5cf08f
[Bugfix] Chat Completions Harmony Refactor Clean up ( #45464 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-15 14:45:19 -04:00
7e612a0f06
[KV Offloading] Implement reset_cache for TieringOffloadingManager ( #44541 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 18:42:53 +00:00
+1
0a1c5034f5
[Model] Add MiniMax M3 support ( #45381 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-16 01:01:25 +08:00
RoyWang and GitHub
a3195fab7b
[AMD][Bugfix][Quantization] Honor fused-name match in is_layer_skipped ( #43981 )
2026-06-15 09:37:52 -07:00
Flora Feng and GitHub
0d80979644
[Chore] Consolidate reasoning/tool parser attributes into unified Parser in chat serving ( #45548 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 11:16:45 -04:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
fa63bb9db6
Remove redundant Triton KV cache dtype asserts and enforce architectural support (fp8 >= sm89) ( #43914 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-15 06:49:57 -07:00
5ed15f42b9
Fix the E8M0 scale computation in the MXFP4 (W4A4) MOE CUTLASS kernel ( #43557 )
...
Signed-off-by: Xin He <xin3.he@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-15 06:04:54 -07:00
Juan Pérez de Algaba and GitHub
b997071ec4
(security) Enforce audio upload size limit before full file materialization ( #45510 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-15 10:25:24 +00:00
Martin Kukla and GitHub
6c5872efc5
[Bugfix] Unset HF's default max_new_tokens for DiffusionGemma ( #45417 )
...
Signed-off-by: Martin Kukla <martin.kukla@cantab.net >
2026-06-15 17:31:57 +08:00
wang.yuqi and GitHub
1d88c4dadd
[Docs] Update the online serving docs. ( #45676 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-15 17:23:36 +08:00
vllmellm and GitHub
25c53d1293
[ROCm][Doc] Add installation notes about python version requirement ( #45671 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-15 17:22:55 +08:00
Yejing Lai and GitHub
9872921c5f
[XPU] skip UT test_with_ngram_gpu_spec_decoding ( #44423 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-06-15 08:46:30 +00:00
Reid and GitHub
c17e2f7c84
[Bugfix][Rust Frontend] Make metrics respect --served-model-name ( #45465 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-15 08:05:10 +00:00
FAUST and GitHub
40eac9a9d9
[Rust Frontend] Support parallel_tool_calls = false ( #44760 )
...
Signed-off-by: zhoujinyu <2319109590@qq.com >
2026-06-15 07:50:48 +00:00
b5adb027ad
[Models] Fix MiMo v2.x QKV TP sharding + FP4 support ( #45200 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 15:13:34 +08:00
Sahil Singh and GitHub
64833f8158
[Rust Frontend] Add external→internal request-id map for abort() ( #45137 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-15 06:51:24 +00:00
ddad5dbda2
[Bugfix][Rust] Sync EngineCoreReadyResponse with the Python dataclass ( #45557 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-15 06:49:42 +00:00
Peter Pan and GitHub
ebb0a71ad0
[Bugfix] Reject out-of-range temperature values in SamplingParams ( #44965 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
2026-06-14 23:12:44 -07:00
Ting SUN and GitHub
48df95c43e
[Feature][Frontend] Report multimodal token counts in usage.prompt_tokens_details ( #45458 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-15 05:20:58 +00:00
7df4fe1bd7
[Model] Remove XverseForCausalLM ( #45638 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-14 22:09:00 -07:00
b8336c3c7c
[Bugfix][V1] Split V2 model-runner attention groups on num_heads_q ( #45564 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-14 21:49:46 -07:00
e8d3e22c88
Fix included router missing path for FastAPI >=0.137 ( #45629 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-15 04:28:52 +00:00
c4a3f9d137
[Frontend] Add Streaming Parser Engine and new Qwen3 Parser ( #45413 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 11:59:05 +08:00
Flora Feng and GitHub
e3e3cd5458
[Bugfix][CI] Update Dockerfile dependency graph PNG ( #45602 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 10:35:24 +08:00
Li, Jiang and GitHub
8760f972ca
[CPU] Refine CPU attention frontend ( #45391 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-14 19:26:54 -07:00
b675cb7d0f
[Bugfix][CPU] Honor cgroup memory limit when computing KV cache size ( #45086 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-14 19:26:50 -07:00
Chaojun Zhang and GitHub
2725c84aae
[XPU] Enable sequence parallel support for XPU ( #38608 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-06-14 19:26:46 -07:00
Noa Neria and GitHub
1801fad0ba
[Bugfix] Stream Llama4 weight loading to avoid host-OOM with copy-returning loaders ( #44645 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-06-14 19:23:44 -07:00
Ting SUN and GitHub
3d6ce816f0
[Bugfix][Model] Validate runai_streamer model_loader_extra_config ( #45291 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-14 19:23:30 -07:00
Taneem Ibrahim and GitHub
2c764c089a
Added real /v1/embeddings support for messages + chat_template_kw ( #45173 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-15 09:08:10 +08:00
Michael Ma and GitHub
c621af1690
[BugFix] Fix prompt_embeds for multimodal models ( #45383 )
...
Signed-off-by: ruinan ma <r7ma3088@gmail.com >
2026-06-14 01:44:56 -07:00
Roger Wang and GitHub
e2bf2b3d84
[Perf] Use bisect for mm feature lookup in model runner v2 ( #45566 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-14 00:22:53 -07:00
Amanzhol Salykov and GitHub
725c3bc808
[ROCm][Perf] Enable W4A16 FlyDSL MoE ( #44400 )
...
Signed-off-by: amd-asalykov <asalykov@amd.com >
Signed-off-by: Amanzhol Salykov <asalykov@amd.com >
2026-06-14 00:14:39 -07:00
9548a1887f
[XPU] Support int4 group_size=32 W4A16 MoE ( #45136 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-14 00:14:35 -07:00
Jeff (Junze) Ma and GitHub
9fd737badc
[Bugfix][DCP] Fix illegal memory access in DCP a2a decode under full CUDA graphs ( #45487 )
2026-06-14 00:14:31 -07:00
4ef4492e9b
[V1][Spec Decode] Add Dynamic SD ( #32374 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-14 00:14:27 -07:00
78e7293bb1
[Build] Fix CUDA arch build coverage gaps ( #45277 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Xin Li <xinli-sw@users.noreply.github.com >
Co-authored-by: ShawRong <ShawRong@users.noreply.github.com >
Co-authored-by: Change72 <Change72@users.noreply.github.com >
2026-06-13 22:09:20 -07:00
54bbf51668
[Bugfix] nightly Docker images crash with ImportError: AnthropicOutputConfig since May 28 ( #44795 )
...
Signed-off-by: achyuthan.s <113010327+Achyuthan-S@users.noreply.github.com >
Signed-off-by: Achyuthan S <achyuthan.sivasankar@gmail.com >
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-13 21:45:29 -07:00
Nick Hill and GitHub
cf027b86af
[Core] Simplify MRV2 async output handling ( #45442 )
2026-06-13 18:15:36 -07:00
71b961dd35
[Perf] SM90 cutlass fp8 mm supports odd M by swap_ab, 180~290% kernel performance improvement ( #44572 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 12:05:45 -07:00
521b88c29e
[Bugfix] Reject structured outputs for diffusion decoders with a clear error ( #45468 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-13 12:04:01 -07:00
Harry Mellor and GitHub
b3f0a0a0df
Fix docs build on main ( #45536 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 08:53:23 -07:00
Juan Pérez de Algaba and GitHub
470229c37e
[Security] Fix DoS via prompt_embeds on M-RoPE models ( #45252 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-13 10:17:38 +00:00
2b3006076c
[Security] Add timeout guard for regex compilation in structured outp… ( #45118 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 09:52:56 +00:00
Wentao Ye and GitHub
96fa5cdd9e
[CI Bug] Fix ValueError: There is no module or parameter named 'model.vision_tower.vision_model' ( #45478 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 02:38:37 -07:00
Andreas Karatzas and GitHub
9261dbbc55
Treat null completion max_tokens like the default ( #45491 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 09:34:09 +00:00
Wentao Ye and GitHub
2ecf7d0eb4
[Model Runner V2] Fix openai.InternalServerError: Error code: 500 - 'list index out of range' ( #45467 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 01:44:16 -07:00
midas and GitHub
0d29612292
[Doc] Fix uv dependency resolution failure for setuptools during CPU source builds (x86 & ARM) ( #45412 )
...
Signed-off-by: midas <the.anon.github@gmail.com >
2026-06-13 06:18:58 +00:00
WEI CHENG CHIU and GitHub
5b2943f5a6
[Bugfix] Return the tokenizer from maybe_make_thread_pool so it survives pickling ( #45460 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 06:01:35 +00:00
43f0e024bc
[Render] Add /derender endpoints for disaggregated postprocessing ( #43606 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-13 13:55:33 +08:00
Andreas Karatzas and GitHub
1033ffac2e
[CI] Wait for SSL cert refresher events in the test ( #45489 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 04:57:18 +00:00
ff5a30cfac
[Bugfix] Replace deprecated Qwen2VLImageProcessorFast with Qwen2VLImageProcessor ( #42700 )
...
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-12 21:04:31 -07:00
WEI CHENG CHIU and GitHub
17ee5b1ac5
[Bugfix] Set type/role explicitly in streaming message_start event ( #45376 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 01:40:50 +00:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
Kevin H. Luu and GitHub
e3e31e54b0
[Bugfix][CPU] Don't build triton-cpu on arm64 release image ( #45401 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-06-12 14:51:45 -07:00
badddd254f
[ROCm][DSV4][Perf] Fuse inverse-RoPE and cache bf16 wo_a in o-projection ( #45103 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 15:57:09 -05:00
c90650088d
Add the QuantizedActivation linear-kernel contract ( #44260 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-12 13:48:15 -07:00
Michael Goin and GitHub
9eaacb23ec
[Kernel] Consolidate Marlin thread-tile padding across all dense Marlin paths ( #45295 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-12 13:46:21 -07:00
78739c1946
[Model Runner v2] Migration from v1 to v2, with Qwen and DSv2 MOE models [3/N] ( #42667 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 20:44:52 +00:00
Matthew Bonanni and GitHub
cf567cbc71
[Attention] Improve attention benchmarks: configs and profiling ( #39336 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-12 16:24:25 -04:00
Micah Williamson and GitHub
39cb9bf292
[ROCm] Bump Torch to 2.11 ( #45362 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-12 15:22:26 -05:00
Flora Feng and GitHub
6e4a547176
[Refactor] Deprecate ResponsesParser wrapper, inline parsing into ParsableContext ( #45431 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-12 16:15:41 -04:00
aab639c705
[Core][AMD] Propagate shutdown timeout to MultiprocExecutor ( #43154 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-12 15:13:31 -05:00
efe7adb5e1
[Perf] Use native DSA indexer decode path for next_n > 2 on SM100 ( #45322 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-12 12:54:00 -07:00
Isotr0py and GitHub
6635279d8a
[Migration] Migrate GGUF quantization support to plugin ( #39612 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 12:02:21 -07:00
Jonas I. Liechti and GitHub
d6fd7ce8da
[Model][Dflash] Enable Dflash support for Qwen3NextForCausalLM targets ( #45319 )
...
Signed-off-by: Jonas I. Liechti <j-i-l@t4d.ch >
2026-06-12 10:30:09 -07:00
272c16953e
[Kernel][Helion][1/N] Add Helion kernel for dynamic_per_token_scaled_fp8_quant ( #33790 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-12 12:50:06 -04:00
Yi Zhong and GitHub
053e7daa79
[Model] Add encoder CUDA graph support to Lfm2VL ( #44930 )
...
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com >
2026-06-12 09:17:26 -07:00
5af4aec141
[Rust Frontend] Add standalone granite4 tool parser ( #45216 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 00:16:36 +08:00
Sai Sridhar Tarra and GitHub
a30addc754
[Docs][KV Connector][NIXL] document KV Transfer stat logging and Prometheus metrics ( #44055 )
...
Signed-off-by: Sai Sridhar <tarrasridhar1154@gmail.com >
2026-06-12 15:39:11 +00:00
Chauncey and GitHub
3b8fc3fe6d
[Frontend] Support strict mode for tool calling with ResponsesAPI ( #45396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-12 10:59:59 -04:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
Guan-Ming (Wesley) Chiu and GitHub
c7aa3d2630
[Core] Support structured outputs for beam search ( #35022 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <guanmingchiu@gmail.com >
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-06-12 06:56:25 -07:00
fbc3a1907a
[Bug] Migrate Reset cache for both v2 and v1 model runner ( #42759 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 09:38:12 -04:00
4171ae406c
[V1][Metrics] Add MLA attention metrics for DeepSeek MFU estimation ( #39457 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-06-12 14:28:40 +01:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
8af550b399
[BUGFIX][XPU] Update fa interface for compatibility ( #45394 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-12 11:45:01 +00:00
f1e13f7df9
[Model] Remove Mono-InternVL (InternLM2VEForCausalLM) ( #45129 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 10:41:09 +00:00
88ed636218
[KV Connector]: Support KV push from Prefill to Decode node using Nixl KV Connector ( #35264 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-06-12 10:38:41 +00:00
a014dddbaa
[11b/n] Migrate Machete kernels to torch stable ABI ( #45304 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-12 10:36:49 +00:00
Thomas Parnell and GitHub
a37b4a940e
[Doc] AGENTS.md: add section about coding style ( #45301 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-12 06:23:04 -04:00
Juan Pérez de Algaba and GitHub
f715f25f29
Fix misleading error for audio duration limit rejection ( #45113 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-12 09:58:08 +00:00
Fynn Schmitt-Ulms and GitHub
462ef83d58
Update hidden states extraction integration test triggers ( #45294 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-06-12 01:05:19 -07:00
1ae1051b4b
[Bugfix][Rust Frontend] Return 400 for prompt-validation submit errors ( #45286 )
...
Signed-off-by: xiaguan <751080330@qq.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 07:53:11 +00:00
2043258dec
[Frontend] Support strict mode for tool calling ( #45003 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: cjackal <44624812+cjackal@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 07:51:48 +00:00
bd59c913bc
[CI] ci-fetch-log.sh: fetch all failed jobs from a build URL or PR number ( #45274 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 00:42:18 -07:00
04cec9e4d8
[XPU][DeepSeek-V4] Fix MTP: sync with upstream fixes #44821 and #43746 ( #45240 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 15:41:36 +08:00
Will Eaton and GitHub
87b98d6d6c
[Rust Frontend][Bugfix] Forward --shutdown-timeout and --disable-log-stats to the managed Python engine ( #45300 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-12 07:39:27 +00:00
Yuwen Zhou and GitHub
0cd9b7af25
[CPU] Support CPU W4A16 INT4 MoE ( #43409 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-12 07:12:37 +00:00
Isotr0py and GitHub
a2c72d4388
[Bugfix] Fix Dockerfile dependency graph pre-commit error ( #45374 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 07:10:18 +00:00
Rohan Potdar and GitHub
fe04238292
[ROCm][gpt-oss] Pass GateMode.INTERLEAVE for MXFP4 W4A16 fused MoE ( #44893 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-12 01:02:04 -05:00
39dee1114a
[MM][Perf][CG] Support ViT full cudagraphs for mllama4 ( #40660 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-11 22:17:55 -07:00
+1
eb28452b10
[Model] Add DiffusionGemma Support ( #45163 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Martin Kukla <martin.kukla@cantab.net >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Dipika Sikka <dsikka@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
Co-authored-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Alec Kohlhoff <134344302+aleckohlhoff@users.noreply.github.com >
Co-authored-by: Porras Huang <20535584+porrashuang@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: scoootscooob <167050519+scoootscooob@users.noreply.github.com >
2026-06-11 22:17:35 -07:00
Divakar Verma and GitHub
1ce3cdc5c1
[ROCm][CI] fix fp8 support for test_deepep_moe ( #45302 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-12 00:16:14 -05:00
Dao007forever and GitHub
6fbfdd1831
[NIXL] Per-region KV transfer classification for mixed full-attn + MLA groups ( #44583 )
2026-06-11 21:42:41 -07:00
Chris Leonard and GitHub
7021be66e8
[11a/n] Migrate Marlin kernels to torch stable ABI ( #45176 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-06-11 21:22:37 -07:00
Ekagra Ranjan and GitHub
226ba9fc9e
[ASR] Add Long Audio benchmark and correctness test ( #44587 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-06-12 04:11:16 +00:00
b927004c44
[Bugfix] Mamba CPU Offloading ( #44599 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-11 21:07:35 -07:00
e0b9fb1290
[ASR] Optimize CPU preproc to get 2.5x RTFx via multi-threading ( #44612 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 21:05:11 -07:00
42ae5e7ac6
[Bugfix] Fix --enable-prompt-tokens-details omitting zero cached tokens ( #44383 )
...
Signed-off-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-11 20:37:42 -07:00
Nick Hill and GitHub
2263f8a3de
[CI][BugFix] Fix broken test_mamba_prefix_cache.py due to stale mock ( #45345 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 03:26:17 +00:00
Ting SUN and GitHub
c1076839c9
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads ( #45308 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-11 20:21:46 -07:00
fcf5115c45
[ROCm][DSv4][Perf] Flash-decode split-K decode attention kernel ( #44899 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-12 03:17:52 +00:00
4bc83323f2
[Bugfix] OffloadingConnector: respect skip_reading_prefix_cache flag ( #44592 )
...
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: littlecircle0730 <43994952+littlecircle0730@users.noreply.github.com >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-06-12 02:20:39 +00:00
yzong-rh and GitHub
e0871ad225
[Refactor] Chat Completions Streaming Harmony Refactor and Bugfixes ( #45104 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-12 01:09:47 +00:00
6f573f486b
[Bugfix] Initialize missing attributes in mistral eagle ( #45217 )
...
Signed-off-by: jpwang <jpwang@smail.nju.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 08:21:01 +08:00
Neil Schemenauer and GitHub
9bbf42be26
Make mistral_common optional by deferring MistralToolCall import ( #45305 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
2026-06-11 22:59:11 +00:00
8a91228dbe
[Bugfix][KVConnector][Mooncake] Close MooncakeDistributedStore on connector teardown ( #45206 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 14:33:48 -07:00
yzong-rh and GitHub
f712fd0d7d
[Refactor] Chat Completions Harmony Refactor, non-streaming path. ( #45171 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-11 21:18:30 +00:00
Wentao Ye and GitHub
5a6c7b7ab5
[Bug] Fix test flashmla for DSv4 ( #45052 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 16:22:26 -04:00
c9340e6f35
[Model] Remove InternLMForCausalLM registry alias ( #45128 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 20:02:51 +00:00
Ben Browning and GitHub
235b63c004
[Bugfix] Fix Anthropic tool_use content handling dropping args ( #45287 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-11 20:01:29 +00:00
3b03a2cf47
[Rust Frontend] Support continuous_usage_stats stream option ( #43965 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 17:50:59 +00:00
wentian-byte and GitHub
b8142294b7
[Bugfix] Restrict FlashInfer cuDNN FP8 ViT attention gate to Blackwell (SM 100) ( #45251 )
...
Signed-off-by: Wentian Byte <3400259131@qq.com >
2026-06-11 16:39:24 +00:00
2ec6594db9
[Kernel][Helion][1/N] Add Helion kernel for per_token_group_fp8_quant ( #36902 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 08:59:08 -07:00
vraiti and GitHub
79f8c5bd8c
[Metrics] Scope unregister_vllm_metrics() to strictly "vllm:" metrics ( #42331 )
...
`unregister_vllm_metrics()` currently uses "vllm" in `collector._name` to decide
which collectors to remove from the Prometheus registry, removing every even
metrics registered by other subsystems or downstream extensions like "vllm_omni:"
Signed-off-by: vraiti <vraiti@redhat.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-11 15:43:14 +00:00
Jiangyun Zhu and GitHub
f81daf8880
[Attention] add triton diff-kv backend for mimo ( #41797 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-11 11:36:31 -04:00
4085ff7cb4
[Core] Add kvcache watermark to reduce preemptions ( #44594 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-11 08:27:31 -07:00
23eb7c8fbb
[Bugfix] Fix NixlEPAll2AllManager's dependency on --enable-elastic-ep to function ( #44422 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-11 08:14:49 -07:00
wineandchord and GitHub
c2b4cd39ac
[Doc][Attention] Fix MLA top-of-file comments ( #37047 )
...
Signed-off-by: wineandchord <guoqizhou19@gmail.com >
2026-06-11 08:14:45 -07:00
Kai K. and GitHub
f1d8d99717
[Bugfix] CohereModel.load_weights: skip modelopt _quantizer.* keys ( #43495 )
...
Signed-off-by: Kai Köhler <kai.koehler@web.de >
2026-06-11 08:14:21 -07:00
Nicolò Lucchesi and GitHub
750aab5b8e
[Bugfix] Fix CPU memory leak related to not cleaning up old remotes data ( #44424 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-11 07:54:52 -07:00
5edf7ff489
[Core] Release cached device memory under pressure on UMA GPUs during weight loading ( #45179 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 17:49:50 +03:00
b78fc47f05
[Docs] Add redirect for moved lmcache examples page ( #45218 )
...
Signed-off-by: nataliepjlin <nataliepjlin@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:41:08 -04:00
Harry Mellor and GitHub
03878d1c22
Deprecations for v0.23 and v0.24 ( #44992 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 14:35:38 +00:00
55911db580
[PD][Core] Fix Mamba prefix cache hit rate in PD disaggregation ( #44243 )
...
Co-authored-by: lHrHenry233 <2381623149@qq.com >
Co-authored-by: underfituu <hzhucong@163.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-06-11 14:10:25 +00:00
cc640ee8bc
[Rust Frontend][Metrics] Export vllm:lora_requests_info from frontend ( #45030 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 06:45:03 -07:00
ebc6ef971a
Hidden states extraction improvements ( #43805 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 09:44:45 -04:00
tc-mb and GitHub
ab3a1fd2e6
minicpmv4_6: fix ImageSize (W,H) order for placeholder token calculation ( #45244 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-11 13:43:56 +00:00
c3662b36ea
[KV offload] Parallel-agnostic fs-tier cache for single full-attention group ( #44733 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-06-11 15:48:37 +03:00
Juan Pérez de Algaba and GitHub
e62d00ab73
docs: add fix disclosure policy to SECURITY.md ( #45253 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 12:48:00 +00:00
1f60771c74
fix: guard flash-attn rotary import ( #42679 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-11 08:43:31 -04:00
05d9848267
[Build] Upgrade CUDA Dockerfiles from GCC 10 to GCC 12 for C++20 compatibility ( #44923 )
...
Signed-off-by: Richard Barnes <rbarnes@meta.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:26:52 +00:00
jasen and GitHub
ef67071b21
[Build] Skip spinloop extension on Python < 3.11 ( #44783 )
...
Signed-off-by: Jasen2201 <yajizhan@amd.com >
2026-06-11 11:23:21 +00:00
x41lakazam and GitHub
3508cb78d4
[Bugfix] Fix broken profile_modular_kernel.py ( #43300 )
2026-06-11 12:17:23 +01:00
Harry Mellor and GitHub
432905d5d6
Only enable PR docs builds manually ( #45262 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 03:14:29 -07:00
1f9dd7900d
[Bugfix][Rust Frontend] Validate out-of-vocab token ids in request params ( #44680 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 03:14:11 -07:00
9492362972
[Security] Apply sanitize_message to Anthropic and STT error paths ( #45119 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 10:05:34 +00:00
7852e50e4d
[docs] Document --scheduler-cls base class requirement (extend AsyncScheduler, not Scheduler) ( #43724 )
...
Signed-off-by: Georgii Kliukovkin <kliukovkin@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:49:51 +01:00
Reid and GitHub
0d657e44dc
[Rust Frontend] Fix DeepSeek V3.2 continue_final_message rendering ( #45155 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-11 09:34:19 +00:00
aa1df36c53
Fix/minicpmv46 missing version ( #44980 )
...
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-11 09:20:45 +00:00
f06aefb4e3
[CPU] Add missing scalar fallback for CPU W4A8 INT4 GEMM ( #44523 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
2026-06-11 08:52:01 +00:00
Julien Denize and GitHub
1c3a72b8b2
[Bugfix] Add fetch_images to MistralCommonImageProcessor ( #45180 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-11 16:13:01 +08:00
Juan Pérez de Algaba and GitHub
d598d23973
[Security] Reject non-finite temperature and repetition_penalty values ( #45116 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 01:12:14 -07:00
Juan Pérez de Algaba and GitHub
f219788f91
[Security] Fix info disclosure via int32 truncation in GGUF dequantize kernels ( #44971 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 08:05:14 +00:00
6e64c1bab1
[10c/n] Migrate MoE kernels to torch stable ABI ( #44565 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-10 23:02:26 -07:00
Kevin H. Luu and GitHub
2f2c5cf4f1
[release] Always block release images to dockerhub ( #45236 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-10 22:53:04 -07:00
Mohammad Miadh Angkad and GitHub
40e065e86a
[Docker] Fix CUTLASS DSL cu13 install order in Dockerfile ( #45204 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-11 05:19:36 +00:00
0b995f8609
Use std::bit_cast for type punning in CPU kernels ( #45089 )
...
Signed-off-by: Yuanyuan Chen <cyyever@outlook.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-10 22:07:44 -07:00
Bugen Zhao and GitHub
43914dd743
[Rust Frontend] Add Python bridge for Rust tool parsers ( #44624 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 04:51:06 +00:00
3501324957
[Build] fix self-contradictory precompiled-flag orthogonality test ( #44942 )
...
Signed-off-by: pjdurden <prajjwalchittori1@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:49:08 +08:00
Flora Feng and GitHub
3a04061701
[Refactor][Parser] Unify Response API to use parser.parse() like Chat Completion API ( #45190 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-11 04:37:51 +00:00
Yifan Qiao and GitHub
f272dfdce1
[KV Connector] Mooncake store: prefix-cache retention interval for sparse attention ( #44774 )
2026-06-10 21:36:34 -07:00
velonica0 and GitHub
f31bc2ea60
[CPU][RISC-V] Enable oneDNN W8A8 INT8 to run on RISC-V ( #44478 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-06-11 04:09:05 +00:00
248e33c40d
[Bugfix][Responses API] Set id on function_call item in streaming done event ( #44608 )
...
Signed-off-by: Aniruddh Krovvidi <aniruddh.krovvidi@oracle.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-11 03:52:42 +00:00
Bugen Zhao and GitHub
5d5591d99b
[Rust Frontend] Populate cached_token_count in responses ( #44887 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 20:50:05 -07:00
Wentao Ye and GitHub
85a0ffae42
[CI Bug] Remove qwen test ValueError: No example model defined for Qwen/Qwen-7B-Chat ( #45194 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-10 20:11:00 -07:00
Harry Mellor and GitHub
18d87a87dc
Deprecate Transformers v4 support ( #45161 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 11:04:01 +08:00
Flora Feng and GitHub
b038a2f73b
[CI][Bugfix] Update Dockerfile dependency graph PNG ( #45209 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 19:40:25 -07:00
Ting SUN and GitHub
2d481f8a94
[Bugfix][Rust Frontend] Stop unescaping XML-style tool-call parameter values ( #45025 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-10 19:05:23 -07:00
7920ccb97c
[Bugfix]: Fix Quark gpt-oss weight loading broken by FusedMoe refactor ( #45067 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 18:17:46 -07:00
Wentao Ye and GitHub
86111c00c7
[Chore] Add Github notification for MRv2 for @yewentao256 ( #45191 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 09:01:49 +08:00
qizixi and GitHub
e2db0222e9
[Perf][Attention] Pin MLA chunked-context metadata tensors so H2D copies are truly non-blocking ( #45074 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-10 15:56:49 -07:00
82d6b59f04
[CI/Build] Skip test_use_trtllm_attention on non-CUDA platforms ( #44687 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 18:18:42 -04:00
Andreas Karatzas and GitHub
16282a9c4e
[ROCm][CI] Moving MI300 tests to MI325 until cluster is stabilized ( #45170 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 20:26:17 +00:00
5b6b536fdc
[ROCm][Bugfix] Make intermediate_pad TP-aware in rocm_aiter_fused_experts ( #44679 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 15:10:50 -05:00
12f3f19c19
feat(qwen3-asr): support prompt parameter in v1/audio/transcriptions ( #35415 )
...
Signed-off-by: Nathan Price <nathan@abridge.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 19:54:59 +00:00
Ilya Markov and GitHub
6471ec75bd
[EPLB] Reject NCCL-based EPLB communicators with async EPLB ( #44978 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
2026-06-10 19:51:27 +00:00
3d300aecb1
[Doc] Switch K8S examples to default MP mode ( #39400 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
Signed-off-by: Peter Pan <peter.pan@daocloud.io >
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-10 18:17:11 +00:00
ffce72c041
[Model Runner V2] Fix v2 AttributeError: 'CohereASRDecoder' object has no attribute 'embed_input_ids' ( #44568 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 11:06:01 -07:00
TJian and GitHub
bfe1001ab6
[Bugfix] [DSV4] [ROCm] Pin apache-tvm-ffi version to 0.1.10 ( #45169 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-10 17:41:15 +00:00
fa8c868a3c
[Bugfix] Fix Llama4 weight loading ( #45047 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-10 13:40:45 -04:00
Ben Browning and GitHub
d1bcb4b44c
[Bugfix] Fix tool parsing crash with non-function tool types (e.g. WebSearchTool) ( #45147 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-10 17:17:16 +00:00
bnellnm and GitHub
29026682cb
[Bugfix] Fix nemotron accuracy drop introduced by #41184 ( #45037 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 13:16:25 -04:00
Stan Wozniak and GitHub
dc66e01a70
[Hybrid] Marconi-style admission policy for hybrid cache ( #37898 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-10 10:03:13 -07:00
Yongye Zhu and GitHub
2ba68d9bf7
[Test] Fix one-sided MNNVL alltoall test workspace under-reservation ( #44946 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-11 00:43:12 +08:00
Julien Denize and GitHub
2131b597b1
[CI] Ping Mistral team for ministral/voxtral/mixtral/pixtral changes ( #45153 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-10 08:48:00 -07:00
0bae1d3848
[MRV2][Spec Decode] DFlash ( #44586 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-10 08:47:46 -07:00
Yufeng He and GitHub
4673ca1d78
fix: prefix DeepSeek V4 MTP projections ( #44821 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
2026-06-10 08:47:04 -07:00
Angela Yi and GitHub
de900fa7e5
fix: AOT compile cache collision for dataclass-based HF configs ( #45059 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-06-10 08:05:29 -07:00
Divakar Verma and GitHub
166d14e9bf
[bugfix] skip conch kernel for g_idx reordering ( #45072 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-10 23:04:19 +08:00
af65e08fc5
KV-Cache multi-tier offloading async batched lookup ( #44193 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 14:59:30 +00:00
Harry Mellor and GitHub
3cc9fecd58
Deprecated 1st generation Qwen and QwenVL models ( #45131 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 14:55:33 +00:00
ccc05de038
[Bugfix] Fix missing sequence_lengths in EXAONE-4.5 vision encoder ( #45073 )
...
Signed-off-by: Jongsu Liam Kim <jongsukim8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 15:44:34 +01:00
6ec7dcd641
[Frontend][Metrics] Add vllm:tool_call_parser_invocations_total Prometheus metric ( #44448 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 10:29:11 -04:00
c9e5bf8135
[Bugfix] Fix layerwise reload dropping params after a composed weight loader ( #44814 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-06-10 06:42:05 -07:00
Roberto L. Castro and GitHub
6850839c6f
[Perf] Fix dsv3_router_gemm heuristic ( #44217 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-10 06:08:41 -07:00
87c15d46e3
[Bugfix] Lazily import the humming quantization backend ( #44921 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 06:06:17 -07:00
4882fd7632
[Bugfix][Reasoning] Nemotron V3: surface reasoning as content when thinking is unterminated ( #39091 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-06-10 05:58:19 -07:00
77f42d9725
[Model] Remove obsolete ERNIE models ( #45127 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 20:54:30 +08:00
9dfc313bdc
Feature/offloading manager stats ( #35669 )
...
Signed-off-by: Sriusa4414@gmail.com
Signed-off-by: srinivas_oo7 <Sriusa4414@gmail.com >
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 12:44:55 +00:00
9ad08c4d15
[Bugfix][Rust Frontend] Fix missing added tokens in hf/fastokens tokenizer ( #44683 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 03:52:41 -07:00
Shantipriya Parida and GitHub
a1ec011a83
[Bugfix] Add deepseek_v32 to Quark dynamic MXFP4 model type check ( #39498 )
...
Signed-off-by: Shantipriya Parida <shantipriya.parida@amd.com >
2026-06-10 02:52:33 -07:00
Bugen Zhao and GitHub
fdfb2566c0
[Rust Frontend] [CI] Unify Rust artifact builds with setuptools-rust ( #44981 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 17:48:34 +08:00
Juan Pérez de Algaba and GitHub
8a5cf1ccd6
[Security] Fix remote DoS via invalid recovered token reinjection ( #44744 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 02:31:43 -07:00
Kunshang Ji and GitHub
fe1d923afc
[BUGFIX][XPU] fix xpu flash_attn_varlen_func interface ( #45110 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-10 17:07:40 +08:00
32daf56b42
[Refactor] Rename rocm_moe.py to rocm_moe_rdna.py ( #45011 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 17:02:09 +08:00
Andreas Karatzas and GitHub
82a42234be
[ROCm][CI] Defer AITER sampler import and isolate server test PYTHONPATH ( #44823 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 08:56:11 +00:00
Harry Mellor and GitHub
af9f583344
Revert "[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures" ( #45029 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 01:37:03 -07:00
yiheng and GitHub
bd2d83ff31
[SpecDecode] Reduce TP communication for large-vocab draft models speculative decoding ( #39419 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
2026-06-10 07:59:24 +00:00
xiaohuguo2023 and GitHub
bb78168b21
[ROCm][gpt-oss] Hybrid CDNA4 swizzle gate for A8W4 MoE ( #44804 )
...
Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com >
2026-06-09 23:59:44 -07:00
89c6a41001
[Bench] Add BFCL dataset for vllm bench serve tool-calling workloads ( #42457 )
...
Signed-off-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-09 23:59:18 -07:00
7fdfa6441d
Model/colbert autoweightsloader ( #44999 )
...
Signed-off-by: Furkan Fidan <dev@yufufi.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 23:58:50 -07:00
Harry Mellor and GitHub
e9b728de8a
Change from owning configs to owning config utils ( #45058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 06:40:25 +00:00
5828a205ef
Fix Harmony tool descriptions for optional fields ( #44686 )
...
Signed-off-by: Varun Shenoy <varun.vinayak.shenoy@oracle.com >
Co-authored-by: Codex <codex@openai.com >
2026-06-09 23:29:22 -07:00
Yaoming Zhan and GitHub
7a74f31d2e
[Rust Frontend] Add seed_oss and step3p5 reasoning parsers ( #44552 )
...
Signed-off-by: yzhan1 <zhanyaoming2014@gmail.com >
2026-06-09 23:01:33 -07:00
47930b59ca
[Bugfix] Handle HWC images in ImageProcessorItems.get_image_size ( #45057 )
...
Signed-off-by: YellowFoxH4XOR <yellowfoxh4xor@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 05:35:50 +00:00
Flora Feng and GitHub
6aec99f030
[Refactor] Remove dead states from chat completion serving ( #45081 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-09 22:20:15 -07:00
bnellnm and GitHub
f4966f8b3d
[Bugfix] Fix weight loading issues caused by #41184 ( #45054 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 01:20:13 -04:00
Mohammad Miadh Angkad and GitHub
2c9c07c85e
[Bugfix][CI/Build] Fix Rust frontend build after chat conversion refactor ( #45085 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 20:04:41 -07:00
Change72 and GitHub
320c52b134
[Bench] benchmark_serving_multi_turn: make non-standard conversation_id payload opt-in ( #43756 )
...
Signed-off-by: Change72 <cguo51@asu.edu >
2026-06-09 19:41:56 -07:00
6deb05e0e4
[Core][Model] Gemma4: Unified FA4 for all layers + FlashAttention mm_prefix support ( #42175 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-09 17:45:39 -07:00
Flora Feng and GitHub
d82ac00923
[Refactor][Mistral] Extract parsing logic into MistralParser ( #44596 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 00:12:23 +00:00
Bugen Zhao and GitHub
dac9e9a640
[Rust Frontend] Extract shared options in route helper params ( #44884 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 17:02:35 -07:00
Wentao Ye and GitHub
d7607ad273
[Bug] Fix deepseek v4 OOM issue ( #44914 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-09 15:47:06 -07:00
d955745d58
[ROCm][CI] fix test_rope_kvcache_fusion.py ( #44678 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-09 21:53:46 +00:00
Micah Williamson and GitHub
e1ed89dbee
Revert "[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuf… ( #45066 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 14:12:06 -07:00
1c2ffc6f88
feat(multi-turn-bench): add api_key and custom headers for multi turn benchmark ( #44516 )
...
Signed-off-by: Jimmy <jinmingyi1998@sina.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: simon-mo <simon.mo@hey.com >
2026-06-09 14:00:07 -07:00
Jiangyun Zhu and GitHub
ca4cfd8731
[Bugfix] fix qwen3.5 ep weight loading ( #45002 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 13:55:30 -07:00
Micah Williamson and GitHub
c9c1540e61
[ROCm][V2] Fix failed assertion in Llama models when using EAGLE with ROCM_AITER_FA ( #44936 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 13:30:52 -05:00
c1d754d681
[Mooncake] Use all HCAs on multi-NIC hosts instead of GPU-indexed RNIC selection ( #43799 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-06-09 11:05:36 -07:00
01d8cd92dd
[ROCm][Perf] Use fused softplus-sqrt-topk router under AITER fused-MoE ( #44945 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-09 17:53:05 +00:00
a4b14b98c6
[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuffle reduction + vectorized I/O ( #44173 )
...
Signed-off-by: SII-yangdian <yangdian@sii.edu.cn >
Co-authored-by: SII-yangdian <yangdian@sii.edu.cn >
2026-06-09 10:41:26 -07:00
Juan Pérez de Algaba and GitHub
cf1c906724
[Security] Fix image EXIF orientation and tRNS transparency handling ( #44974 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-09 09:34:44 -07:00
766ce2bb6b
Fix MiDashengLM TP>1 crash in audio encoder attention ( #44408 )
...
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-09 09:29:14 -07:00
3d119f78f7
[Docs] Add KV offloading usage guide (single- and multi-tier) ( #44415 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 19:20:23 +03:00
Juan Pérez de Algaba and GitHub
1b1359c332
[Security] Fix DoS via audio decompression bomb in speech-to-text endpoint ( #44970 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 00:18:53 +08:00
Tyko Niemi and GitHub
cad4ca12b8
[Bugfix] Add X-Session-ID from conversation_id in multi-turn benchmark ( #44663 )
...
Signed-off-by: Tyko Niemi <tyko.niemi@amd.com >
2026-06-09 08:57:00 -07:00
Andreas Karatzas and GitHub
b697119800
[ROCm][CI] Stabilize ModernBERT token-classification parity against Hugging Face ( #44040 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:52:36 +01:00
Kunshang Ji and GitHub
b4c6dc6454
[WIP][XPU] upgrade torch-xpu to 2.12 ( #42262 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-09 15:51:39 +00:00
Raushan Turganbay and GitHub
2ee5106372
Remove raw_inputs from transformers backend ( #39425 )
...
Signed-off-by: raushan <raushan@huggingface.co >
2026-06-09 15:01:04 +00:00
Jiangyun Zhu and GitHub
7a89b72564
[Perf] fuse qk rmsnorm rope gate for qwen3.5 ( #44176 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 22:12:17 +08:00
Jee Jee Li and GitHub
dc10e467a9
[Bugfix] Fix minimax_qk_norm_fusion ( #44983 )
2026-06-09 06:43:46 -07:00
Terrence Zhao and GitHub
ee4d7df2b5
[Cohere] Cohere2 moe parser fix ( #44907 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:32:18 -07:00
Terrence Zhao and GitHub
3e8afdf785
[Cohere] Fix Cohere2MoE weight loading when using Transformers ≥5.10 ( #44747 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:27:40 -07:00
Nicolò Lucchesi and GitHub
6690a0c4de
[PD][Bugfix] Fix KV Cache sharing with HMA ( #44629 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 06:10:06 -07:00
Maria Guevara and GitHub
1c23c42030
[Rust Frontend] Support Kimi K2 tool call IDs ( #44901 )
2026-06-09 05:31:26 -07:00
xiangdong and GitHub
b12e42d132
[XPU][CI] Refine docker image build and pull/create lock mechanism in Intel GPU CI ( #44481 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-09 20:20:32 +08:00
69fdaffbcd
[Rust Frontend] Add /tokenize and /detokenize endpoints ( #44222 )
...
Signed-off-by: Tan Ngoc Do <darkknightkhtn2008@gmail.com >
Signed-off-by: TanNgocDo <darkknightkhtn2008@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 05:11:37 -07:00
80e2c4462d
[ROCm][Compile] Fuse AR + RMSNorm + per-group FP8 quant (+ DSv3.2 indexer fan-out) ( #42864 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
2026-06-09 12:06:56 +00:00
Sage and GitHub
5b3807e862
[KV Events] Switch event structs from array to map encoding ( #42892 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-06-09 11:39:52 +00:00
Qiuyang Yue and GitHub
59401ac9f1
[Kernel][Perf] Tune fused_moe FP8 config for Qwen3-Next-80B tp=4 on H100 (+25% at batch 96-512) ( #44830 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-09 04:15:51 -07:00
d841386d27
[Rust Frontend] Support API key authentication ( #44321 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 10:15:20 +00:00
Mohammad Miadh Angkad and GitHub
fff9210b2a
[CI/Docs] Remove stale disagg prefill links ( #44918 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 03:05:53 -07:00
Ma Jian and GitHub
70db1488c5
[DSV4][XPU] Add MHC fused_post_pre support ( #44144 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-09 17:23:17 +08:00
Andreas Karatzas and GitHub
2385e140d6
[ROCm][CI] Stabilize sleep-mode memory release ( #43022 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:51:12 +08:00
Nicolò Lucchesi and GitHub
dab60fc658
[Bugfix][CI] Fix test_offloading_connector.py::test_fs_tiering_offloading ( #44903 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 00:57:34 -07:00
wang.yuqi and GitHub
996222f4bf
[CI] Reorganize entrypoints CI ( #44947 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 00:46:11 -07:00
e6fc848d4f
[Bugfix][MiniCPM-o] Fix cuda/cpu device mismatch in Resampler2_5 pos_embed ( #43844 )
...
Signed-off-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 23:28:26 -07:00
Andreas Karatzas and GitHub
f843ac1a1c
[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures ( #44952 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 05:49:30 +00:00
7c2aa3108a
fix: prevent MM cache hang from stale LRU order keys ( #43595 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:48:31 -07:00
ebf53ba373
[Bugfix][Rust Frontend] Set a structured-output backend so requests do not 500 ( #44729 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 22:30:54 -07:00
baacbfcebf
[ROCm][MLA][Bugfix] Reserve FP8 prefill workspace before lock for Kimi-K2.5 ( #42978 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 22:25:52 -07:00
d8218b1ee7
[Bugfix] Propagate ImportError from load_audio_pyav when vllm[audio] … ( #44750 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
2026-06-09 04:24:52 +00:00
9f153aa781
[MM][Perf][CG] Support ViT full CUDA graph for glm4_1v image and video inference ( #40576 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-09 11:13:56 +08:00
Kunshang Ji and GitHub
d3de61502f
[XPU][CI] fix test case path ( #44940 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-08 20:02:31 -07:00
Lanze Liu and GitHub
540aaf2140
[Bugfix][Model] Qwen3-Omni: move cu_seqlens to GPU before VIT attention ( #44264 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-08 20:02:27 -07:00
Lanze Liu and GitHub
4128605ad4
[Docs] Remove broken link to deleted disaggregated_prefill.sh ( #44929 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-09 01:40:06 +00:00
e2f993dc41
[WideEP] Integrate DeepEP v2 ( #41183 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-08 18:07:29 -07:00
Andreas Karatzas and GitHub
05cb606cad
[ROCm][CI] Re-route NixlConnector jobs ( #44809 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 18:57:11 -05:00
3f627ebef7
[Misc] usage_stats: report more engine, spec-decode, and EP config ( #44595 )
...
Signed-off-by: Zach Xi <zachary.xi@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-08 15:20:00 -07:00
Bugen Zhao and GitHub
bc941f375d
[Rust Frontend] [Refactor] Refine utility call interfaces ( #44856 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 15:13:08 -07:00
Michael Goin and GitHub
6afa25000c
[Bugfix] Canonicalize FP8 weight layout to (K, N) at the source ( #44735 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-08 14:37:36 -06:00
Mohammad Miadh Angkad and GitHub
823a0ab754
[Bugfix][MoE] Fix fused MoE expert mapping helper call sites ( #44897 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-08 13:35:04 -07:00
Wentao Ye and GitHub
2c27c294c0
[Model Runner V2] Fix mrv2 mm lora issue ( #44450 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-08 14:30:09 -04:00
ba94a3b998
[Attention] Extract KV-cache update from CPU attention backend ( #40470 )
...
Signed-off-by: Diego Maniloff <diego.maniloff@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-08 15:43:05 +00:00
bnellnm and GitHub
dc68bd8c41
[MoE Refactor] FusedMoE/MoERunner inversion refactor ( #41184 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-08 10:42:58 -04:00
753e9d55e6
[Quantization] add online fp8 ptpc ( #44132 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:42:11 +08:00
akii96 and GitHub
ac3409d162
[Benchmark] Auto-detect and correct client/server tokenizer mismatch for random dataset ( #44708 )
2026-06-08 06:10:20 -07:00
wang.yuqi and GitHub
93ee4cd47f
[CI] Consolidate multimodal entrypoint tests. ( #44819 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-08 04:48:08 -07:00
Li, Jiang and GitHub
980796cd07
[CI/Build][CPU] Fix flaky CI image build failure and unexpected warnings ( #44852 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-08 11:10:06 +00:00
Nicolò Lucchesi and GitHub
5add018beb
[Connector] Remove P2pNcclConnector ( #44854 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-08 18:58:29 +08:00
d5fe994e79
[CPU][Spec Decode] Warn about throughput loss when libiomp5 is not preloaded ( #44419 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <bigpyj64@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-08 03:45:08 -07:00
Chaojun Zhang and GitHub
fa662b1a8b
[XPU] Cap topk/topp Triton BLOCK_SIZE to 4096 to fix Top-p mask difference failures ( #44470 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-08 09:36:51 +00:00
3c0b4432be
[Rust Frontend] Add /pause, /resume, /is_paused endpoints ( #44499 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 17:28:37 +08:00
Sungjae Lee and GitHub
469f3dcf1d
[BugFix] Use served model name in gemma4 audio-tower error message ( #44828 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
2026-06-08 06:58:31 +00:00
xiangdong and GitHub
94fcdd007f
[XPU][CI] Add more test cases in Intel GPU CI ( #43663 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-08 06:21:24 +00:00
Andreas Karatzas and GitHub
d9ff7e4e9a
[ROCm][CI] Stabilizing teardown and timeout of flaky tests to prevent rare OOMs ( #44761 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 14:11:17 +08:00
Andreas Karatzas and GitHub
967c5c3bc3
[ROCm][CI] Stage C mirrors ( #42793 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 23:00:59 -07:00
Yan Ma and GitHub
54c660c3a6
[XPU][Minor] format moe kernel name and add in kernel list ( #44771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-08 13:58:16 +08:00
Shanshan Shen and GitHub
8fb0274415
[MM][CG] Simplify ViT CUDA graph interfaces ( #44484 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-08 05:57:06 +00:00
Ma Jian and GitHub
eebce65756
[XPU]feat: add DeepSeek-V4 XPU attention decode path ( #42953 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-08 13:27:12 +08:00
303916e93d
[Bugfix]: Fix assertion in MambaManager.allocate_slots() ( #39562 )
...
Signed-off-by: Holworth <kangqihan17@mails.ucas.ac.cn >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-08 00:34:37 -04:00
Taneem Ibrahim and GitHub
5633405964
Added extra_repr() to pooler classes to improve debuggability ( #44805 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-08 03:19:31 +00:00
6124a98a9b
[Bugfix] Fix FunASR-Nano crash during initialization ( #44215 )
...
Signed-off-by: SunskyXH <sunskyxh@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-07 20:00:02 -07:00
Agata Dobrzyniewicz and GitHub
2ed0a9627b
[Kernel][Test] Make kernel tests for mamba dual-HW (CUDA + XPU) ( #42736 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-08 08:22:47 +08:00
4dcd10eb0d
[1/N][KV-Cache Layout Refactor] Refactor DSV4 KV cache config construction ( #44454 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-07 14:53:37 +00:00
Charlie Fu and GitHub
228bcc436b
[ROCm][Kernel] Enable permute_cols for ROCm ( #44674 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-07 09:50:03 +00:00
3d3ba460a2
Modify torch dependency in xpu.txt ( #43087 )
...
Signed-off-by: Bram Vanroy <2779410+BramVanroy@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-07 16:33:50 +08:00
Mohammad Miadh Angkad and GitHub
66ecfd0568
[Dependency] Remove stale cuDNN frontend upper bound ( #42599 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-07 16:09:25 +08:00
Andreas Karatzas and GitHub
f0f6805d8a
[CI] Stabilize the multi-audio OpenAI server path ( #44051 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 15:54:32 +08:00
15652a6b70
[Doc] Fix multimodal torch.compile troubleshooting to not use removed VLLM_TORCH_COMPILE_LEVEL ( #44378 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-07 00:34:07 -07:00
Yifan Qiao and GitHub
51ef688831
[Bugfix][Mooncake] Fix per-group block_size/block_hash and group_idx in MooncakeStoreConnector KV events ( #44103 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-07 07:12:43 +00:00
Jared Wen and GitHub
6ac69203e8
[videoloader] implement glm46v video loader ( #44417 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-07 06:27:20 +00:00
1505b3d8a1
[Cohere] Enable Cohere Mini Code model and update Command A-plus test registry ( #44707 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:44:16 -07:00
32f34d3935
[feature] add index share feature for DSA MTP ( #44420 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:04:14 -07:00
Qiuyang Yue and GitHub
9c7f7741d4
[Bugfix] Fix benchmark_moe.py after inplace mechanism removal ( #44041 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-07 00:32:00 -04:00
6181e80fe0
[XPU] add xpu branch in compressed_tensors_moe_w4a4_mxfp4 ( #44540 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Co-authored-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-07 12:27:34 +08:00
Yan Ma and GitHub
3bb46975bd
[XPU][Feature] transparent sleep mode support for XPU platform ( #37149 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-07 10:45:31 +08:00
Chaojun Zhang and GitHub
810966453a
[XPU] Support cpu kv offloading and tiering offloading on XPU platform ( #36423 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-07 09:59:28 +08:00
Woosuk Kwon and GitHub
2a983c79ac
[DSV4] Decouple DS V4 Sparse MLA Metadata from DS V3.2 ( #44699 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-06 20:37:56 -04:00
bc5745a00f
[ROCm][MLA] Replace torch.cat in sparse-MLA forward_mqa with fused concat_mla_q ( #42838 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 18:20:50 -05:00
Nick Hill and GitHub
3b3d5287fa
[BugFix] Resolve multiple async kv load deadlock ( #44560 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-06 23:05:47 +00:00
062b05ff3a
[ROCm][Perf] Fused MoE W4A16 HIP kernel for AMD RDNA3 (gfx1100) ( #44075 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 15:30:39 -05:00
fa27d4e9cf
[PERF] [Qwen3.5] Split mixed prefill+decode batches: route decodes to the recurrent kernel ( #44700 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:13:50 +08:00
Vadim Gimpelson and GitHub
67d3792d99
[Bugfix] Fix Qwen3.5-FP8 nightly fail. Guard fused_add_rms_norm input/weight dtype mismatch in RMSNorm + quant fusion ( #44694 )
2026-06-06 08:46:14 -04:00
00d1fb7747
[Bugfix][ROCm] ApplyRotaryEmb: fall back to native when flash_attn rotary grid would exceed the HIP per-dim limit ( #43684 )
...
Signed-off-by: vLLM ROCm fix <noreply@example.com >
Signed-off-by: amd-fuweiy <fuweiy@amd.com >
Co-authored-by: vLLM ROCm fix <noreply@example.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 02:29:16 -07:00
c9b4b184b4
[Bugfix][Voxtral] Add fetch_audio to MistralCommonFeatureExtractor (transformers>=5.10 compat) ( #44559 )
...
Signed-off-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 07:58:09 +00:00
f87df1df9e
[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig ( #44613 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-05 23:14:41 -07:00
Taneem Ibrahim and GitHub
eafbb06331
[Misc] Replaced asserts with proper exceptions to improve UX for pooling ( #44593 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-06 05:57:26 +00:00
ec0a31d4aa
[Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 ( #44692 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 10:44:21 +08:00
Devin Lai and GitHub
c8beda4cc3
[Rust Frontend] Add Phi-4 mini JSON tool parser ( #44213 )
2026-06-06 10:40:00 +08:00
2f27c9a150
Preserve layout-changing clones ( #44574 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-05 20:45:24 -04:00
4765f0f189
[Bugfix] Fix sequence_parallel_chunk_impl custom op aliasing its input ( #44130 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-05 23:56:36 +00:00
Terrence Zhao and GitHub
a50e675b0d
[Cohere] fix RoutingMethodType ( #44021 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-05 16:25:53 -07:00
Daoyuan Li and GitHub
f6a708ab2b
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models ( #44435 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-05 16:04:32 -07:00
4200f62147
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path ( #42832 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 16:22:19 -05:00
Walter Beller-Morales and GitHub
c73b0d0db9
[Core][Engine] allow DP ray placement groups to be set on specific nodes ( #44669 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-06-05 20:07:47 +00:00
Harry Mellor and GitHub
e28e369f78
Male Mergify comment less spammy ( #44666 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 10:56:52 -07:00
yzong-rh and GitHub
703fb17b13
[Bugfix] GPT-OSS instruction rendering ( #44330 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-05 13:52:32 -04:00
Sting Lin and GitHub
b593396c7a
Upgrade tpu-inference to v0.21.0 ( #44621 )
...
Signed-off-by: StingLin <sting.lin@cienet.com >
2026-06-05 16:12:49 +00:00
Flame and GitHub
91e17d4315
Fix sarvam forward compatibility with transformers v5 ( #38804 )
...
Signed-off-by: vikrantpalle <vikrantpalle@gmail.com >
2026-06-05 11:51:44 -04:00
TJian and GitHub
aa6fb8a329
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm ( #44648 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-05 15:51:17 +00:00
Effi Ofer and GitHub
6a894574bf
Add objectstore as a secondary tier to multi-tier kv cache offloading ( #41968 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
2026-06-05 18:05:41 +03:00
Yan Ma and GitHub
7f003a1285
Support MiniCPMV batched preprocessing ( #44609 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-05 15:05:31 +00:00
Harry Mellor and GitHub
ef0df7dbd6
[CI] Bump mypy version 1.19.1 -> 1.20.2 ( #44647 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:56:27 +00:00
Harry Mellor and GitHub
a80af24356
Speed up docs build ( #44635 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:51:44 +00:00
Harry Mellor and GitHub
c66b19800b
[CI] Bump mistral-common ( #44649 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:18:50 +00:00
6a11d72df7
[Reasoning][Structured Outputs] Add Command A plus tags for structural tags ( #44588 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-05 06:51:20 -07:00
Woosuk Kwon and GitHub
02d2da0748
[DSV4] Move more ops out of eager breakpoint ( #44561 )
2026-06-05 06:42:41 -07:00
adhithyamulticoreware and GitHub
bbb6c274c8
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call ( #44615 )
...
Signed-off-by: ADHITHYA BALAKRISHNAN <adhithya.balakrishnan@multicorewareinc.com >
2026-06-05 12:47:56 +00:00
62215e72c6
Remove KV cache scale boilerplate from model weight loading methods ( #43167 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 05:19:04 -07:00
7fe7800fa4
[BUG] Fix FP64 Gumbel precision coverage ( #43150 )
...
Signed-off-by: tianyu-z <zhangtianyupro@gmail.com >
Signed-off-by: Tianyu Zhang <53099276+tianyu-z@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-05 19:04:14 +08:00
8a83e6f2d7
[Rust Frontend] Batch auto-abort requests by engine ( #44591 )
...
Signed-off-by: Hugh Ryan <197298026+HueCodes@users.noreply.github.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-05 02:59:09 -07:00
Chunyang Wen and GitHub
efc347f1b2
docs: fix tokenizer optimization typo ( #44066 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-06-05 02:12:49 -07:00
Nicolò Lucchesi and GitHub
d98b8f371c
[NixlConnector] Initiate deprecation cycle for kv_both role ( #43874 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-05 11:08:17 +02:00
Chao-Ju Chen and GitHub
e64237ae82
[Rust Frontend] Support include_reasoning=false ( #44391 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-06-05 16:47:50 +08:00
d61d8566ec
[Bugfix] Update mistral tokenizer test for continue_final_message fix ( #44622 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 16:13:26 +08:00
Uranus and GitHub
d2f70da116
fix: pad dummy run query_start_loc ( #44603 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-05 00:43:04 -07:00
6542d48964
[Bugfix] Fix test_invocations flaky failure with newer openai SDK ( #44618 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 07:36:20 +00:00
Ting SUN and GitHub
ca73293fa6
[Bugfix][Rust Frontend] Fix UTF-8 char-boundary panic in incremental detokenizer ( #44620 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-05 07:36:17 +00:00
Vic Wen and GitHub
ef3af56a97
Fix LLM.wait_for_completion output type docstring ( #44617 )
...
Signed-off-by: viiccwen <viiccwen@gmail.com >
2026-06-05 00:16:38 -07:00
b4a6f26c90
[ROCm][perf] Use workspace manager for sparse indexer allocations ( #41002 )
...
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:46:29 -07:00
165b7864d0
[ROCM] [FEAT] Integrate Aiter hipBLASLt GEMM online tuning ( #40426 )
...
Signed-off-by: hanlin12 <hanlin12@amd.com >
Signed-off-by: Han Lin <hanlin12@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:45:36 -07:00
Li, Jiang and GitHub
c505cd93ef
[CI/Build] Disable CPU-Compatibility Tests ( #44605 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-05 13:14:26 +08:00
qizixi and GitHub
96229fa99e
[KVConnector][1/N] PP-aware handshake aggregation and intermediate-PP output plumbing ( #43720 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-04 22:04:19 -07:00
da1daf40bf
[Bugfix] Exclude vision embedder from quantization in Gemma4 Unified ( #44571 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-04 20:47:38 -07:00
Woosuk Kwon and GitHub
4efd6ffde0
[DSV4] Refactor DeepseekV4Attention ( #44569 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-04 20:23:07 -07:00
Chris Leonard and GitHub
56aff0dd15
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to torch stale ABI. ( #44334 )
2026-06-04 20:14:43 -07:00
zofia and GitHub
063ce98fb7
[XPU][MoE] support block_fp8_moe on xpu ( #42139 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
2026-06-05 08:36:58 +08:00
Bugen Zhao and GitHub
62d6f06e3d
[Rust Frontend] Skip loading multimodal processor if --language-model-only is specified ( #44500 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-04 17:02:54 -07:00
Schwinn Saereesitthipitak and GitHub
b7c5baf63d
fix: keep DeepSeek V4 RoPE cache on inv_freq device ( #43926 )
...
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com >
Signed-off-by: Schwinn Saereesitthipitak <17022745+galletas1712@users.noreply.github.com >
2026-06-05 02:30:29 +04:00
Jiangyun Zhu and GitHub
a55fccfc7c
[mamba] unify KDA conv states into one cache to match 2-state SSM layout ( #44539 )
2026-06-04 20:38:05 +02:00
Wentao Ye and GitHub
41a4829f22
[Logs Refactor] Optimize shutdown logs, easier to follow and consistent ( #43707 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-04 14:36:32 -04:00
38fd2405f3
use split_group for pytorch process group creation ( #41980 )
...
Signed-off-by: Tushar Jain <tushar00jain@users.noreply.github.com >
Co-authored-by: Tushar Jain <tushar00jain@users.noreply.github.com >
2026-06-04 14:36:07 -04:00
Agata Dobrzyniewicz and GitHub
a947f7a420
[Kernel][Test] Extend lightning_attn and awq_triton kernel tests to XPU ( #43307 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-04 14:25:59 -04:00
bnellnm and GitHub
439203d32c
[Bugfix] Fix test_cutlass_moe.py ( #44380 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-04 14:18:52 -04:00
Taneem Ibrahim and GitHub
8d9536a775
[Misc] Add unit tests for pooler head classes ( #44471 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-04 17:59:25 +00:00
Fadi Arafeh and GitHub
3da29aa4a5
[DOC] Add INT8 W4A8 docs and Arm's supported quantization schemes ( #34894 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-04 16:27:17 +00:00
06f94633e7
[ROCm][CI] Add test for Aiter unified attn kernel ( #44436 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 16:15:05 +00:00
99ef652907
[Bugfix] Reject non-positive values for ParallelConfig int knobs ( #44057 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-04 11:46:50 -04:00
Tyler Michael Smith and GitHub
4cc78c9d5d
[Core] Freeze garbage collector in workers after model initialization ( #44363 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 08:39:04 -07:00
tc-mb and GitHub
3dbb4e0ace
[Bugfix] MiniCPM-V-4.6 video inference crash: placeholder count mismatches visual embedding count ( #44509 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-04 08:22:30 -07:00
b21443e23c
Add model support for granite speech plus ( #43519 )
...
Signed-off-by: Zvi Kons[WSL] <zvi@il.ibm.com >
Signed-off-by: Zvi Kons (BlueVela) <zvi@il.ibm.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
2026-06-04 14:47:48 +00:00
Michael Goin and GitHub
06ee2d8433
[Quant] Support compressed-tensors WNA8O8Int linears and WNInt embeddings ( #44340 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-04 07:40:33 -07:00
Yongye Zhu and GitHub
b5235fca2e
[DSv4] Adding TRTLLM gen attention kernel ( #43827 )
2026-06-04 07:35:09 -07:00
Andreas Karatzas and GitHub
3e77036768
[ROCm][CI] Specifying time outs for the lm eval models ( #44255 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:35:00 +08:00
Andreas Karatzas and GitHub
6f68ca3e91
[ROCm][CI] Stabilize memory-release in the Hybrid model generation tests ( #44046 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:34:24 +08:00
Turner Jabbour and GitHub
0c96dd64fb
[ROCm] Bump fastsafetensors to v0.3.2 from PyPI, remove git source build ( #43625 )
...
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com >
2026-06-04 07:30:57 -07:00
Nicolò Lucchesi and GitHub
68f5e565c9
[PD][Nixl] Mamba prefix caching mode support ( #42554 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 06:41:46 -07:00
QiliangCui2023 and GitHub
9354fb1ba5
[Bugfix][Compile] Guard per_token_group_fp8_quant lookup on non-CUDA platforms ( #44476 )
2026-06-04 09:31:50 -04:00
Harry Mellor and GitHub
f35b557239
Add GH token to docs build pre run check ( #44534 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-04 05:43:49 -07:00
Dipika Sikka and GitHub
e68988a248
Refactor CT NVFP4 linear to use a single class ( #42443 )
2026-06-04 08:25:08 -04:00
4b87b3e845
[Bugfix] fix EVS for qwen3-vl ( #44205 )
...
Signed-off-by: Rui "Garry" Gao <garrygaogg@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-04 11:06:51 +00:00
90619351e3
[Attention] Mamba attention module refactor - LINEAR ( #43556 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 18:45:29 +08:00
d0975a4b50
[perf] Add gemma RMS AR fusion ( #42646 )
...
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-04 01:33:59 -07:00
Kevin_Xiong and GitHub
1bdc60ed53
Fix Kimi-K2.5 FlashInfer ViT metadata ( #44493 )
...
Signed-off-by: Kevin-XiongC <kevin_xiong1997@outlook.com >
2026-06-04 08:14:35 +00:00
a6183563b6
[Prefix Caching] DeepSeekv4 - Support selective prefix-cache retention for sliding-window KV cache ( #43447 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-04 00:48:31 -07:00
Andreas Karatzas and GitHub
22c2e87555
[CI] Reverted gitignore changes ( #44497 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 00:37:44 -07:00
wang.yuqi and GitHub
d01d0b4646
[Frontend] Consolidate online serving utils. ( #44479 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-04 06:49:31 +00:00
b4b4aaa70e
[Inductor] Fast-path Inductor fallback for vllm::*/vllm_aiter::* custom ops ( #42129 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-04 00:03:52 -05:00
Andreas Karatzas and GitHub
5e2af28838
[CI] Resolve release V2 docker build after ROCm CI wheels change ( #44463 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-03 21:35:40 -07:00
4f423bd5bc
[EPLB] Nixl communicator optimization. Zero-copy transfers ( #41633 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 03:40:34 +00:00
f0cd590d62
optimize the compressor 128 split cutedsl kernel ( #44230 )
...
Signed-off-by: Jie Fang <jief@nvidia.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-03 20:22:57 -07:00
e6018c644a
[Refactor] Remove dead code in tests and parallel_state ( #41471 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 19:32:39 -07:00
f25952e59b
[MM][Perf][CG] Support ViT full CUDA graph for InternVL ( #41759 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-04 10:24:25 +08:00
maobaolong and GitHub
b58e082d95
[KV Connector] Update lmcache kv_offloading_backend to use LMCacheMPConnector ( #42865 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-06-03 19:23:55 -07:00
Ted Mostly and GitHub
0c1e6f63f5
[Bugfix] Fix VLLMNotFoundError when using LoRA adapter name in poolin… ( #44410 )
...
Signed-off-by: Ted Mostly <wanghenshui@qq.com >
2026-06-04 02:22:03 +00:00
Giancarlo Delfin and GitHub
ceb0111a90
[Model Runner V2][Spec Decode] Add Gemma4 MTP support ( #43241 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-04 00:51:06 +00:00
0414d75410
[XPU] skip unapplied UT in test_gpu_model_runner.py ( #44289 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 08:48:17 +08:00
128adabfe0
[Bugfix] Fix Gemma4 MTP block_table batch_size mismatch under concurrent load ( #43982 )
...
Signed-off-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
Co-authored-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
2026-06-03 17:11:10 -07:00
bdbf08fc02
Bump actions/stale from 10.1.1 to 10.2.0 ( #35078 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 14:14:41 -07:00
Woosuk Kwon and GitHub
6bad553f4e
[Minor] Remove FlashInfer version check in topk_topp_sampler ( #44442 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 21:06:00 +00:00
91945b6e4a
[Bug Fix][Model Runner V2][Spec Decode] Warmup & capture with different attention states for speculator prefill ( #44253 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 13:32:40 -07:00
2b237c7a41
[Bugfix] Honor tool_choice="none" in Chat Completions streaming ( #42752 )
...
Signed-off-by: hoobnn <111053672+hoobnn@users.noreply.github.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 13:27:45 -07:00
Wentao Ye and GitHub
dad95e34d8
[Feature] Support batch invariant rms norm with residual ( #42453 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 15:22:01 -04:00
a248b45d05
[Model] Add Gemma4 Unified (encoder-free) support ( #44429 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-03 12:01:39 -07:00
linitra24 and GitHub
271328e256
[LoRA] Fix dedup for post-replacement module aliases ( #44413 )
...
Signed-off-by: bk-201 <joy25810@foxmail.com >
2026-06-03 18:23:23 +00:00
Wentao Ye and GitHub
2b91012650
[Refactor] Remove dead code fp quant ( #44122 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 14:22:23 -04:00
JartX and GitHub
5b2a2beade
[ROCm][CI] Move Model Executor test step from MI250 to MI300 (gfx942) ( #44370 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-03 12:23:51 -05:00
59d0236193
[10b/n] Migrate custom all-reduce, DeepSeek V4 fused MLA, MiniMax reduce-RMS, and MXFP8 MoE to libtorch stable ABI ( #44365 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-04 00:29:46 +08:00
0a5cbf633e
Handle spinloop ext load failure gracefully ( #43659 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-03 16:09:52 +00:00
Willow Lopez and GitHub
51e0c579b0
fix(config): validate max_num_scheduled_tokens >= 0 on all paths ( #44207 )
...
Signed-off-by: Oxygen56 <1391083091@qq.com >
2026-06-03 16:06:45 +00:00
0c6631f02a
[KVCache] Support Pluggable KVCacheSpec ( #37505 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 09:05:16 -07:00
Nicolò Lucchesi and GitHub
df7252c343
[CI] Align PD tests to HMA on by default ( #44174 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 00:04:30 +08:00
Jee Jee Li and GitHub
4d1fd13613
[CI/Build] Fix LoRA testing ( #44425 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-03 08:58:06 -07:00
Nick Hill and GitHub
ec8d60bea8
[Model Runner V2] Use FlashInfer sampler ( #42472 )
2026-06-03 07:59:31 -07:00
27f1d34a23
[Frontend][Responses API] Move developer-to-system conversion into HF renderer ( #43590 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: kdcyberdude <kdsingh.cyberdude@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-03 14:52:24 +00:00
Flora Feng and GitHub
e3e132d2dd
[Refactor] Suppress SyntaxWarning from ast.literal_eval in tool parsers ( #44346 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 10:42:19 -04:00
e5232679a3
[XPU] Add XPU block-scaled W8A8 fp8 path ( #39968 )
...
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com >
Co-authored-by: Yuxiang <yuxiang.liang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 20:16:19 +08:00
309385a359
[Rust Frontend] Add /server_info to Rust frontend ( #43942 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 04:30:47 -07:00
3d76f395e3
[SharedOffloadRegion] Align blocks to page-size ( #43689 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 14:25:57 +03:00
Li, Jiang and GitHub
823d271c0d
[Attention][CPU] Standardize kv layout to blocks first ( #44393 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-03 19:03:09 +08:00
Andy Lo and GitHub
95b1615ec9
[Perf] Improve multimodal item handling from O(n) to O(log n) per step ( #44212 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-06-03 11:00:26 +00:00
1fa9ea09f6
[Perf] Triton fast path for small CPU→GPU swap_blocks_batch in the offloading connector ( #42212 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 13:38:17 +03:00
02564b4de0
[XPU]fallback to TRITON_ATTN for vit attn on xpu when use float32 dtype ( #43759 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 03:20:21 -07:00
Flora Feng and GitHub
209709a8c1
[Bugfix] Fix unstreamed tool call args dropped in Responses API streaming ( #44348 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 03:19:08 -07:00
ace95c9cf8
[Bugfix] Update TrtLLM MoE routing methods ( #44347 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 02:56:43 -07:00
Shanshan Shen and GitHub
0e2b13103b
[Doc] Update ViT CUDA graph interfaces ( #44388 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-03 01:20:59 -07:00
Bugen Zhao and GitHub
449be4f934
[Rust Frontend] Fix several hf chat template rendering issues ( #44311 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 01:04:43 -07:00
6550ff12f2
[Rust Frontend] Add dynamic LoRA endpoints ( #43778 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:55:29 +00:00
4aaed4ca22
[Rust Frontend] Add server router extension hook ( #43774 )
...
Signed-off-by: NolanHo <kujyo.eia.serias@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:45:31 +00:00
7268457999
[KV Offloading] Enable HMA models for Tiering Offloading ( #44287 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 10:03:00 +03:00
9af53a3c13
[Perf] Add tuned selective_state_update configs for H200 and RTX PRO … ( #44251 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-06-02 23:59:01 -07:00
Andreas Karatzas and GitHub
87954eb50e
[ROCm][CI] Optimize ROCm Docker build: registry cache, DeepEP, and ci-bake script ( #36949 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 23:43:07 -07:00
Charlie Fu and GitHub
71df063c49
Enable perf_token_group_quant/_C_stable_libtorch for ROCm ( #42758 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-02 23:23:28 -07:00
Albert Cheng and GitHub
e0081ef8cf
[Benchmark] Enable reasoning-model (thinking) benchmarking via --chat-template-kwargs for client-rendered datasets ( #44244 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
2026-06-02 22:49:51 -07:00
f0204358d9
[Bugfix] fix crash in postprocess for null tool args ( #43862 )
...
Signed-off-by: William-Rom <william.rom@intility.no >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 22:17:26 -07:00
Willow Lopez and GitHub
597bc15936
fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init ( #44236 )
...
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-03 01:07:10 -04:00
Rotem Shavitt and GitHub
3f0a91bb96
Nit Changes in Tiered KV Offload ( #44293 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-06-02 21:53:21 -07:00
Flora Feng and GitHub
e67063826b
[CI] Add missing vllm/parser/ CI trigger and fix test_parse.py ( #44352 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 21:05:19 -07:00
Andreas Karatzas and GitHub
53b88d1dfc
[CI] Reject out-of-vocabulary before they reach the GPU logprob path ( #44042 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 22:27:52 -05:00
JartX and GitHub
7b476c8f14
[ROCm][CI] Skip fp8 reload tests on gfx90a (MI250) ( #44369 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:27:14 -05:00
JartX and GitHub
4454a18695
[ROCm][CI] Fix stale wvSplitK GEMM fallback test for N=5 ( #44368 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:00:25 -05:00
02a01496fc
[Platform] Add is_cumem_allocator_available ( #43838 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 10:54:50 +08:00
Kevin H. Luu and GitHub
27a93cd426
[docker] Stop using extra-index-url for flashinfer-jit-cache ( #44366 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-02 18:58:22 -07:00
Wei Zhao and GitHub
969aec4bc8
[Bugfix] Fix Deepseek v4 non-mega-moe model init error ( #44356 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-02 18:26:30 -07:00
ca17b6b17d
[Perf] Apply single-pass min_larger finding and binary search in Triton Top-p path. ( #42191 )
...
Signed-off-by: js_park <cakeng@naver.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 17:57:26 -07:00
Woosuk Kwon and GitHub
b254e0456c
[DSV4] Minor cleanup for DeepseekV4MegaMoEExperts ( #44367 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 17:54:27 -07:00
Daoyuan Li and GitHub
bd98e97557
[Misc] Remove dead VLLM_RPC_TIMEOUT env var and fix profiling doc that references it ( #44128 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-03 00:22:10 +00:00
a4ac746405
[MoE/b12x] Accept W4A16 (kNvfp4Static, None) in FlashInferB12xExperts supports check ( #43332 )
...
Signed-off-by: Junhao Shen <junshen@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-02 15:20:37 -07:00
8b3b71ee9d
[CI/Build] Bump flashinfer to v0.6.12 ( #44036 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-02 15:19:05 -07:00
Siddharth Bedekar and GitHub
0917a009d3
Fix sparse NCCL weight transfer test construction ( #44345 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
2026-06-02 21:51:21 +00:00
3099de3617
[Kernel][MoE] Add GELU_TANH to CPU, CUTLASS, and WNA16 MoE backends ( #42027 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
2026-06-02 17:12:08 -04:00
Nick Hill and GitHub
e15f20258b
[ModelRunnerV2] Avoid pipeline parallel bubbles ( #42187 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 14:02:01 -07:00
557781131a
[Misc] Remove stray empty file ( #44350 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-02 12:53:03 -07:00
Yifan Qiao and GitHub
e9e08c49b9
[Bugfix] Cache the EAGLE/MTP lookahead block in the SWA prefix-cache mask ( #44082 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 12:21:07 -07:00
Woosuk Kwon and GitHub
e4a2e584e5
[MRV2] Remove assignment of graph_pool in cudagraph_utils ( #44338 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 11:50:27 -07:00
b8b49e2395
Bump actions/github-script from 8.0.0 to 9.0.0 ( #39667 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-02 11:26:57 -07:00
da107a59e5
[MRV2] Also enable MRV2 for Llama and Mistral dense models ( #43458 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: yewentao256 <zhyanwentao@126.com >
2026-06-02 11:18:46 -07:00
ed9a7526b6
[Anthropic] Support system role messages inside messages array ( #44283 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
Co-authored-by: Ang Kah Min, Kelvin <syraxius@hotmail.com >
2026-06-02 18:13:54 +00:00
2427094152
[Feature] Support EPLB for DeepSeek v4 Mega Moe ( #43339 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-lyris01.lyris.clusters.nvidia.com >
2026-06-02 10:56:44 -07:00
Kartavya sonar and GitHub
fe32e7830b
[Bugfix] flashinfer: fail fast when --kv-cache-dtype nvfp4 used on unsupported arch ( #43669 )
...
Signed-off-by: Kartavya Sonar <sonarkartavya@gmail.com >
2026-06-02 10:50:00 -07:00
afcb580715
[BugFix] Fix Humming MoE deploy error ( #43100 )
...
Signed-off-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Signed-off-by: Alireza Dadgarnia <49554709+adotdad@users.noreply.github.com >
Co-authored-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Co-authored-by: Jinzhen Lin <linjinzhen@hotmail.com >
2026-06-02 09:32:50 -07:00
3f3e2702c2
[XPU] Enable rms_norm/act quant fusions ( #43963 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 16:14:41 +00:00
Flora Feng and GitHub
478b49ddec
[Refactor] Remove dead code from parser infrastructure ( #44279 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 12:08:27 -04:00
Nick Hill and GitHub
cab5c9a2a9
[Core] Move max_concurrent_batches to VllmConfig ( #44274 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 08:57:25 -07:00
Brian Dellabetta and GitHub
774e552397
[compressed-tensors] Asymmetric support for MoE WNA16 marlin ( #44025 )
...
Signed-off-by: Brian Dellabetta <bdellabe@redhat.com >
2026-06-02 08:51:45 -07:00
XiaoZ and GitHub
53fa09d085
[Misc] Support local image encoding in benchmarks ( #43843 )
...
Signed-off-by: xiaoz <Sukra1@outlook.com >
2026-06-02 15:15:06 +00:00
Chris Leonard and GitHub
4d93bc35c9
Migrate header files to torch stable abi ( #44013 )
2026-06-02 08:09:52 -07:00
Bugen Zhao and GitHub
586201ebdc
[Rust Frontend] Cover different thinking modes in roundtrip tests ( #44320 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:51:25 -07:00
pschlan-amd and GitHub
88f172188b
[ROCm] Fix AITER RMSNormQuantFusion for Kimi-Linear ( #44308 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-06-02 14:50:21 +00:00
Bugen Zhao and GitHub
880fc032f4
[Rust Frontend] Support recursive tool parameter conversion ( #44299 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:45:35 -07:00
6314de8bad
[XPU] [Bug] remove xpuw4a16 output size check ( #44168 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 22:26:20 +08:00
IdoAtadTD and GitHub
c91a87f01a
[BugFix] [GDN] Read linear_key_head_dim from hf_text_config for multimodal models ( #43978 )
...
Signed-off-by: IdoAtadTD <ido.atad@twodelta.com >
2026-06-02 17:17:55 +03:00
Matthew Bonanni and GitHub
ea0d045a05
[FlashAttention] Sync FA with upstream ( #44065 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-02 07:15:37 -07:00
0bdfd5eb84
[Bugfix] Vendor MiniCPMV/MiniCPMO processors to unblock Transformers v5 ( #44282 )
...
Signed-off-by: guanwei-wu <b08901019@ntu.edu.tw >
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: guanwei-wu <b08901019@ntu.edu.tw >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-02 07:14:38 -07:00
0cbc48c4f9
Support ModelOpt MXFP8 non-gated MoE ( #42958 )
...
Signed-off-by: tbarnatan <tbarnatan@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-02 13:56:03 +00:00
2fd0e52252
[Bugfix] Fix Gemma4 startup crash with recent transformers multimodal processor ( #44232 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-02 13:42:40 +00:00
654bd2bca4
[Bugfix] Sync block_size from EngineCore to frontend for hybrid Mamba… ( #42967 )
...
Signed-off-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-02 13:41:00 +00:00
wang.yuqi and GitHub
b623f7ea95
[Frontend] Consolidate dev entrypoints. ( #44170 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-02 06:30:21 -07:00
Shreyas Kulkarni and GitHub
0eeba5eec1
Fix DFlash prefix cache corruption due to missing lookahead block ( #42971 )
...
Signed-off-by: Shreyas Kulkarni <shreyas.gp269@gmail.com >
2026-06-02 12:06:33 +00:00
f69ede495b
[XPU][Mamba] Triton-based selective scan forward op for XPU ( #43421 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:50:26 -07:00
Ronen Schaffer and GitHub
2a2b5ca791
[KV Offload] Add on_schedule_end() hook to separate step lifecycle from event draining ( #44206 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 13:42:52 +03:00
689b0eeb9e
[HARDWARE][POWER] Enable SHM communicator support for PowerPC ( #43754 )
...
Signed-off-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
Co-authored-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Co-authored-by: Akash kaothalkar <61960177+Akashcodes732@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 18:06:32 +08:00
Isotr0py and GitHub
f8e9c56d15
[Multimodal] Automatically select registered video loader for VLM ( #44126 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-02 09:09:47 +00:00
alberto and GitHub
e30313220c
[Parser] Migrate ResponsesParser to unified Parser interface ( #42977 )
...
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com >
2026-06-02 08:50:05 +00:00
d247a9dc13
[EC Connector] Non blocking EC Connector lookup ( #41627 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 08:48:25 +00:00
Yifan Qiao and GitHub
7c37096620
[Core][Refactor]: thread scheduler_block_size into KVCacheManager and KVCacheCoordinator ( #44165 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 01:14:44 -07:00
Maria Guevara and GitHub
b817b23f7b
[Rust Frontend] add --enable-request-id-headers flag support. ( #43883 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-06-02 16:08:37 +08:00
Ronen Schaffer and GitHub
93da882e73
[kv_offload] Add @override decorators to subclass method implementations ( #44177 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 08:07:47 +00:00
0b25cf4419
[CPU][Perf] Enable fused kernels for GDN's gated delta rules ( #43534 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 08:00:48 +00:00
Jiangyun Zhu and GitHub
dcdfe66bfa
[Perf] use triton moe backend on hopper by default ( #44220 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-02 15:52:30 +08:00
Flora Feng and GitHub
68dafcca75
[Refactor] Unify reasoning + tool-call parsing behind Parser.parse() ( #44267 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 15:11:42 +08:00
zhrrr and GitHub
1edfd09ffd
[Model Runner V2] Use actual batch max_seq_len for attn metadata ( #43991 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 06:07:56 +00:00
zhrrr and GitHub
8a9eb40808
[Model Runner V2] Support zeroing freshly allocated KV blocks for hybrid + fp8 KVCache ( #43990 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 05:56:53 +00:00
f91fb2fcf3
[Bugfix] Convert Gemma4-MM ViT linear layers to vllm native impl ( #43798 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: ZiTian Zhao <zitian.zhao@tencentmusic.com >
Co-authored-by: B-201 <Joy25810@foxmail.com >
2026-06-01 21:41:16 -07:00
JooHo Lee and GitHub
a045c7425f
[MM][CG] Profile encoder CUDA graph pool memory ( #41714 )
...
Signed-off-by: JooHo Lee <jooho414@gmail.com >
2026-06-02 12:27:34 +08:00
a3a5a5ece5
[XPU][Bugfix] Fix per_token_group_fp8_quant missing dummy args on XPU ( #43930 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:09:21 +00:00
Or Ozeri and GitHub
480fadab1b
[BugFix][kv_offload]: Prevent offloading stale sliding window blocks ( #42959 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-06-02 05:59:48 +03:00
279d25f5cb
[BugFix] Fix TypeError in MiniCPM-O audio feature unpadding ( #38053 )
...
Signed-off-by: Krishna Chaitanya Balusu <krishnabkc15@gmail.com >
Signed-off-by: wjinxu <1299461899@qq.com >
Signed-off-by: Kc Balusu <kcbalusu@users.noreply.github.com >
Co-authored-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Kc Balusu <kcbalusu@users.noreply.github.com >
2026-06-01 19:57:28 -07:00
Andreas Karatzas and GitHub
54d0c36fff
[CI] Stabilize OpenAI schema fuzzing for malformed structural tags ( #44131 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-01 19:56:15 -07:00
Flora Feng and GitHub
9affc17a05
[Refactor] Move unstreamed tool-arg flush from serving layer to parser ( #44017 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 10:37:43 +08:00
Alec and GitHub
816cc73a9b
[Bugfix][CI] Normalize NIXL connector CUDA wheel installs ( #44266 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
2026-06-01 19:34:05 -07:00
Micah Williamson and GitHub
2588ec4f0a
[ROCm] Upgrade AITER to v0.1.13.post1 ( #44265 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-02 01:48:59 +00:00
d68f0b220e
[Bugfix][Mooncake] Release GPU pin on failed store in MooncakeStoreConnector ( #43742 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-01 18:29:18 -07:00
Woosuk Kwon and GitHub
517e74a964
[DSV4] Refactor RoPE initialization ( #44262 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 01:26:58 +00:00
JartX and GitHub
48c0d13e65
[ROCm][CI] Skip unbacked dynamic shapes tests on PyTorch < 2.11 ( #44256 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-01 19:09:01 -05:00
Woosuk Kwon and GitHub
8c3cc98cff
[DSV4] Remove unncessary classes & functions ( #44246 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-01 14:43:00 -07:00
Nick Hill and GitHub
e4cbc4385d
[Test][BugFix] Fix double-BOS in PD+specdec acceptance test ( #44234 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:31:12 -07:00
Nick Hill and GitHub
6f8b40a23f
[BugFix][CI] Fix added _has_module tests ( #44248 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:23:12 -07:00
266b9d9c64
[Frontend][Core] Add sparse NCCL weight transfer support for in-place updates ( #40096 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-01 15:37:30 -04:00
182c67daf1
[Rust Frontend] Support streaming generate endpoint ( #43779 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 19:30:55 +00:00
fd9e91d7e4
[ROCm][CI] Fix and stabilize EAGLE3 acceptance tests ( #41294 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
2026-06-01 12:40:01 -05:00
Yongye Zhu and GitHub
035733515f
[Kernel][DSv4] Optimize sparse FP8 compressor kernels ( #44161 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-02 00:18:32 +08:00
023808c23d
[Feature] Add support for JetBrains' Mellum v2 code generation model ( #43992 )
...
Signed-off-by: Madeesh Kannan <madeeswaran.kannan@jetbrains.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-01 10:11:35 -04:00
985c97a6a8
[Perf] Optimize cutlass fp8 scaled mm bypassing padding, 20% kernel performance improvement ( #43706 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-01 09:05:21 -04:00
Chaojun Zhang and GitHub
bd0aecdc08
[XPU][CI] Fix test_audio_in_video flake by using module-scoped server fixture ( #44146 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-01 11:21:36 +00:00
8796838910
[Bugfix] fix wrong partial_rotary_factor calculation for bailing_moe model. ( #43770 )
...
Signed-off-by: zzt <zengzetang.zzt@antgroup.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-01 02:42:49 -07:00
de21863419
[Rust Frontend] Add InternLM2 tool parser ( #43481 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 08:58:46 +00:00
wang.yuqi and GitHub
0910f7e0e1
[Frontend] Resettle generative scoring entrypoint. ( #44153 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-01 07:54:59 +00:00
Uranus and GitHub
1f6048abe5
fix: glm5.1 pp model loading ( #42944 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-01 15:14:47 +08:00
98f1279815
[CPU][RISC-V] Add missing RVV cpu_types helpers for WNA16 ( #42730 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-01 14:56:41 +08:00
Isotr0py and GitHub
1fd8bd02a4
[Docs] Replace broken video url in examples ( #44159 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-01 06:01:10 +00:00
29d69332aa
[BugFix] Fix _has_module to verify native deps via trial import ( #44035 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: esmeetu <jasonailu87@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-31 22:06:33 -07:00
Lucas Wilkinson and GitHub
4721bb3aa4
[MRV2] Remove Eagle's dedicated CUDA graph pool ( #44078 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-05-31 22:00:33 -07:00
Umut Polat and GitHub
f46e6be169
[Misc] Use VLLMValidationError consistently in chat completion and completion protocol validators ( #36254 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
2026-06-01 04:04:11 +00:00
8b8546da1c
docs: fix MLA attention docstring examples ( #44118 )
...
Co-authored-by: nightcityblade <nightcityblade@gmail.com >
2026-05-31 12:28:38 -07:00
Jee Jee Li and GitHub
6bdabbad5b
[CI/Build] Enable Step3p7ForConditionalGeneration testing ( #43956 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-31 05:16:12 +00:00
3fd9d2d357
[CPU][Zen] Route W8A8 and W4A16 linear inference through zentorch on AMD Zen CPUs ( #41813 )
...
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Harshal Adhav <harshal.adhav@amd.com >
Signed-off-by: Aakar Dwivedi <aadwived@amd.com >
Co-authored-by: R <Ganesh.R@amd.com >
Co-authored-by: Harshal Adhav <harshal.adhav@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-30 14:17:21 -05:00
Woosuk Kwon and GitHub
27fa5aa3b9
[MRV2] Support breakable CUDA graph ( #44050 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-30 09:40:52 -07:00
e1105064b2
[Bug] Fix gemma4 MTP IMA issue when TP>1, CUDA error: an illegal memory access was encountered ( #43909 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-30 10:34:33 -04:00
Bugen Zhao and GitHub
50c80d7923
[Governance] Add @BugenZhao as Rust frontend code owner ( #44047 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-30 22:23:54 +08:00
3becc5db40
[ROCm] Add attention sink support to AITer flash attention backend ( #43817 )
...
Signed-off-by: Xiaoran Chen <xiaoran@fb.com >
Co-authored-by: Xiaoran Chen <xiaoran@fb.com >
2026-05-30 18:13:18 +08:00
124fac10cb
[Bugfix] Fix RMSNorm kernels to multiply in weight's native dtype ( #42379 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 23:16:53 -07:00
e9499996df
[BugFix][Platform] Fix import vllm.platforms.rocm error on non-CUDA test_gpt_oss.py ( #43571 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 23:16:49 -07:00
c0056b19bf
[ROCm] cmake: support PYTORCH_FOUND_HIP for torch 2.13 native HIP language support ( #43881 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 22:16:57 -07:00
Andreas Karatzas and GitHub
ef8840adc7
[ROCm][CI] Fix failure in the Phi3V pooling test ( #44028 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-30 12:14:37 +08:00
Flora Feng and GitHub
1a096d8208
[Refactor] Remove dead current_tool_name_sent assignments from tool parsers ( #43997 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 21:45:15 -04:00
Gagan Dhakrey and GitHub
1e2ce5d11a
offload prompt_embeds decode in render_prompts_async to avoid blocking ( #43792 )
...
Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com >
2026-05-30 01:36:34 +00:00
559d6710bf
[PERF]MiniMax-M2 gate kernel ( #38445 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Co-authored-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
2026-05-29 18:28:34 -07:00
bnellnm and GitHub
187457a952
Revert "[MoE Refactor] Migrate MoeWNA16Method quantization to MK orac… ( #44033 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 16:45:29 -07:00
8fad266507
[CI] Fix smoke test step key to bypass block gate ( #43974 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 16:28:32 -07:00
Flora Feng and GitHub
8c6daf6e2f
[CI] Remove duplicate Harmony test coverage ( #44023 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 22:52:46 +00:00
bnellnm and GitHub
7b98f498cd
[MoE Refactor] Remove supports_expert_map ( #43108 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 17:26:56 -04:00
106aa92f04
[MoE Refactor] Migrate MoeWNA16Method quantization to MK oracle ( #42647 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 17:19:31 -04:00
yzong-rh and GitHub
46409fd2a1
[Fronten] Clean up stop_token_ids override for Harmony ( #44009 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-29 13:28:06 -07:00
38b864d81d
[Metrics] Exclude KV transfer tokens from iteration_tokens_total ( #43346 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 19:56:44 +00:00
Wentao Ye and GitHub
5dbf1605a0
[Feature] SSL support for dp supervisor ( #43688 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 19:28:12 +00:00
Kevin H. Luu and GitHub
acbc203340
Add @khluu to CODEOWNERS ( #44019 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-29 12:24:29 -07:00
Flora Feng and GitHub
6de08e8b46
[CI] Remove redundant test_chat_with_tool_reasoning.py ( #44011 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 19:23:56 +00:00
6aabe221a5
[CI] Make Model Executor test hangs fail fast with a traceback ( #43971 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 11:58:25 -07:00
Wentao Ye and GitHub
739096a028
[Bug] Fix torch device issue for MOE permute ( #44005 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 18:55:00 +00:00
czhu-cohere and GitHub
8b9deeec4b
[Bugfix] Fix Ray placement group allocation with grouped nodes ( #43998 )
...
Signed-off-by: <conway.zhu@cohere.com >
Signed-off-by: root <conway.zhu@cohere.com >
2026-05-29 12:51:05 -06:00
d07ad0693b
[Bugfix] Use storage_block_size in KV cache reshape for compressed specs (DeepSeek V4) ( #43988 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-05-29 11:14:25 -07:00
4aaba00f92
[EPLB] Make async EPLB default ( #43219 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-29 18:07:16 +00:00
84b2a8a7e7
[MoE Refactor] WNA16 MoE backend selection into oracle module ( #42553 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 13:11:17 -04:00
4ff865c38e
[Bugfix] Disable allreduce_rms_fusion when pipeline_parallel_size > 1 ( #43616 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 22:57:43 +08:00
5502c3b52d
[Misc] added unit tests for the core pooling methods ( #43818 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-29 14:40:31 +00:00
Chunyang Wen and GitHub
f191d5630e
docs: clarify ITL acronym in optimization docs ( #43922 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-29 07:40:05 -07:00
11dfa3169d
Add vLLM library info to Hugging Face Hub requests ( #43857 )
...
Signed-off-by: Wauplin <lucainp@gmail.com >
Signed-off-by: Lucain Pouget <lucain@huggingface.co >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-29 14:04:58 +00:00
Li, Jiang and GitHub
3f6f508e14
[Bugfix][CPU] Remove invalid extra deps ( #43977 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-29 22:02:09 +08:00
Harry Mellor and GitHub
0585b5ba2e
Skip docs build if PR doesn't affect docs ( #43972 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-29 12:09:52 +00:00
Thien Tran and GitHub
d2889722ff
[Bugfix] Corrupted MLA + linear attention ( #43961 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-05-29 05:00:51 -07:00
0b56815a24
[ROCm][Perf] DSv3.2 MI355X TP4 decode-step orchestration cleanup (3 micro-opts) ( #42982 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-29 04:26:57 -07:00
ab12aab127
[Bugfix] [ROCm] [DSV4] Fix AITER MXFP4 MoE weight loading and shuffle… ( #42595 )
...
Co-authored-by: MHYangAMD <MHYangAMD@users.noreply.github.com >
2026-05-29 04:08:33 -07:00
JartX and GitHub
0cff0741ff
[Kernel][ROCm] Native W4A16 kernel for AMD RDNA3 (gfx1100) — fp16 + bf16 ( #41394 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-05-29 11:04:40 +00:00
60a7a2214f
[Bugfix] Fix Step3 pipeline parallel KeyError for residual tensor ( #37622 )
...
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-29 03:04:02 -07:00
Nicolò Lucchesi and GitHub
7ebc0ec104
[CI] Nixl+SimpleCPUOffloadingConnector unit tests ( #43871 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-29 02:40:42 -07:00
e8b5199973
[XPU] support MTP of gdn attention ( #43565 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 17:10:24 +08:00
Simon Danielsson and GitHub
b7fb747d8d
[CI][ROCm] Don't skip MoRI-IO Connector tests ( #43703 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-29 17:06:23 +08:00
Kunshang Ji and GitHub
30c6289b8e
[XPU] fix xpu install document triton-xpu version ( #43947 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 02:05:12 -07:00
Andreas Karatzas and GitHub
ff990d0d32
[ROCm][CI] Fix AITER unified attention for encoder-decoder cross-attention ( #43945 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-29 16:43:39 +08:00
Chauncey and GitHub
87f12e5c7c
[Frontend]Responses API supports chat_template_kwargs ( #43761 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-29 07:58:19 +00:00
kliuae and GitHub
ab7521d77c
[ROCm][DSv4] Remove device pipeline stall in sparse attention ( #43898 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
2026-05-29 15:42:40 +08:00
94d3f4d205
[CPU Backend] CPU top-k and top-p sampling kernels using Triton ( #43633 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 15:02:39 +08:00
04516eabc8
[XPU] add gelu_tanh to xpu moe backend supported activations ( #42822 )
...
Signed-off-by: yintong-lu <yintong.lu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 14:37:20 +08:00
648c3ebee6
[CI] Separate non-root smoke tests from image build step ( #43712 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 23:34:16 -07:00
22a58640b4
[9/n] Migrate attention and cache kernels to torch stable ABI (continued) ( #43717 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 04:44:45 +00:00
710f077617
[Refactor] Remove dead code ( #43234 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 00:29:56 -04:00
d63108fb18
[kv_offload] Skip decode-phase blocks in CPU offload ( #43797 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-05-29 06:39:43 +03:00
9636709372
[XPU] add scale transpose to prepare_fp8_moe_layer_for_xpu and bump up kernels ( #43277 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 03:22:51 +00:00
Weida Hong and GitHub
dfe8ba7c80
Adjust design around encoder_cudagraph_forward ( #42288 )
...
Signed-off-by: Weida Hong <wdhongtw@google.com >
2026-05-29 03:02:52 +00:00
212deff2ec
[feat] add GlmgaProcessor specific logits in glm4_1v.py ( #43575 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-05-29 02:56:02 +00:00
Woosuk Kwon and GitHub
7bd45da585
[DSv4] Move mHC tilelang kernels & Don't use CustomOP in dsv4/nvidia ( #43905 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-29 10:25:02 +08:00
bf18d7e0b4
[Misc][NUMA] Auto-bind to PCT priority cores on DGX B300 + widen EngineCore across shard NUMA nodes ( #43270 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <noreply@cursor.com >
2026-05-29 10:07:44 +08:00
Bugen Zhao and GitHub
1521173c17
[Rust Frontend] Add /version endpoint using engine-reported value ( #43854 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-29 00:32:27 +00:00
b690b2bb67
[Model]Support Step-3.7-Flash ( #43859 )
...
Signed-off-by: luotingdan <luotingdan@stepfun.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: luotingdan <luotingdan@stepfun.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Yu Huang <yuhuang@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-28 17:01:48 -07:00
yzong-rh and GitHub
325a1ec4fb
[CI] Enable prefix caching in BFCL benchmark ( #43925 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-28 23:36:31 +00:00
69c9f19957
fix(frontend): Add multimodal placeholders to Gemma4 tool message template ( #41459 )
...
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-05-28 14:48:12 -07:00
rasmith and GitHub
9769e2df2a
[AMD][CI][BugFix] Fix Distributed Compile Unit Tests (2xH100-2xMI300) group ( #43120 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-28 14:39:01 -07:00
Michael Goin and GitHub
03f03f9630
Refactor output filename handling in ci-fetch-log.sh ( #43901 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-05-28 14:20:12 -07:00
Benjamin Chislett and GitHub
9202ea6fda
[Spec Decode] Allow causal DFlash ( #43445 )
2026-05-28 21:18:44 +00:00
Woosuk Kwon and GitHub
69b8956dcd
[Model Refactoring] Remove unncessary torch op registration for DSv4 ( #43891 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:04:55 -07:00
a3ed5ab10c
[KV Offload] Add per-request offloading policy via on_new_request lifecycle hook ( #43205 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 20:45:18 +00:00
7e53283b1c
[Core] Cleanup KVConnector handling with PP + fix MRV2 ( #43732 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 13:12:03 -07:00
9090368b65
[Feat] Add support for per GPU worker RDMA NIC selection ( #42083 )
...
Signed-off-by: Raj Joshi <rajjoshi@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-28 12:45:23 -07:00
Harry Mellor and GitHub
085ac221a3
Deprecate JAISLMHeadModel ( #43784 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 18:29:12 +00:00
Hua Huang and GitHub
9006204e90
[MM][CG] Avoid over-padding Qwen2.5-VL encoder cudagraph window metadata ( #42796 )
...
Signed-off-by: Hua Huang <huah@nvidia.com >
2026-05-28 11:22:56 -07:00
ed7fe831da
[ROCm] Enable the aiter top-k/top-p sampler by default ( #43331 )
...
Signed-off-by: John Qin <yanyuan.qin@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-28 13:19:59 -05:00
Nicolò Lucchesi and GitHub
5b115bb8a3
[Attention][AMD] Standardize kv layout to blocks first for AMD ( #43660 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-28 12:28:50 -05:00
53a2088675
Allow native KV cache dtype in Triton cache update ( #43330 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-05-28 16:51:40 +00:00
Chao-Ju Chen and GitHub
099024762c
[Rust Frontend] Optimize multimodal prompt expansion ( #43670 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-05-28 09:46:18 -07:00
9aa131f944
Add Cosmos3 Reasoner model ( #43356 )
...
Signed-off-by: Maciej Bala <mbala@nvidia.com >
Signed-off-by: MaciejBalaNV <mbala@nvidia.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-28 09:43:55 -07:00
Micah Williamson and GitHub
1b5437cec8
[ROCm] Bump ROCm to 7.2.3 ( #43136 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-28 09:42:43 -07:00
3207e7680e
[XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) ( #41426 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-28 16:30:48 +00:00
Matthias Gehre and GitHub
a9ec46d4b7
[ROCm][Perf] Support N=5 in wvSplitK skinny GEMM kernels for speculative decoding ( #40687 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-05-28 16:28:21 +00:00
Ronen Schaffer and GitHub
4bfa0f2b14
[KV Offload] Rename SecondaryTierManager.get_finished() to get_finished_jobs() ( #43870 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-28 16:00:18 +00:00
Vadim Gimpelson and GitHub
5d126dd155
[Bugfix] Exclude Ray DP from #42585 's deferred port allocation ( #43864 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-28 15:55:14 +00:00
c08ebebf30
[Perf] Add do_not_specialize to Mamba SSD chunk kernels ( #43803 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
2026-05-28 15:40:02 +00:00
Wentao Ye and GitHub
be4062fd6c
[Bug] Fix tests/distributed/test_elastic_ep.py - assert False ( #43813 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 11:00:56 -04:00
577d693838
[rust] fix: aggregate is_sleeping and reset_prefix_cache across DP engines ( #43429 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-28 07:56:56 -07:00
Bugen Zhao and GitHub
61a1e30473
[Rust Frontend] Reduce Gemma4 tool parser args scan complexity ( #43850 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:52:29 +00:00
Bugen Zhao and GitHub
3a282230ee
[Rust Frontend] Add hy_v3 tool parser ( #43872 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:42:47 +00:00
Li, Jiang and GitHub
20d69d100a
[CPU] Migrate cpu_awq into awq_marlin ( #43841 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-28 22:36:31 +08:00
Simon Danielsson and GitHub
552eb81918
[Bugfix][ROCm] Resolve MoRI connector hangs at high concurrency ( #40344 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-28 14:30:21 +00:00
Woosuk Kwon and GitHub
9957e4d240
[Model Refactoring] Remove torch compile dependency in DSv4 ( #43746 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:26:25 +00:00
864990e8d9
Add token-offset based selective offload in OffloadConnector ( #39983 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-05-28 14:11:02 +00:00
f3b2a819f7
[Perf][KDA] Fuse gate softplus, chunk-local cumsum, and RCP_LN2 scaling ( #43667 )
...
Signed-off-by: haojiangzheng <justineric096@gmail.com >
Co-authored-by: haojiangzheng <justineric096@gmail.com >
2026-05-28 13:47:08 +00:00
Wentao Ye and GitHub
64e1218673
[Perf] Optimize moe permute by pre-allocate buffer, 9~14% kernel performance improvement ( #43014 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 06:18:26 -07:00
Julien Denize and GitHub
02606b0b09
[BUGFIX] Multimodal benchmark with MistralTokenizer ( #42965 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-05-28 05:36:24 -07:00
19af4e6dd4
Fix OlmoHybridForCausalLM not initialising ( #43846 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 05:33:31 -07:00
omerpaz95 and GitHub
811d805195
[EC Connector] Add shutdown API to EC Connector. ( #42423 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
2026-05-28 12:28:01 +00:00
Vadim Gimpelson and GitHub
c1c4db8b4b
Log dummy DP step in iteration details ( #41406 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-28 12:18:39 +00:00
Chauncey and GitHub
d692b89c2c
[Feature] Add structured output and effort support to Anthropic Messages API ( #42396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-28 12:06:48 +00:00
Bugen Zhao and GitHub
8e0580f4ee
[CI] Auto-apply rust label to relevant PRs ( #43866 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 11:57:22 +00:00
61288b5458
[Bugfix] Fix HyperCLOVAX CI failure after upstream removed remote code ( #43860 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 03:37:36 -07:00
a583c84e2b
[Bugfix][ROCm] Fix Accuracy Drop in Sparse Indexer on gfx950 ( #43781 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-28 03:37:15 -07:00
4ec2817313
[Model][Bugfix] Rename weight_mapper to hf_to_vllm_mapper in LlamaNemotronVL pooling models ( #43581 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
Co-authored-by: opencode <noreply@opencode.ai >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-05-28 03:32:22 -07:00
Wei Zhao and GitHub
f2caefe226
[UX] Increase DP Coordinator startup timeout from 30s to 120s ( #42343 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-28 03:31:45 -07:00
Animesh Trivedi and GitHub
bfb9ebc211
[Feature] Add support for timed trace replay in vllm bench serve to replay Moonshot and Alibaba workload traces ( #39795 )
...
Signed-off-by: Animesh Trivedi <Animesh.Trivedi@ibm.com >
2026-05-28 03:31:34 -07:00
Andreas Karatzas and GitHub
a9bc0ad8e4
[ROCm][CI] Move workload from MI300 to MI325 ( #43824 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 03:31:29 -07:00
b372ad3e90
[Bugfix] Stream DeepSeek DSML tool-call argument deltas incrementally ( #42879 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-28 17:50:23 +08:00
Harry Mellor and GitHub
2a781756a1
Restore Literal for WeightTransferConfig.backend ( #43183 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 09:39:41 +00:00
Woosuk Kwon and GitHub
a04afd76aa
[DSV4] Remove AMD/XPU path in deepseek_v4/nvidia ( #43829 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 08:00:52 +00:00
6cc8577421
[Kernel] Marlin MoE: include SM 12.x in default arch list ( #40923 )
...
Signed-off-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-28 15:30:26 +08:00
d6b48f928f
[BugFix] Fix hard-coded timeout for multi-API-server startup ( #43768 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:09:13 -07:00
Rotem Shavitt and GitHub
1b16f2ddc9
change name of fs_python secondary tier to fs. ( #43600 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-05-28 07:05:48 +00:00
TJian and GitHub
0ba46d4b11
[ROCm][DSV4] Enable Tilelang MHC replacing torch/triton mhc ( #43679 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-28 07:05:28 +00:00
JINO ROHIT and GitHub
e1814f822d
minor docs: fix incorrect example path ( #43830 )
...
Signed-off-by: JINO-ROHIT <find.jinorohit@gmail.com >
2026-05-27 22:58:43 -07:00
7909f82a45
[Bugfix][Frontend] streaming tool-call serializer drops first args chunk when name and args share a DeltaMessage ( #42683 )
...
Signed-off-by: ignaciosica <mignacio.sica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-28 05:20:55 +00:00
Nick Hill and GitHub
626fa9bba5
[BugFix] Fix blocked reasoning parsing with MRV2 ( #43808 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 04:59:34 +00:00
Thien Tran and GitHub
e54eff769d
[Bugfix] Pass routed_scaling_factor to FlashInfer TRTLLM BF16 MoE ( #43769 )
2026-05-27 21:29:14 -07:00
05ac829629
fix: parse Qwen3 XML JSON arguments first ( #43243 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-28 03:35:59 +00:00
Andreas Karatzas and GitHub
33e94fc3ad
[ROCm][CI] Stabilize Cargo cache and pre-test image checks ( #43815 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 11:24:44 +08:00
413ac5c070
[Misc][Rocm] Remove redundant AiterUnifiedAttentionBackend block size log ( #43664 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 22:19:11 -05:00
Yongye Zhu and GitHub
2d2c660104
[MoE] Remove inplace fused experts mechanism ( #43727 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 20:00:19 -07:00
Benjamin Bartels and GitHub
05eec7120e
Fix RunAI streamer tensor buffer reuse during weight loading ( #43464 )
...
Signed-off-by: bbartels <benjamin@bartels.dev >
2026-05-27 19:16:52 -07:00
Bugen Zhao and GitHub
c87f62ccf8
[Rust Frontend] Introduce mock engine for benchmark baseline ( #43469 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 01:40:35 +00:00
1223732dda
[ModelRunnerV2][Hybrid model] Support kernel block size in hybrid model ( #38831 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:55:55 +00:00
amitz-nv and GitHub
381edde1b9
[Bugfix][Kernel] TRTLLM NVFP4 MoE chunking ( #43599 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-28 00:36:21 +00:00
Andreas Karatzas and GitHub
094124af15
Add @AndreasKaratzas to CODEOWNERS ( #43740 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-27 16:14:50 -07:00
Dakai An and GitHub
5963c19478
Fix Qwen3-VL and Qwen3-omni-thinker accuracy degradation from deepstack inputs under torch.compile ( #43617 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-05-27 15:34:08 -07:00
7fb9c0197a
[Bugfix][DFlash]allocate the proper number of lookahead slots ( #43733 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@gmail.com >
2026-05-27 21:45:34 +00:00
Harry Mellor and GitHub
2c2c966669
Validate against some config fields being set to 0 ( #43794 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 21:14:49 +00:00
Harry Mellor and GitHub
2616f67faa
Remove Transformers forward/backward compatibility tests ( #43785 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 12:46:36 -07:00
206b72c982
[Quantization] Fix Humming RoutedExperts import ( #43540 )
...
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 10:51:56 -07:00
284e6f543d
[8/n] Migrate merge_attn_states, mamba, sampler to torch stable ABI (continued) ( #43361 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-27 09:35:24 -07:00
jatseng-ai and GitHub
05c50c721e
[ROCm] mori: add InterNodeV1LL inter-node kernel selection via VLLM_MORI_INTERNODE_KERNEL ( #41751 )
...
Signed-off-by: jatseng-ai <jatseng@amd.com >
2026-05-28 00:33:32 +08:00
Harry Mellor and GitHub
41688e2dc7
Fix early CUDA init ( #43791 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 09:30:11 -07:00
Chunyang Wen and GitHub
49a3510266
[Docs] Fix the duplicate doc icon issue ( #43546 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-27 16:09:58 +00:00
Injae Ryou and GitHub
165460941f
[BugFix] HFValidationError with cloud storage URIs when HF_HUB_OFFLINE=1 ( #39155 )
...
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
2026-05-27 10:53:32 -05:00
Yongye Zhu and GitHub
03d9cc2fe2
[misc] Bump cutedsl version to 4.5.2 ( #43745 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 08:25:36 -07:00
52a31ccecc
[Bugfix] Map reasoning_effort to enable_thinking in chat template kwargs ( #43401 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 05:39:49 -07:00
2272062471
[Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes ( #43731 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 18:36:27 +08:00
Mohammad Miadh Angkad and GitHub
158289e0fc
[Docs] Fix MLA prefill backend default docs ( #43697 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-27 10:13:22 +00:00
Bugen Zhao and GitHub
396c8fee50
[Rust Frontend] Align tool parser fallback behavior between streaming & non-streaming paths ( #43662 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 10:13:12 +00:00
ad464e16c0
[Doc] Add Ascend NPU tab to the quickstart installation guide ( #43550 )
...
Signed-off-by: Aditya Singh <adisin650@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 08:41:29 +00:00
akii96 and GitHub
de12f5ca0b
[ROCm][GPT-OSS] Avoid repeated compile-time cos_sin_cache.to(bf16) casts in rotary path ( #42833 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-27 16:22:27 +08:00
683033d4ba
[Frontend] Add MiniCPM5 XML tool call parser ( #43175 )
...
Signed-off-by: zhangtao <zhangtao2@modelbest.cn >
Signed-off-by: zhangtao2 <zhangtao2@modelbest.cn >
Co-authored-by: zhangtao <zhangtao2@modelbest.cn >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 00:39:35 -07:00
8c94938cfb
[MRV2][BugFix] Fix KV connector handling in spec decode case ( #43719 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 06:37:56 +00:00
Nico Holmberg and GitHub
7b54690244
[ROCm][Perf] Expose AITER MoE sorting dispatch policy via env var ( #39177 )
...
Signed-off-by: nholmber <nholmber@users.noreply.github.com >
2026-05-27 13:11:02 +08:00
1fc2cee50a
[KVConnector][Mooncake] Wire reset_cache cascade end-to-end ( #42694 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Signed-off-by: Ao Shen <aoshen@inferact.ai >
Co-authored-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 20:52:35 -07:00
Angela Yi and GitHub
0fa3114ae1
Fix test_aot_compile for torch 2.12 ( #43695 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-05-26 23:12:49 -04:00
Woosuk Kwon and GitHub
adaa5e455a
[DSv4] Refactor compressor & Fix ROCm compatibility ( #43710 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 19:56:46 -07:00
c02c758ea4
[Deprecation] Deprecate functions as scheduled for v0.21.0 ( #43358 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 19:56:21 -07:00
Matthew Bonanni and GitHub
aa6138169f
[MLA][Attention] Add OOT MLA prefill backend registration mechanism ( #43325 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-26 19:56:09 -07:00
7e33081cee
[Attention] Make FlexAttention and FlashAttention use num-blocks first layouts ( #42095 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-26 19:55:56 -07:00
Xin Yang and GitHub
d8eebe6d97
[Perf] Optimize Fp8BlockScaledMMLinearKernel input_scale tensor using new_empty() ( #43677 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-05-26 19:55:52 -07:00
Andreas Karatzas and GitHub
5bdb181df5
[ROCm][CI] Fix ROCm multimodal Qwen2.5-VL activation compile and Phi4MM ragged image mask handling ( #43647 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 19:53:34 -07:00
Bugen Zhao and GitHub
0b68f21e7c
[Rust Frontend] Add reasoning/tool parser & renderer roundtrip tests ( #43582 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 00:49:30 +00:00
dede691c95
[Bugfix] Split attention groups by num_heads_q for spec-decode drafts ( #43543 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 00:11:01 +00:00
e19b9b1045
[ci] Add arm64 ci image ( #41303 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 14:38:09 -07:00
812e7e7364
[Bugfix][V1] Fix TOCTOU race causing intermittent EADDRINUSE on multi-API-server DP startup ( #42585 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 14:06:00 -07:00
d98cbf472b
[KV Connector] MooncakeStore: drop dead discard_partial_chunks parameter ( #43627 )
...
Signed-off-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 13:40:21 -07:00
Jee Jee Li and GitHub
6e503868ca
[Kernel] Porting fuse_minimax_qk_norm to manual fusion ( #43410 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-26 13:16:03 -07:00
49b4882779
[CI] Soft-fail AMD entrypoints mirror tests ( #43709 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 13:08:48 -07:00
Woosuk Kwon and GitHub
193ce8812e
[DSv4] Drop _get_compressed_kv_buffer in DeepseekCompressor ( #43690 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 10:11:25 -07:00
3aea37d28e
[Doc] Add line limit to AGENTS.md ( #43635 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-05-26 09:31:23 -07:00
Wei-Ming Chen and GitHub
6f5b533241
Add LM head quantization support for ModelOpt ( #42124 )
...
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com >
2026-05-26 09:21:05 -07:00
Woosuk Kwon and GitHub
c8414a8271
[ROCm] Remove MegaMoE integration in deepseek v4 ( #43629 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 08:56:04 -07:00
f51bbc694d
[MoE Refactor] W4a8 int8 oracle ( #42789 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:15:42 -04:00
b226ddacfd
[MoE Refactor] Migrate ModelOptMxFp8FusedMoE to oracle ( #42768 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:14:14 -04:00
Yongye Zhu and GitHub
6ab6ffb428
[Feat][DSV4] Fuse q pad into deepseek v4 fused kernel ( #43162 )
2026-05-26 05:12:54 -10:00
Andreas Karatzas and GitHub
445ded18c1
[ROCm][CI] Extend ROCm quick reduce coverage ( #40990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 21:57:13 +08:00
d565357a90
[Docs][ROCm] MoRI-IO Connector Usage Guide ( #43603 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Signed-off-by: Simon Danielsson <70206058+simondanielsson@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 21:52:30 +08:00
Mohammad Miadh Angkad and GitHub
a970fb5a1a
Fix CuPy runtime deps and restore humming ( #43530 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-26 05:59:40 -07:00
Chaojun Zhang and GitHub
861b97765d
[XPU] Fix fused MoE LoRA kernel crash on XPU by using platform-agnos num_compute_units ( #43646 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-05-26 03:40:32 -07:00
ebd0692f80
[Model] Use AutoWeightsLoader for InternLM2 ( #38278 )
...
Signed-off-by: Jesus De Jesus <dejesus.9297@gmail.com >
Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-26 03:39:26 -07:00
739af5c7e1
[Reasoning] [Bugfix] Reject invalid thinking_token_budget values ( #43402 )
...
Signed-off-by: linzm1007 <linzm1007@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 03:37:30 -07:00
Thibault Castells and GitHub
5d09f471f4
[Misc] Support interleaved custom image benchmark datasets ( #43636 )
...
Signed-off-by: ThibaultCastells <thib.castells@icloud.com >
2026-05-26 03:37:25 -07:00
681d7dd38b
[Misc][Refactor][ROCm] Convert MoRI-related envvars to extra config args ( #43303 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-26 03:33:35 -07:00
Ethan Feng and GitHub
755043cf3c
[KV Transfer] Enable HMA by default for connectors that support it ( #41847 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-26 12:28:51 +02:00
97e4022c6c
[Bugfix] Apply fc_norm in Eagle3DeepseekV2 combine_hidden_states ( #43482 )
...
Signed-off-by: Yubo Wang <yubowang2019@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-26 00:46:10 -07:00
Hank_ and GitHub
b3269454b1
[chores][log] change registry log from warning to debug ( #43045 )
...
Signed-off-by: Hank <hcc.mayday@gmail.com >
2026-05-26 00:13:46 -07:00
a37e47100c
Add CuTe DSL sparse compressor support ( #43584 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-26 00:11:12 -07:00
Sting Lin and GitHub
e6adbd7834
Upgrade tpu-inference to v0.20.0 ( #43394 )
2026-05-25 20:26:25 -10:00
zhao, zhenhui and GitHub
771e1e48b1
[CPU] Enable non-divisible GQA for decode workitems in mixed batches ( #43032 )
...
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com >
2026-05-26 14:15:47 +08:00
Thien Tran and GitHub
d56612c621
[GDN] GDN Prefill kernel for SM100 ( #43273 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-05-26 14:02:11 +08:00
6f955986e1
[Bugfix][Model] Fix GPT2ForSequenceClassification sub-module prefix ( #43579 )
...
Signed-off-by: QingZhou-YangHY <3868850350@qq.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-25 22:43:19 -07:00
d5cf7b4a2c
[Frontend] Split the offline inference APIs and utils. ( #43553 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 05:20:24 +00:00
Yan Ma and GitHub
f815c99954
[Bugfix] fix device mismatch in MiniCPM-o-4_5 resampler ( #43194 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-05-26 13:12:50 +08:00
Dao007forever and GitHub
c2a4005c70
[KV Connector] Propagate MooncakeStore load failures ( #42788 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
2026-05-25 22:12:15 -07:00
7966fc7233
[KV Connector][Bugfix] MooncakeStore: don't double-apply Eagle prune in load_mask ( #43516 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-25 22:11:57 -07:00
Woosuk Kwon and GitHub
aa2b56ffb0
[DeepSeek V4] Move MegaMoE input prep kernel to nvidia/ops ( #43632 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-25 21:08:29 -07:00
Jee Jee Li and GitHub
ec5de7fa7d
[LoRA] Add one shot triton kernel For MoE LoRA ( #42290 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-25 19:47:04 -07:00
71d810bbf4
[XPU] Ensure RNG offset alignment with PyTorch requirements in XPU sampler ( #43028 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 02:01:30 +00:00
Jee Jee Li and GitHub
d4004455d2
[Kernel] Remove NormGateLinear ( #43554 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-25 09:49:19 +00:00
Nicolò Lucchesi and GitHub
716d5294e6
[Misc] Print accuracy value for PD tests even on success ( #43583 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-25 02:10:01 -07:00
873758c13a
[KV Connector] Handle Mooncake finish after preemption ( #43281 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-25 01:58:38 -07:00
5c1aec3dc0
Reduce memory usage for granite_speech. ( #42933 )
...
Signed-off-by: Yihuki <wangbovbvb@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-25 14:12:57 +08:00
Roy Wang and GitHub
0c942c69d6
[Doc] Add section on escalating stalled contributions ( #43568 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-05-25 14:11:01 +08:00
Yifan Qiao and GitHub
81252d4e24
[Feat][KVConnector] Support DSV4 in SimpleCPUOffloadBackend ( #42296 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-25 14:04:30 +08:00
3df1c7c43e
[Docker] Non-root support for vllm-openai; add opt-in vllm-openai-nonroot target ( #40275 )
...
Signed-off-by: TheDuyIT <nduy250299@gmail.com >
Signed-off-by: dtnguyen <dtnguyen@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-25 13:45:31 +08:00
1b26fa361e
[Docs] Reorganize offline inference docs. ( #43552 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-25 13:44:39 +08:00
weizhoublue and GitHub
6cbe448eed
fix: MoE model using shared routed experts crashes on AMD GPUs ( #42373 )
...
Signed-off-by: weizhou.lan@daocloud.io <weizhou.lan@daocloud.io >
2026-05-25 12:03:05 +08:00
Jee Jee Li and GitHub
b06813e872
[Kernel] Add mhc_pre_big_fuse_with_norm_tilelang ( #43474 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-25 01:19:45 +00:00
d0a100c87a
File system secondary tier implemented in python ( #41735 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-05-24 18:14:44 +00:00
d56285c747
Tuning script and configs for Triton Mamba SSU kernel ( #43083 )
...
Signed-off-by: Banani Ghosh <bg2502@nyu.edu >
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
Co-authored-by: Banani Ghosh <bg2502@nyu.edu >
2026-05-24 20:12:44 +03:00
TJian and GitHub
1806d1adfc
[ROCm] [DSv4] [Perf] Support DeepSeek v4 MTP ( #43385 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-24 18:43:08 +08:00
Andreas Karatzas and GitHub
5940590855
[ROCm][CI] Stabilize 400 error return code for invalid schema inputs ( #43016 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-24 10:06:49 +00:00
Or Ozeri and GitHub
357fddf614
[kv_offload]: Add DSv4 support ( #43142 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-24 11:10:12 +03:00
0902d8e62f
[KV Connector] Keep MooncakeStore full hits block-aligned ( #43494 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-23 23:15:03 -07:00
Wentao Ye and GitHub
33d7cbe02c
[Model Runner v2] Force v1 runner for tests ( #43233 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-23 16:37:24 -07:00
Flora Feng and GitHub
b32fe416ea
[Bugfix] Fix reasoning dropped on streaming boundary deltas ( #42691 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-23 16:18:30 -07:00
Michael Goin and GitHub
10d264a2b9
Revert "[Misc] add humming to dependencies" ( #43492 )
2026-05-23 14:21:13 -07:00
TJian and GitHub
46f95b2ec2
[ROCm][Critical] Fix the GDN import bug ( #43486 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-23 21:12:58 +00:00
Dao007forever and GitHub
819c610f9b
[Mooncake] Add metrics for MooncakeStoreConnector operations ( #43392 )
2026-05-23 13:34:40 -07:00
4438b6e7dc
[MoE] Migrate W4A8 CT to oracle kernel setup ( #42680 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-05-23 13:56:01 -04:00
Holegots and GitHub
8737e4a857
[Docs] Fix stale version number in token_classify.md ( #43489 )
...
Signed-off-by: holegots <ikun3.1415927@gmail.com >
2026-05-23 10:42:20 -07:00
Holegots and GitHub
7c2ff1f819
[Docs] Fix stale version number in token_embed.md ( #43488 )
...
Signed-off-by: holegots <ikun3.1415927@gmail.com >
2026-05-23 10:06:56 -07:00
a0be71ee47
[MM] Enable FlashInfer metadata support for Qwen2.5-VL vision attention ( #42787 )
...
Signed-off-by: Hua Huang <huah@nvidia.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-23 16:08:40 +00:00
d8b385b7ea
[Bugfix][Frontend] Fix input_audio parsing when uuid is present ( #43414 )
...
Signed-off-by: ffggs <314137448@qq.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-23 09:03:19 -07:00
Andreas Karatzas and GitHub
2a7d5b7324
[ROCm][CI] Remove benchmarks test group and shard long test groups ( #41669 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 23:31:46 +08:00
5bb8d2767a
[Kernel] Batch invariant NVFP4 linear using cutlass ( #39912 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-23 09:41:12 -04:00
GuangYaoZheng and GitHub
3f3e862681
fix(eagle3): read norm_before_fc from eagle_config for NVIDIA checkpoint ( #42143 )
...
Signed-off-by: FERRARIZHENG <popkart06@gmail.com >
2026-05-23 08:21:34 +00:00
Gabriel Wu and GitHub
82536acc54
Keep scheduler alive for delayed KV connector frees ( #43433 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-05-23 06:23:32 +00:00
Wei-Ming Chen and GitHub
09a219c075
[ModelOpt] Support Qwen3.5/3.6 VLM quantized prefix mapping ( #42546 )
...
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com >
2026-05-23 06:23:31 +00:00
d19db10974
[Bugfix] Fix native Triton top-k/top-p kernel assumes contiguous logi… ( #42739 )
...
Signed-off-by: xiaogang.zhou <xiaogang.zhou@bytedance.com >
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com >
2026-05-22 22:56:16 -07:00
Taneem Ibrahim and GitHub
3a1c062151
[Misc] Added missing return type annotations to improve mypy and IDE tooling ( #43383 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-23 13:28:22 +08:00
a7be0f342d
[7/n] Migrate pos_encoding and norm kernels to libtorch stable ABI (continued) ( #43209 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-23 13:20:00 +08:00
54d153637b
[XPU] reudce host overhead of XPU MOE ( #42915 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 13:09:34 +08:00
a5bbd81e2e
[XPU]feat: enable FP8 block-scaled quantization on XPU ( #42952 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 12:33:18 +08:00
Andreas Karatzas and GitHub
d28bdf9344
[ROCm][CI] Fix ROCm LoRA Transformers fallback with full CUDA graphs ( #41577 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 04:31:32 +00:00
84e351555a
[Bugfix] Auto-raise max_num_batched_tokens for prefix-LM multimodal models ( #43051 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-05-22 21:23:50 -07:00
Andreas Karatzas and GitHub
76ea1d5d2f
[ROCm][CI] Stabilize Granite tool-use and test URL construction ( #43017 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 12:21:11 +08:00
Andreas Karatzas and GitHub
6a4723a2e0
[ROCm][CI] Stabilize runner teardown between sampler tests ( #43023 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 12:19:54 +08:00
Yongye Zhu and GitHub
367cb81966
[DSV4] More multi-stream enablement for c4a ( #42925 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-23 09:22:27 +08:00
3cb83c9592
Add model to WeightTransferEngine.__init__ ( #42922 )
...
Signed-off-by: SumanthRH <sumanthrh99@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-22 17:52:15 -07:00
Duncan Moss and GitHub
552bbe6f4e
[Attention] Add head_dim=512 support for FlashInfer trtllm attention backend ( #38822 )
2026-05-22 20:27:35 -04:00
Itay Alroy and GitHub
6d30655b13
elastic_ep: stage/commit MoE quant method on reconfigure ( #40881 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-05-22 18:57:26 -04:00
8de5cabeb7
[XPU]fix: add XPU platform guards to DeepSeek-V4 ops ( #42950 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 06:29:45 +08:00
4e2eba28be
[Perf] Optimize hidden state extraction logic ( #37374 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-22 18:23:08 -04:00
gnovack and GitHub
f743254143
DSv4 fused Q-norm kernel grid refactor ( #42353 )
2026-05-22 15:21:33 -07:00
Nick Hill and GitHub
47d4407d7c
[Model Runner V2] Support sharing kv cache layers ( #35045 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 22:18:23 +00:00
Juhi Mittal and GitHub
e203006a8b
[Quantization][ModelOpt] W4A16 NVFP4 fused MoE + mixed-precision dispatch ( #42566 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
2026-05-22 20:51:49 +00:00
08cb46789d
mhc_post - remove sts & add vectorized copies ( #43437 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-22 13:44:29 -07:00
4e597b7491
[Bugfix] Clear error message for FP8 torchao quantization on unsupported GPUs ( #36854 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-22 20:09:17 +00:00
Artem Perevedentsev and GitHub
23f7b11bf4
[Bugfix] Detect wrong libcute_dsl_runtime.so variant in FlashInfer GDN ( #43427 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-22 19:33:33 +00:00
977703aa94
[RFC][EPLB][ #32028 ] Remove dead torch.accelerator.synchronize() from sync path ( #40733 )
...
Signed-off-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com >
Co-authored-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com >
2026-05-22 15:19:24 -04:00
2b94d1c0ca
[Frontend] Simplify AuthenticationMiddleware path extraction ( #43426 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-22 11:59:14 -07:00
Yongye Zhu and GitHub
843715739b
[Refactor] Extract DeepSeek V4 sparse MLA impl into model folder ( #43149 )
2026-05-22 10:06:31 -07:00
b21f3d56d4
[KV Connector] MooncakeStore: don't co-queue save with load to avoid double delayed-free ( #43371 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-22 16:14:11 +00:00
c7624bea5e
[Bugfix] Source num_qo_heads from Attention layers in Flashinfer/Triton metadata builders ( #42650 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
2026-05-22 16:10:03 +00:00
Bugen Zhao and GitHub
91f5b92438
[Rust Frontend] [Refactor] Extract a newtype for utility call ID ( #43405 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-22 08:22:11 -07:00
Isotr0py and GitHub
f0feb15e7f
[Multimodal] Simplify ViT CUDA graph interfaces ( #41234 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-22 22:31:00 +08:00
sychen52 and GitHub
fb21d8b4f9
Add NVFP4 MOE support for Deepseek V4. ( #42209 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-22 07:21:51 -07:00
haosdent and GitHub
a377631d21
[CI] Fix AMD docker build tests ( #43329 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 14:06:24 +00:00
d3a563501b
[EPLB] Change default EPLB communicator ( #43110 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-05-22 09:43:27 -04:00
Jee Jee Li and GitHub
15f7cd33dc
[LoRA] Reduce memory of 2D weights when EP is set ( #42737 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-22 06:41:56 -07:00
79ff0ffa98
[BugFix] wire make_empty_intermediate_tensors on AyaVision and Voxtral ( #43118 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
2026-05-22 05:26:41 -07:00
Tobias Wasner and GitHub
4658bf882b
[Bugfix] Clear P0 mm sender cache on sleep/pause to fix mm_hash desync ( #43001 )
...
Signed-off-by: Tobias Wasner <wasnertobias@gmail.com >
2026-05-22 03:54:29 -07:00
b3c7ffcab8
[Misc] Replace assert with proper exceptions for security and validation in pooling ( #43286 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 18:43:33 +08:00
d3d1cf6972
[XPU]feat: add XPU fallback for MoE topk routing and MXFP4 backend ( #42951 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-22 10:22:45 +00:00
wangxiyuan and GitHub
7e1b45a092
[Attention] Mamba attention module refactor ( #41126 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-05-22 17:13:12 +08:00
Li, Jiang and GitHub
65b7a812a2
[CPU] Experimentally enable Triton and MRV2 ( #43225 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-22 01:48:17 -07:00
2380bfc210
[Docs] Note image preprocessing difference between qwen_vl_utils and vllm. ( #43393 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-22 01:43:14 -07:00
mrjunwan-lang and GitHub
a761697717
Fix the docker build failure in tpu-inference ( #43360 )
...
Signed-off-by: mrjunwan-lang <mrjunwan@google.com >
2026-05-22 01:36:17 -07:00
Nick Hill and GitHub
694d9a81bb
[BugFix] Fix setuptools-rust dep in requirements files ( #43377 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 15:25:10 +08:00
Weida Hong and GitHub
6bb8753db1
Correcting the mock classes for MM GC tests ( #43321 )
...
Signed-off-by: Weida Hong <wdhongtw@google.com >
2026-05-22 15:21:35 +08:00
haosdent and GitHub
025d4f5cd2
[CI] Fix "test_awq_load[gemma4-moe-*]" failure ( #43296 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 07:13:59 +00:00
5ea76fa89a
[CI] Fix test_lora_with_spec_decode on V2 model runner ( #43314 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-22 14:24:18 +08:00
tc-mb and GitHub
fa1ff88b31
[Model] Fix MiniCPM-V 4.6 vit_merger qkv weight loading ( #43213 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-05-21 22:44:06 -07:00
Furkan F and GitHub
e746a2eebf
[Model] Use AutoWeightsLoader for Voyage ( #42972 )
...
Signed-off-by: Furkan Fidan <dev@yufufi.com >
2026-05-22 05:28:23 +00:00
haosdent and GitHub
1fe3303983
[CI] De-flake renderers/test_hf.py::test_resolve_content_format_fallbacks[Qwen/Qwen-VL-string] ( #43064 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 12:15:22 +08:00
8c8b1825eb
[XPU] Enable multiple key kernels for sparse attention ( #37888 )
...
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com >
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-22 12:02:51 +08:00
18a27cc9a3
[Bugfix] Make CuMemAllocator free callback stream-aware ( #43020 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-22 03:36:22 +00:00
0ddd7dd656
[Frontend] DP Supervisor ( #40841 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: robertgshaw2-redhat <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 20:33:16 -07:00
60af5c16ee
[Frontend] Add truncation side to OpenAI endpoints ( #43260 )
...
Signed-off-by: Rui Zhang <rza21.bc@gmail.com >
Signed-off-by: Rui Zhang <rui.zhang@globalrelay.net >
Co-authored-by: Rui Zhang <rui.zhang@globalrelay.net >
2026-05-21 20:32:31 -07:00
Divakar Verma and GitHub
35d0141a0b
[ROCm][CI] add warmup to mem_util test before measurement ( #43236 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-22 03:17:54 +00:00
Simon Danielsson and GitHub
86ccef7d44
[ROCm] Add XGMI backend for MoRI Connector ( #41753 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-22 03:06:40 +00:00
2998a047aa
[Bugfix] Fix DSV4 Base model swiglu limit issue in FP8 path ( #42855 )
...
Signed-off-by: Chengze Fan <chengze@meta.com >
Signed-off-by: Chengze Fan <fancz2002@gmail.com >
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com >
2026-05-21 19:43:01 -07:00
Isotr0py and GitHub
ba369b7eb5
[CI] Fix dockerfile dependency graph failure for pre-commit ( #43378 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-22 10:26:05 +08:00
39910f2b25
[Rust Frontend] Move code from vllm-frontend-rs ( #43283 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Eric Curtin <eric.curtin@docker.com >
Signed-off-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com >
Signed-off-by: Will.hou <1205157517@qq.com >
Signed-off-by: Will.hou <willamhou@ceresman.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Eric Curtin <eric.curtin@docker.com >
Co-authored-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com >
Co-authored-by: Will.hou <1205157517@qq.com >
Co-authored-by: Will.hou <willamhou@ceresman.com >
Please see https://github.com/Inferact/vllm-frontend-rs for full original commit history.
2026-05-21 17:21:48 -07:00
Lanze Liu and GitHub
39d5fa96a7
[Bugfix] Zero stale is_prefilling in padded CUDA graph rows for Mamba ( #41873 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-05-21 15:42:42 -07:00
Nick Hill and GitHub
565b745ec5
[BugFix] Use correct logprobs for logprob_token_ids ( #43125 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 15:42:20 -07:00
e26e1f0928
[Feature] Add --cpu-distributed-timeout-seconds CLI Option for CPU Process Group Timeout ( #42968 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Signed-off-by: zWaNg3 <389750525@qq.com >
Co-authored-by: zWaNg3 <389750525@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-21 15:42:07 -07:00
Nick Hill and GitHub
0f66623b0d
[Frontend] Rework fastokens integration ( #43168 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 15:36:58 -07:00
0b59fc45dd
Disable build isolation to bypass CUDA related deps for vllm-tpu ( #43038 )
...
Signed-off-by: Ylang Tsou <ylangt@google.com >
Co-authored-by: Ylang Tsou <ylangt@google.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-21 18:00:52 -04:00
17b69828a0
[Core] Add native ModelExpress load format ( #43105 )
...
Signed-off-by: Zheng Luo <zheluo@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-21 16:05:01 -04:00
Wentao Ye and GitHub
b29cbf0652
[Perf] zeros -> empty to remove additional fill ( #42988 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-21 16:00:29 -04:00
Michael Goin and GitHub
9b54e50e2c
[Deprecation] Mark env vars covered by --moe-backend / --linear-backend ( #43148 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-05-21 12:51:12 -07:00
1c78f76c29
[Bugfix] Add early validation to reject incompatible runner types for embedding models ( #43079 )
...
Signed-off-by: anish <anishesg@users.noreply.github.com >
Signed-off-by: Your Name <ak8686@princeton.edu >
Signed-off-by: anish <145943060+anishesg@users.noreply.github.com >
Co-authored-by: anish <anishesg@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-21 11:07:46 -04:00
haosdent and GitHub
9b9d5dbaab
[CI] Fix CPU tests failing on tl.exp2 import ( #43311 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 14:28:34 +00:00
b730c46352
[Perf] [Hybrid] Fused Triton kernel for GPU-side Mamba state postprocessing ( #40172 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-21 04:50:54 -07:00
c68c55d43e
[CPU][RISC-V] Add VLEN=256 support to RVV attention kernels ( #42943 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
Signed-off-by: velonica0 <47554626+velonica0@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-21 04:50:49 -07:00
5ecd8e9c70
[XPU][CI]Fix Docker image pull-to-run race in Intel GPU CI ( #43266 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-21 10:41:38 +00:00
haosdent and GitHub
caf69823d6
[CI] Pin protoc binary in rust-build stages ( #43292 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 03:38:07 -07:00
68e07d5916
[Bug] Fix ci issue assert output_size is not None AssertionError ( #43261 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-05-21 16:58:09 +08:00
ebbfb34e3e
[Test] Replace zephyr-7b-beta (7B) with SmolLM2-135M in tokenization test ( #43085 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-21 01:57:47 -07:00
zhangxin81 and GitHub
edafea3555
Fix FlashInfer TRTLLM NvFP4 monolithic MoE routing ( #43223 )
...
Signed-off-by: zhangxin81 <115389973+zhangxin81@users.noreply.github.com >
2026-05-21 01:17:12 -07:00
b719b1635b
Update KDA chunk prefill decay to use exp2 semantics ( #43195 )
...
Signed-off-by: zexplorerhj <19794632+zexplorerhj@users.noreply.github.com >
Co-authored-by: zexplorerhj <19794632+zexplorerhj@users.noreply.github.com >
2026-05-21 01:16:27 -07:00
Kunshang Ji and GitHub
0a54df2847
[XPU] add setuptools-rust for xpu dependency ( #43287 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-21 00:14:13 -07:00
haosdent and GitHub
a950e9447e
[CI] De-flake test_models for bigscience/bloom-560m ( #43197 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 06:30:14 +00:00
050611a3dd
[Bugfix] Fix glm4_moe_tool_parser._is_string_type for /v1/responses FunctionTool format ( #39601 )
...
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 22:58:59 -07:00
yzong-rh and GitHub
905b97adfa
[Benchmark] Add num-warmup to vllm bench throughput ( #43245 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-21 05:13:15 +00:00
Daoyuan Li and GitHub
a6682d1d25
[Bugfix] Warn when renderer_num_workers has no effect on offline LLM ( #42905 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-05-20 21:35:08 -07:00
f2ace1d57d
[Frontend][RFC] Rust front-end integration ( #40848 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-05-21 12:24:48 +08:00
d97ba29fdc
[ToolParser][Bugfix] Re-land: Fix anyOf/oneOf/$ref type resolution in Qwen3CoderToolParser ( #37831 ) ( #38973 )
...
Signed-off-by: AAISSJ <maze0717@g.skku.edu >
Signed-off-by: <>
Signed-off-by: sejung-son <sejung.son@nhn.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: 세덩 <saison@sedeong-ui-MacBookAir.local >
Co-authored-by: sejung-son <sejung.son@nhn.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-21 12:24:08 +08:00
Flora Feng and GitHub
6441cf4a44
[Refactor] Use shared coerce_to_schema_type in Seed-OSS tool parser ( #43140 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 21:24:06 -07:00
346cf163a1
[Frontend] Normalize reasoning_content to reasoning for client compatibility ( #42664 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-20 21:23:47 -07:00
haosdent and GitHub
7e5070934e
[CI] Fix "test_vit_cudagraph_[image|video][step3_vl]" failure ( #43082 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-20 21:22:10 -07:00
2b75a73b8e
[Perf][Gemma4] Batch vision encoder calls for image and video processing ( #43169 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-20 21:22:06 -07:00
e45df8c3f7
[Bugfix] Fix Qwen3.5 GatedDeltaNet in_proj_ba Marlin failure at TP>=2 ( #36329 )
...
Signed-off-by: Adi McM Sonus Flow <biuro@sonusflow.pl >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-20 21:22:01 -07:00
Jee Jee Li and GitHub
ee05e8137e
[Minor] Bigger overlap for FI AR ( #43103 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-20 21:20:57 -07:00
Louie Tsai and GitHub
5d041cc1fe
update GPU json file based on h200 recipes ( #43262 )
...
Signed-off-by: louie-tsai <louie.tsai@intel.com >
2026-05-21 03:57:48 +00:00
9640970de2
[Model Runner V2] Fix lora Triton Error [CUDA]: device-side assert triggered ( #43139 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 01:00:30 +00:00
63ea11709b
[CI] Add composed-schema regression tests for DeepSeek V3.2/V4 parsers ( #43255 )
...
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-21 00:36:16 +00:00
akii96 and GitHub
bde560ed6e
[ROCm] Add QuickReduce min-size override and codec threshold ( #41675 )
...
Signed-off-by: <>
2026-05-20 17:46:51 -05:00
Jiangyun Zhu and GitHub
6dc0a71843
[Misc] downgrade nvidia-cutlass-dsl to 4.5.0 ( #43230 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-20 14:19:50 -07:00
Michael Goin and GitHub
5774aad9c5
[Perf][gpt-oss] Downgrade triton_kernels to v3.5.1 ( #43135 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-20 14:13:12 -07:00
Douglas Lehr and GitHub
452baa860b
Add dllehr-amd to CODEOWNERS and committers list ( #42772 )
...
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com >
2026-05-20 16:10:44 -05:00
Flora Feng and GitHub
2a43b407c5
[Bugfix][CI] Add missing import of pad_nvfp4_activation_for_cutlass in flashinfer ( #43237 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 11:59:12 -07:00
53ff50fcd3
[Perf] Optimize CutlassFP8ScaledMMLinearKernel when padding needed by pre-weight processing, 13.5% TTFT improvement ( #42651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-20 11:57:42 -07:00
363fc84407
Integrate flashinfer b12x MoE and FP4 GEMM kernels for SM120/121 ( #40082 )
...
Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-05-20 17:21:11 +00:00
f2d5e3d3ae
[CI] Lower granite-4.0-h-tiny gsm8k threshold for Hybrid SSM NixlConnector PD accuracy tests (4 GPUs) ( #43186 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-20 17:00:24 +00:00
2d6b3489b9
[R3] Add routed experts to openai entrypoint ( #38939 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-20 09:07:59 -07:00
Vadim Gimpelson and GitHub
9c78c99995
[MISC] Fix symm_mem cap-equal gate; log AR backend selection ( #42993 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-20 08:50:24 -07:00
Flora Feng and GitHub
a10d69116c
[Bugfix] Use shared coerce_to_schema_type in DeepSeekV32 tool parser ( #43019 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 10:21:00 -04:00
644b2a28e7
[Bugfix] Use enable_sm120_family for per-tensor FP8 CUTLASS kernels on SM12.1 ( #41215 )
...
Signed-off-by: j9smith <j.smith9103@outlook.com >
Signed-off-by: Joel Smith <j.smith9103@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-20 14:10:01 +00:00
ded871201a
[Bug][Structured Outputs] Fix bug that leads to unconstrained generations with structural tags ( #42452 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-20 07:08:58 -07:00
Dipika Sikka and GitHub
df84fb07a6
Remove additional dead code as a follow-up to #42889 ( #43144 )
...
Signed-off-by: Dipika Sikka <dipikasikka1@gmail.com >
2026-05-20 10:01:45 -04:00
Benjamin Chislett and GitHub
0a508743d4
[Spec Decode] Support non-MTP speculation for NemotronH ( #43130 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-20 09:15:52 -04:00
Kebe and GitHub
19cf334207
[Feature] Support manually enabling the cumem allocator ( #33648 )
...
Signed-off-by: Kebe <mail@kebe7jun.com >
2026-05-20 08:58:30 -04:00
87e31455b0
[Doc] Sync CLI guide with actual help modes and launch subcommand ( #40326 )
...
Signed-off-by: Rui Wang <raygorous@gmail.com >
Co-authored-by: Rui Wang <raygorous@gmail.com >
2026-05-20 02:32:03 -07:00
cb600d1cdb
[Frontend] Forward X-data-parallel-rank header on /inference/v1/generate ( #42330 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-20 08:58:46 +00:00
xiangdong and GitHub
6f21558da1
[XPU][CI] Add 2 server model test files in Intel GPU CI ( #42499 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-05-20 16:54:58 +08:00
Artem Perevedentsev and GitHub
1cb224430b
[GDN] Enable FI Blackwell GDN prefill kernel ( #40717 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-20 01:46:55 -07:00
Harry Mellor and GitHub
9b343dd4f5
Enable mermaid diagrams in the docs ( #43192 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-20 08:10:00 +00:00
07aeaf9d4d
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) ( #42663 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-20 00:18:12 -07:00
Nicolò Lucchesi and GitHub
40651c0207
[Docs][PD][NIXL] Bidirectional kv-cache transfer ( #43097 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-20 09:02:36 +02:00
Nicolò Lucchesi and GitHub
7e4bc2cecb
[Docs][PD][NIXL] Lease extension mechanism for blocks on P ( #43099 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-20 08:58:25 +02:00
Kevin H. Luu and GitHub
85959567c3
[ci] Revert model executor test back to L4 ( #43188 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-19 23:01:41 -07:00
Ronen Schaffer and GitHub
4f940896a3
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers ( #43076 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-20 03:32:08 +00:00
Michael Goin and GitHub
cd0ff26e7a
[CI] Add DSV4-Flash to gsm8k moe-refactor/config-b200.txt ( #42111 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-19 20:21:01 -07:00
Izik Golan and GitHub
2ae910ed88
[Perf] Avoid forward scan for async output placeholders ( #42938 )
2026-05-19 20:16:07 -07:00
fadf5d332c
add enqueue all option to throughput benchmark ( #42975 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Signed-off-by: pmaybank <113125070+pmaybank@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 20:16:02 -07:00
Benjamin Chislett and GitHub
c628a93a64
[Perf][Bugfix] Update dflash aux layer indexing ( #40727 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-19 20:15:57 -07:00
Terrence Zhao and GitHub
5774aaed0c
[Cohere] Enable Cohere MoE ( #43143 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-05-19 19:32:06 -07:00
Nick Hill and GitHub
39bba710be
[MRV2][BugFix] Fix default-stream CG capture in P/W LoRA case ( #43160 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-19 19:19:05 -07:00
Aaron Hao and GitHub
73dd2f33b7
[bug] fix WeightTransferConfig.backend to allow for all strings ( #43121 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
2026-05-19 21:01:29 -04:00
Fadi Arafeh and GitHub
be16785998
[CPU][DOC] Fix installation commands for Arm CPUs ( #43115 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-05-19 23:31:15 +00:00
117afeea46
Fix error in Dynamic NTK scaling ( #41277 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
Signed-off-by: Max de Bayser <maxdebayser@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-19 17:27:54 -04:00
Doğaç Eldenk and GitHub
1242196295
[Model] Support post-norm architecture for EAGLE-3 supeculators ( #42764 )
...
Signed-off-by: Doğaç Eldenk <dogacel@gmail.com >
2026-05-19 13:39:00 -07:00
Kevin H. Luu and GitHub
a65093c1a3
[ci] Move language models tests (hybrid) back to L4 ( #43129 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-19 11:51:34 -07:00
9aaf83ef50
[CI failure] Temporarily disable using persistent cache for flashinfer autotune ( #43119 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-19 11:44:32 -07:00
tomeras91 and GitHub
f54721bcc3
[Bugfix][MoE] FlashInfer one-sided: workspace union across heterogeneous layers ( #42976 )
...
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com >
2026-05-19 14:43:04 -04:00
aed2eb355a
[Docs] Fix MooncakeStoreConnector role in disaggregated example ( #42994 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-19 11:14:43 -07:00
Dom Brown and GitHub
d247a931cc
[feat] Add FP8 per-tensor Q scale support to Triton attention backend ( #42080 )
...
Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com >
2026-05-19 09:02:05 -07:00
Jinzhen Lin and GitHub
8200fbe1ac
[Misc] add humming to dependencies ( #42540 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
2026-05-19 08:36:47 -07:00
Flora Feng and GitHub
42b4f1fdf7
[Refactor] Extract extract_types_from_schema utility from Minimax M2 tool parser ( #43025 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-19 11:21:12 -04:00
Wang Yiwen and GitHub
1c6158083a
[Model] Openvla support ( #42654 )
...
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com >
2026-05-19 08:17:42 -07:00
Xinyu Chen and GitHub
d740e2c029
[XPU] update xpu graph usage ( #43043 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
2026-05-19 23:09:07 +08:00
Nick Hill and GitHub
b82e908b4c
[Perf][4/n] Eliminate various GPU<->CPU syncs ( #42347 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-19 10:35:54 -04:00
Sage and GitHub
a78b842d0e
[Bugfix] Fix top logprobs token placeholders in /inference/v1/generate ( #42887 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-05-19 10:21:49 +00:00
129019f334
[CI] Add MTP + PD disagg test for Qwen3.5 ( #42677 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-19 11:44:33 +02:00
Shanshan Shen and GitHub
ef54a4d604
[Misc][MM] Remove redundant code in CLIPAttention ( #43046 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-05-19 08:43:16 +00:00
Woosuk Kwon and GitHub
07beaed842
[Model Refactoring] Rename deepseek_v4.py to model.py [4/N] ( #43077 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 01:12:46 -07:00
Yifan Qiao and GitHub
056bc2e166
[KVConnector][DSV4] HMA support for Mooncake store connector ( #42828 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-19 01:07:46 -07:00
f34623bf3c
[bug] AsyncScheduler drops first post-resume token after pause_generation + clear_cache ( #42117 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 01:06:21 -07:00
Woosuk Kwon and GitHub
b14be81c1f
[Model Refactoring] Move deepseek_v4_ops to models/deepseek_v4 [3/N] ( #43073 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 00:52:54 -07:00
wang.yuqi and GitHub
301d986473
[Frontend] Consolidate beam search by BeamSearchMixin. ( #42946 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-19 07:37:40 +00:00
257af77bc2
[Docs] Reorganize online serving docs. ( #41907 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-19 14:43:18 +08:00
Taneem Ibrahim and GitHub
4a4fdabe28
[Misc] Aligning tokwise pooler heads for consistency ( #43041 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-19 06:16:42 +00:00
f1e3f0e6d6
[XPU] Use custom op collective behavior ( #41354 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-19 14:14:59 +08:00
9fd8487d2f
[Docs] Add SVG images for pooling models. ( #42626 )
...
Signed-off-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-18 22:50:38 -07:00
27f4ba9481
fix: use keyword arguments for shard_id and expert_id in weight_loade… ( #42671 )
...
Signed-off-by: junyanxu <junyanxu5513@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 05:29:04 +00:00
6e889b582b
[ci] Route 28 gpu_1_queue tests to h200_35gb queue ( #43030 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-18 21:58:36 -07:00
fab07e4d0f
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B ( #42289 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
Co-authored-by: gemini-code-assist <noreply@google.com >
2026-05-18 21:22:33 -07:00
3ca8db2ef8
add cutedsl dsv4 indexer fp8 kernel ( #42899 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-18 21:17:56 -07:00
Woosuk Kwon and GitHub
87b08c5f64
[Model Refactoring] Move DeepSeek V4 layers to models/deepseek_v4/ [2/N] ( #43039 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-18 21:00:58 -07:00
fba010dd74
[Bugfix][MRV2] Fix KVCache tensor explicit kernel_block_size dim ( #42766 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-18 20:25:41 -07:00
Mohammad Miadh Angkad and GitHub
da03e549b3
[UX] Add a persistent cache for FlashInfer autotuning ( #42537 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-18 20:25:37 -07:00
Kunshang Ji and GitHub
36dcaf25d8
[XPU] add gptq(int4) support ( #37844 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-19 11:17:09 +08:00
Ofir Zafrir and GitHub
8f16c4a5c0
[BugFix][CPU][Spec Decode] Fix Eagle implementation on CPU backend ( #42468 )
...
Signed-off-by: Ofir Zafrir <ofir.zafrir@intel.com >
2026-05-19 03:16:07 +00:00
afd7b1dce9
[Bugfix] Use platform-agnostic device in example_connector load ( #42926 )
...
Signed-off-by: Revital Sur <eres@il.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-19 03:12:04 +00:00
Woosuk Kwon and GitHub
287471b994
[Model Refactoring] Migrate DeepSeek V4 to vllm/models/ [1/N] ( #43004 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-18 19:50:02 -07:00
239b5ff30c
[Frontend] Add --spec-method/--spec-model/--spec-tokens CLI aliases ( #42476 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-18 17:22:27 -07:00
Artem Perevedentsev and GitHub
f85c76d701
[CI/Build] Bump nvidia-cutlass-dsl to 4.5.1 ( #42991 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-18 16:58:15 -07:00
shanjiaz and GitHub
a171e6b52d
Add parallel drafting to v2 model runner unsupported features ( #43010 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-05-18 16:39:09 -07:00
Wentao Ye and GitHub
37ece593c1
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement ( #42774 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 16:38:12 -07:00
Flora Feng and GitHub
57fef4e0bf
[Refactor] Extract shared coerce_to_schema_type utility from Minimax M2 tool parser ( #43006 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-18 17:55:39 -04:00
haosdent and GitHub
0191354827
[Perf][MLA] Enable FULL cudagraph capture for TRITON_MLA decode ( #42885 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-18 14:29:10 -07:00
Wentao Ye and GitHub
cd49a05d5a
[Refactor] Remove dead code ( #42889 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 16:41:22 -04:00
Ronen Schaffer and GitHub
84747489de
Tier offload followup ( #42529 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-18 19:41:58 +00:00
Tuukka Sarvi and GitHub
8fc1c284b9
[ROCm] Guard AITER GDN decode fast path by layout ( #42880 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-05-18 11:56:22 -07:00
Amit Portnoy and GitHub
ce88f01c9a
[Docs] update attribution to reflect EDEN foundation ( #41666 )
...
Signed-off-by: amitport <1131991+amitport@users.noreply.github.com >
2026-05-18 11:22:56 -07:00
Wentao Ye and GitHub
00e20e76f7
[Refactor] Remove dead cuda kernels ( #42767 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 11:14:21 -07:00
czhu-cohere and GitHub
9758a6e5c5
[BugFix] support PP for Cohere vision model ( #42819 )
...
Signed-off-by: <conway.zhu@cohere.com >
Signed-off-by: root <conway.zhu@cohere.com >
2026-05-18 11:12:06 -07:00
Bowen Bao and GitHub
a2c8fc6657
[ROCm][Quantization][3/N] Refactor quark_moe w4a4 w/ oracle ( #41436 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-18 13:46:13 -04:00
6859ca7615
[Bugfix] fix swiglu limit issue for humming backend + deepseek v4 ( #42541 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-18 17:32:26 +00:00
Mohammad Miadh Angkad and GitHub
67f58ce23f
[Bugfix] Fix DSV4 MTP after ROCm mHC integration ( #42930 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-18 17:02:01 +00:00
Wei Zhao and GitHub
8c296de63b
[Perf] Re-enable flashinfer autotune by default and cleanup ( #42857 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-18 09:12:27 -07:00
Harry Mellor and GitHub
b12745e4f3
Fix --convert passed without --runner on causal models ( #42935 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 15:56:09 +00:00
Wentao Ye and GitHub
e26736973a
[Model Runner V2] Fix prompt logprobs calculation Sizes of tensors must match error ( #42778 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 08:27:21 -07:00
Netanel Haber and GitHub
47829b1159
[Bugfix] mamba: run single-token extends as decodes ( #42430 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-18 15:26:00 +00:00
Blanc Swan and GitHub
4a39b4f553
[Model] Add Apertus Tool Parser ( #41154 )
...
Signed-off-by: Blanc <swan.blanc@infomaniak.com >
2026-05-18 11:20:04 -04:00
78e7a7b9b0
Refactor AWQ Marlin MoE onto modular WNA16 oracle ( #42483 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Signed-off-by: Siddharth Bedekar <104613085+bedeks@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 08:02:43 -07:00
f5d3dc7115
[Model Runner v2] Support update_config ( #42783 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-18 10:26:07 -04:00
1ac10f159a
Revert "[torch.compile] Add patch for fullgraph compilation" ( #42686 ) ( #42913 )
...
Co-authored-by: Luka Govedič <luka.govedic@gmail.com >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-18 09:02:51 -04:00
e5417657e5
[KV Connector][Offloading] Flush all pending jobs on last step ( #42611 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 12:59:42 +00:00
xiangdong and GitHub
2e40faf08b
[XPU][CI] Temporarily skip test_moe_lora_align_block_size_mixed_base_and_lora[1] in Intel GPU CI ( #42954 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-05-18 20:34:48 +08:00
Nicolò Lucchesi and GitHub
69c91d010a
[MRv2] Default to MRv1 when a connector is present ( #42955 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-18 20:34:16 +08:00
roikoren755 and GitHub
737bfa3a43
[Bugfix][Hybrid][NemotronH] Fix mamba_cache_mode=all + speculative decoding crash ( #41233 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-05-18 14:54:00 +03:00
Kfir Toledo and GitHub
e414e1f1c0
[Bugfix][KV Offload] count appended GPU blocks in store group_sizes ( #42945 )
...
Signed-off-by: Kfir Toledo <kfir.toledo@ibm.com >
2026-05-18 11:36:02 +00:00
df852ed503
fix: remove unused norm for dpskv4 ( #41710 )
...
Signed-off-by: inisis <desmond.yao@buaa.edu.cn >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-18 18:33:29 +08:00
Yuwen Zhou and GitHub
88a860d754
[CPU] Add MXFP4 W4A16 MoE support ( #41922 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-05-18 03:04:45 -07:00
cac81b6eda
[CPU Backend] Improve cpu thread utilization ( #42666 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 03:04:41 -07:00
Li, Jiang and GitHub
b4601ad43f
[CPU] Add fused GDN support for AMX CPU platform ( #42707 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-18 03:04:36 -07:00
Jee Jee Li and GitHub
2267f70070
[Kernel] Pack topk id/weights triton kernel ( #42527 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-18 03:04:31 -07:00
965d076148
[CPU] Specify required KV cache layout for CPU attention backend ( #42740 )
...
Signed-off-by: Tony Lin <tony.lin@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-18 17:38:54 +08:00
c38bed4248
delete xpu ci ( #42582 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 16:36:45 +08:00
Xin Yang and GitHub
998714b21b
[Perf] Add do_not_specialize in fused FP8 RoPE kernel ( #42849 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-05-18 01:32:46 -07:00
Harry Mellor and GitHub
9537542537
Revert checkpoint specific workaround in Transformers modelling backend ( #42923 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 17:31:06 +09:00
Rishapveer Singh and GitHub
5ab6d1b3fd
[Model] [Perf] Use flatten for Qwen3.5's GDN output projection ( #42311 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-05-18 16:14:36 +08:00
7d5b033782
[LoRA] Support 2D and 3D MoE LoRA adapter at the same time ( #42242 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-18 15:22:26 +08:00
e3aeee5ff8
[Bugfix] moe lora align kernel grid ( #40131 )
...
Signed-off-by: TheDuyIT <nduy250299@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: dtnguyen <dtnguyen@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-18 00:17:53 -07:00
Harry Mellor and GitHub
c1f7854342
Improve logging when docs build is skipped ( #42929 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 06:33:32 +00:00
gaozihao-shy and GitHub
23c15acd77
[BugFix] Kimi-K2.5: skip vision tower dtype conversion when using quantization ( #42869 )
...
Signed-off-by: gaozihao-shy <gaozihao-shy@users.noreply.github.com >
Signed-off-by: gaozihao <gaozihao3@huawei.com >
2026-05-18 05:07:16 +00:00
Andreas Karatzas and GitHub
b50646e5ef
[ROCm][CI] Stabilize ROCm pooling and multimodal CI ( #42909 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-18 03:57:59 +00:00
Soyaazz and GitHub
990f49bdcb
[MM][CG] Enable encoder Cudagraph for Step3VL ( #42224 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
2026-05-17 20:19:13 -07:00
107210442d
[CI] Add NIXL EP import canary ( #42567 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-05-17 19:11:46 -07:00
03ddc1c9bc
[Perf] Wire silu_and_mul_per_block_quant into TritonFP8MoE (MiniMax-M2) ( #42497 )
...
Signed-off-by: qianlihuang <yiliu.dong@qq.com >
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Co-authored-by: qianlihuang <yiliu.dong@qq.com >
2026-05-17 21:57:04 -04:00
Luka Govedič and GitHub
966903eb93
[torch.compile] Add patch for fullgraph compilation ( #42686 )
...
Signed-off-by: Luka Govedič <luka.govedic@gmail.com >
2026-05-17 19:49:16 +00:00
TJian and GitHub
599e75f432
[ROCm] [Bugfix] Fix DeepSeek V4 Functionality and Accuracy ( #42810 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-17 12:18:50 -04:00
Taneem Ibrahim and GitHub
1c8e9c0399
Refactor: Pass num_labels explicitly to PoolerClassify instead of reading from global config ( #42851 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-17 14:40:21 +00:00
0fa888465e
[XPU] fix weight scale shape ( #42725 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-17 16:55:10 +08:00
liuzhenwei and GitHub
ff712f6447
[MRV2][XPU] add Model Runner V2 log ( #42710 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-17 04:15:50 +00:00
Qi Zhou and GitHub
504a26ce2b
Support bf16 for mamba ssm cache ( #41680 )
...
Signed-off-by: Qi Zhou <qizzzh@google.com >
2026-05-16 17:54:58 -07:00
weizhoublue and GitHub
a94189295b
Fix Weight loading for Qwen3.5-MTP and Qwen3-VL using runai_streamer ( #42716 )
...
Signed-off-by: weizhoublue <weizhou.lan@daocloud.io >
2026-05-16 17:54:27 -07:00
0867497368
[CI/Build] Bump flashinfer to v0.6.11.post2 ( #41711 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-16 14:55:12 -07:00
36e74c9ea4
[KV Connector] Support disk offloading in MooncakeStoreConnector ( #42689 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-16 13:34:15 -07:00
Taneem Ibrahim and GitHub
787bc0d031
Add unit tests for pooler activation functions ( #42824 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-16 14:58:16 -04:00
weizhoublue and GitHub
d1586e1a12
Fix: Propagate pinned model revisions into Ultravox secondary weight loading ( #42830 )
2026-05-16 17:02:54 +00:00
Jiangyun Zhu and GitHub
8a56da3845
[Experimental] Breakable CUDA graph ( #42304 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-16 22:04:12 +08:00
Andreas Karatzas and GitHub
4db300e95f
[ROCm][CI] Removed problematic command override mechanism ( #42807 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-16 17:35:05 +08:00
657b42b592
[Docker][KVConnector] Build mooncake-transfer-engine from source ( #42114 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: khluu <khluu000@gmail.com >
2026-05-16 00:26:25 -07:00
Jee Jee Li and GitHub
32b7177909
[LoRA][Bugfix] Dedup LoRA wrapping for modules referenced from multiple attribute paths (MoE gate) ( #42757 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-16 11:22:35 +08:00
39c67d714e
fix: add API key authorization to /v2 endpoints ( #42594 )
...
Signed-off-by: DustHunter <dusthunter@126.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-16 01:29:27 +00:00
87a2adcb43
[Misc] Add common random prefix option to structured-output serving benchmark ( #41632 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-16 00:44:48 +00:00
Michael Goin and GitHub
852f567444
[Bugfix] Respect explicit --kv-cache-dtype over checkpoint kv_cache_scheme ( #42782 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-15 17:15:52 -07:00
Michael Goin and GitHub
b2a27b82d9
[Kernel][UX] Add --linear-backend arg for linear kernel selection ( #39538 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-15 17:07:39 -07:00
Keyi Li and GitHub
d0921bafef
[Bugfix] Unwrap VLM wrappers for EPLB on Model Runner V2 ( #42706 )
2026-05-16 07:20:33 +08:00
1ccdf87507
[Bugfix] Fix layerwise reload alias-buffer corruption ( #42481 )
...
Signed-off-by: rasdani <73563550+rasdani@users.noreply.github.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-15 15:20:53 -07:00
Rita Brugarolas and GitHub
bd9dbe6060
[ROCm][Bugfix] Fix fused_mla_dual_rms_norm for AITER API rename _fused_qk_rmsnorm ( #42606 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-05-15 14:50:03 -06:00
de2d76f352
[Build] Switch CUDA 12.9 wheel builds to PyTorch manylinux_2_28 base ( #41668 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-15 13:46:16 -07:00
9a7a273dfe
Add HumanEval and GSM8K benchmarks to datasets ( #42648 )
...
Signed-off-by: southfreebird <yvorott@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-15 13:01:21 -07:00
b2c58ee942
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants ( #42685 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-15 15:34:59 -04:00
frida-andersson and GitHub
4d67d3bde2
[ROCm] Restore fast top_k_per_row kernels for sparse MLA when topk_tokens=2048 ( #42072 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-05-15 19:02:57 +00:00
06d020bb6e
[Bugfix] Fix SM121 (DGX Spark) exclusion from Marlin/CUTLASS FP8 paths ( #35568 )
...
Signed-off-by: Blake Ledden <blake@secondnaturecomputing.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Pavani Majety <pmajety@nvidia.com >
2026-05-15 10:59:00 -07:00
chunxiaozheng and GitHub
f45c210885
[LMCacheMPConnector] Prioritize importing the lmcache_mp_connector from lmcache ( #42596 )
...
Signed-off-by: idellzheng <idellzheng@tencent.com >
2026-05-15 17:46:31 +00:00
akii96 and GitHub
be7a03ea65
[ROCm] Widen AITER fused AR RMSNorm 1-stage gate ( #42409 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-15 17:44:38 +00:00
6147c70224
[Model Runner v2] Support reload weights (sleep mode) ( #42673 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-15 16:41:23 +00:00
0162596603
[Model Runner V2] FP32 gumbel sampling. ( #41775 )
...
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
2026-05-15 09:20:08 -07:00
46a95815d3
[ROCm][MLA] FP8 ASM prefill for AITER dense MLA backend on gfx950 ( #42509 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: clintg6 <clint.greene@amd.com >
Co-authored-by: frida-andersson <frida.andersson@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-15 23:56:58 +08:00
BadrBasowid and GitHub
fb5bd03f51
[Perf] Set IR Op Priority Once at Worker Init ( #42631 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-05-15 15:56:13 +00:00
Mohammad Miadh Angkad and GitHub
ee58665aac
[Bugfix] Fix DeepGEMM context lens contiguity in MLA indexer ( #42135 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-15 23:29:58 +08:00
Wentao Ye and GitHub
491e8d8539
[Perf] Optimize MLA attention _v_up_proj bmm by removing additional copy ( #42561 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-15 08:14:26 -07:00
Wentao Ye and GitHub
af9616d845
[Model Runner V2] Fix kv_connector pre_forward order ( #42676 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-15 08:13:59 -07:00
d792d993c1
[ROCm] Widen OAI Triton MoE capability range to include gfx12 (RDNA4) ( #37826 )
...
Signed-off-by: L.B.R. <lbr@mmonad.com >
Co-authored-by: L.B.R. <lbr@mmonad.com >
2026-05-15 07:59:57 -07:00
Aaron Hao and GitHub
e0a45f1455
[Feat][RL] IPC weight sync optimizations: multigpu support and chunked packed tensors ( #37476 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Signed-off-by: hao-aaron <ahao@anyscale.com >
2026-05-15 22:53:06 +08:00
Benjamin Chislett and GitHub
0fe7550254
[Bugfix] DFlash FP8 KV-Cache ( #42692 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-15 08:29:45 -06:00
Li, Jiang and GitHub
95cfe102a5
[Bugfix] Ensure embeding model compilation on CPU ( #42709 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-15 18:58:19 +08:00
1dc3fe08ea
gemma3 multi-gpu bug-fix ( #42630 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-15 02:32:05 -07:00
d26a28ab03
fix: propagate revision/code_revision pins to all artifact boundaries ( #42616 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-15 02:31:54 -07:00
Andreas Karatzas and GitHub
d735968f6d
[ROCm][CI] Stage B gating ( #42025 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-15 01:49:27 -07:00
ccde9540be
DeepSeekV4-Pro enable cuda graph full and piecewise mode ( #42604 )
...
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-15 01:45:30 -07:00
wang.yuqi and GitHub
75fd68c7a5
[Entrypoints] Split the pooling offline API into PoolingOfflineMixin. ( #42267 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-15 08:05:57 +00:00
Yifan Qiao and GitHub
4b364f810e
[Core][DSV4] Skip caching SWA blocks that can never serve a prefix-cache hit ( #42258 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-15 15:59:18 +08:00
31fa757cf9
[Misc] Make it simpler to replace out-of-tree layer classes with related LoRA layers. ( #42306 )
...
Signed-off-by: paulyu12 <507435917@qq.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-15 15:20:42 +08:00
Cyrus Leung and GitHub
2676ab1e0b
[Deprecation] Remove old locations of get_tokenizer and resolve_hf_chat_template ( #35024 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-15 00:13:32 -07:00
27b85d2084
[Bugfix] Clarify CPU backend memory error messages reference shared flag ( #42479 )
...
Signed-off-by: daniel-devlab <282598346+daniel-devlab@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-15 06:35:05 +00:00
Louie Tsai and GitHub
e30f39c4f1
Update Intel Xeon model list and vLLM Benchmark Suite BKMs ( #42607 )
...
Signed-off-by: louie-tsai <louie.tsai@intel.com >
2026-05-15 05:14:03 +00:00
bf610c2f56
[Bugfix] Fix inverted condition causing thinking_token_budget to be silently ignored ( #41674 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-15 12:48:49 +08:00
faa4b76afa
[Model] Support InternS2 Preview ( #42705 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: zxy <46674730+CUHKSZzxy@users.noreply.github.com >
2026-05-14 21:30:26 -07:00
f351455f0f
[CPU][RISC-V] Add RVV-optimized attention kernels for RISC-V Vector Extension ( #40119 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-15 12:08:23 +08:00
Cyrus Leung and GitHub
56434e8651
[Bugfix] Fix incorrect chat template format for Qwen3.5 ( #42660 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-14 20:52:52 -07:00
0d4d334eaa
Bump llguidance to 1.7 ( #42150 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 20:35:27 -04:00
fa2a33b893
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py ( #38288 )
...
Signed-off-by: Chengyi Nie <cnie@roblox.com >
Co-authored-by: Chengyi Nie <cnie@roblox.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-15 08:25:52 +08:00
Giancarlo Delfin and GitHub
3b6a204789
[Model Runner V2][Bug Fix][DSV4] Ensure lazy attention state initializations happen during cudagraph capture ( #42444 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-05-14 16:16:17 -07:00
f8848b2f2d
[Bugfix] Add swiglu limits to deepgemm fp8 methods ( #41986 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-14 15:43:13 -07:00
Charlie Fu and GitHub
4cfcc0866f
[CI][ROCm] Remove unsupported cases in test_fusion.py ( #38680 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-05-14 17:37:18 -04:00
f887aa1a53
[Aiter][ROCm] RMSNormGated+GroupedQuantFP8 fusion ( #40710 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Tres Popp <trespopp@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-14 15:37:09 -04:00
Matthew Bonanni and GitHub
9898f94abe
[Attention] Remove deprecated MLA prefill arguments ( #42555 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-14 10:34:06 -07:00
ae4f59f0ec
[Model Runner v2] Oracle for model runner v2 - qwen3 dense model by default [1/N] ( #39337 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-14 10:02:33 -07:00
Ranran and GitHub
f3d5360591
[Bugfix][Multimodal] PyAV video backend returns keyframes labeled as targets ( #42586 )
...
Signed-off-by: Ranran <hzz5361@psu.edu >
2026-05-14 08:56:59 -07:00
Baorun (Lauren) Mu and GitHub
a7737cb4f3
[Fix] Misc Fixes in ViT CUDA Graph ( #38040 )
...
Signed-off-by: Baorun Mu <bmu@nvidia.com >
2026-05-14 23:49:06 +08:00
Cyrus Leung and GitHub
b8a25d0e12
[Bugfix] Fix LM detection for Nemotron Parse ( #42641 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-14 23:42:10 +08:00
frida-andersson and GitHub
f07b1da797
[ROCm] Enable gluon paged MQA logits on gfx950 (MI355X) ( #42062 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-05-14 15:39:26 +00:00
f60c6b33a5
[V1][DP][LB] Publish request counts at the start of each engine step ( #41626 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-14 15:39:24 +00:00
24337fb860
PD disagg with NIXL Connector: GDN support (Qwen3.5) ( #41869 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-14 16:33:01 +02:00
c7560af424
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support ( #39568 )
...
Signed-off-by: xhx1022 <1737006628@qq.com >
Signed-off-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com >
2026-05-14 14:12:30 +00:00
Mohammad Miadh Angkad and GitHub
2317682f95
[Bugfix] Fix TRTLLM ragged MLA prefill workspace warmup ( #42112 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-14 09:48:56 -04:00
5bd8c71e79
[kv_offload] Implement reset_cache() for the offloading connector ( #41956 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com >
2026-05-14 16:00:10 +03:00
Wentao Ye and GitHub
6548560496
[Compile] Fix compile warning with topk softplus sqrt ( #41261 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-14 05:12:50 -07:00
0a65d46628
[DSV4] Fuse norm and router for low latency scenario ( #41263 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io >
Co-authored-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 05:11:02 -07:00
Zhenzhong Xu and GitHub
1ea9401364
[Quantization][Autoround][Toolkit] Add W4A16 Support ( #39778 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
2026-05-14 19:18:49 +08:00
9946c38b7f
[XPU] Fix double-transpose in XPUFP8ScaledMMLinearKernel for W8A8 quant method ( #41689 )
...
Signed-off-by: Libin Tang <libin.tang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 17:17:39 +08:00
23c85343fb
[Bug] Fix DeepSeek V4 AttributeError: module 'cutlass.cute.nvgpu' has no attribute 'LoadCacheMode' ( #42342 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-14 02:00:20 -07:00
rasmith and GitHub
768f4a6f26
[CI][AMD][BugFix] Prevent triton compiler error when running test_moe_layer with use_ep = True on ROCm ( #40857 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-14 08:44:22 +00:00
rasmith and GitHub
addef3299c
[CI][AMD] Skip tests where models have problems or fails on both HW types ( #42126 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-14 08:21:06 +00:00
ce29c26b31
Update Dockerfile.rocm for AINIC & Thor NIC ( #40453 )
...
Signed-off-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu >
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu >
Co-authored-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
Co-authored-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-14 15:24:27 +08:00
aoshen02 and GitHub
8c79ad6580
Revert "[Core] Replace routing replay with device cache and async D2H pipeline" ( #39917 ) ( #42434 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-05-13 23:49:01 -07:00
0d2732dd91
[MLA Attention Backend] Add TOKENSPEED_MLA backend for DSR1/Kimi K25 prefill + decode on Blackwell ( #41778 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-13 23:48:02 -07:00
Rebecca Lee and GitHub
fd7d858c8a
Use hidden_pad and intermediate_pad from vLLM #34301 ( #42098 )
...
Signed-off-by: Rebecca Lee <Rebecca.Lee@amd.com >
2026-05-14 14:21:04 +08:00
liuzhenwei and GitHub
b26558d4a3
[CI][XPU] skip ut of offload connector ( #42598 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-14 13:13:53 +08:00
Sarah Salah and GitHub
bf0d2dc6d7
[Misc] Fix mypy error in parser_manager type narrowing ( #42441 )
...
Signed-off-by: Sarah-Salah <11881117+Sarah-Salah@users.noreply.github.com >
2026-05-14 02:48:59 +00:00
ca60a4e84f
[Fix] Weight loading for qwen3_5 using runai_streamer ( #42521 )
...
Signed-off-by: Harsh Shah <iharsh@google.com >
Co-authored-by: Harsh Shah <iharsh@google.com >
2026-05-14 10:36:20 +08:00
Roy Wang and GitHub
77e1421a68
[Bugfix] Fix EPLB initialization for VLM wrapper models ( #39805 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-05-14 02:26:15 +00:00
Kunshang Ji and GitHub
751b9f14bd
[XPU][CT] Support mxfp8 moe model ( #41918 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-14 09:47:10 +08:00
Krish Gupta and GitHub
70c00163ff
[Feature] Add instruction support for score/rerank chat templates ( #42412 )
...
Signed-off-by: KrxGu <krishom70@gmail.com >
2026-05-14 09:41:22 +08:00
Siddharth Bedekar and GitHub
f51f6844f9
[Bugfix][Spec Decode] Wire draft_probs into probabilistic draft_model rejection ( #40269 )
2026-05-13 21:04:03 -04:00
longguo and GitHub
665f9c4253
[Bugfix] Fix Gemma4ToolParser streaming float corruption ( #42128 )
...
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-05-13 18:03:30 -07:00
Flora Feng and GitHub
1087676a90
[Refactor] Use shared utils in hermes tool parser ( #42570 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-13 20:35:45 -04:00
63cc8a55a9
fix(tool-parser): preserve "none"/"nil" strings as valid enum values in minimax_m2 ( #39599 )
...
Signed-off-by: Yiyang Liu <yiyangliu@microsoft.com >
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-13 20:35:34 -04:00
Divakar Verma and GitHub
ca7e4546da
[CI] set max transformers version for skywork model ( #42104 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-13 16:53:49 -07:00
b2198670b1
[Bugfix] V1: support tuple model outputs in ubatch wrapper (dbo + spec decode) ( #40789 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-13 15:47:51 -07:00
Mohammad Miadh Angkad and GitHub
f1cc7aad3c
[Bugfix] Fix DeepSeek V4 MTP HC state handling ( #42320 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-13 15:44:52 -07:00
Lukas Geiger and GitHub
597ed13803
[Core][MM] Do not use urllib3 to parse data URLs ( #42535 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-05-13 22:21:01 +00:00
liangel-02 and GitHub
6b5c389ee3
expose flex block size for batch invariant mode ( #41252 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-05-13 14:11:57 -07:00
Michael Goin and GitHub
8efd508204
[Quantization] Rework quantization_config to use QuantKey and allow for activation override ( #41566 )
2026-05-13 16:58:32 -04:00
ovidiusm and GitHub
cca32d55a2
[PD] Fix broken NIXL EP installation ( #42542 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
2026-05-13 13:55:51 -07:00
Walter Beller-Morales and GitHub
873910d608
[Frontend] add support for thinking_token_budget in completions ( #42116 )
2026-05-13 16:01:52 -04:00
Wentao Ye and GitHub
3f611f6106
[CI] Fix pre-commit issue ( #42563 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-13 12:37:26 -07:00
Nick Hill and GitHub
a505cf807e
[ModelRunner V2] Share identical MTP weights ( #42538 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-13 18:57:04 +00:00
40330967ab
[Quark] Support loading Quark NVFP4 checkpoints in vLLM ( #35859 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Signed-off-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-05-13 11:17:36 -07:00
Fynn Schmitt-Ulms and GitHub
ab1ad0d7a9
Remove verifier model type check in speculative config ( #42536 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-05-13 18:14:39 +00:00
Ben Browning and GitHub
0f69128a37
[Bugfix] Handle real-world gpt-oss tool call output in Harmony parsing ( #42454 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-05-13 17:54:46 +00:00
b3c69595a6
[MM][CG] Support ViT CG for Qwen2-VL ( #41736 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-14 01:52:35 +08:00
2f821faeae
[Spec Decode] Support hybrid attention models in extract_hidden_states ( #39949 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-13 10:45:53 -07:00
5794c65f8c
[Bugfix][Model] Gemma4 MoE routing closure captures per_expert_scale, breaking functional_call substitution ( #42250 )
...
Signed-off-by: Noelia <noeliabentancor1@gmail.com >
Signed-off-by: Noelia Bentancor <71080743+NoeliaBentancor@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 17:43:12 +00:00
CynicDora and GitHub
256dbcaabf
[Feature] Support custom callable proposer backend for speculative decoding ( #39487 )
...
Signed-off-by: 524031910363 <hyzhyzsh@sjtu.edu.cn >
Signed-off-by: CynicDora <hyzhyzsh@sjtu.edu.cn >
2026-05-13 16:53:01 +00:00
Wentao Ye and GitHub
e35c0d4c63
[Feature] Support compile mode for batch invariance on SM80 ( #42456 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-13 11:02:39 -04:00
Ronen Schaffer and GitHub
11f6b545d4
[kv_offload] Add multi-tier KV cache offloading framework ( #40020 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-13 17:21:43 +03:00
a8887c208f
[Bugfix] [ROCm] [DSV4] [Perf] Add aiter mhc support ( #41946 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-13 21:43:15 +08:00
0ddaf6dffa
[XPU] [CT] Enable CT W4A4MxFp4 path and add xpu kernel ( #38896 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-13 06:43:00 -07:00
Marek Wawrzos and GitHub
67671692ac
[CI] Re-enable Nemotron Parse parity test and switch testing to nemotron-parse v1.2 ( #42498 )
...
Signed-off-by: <mwawrzos@nvidia.com >
2026-05-13 21:05:27 +08:00
hissu-hyvarinen and GitHub
0a62f5eec9
[AMD] skip machete tests for rocm ( #42326 )
...
Signed-off-by: Hissu Hyvarinen <hissu.hyvarinen@amd.com >
2026-05-13 12:11:03 +00:00
PikaPikachu and GitHub
3b1ef03be4
[Bugfix][Quark] Fix W8A8 INT8 garbage outputs on Step-3.5-Flash (and other 3-key fused-MoE Quark exports) ( #41892 )
...
Signed-off-by: kangletian <kangletian@hotmail.com >
2026-05-13 11:59:49 +00:00
3c413a5481
Triton attention: add USE_TD constexpr for tensor descriptor Q/K/V load/store ( #40327 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: quinnlp <quinnlp@users.noreply.github.com >
2026-05-13 13:57:41 +02:00
Ronen Schaffer and GitHub
79fd1bc7ed
[kv_offload] Add req_id to ReqContext for per-request tracking ( #42507 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-13 11:11:10 +00:00
SILONG ZENG and GitHub
cee6751e54
[Bugfix][Qwen3-VL] Fix pipeline-parallel deepstack initialization ( #42394 )
...
Signed-off-by: MrZ20 <2609716663@qq.com >
2026-05-13 10:58:42 +00:00
16863072ca
[Bugfix] Fix scipy audio resampling ratio ( #42233 )
...
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
2026-05-13 18:52:41 +08:00
Andreas Karatzas and GitHub
d628a3c5cb
[ROCm][CI] Skip ROCm batch invalid-input test pending torch fix ( #41572 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-13 18:50:47 +08:00
akii96 and GitHub
74dffae666
[ROCm] Run AITER RMSNorm pad fusion before AR RMS fusion ( #42411 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-13 18:35:12 +08:00
97c4317bf5
[Bugfix][Frontend] Default max_tokens server-side on /inference/v1/generate ( #42329 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-13 11:16:46 +02:00
f6e868fbdf
[CI] Use uv with Python 3.12 for PyPI wheel upload ( #42470 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-13 02:12:06 -07:00
13bf242100
[Feat][KVConnector] Add bind_gpu_block_pool() to KVConnectorBase_V1 ( #39654 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 02:10:29 -07:00
Jiangyun Zhu and GitHub
140dc2ec30
[Bugfix] Install nvidia-cutlass-dsl[cu13] extra on CUDA 13 platforms ( #42438 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-13 01:57:21 -07:00
9ce74042d3
[Bugfix][SimpleCPUOffloadBackend] Dedup in-flight CPU offload stores across scheduler steps ( #41289 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 01:53:32 -07:00
sychen52 and GitHub
a8c13d2837
Patch SlidingWindowSpec.real_page_size_bytes for nvfp4 kv ( #42464 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-13 01:46:30 -07:00
Shanshan Shen and GitHub
92def124bc
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3.5 ( #42151 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-05-13 16:00:32 +08:00
85b2fecab7
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) ( #42339 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-05-13 07:24:39 +00:00
503697c9ce
[chore] Refactor pooling metadata token ID accessors ( #42368 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-13 06:08:01 +00:00
Nicolò Lucchesi and GitHub
71bcd02ef3
[Bugfix][PD] Fix multi-node TP (TP>8) ( #39907 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-12 22:20:57 -07:00
Matthew Bonanni and GitHub
dcacdf9a88
[Attention] Sync FA with upstream ( #41052 )
2026-05-12 23:34:18 -04:00
bnellnm and GitHub
18f6bf5a21
[MoE Refactor] Add sequence parallel tests to test_moe_layer.py ( #41299 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-12 21:52:19 -04:00
Alec and GitHub
07534b8782
[PD] Bump NIXL connector dependency to 1.x ( #42364 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
2026-05-12 18:05:01 -07:00
Wentao Ye and GitHub
3d635c58c0
[Perf] Optimize MLA compute_prefill_context memory allocation ( #42460 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-12 16:23:46 -07:00
+4
ebeb09d822
[KV Transfer] Add MooncakeStoreConnector for KV cache offloading via Mooncake distributed store ( #40900 )
...
Signed-off-by: leichao.lc <leichao.lc@antgroup.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: leichao.lc <leichao.lc@antgroup.com >
Co-authored-by: ivanium <yifanqiao@inferact.ai >
Co-authored-by: aoshen524 <aoshen@inferact.ai >
Co-authored-by: Dao007forever <daole@inferact.ai >
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com >
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com >
Co-authored-by: foraxe <1055696449@qq.com >
Co-authored-by: Skywalker-EP <173423846@qq.com >
Co-authored-by: fems14 <1804143737@qq.com >
Co-authored-by: jianzs <zheng.shoujian@outlook.com >
Co-authored-by: baxingpiaochong <771405853@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 16:09:10 -07:00
Michael Goin and GitHub
184577ae46
[Build] DeepGEMM: trim comments, add integration notes + TODOs ( #42429 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-12 15:57:58 -07:00
Kevin H. Luu and GitHub
8c4fc4202a
[CI] Inline build artifact annotations in release pipeline ( #42357 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 15:57:43 -07:00
Nick Hill and GitHub
fe8b42e80c
[CI] Fix test_async_scheduling.py flakiness ( #42455 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-12 21:38:32 +00:00
Giancarlo Delfin and GitHub
fe5b4e0fe7
[Model Runner V2] Apply synthetic mode to probabilistic rejection sampler ( #41035 )
2026-05-12 13:37:03 -07:00
0ce6613b9c
platforms: add uses_cpu_device() hook to Platform for DeviceConfig ( #42313 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-12 12:39:17 -07:00
379f0ec369
[CI] Migrate 6 verified jobs from gpu_1_queue to h200_18gb MIG ( #42446 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-12 11:52:01 -07:00
KaivalyaMDabhadkar and GitHub
67c89fe40a
[Model][Bugfix] Fix Step3-VL image_embeds input path ( #42333 )
...
Signed-off-by: Kaivalya Dabhadkar <kdabhadkar@nvidia.com >
2026-05-12 18:47:55 +00:00
d9b4990783
[MoE Refactor] EPLB refactoring for FusedMoE ( #41055 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-12 14:16:31 -04:00
4d591db470
[MoE Refactor] Introduce RoutedExperts alias for FusedMoE and don't store SharedExperts in MK ( #40735 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-05-12 13:37:44 -04:00
yzong-rh and GitHub
6ff7405b81
[Bugfix] [Frontend] Responses API, fix merging of messages ( #42189 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Yifan <yzong@redhat.com >
2026-05-12 16:09:59 +00:00
Yan Ru Pei and GitHub
bcb9c133ba
feat(kv-events): emit KV cache metadata ( #40984 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com >
2026-05-12 15:58:48 +00:00
c8a6e272e0
[CPU] Fix rotary embedding for CPU without flash-attn ops ( #42225 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-12 15:05:35 +00:00
Wentao Ye and GitHub
a1b2d87498
[Refactor] Clean up pooling models build_tok_params logic ( #42341 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-12 15:05:05 +00:00
Martin Hickey and GitHub
418ba8ef14
[kv_offload][BugFix] Fix store deferral ( #41945 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-12 18:04:44 +03:00
5a6a9fc6f6
[docs] Added one new contact to the Vulnerability Management team ( #42145 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 10:59:59 -04:00
289cee0473
[vLLM IR] Minor improvements ( #39362 ) ( #39558 )
...
Signed-off-by: Avishek Goswami <avishek.goswami@ibm.com >
Co-authored-by: Avishek Goswami <avishek.goswami@ibm.com >
2026-05-12 10:58:36 -04:00
shanjiaz and GitHub
6ccb10d794
Added peagle speculators support ( #41826 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-05-12 07:55:57 -07:00
7a9cc5e7f0
[Model] Support MiniCPM-V 4.6 ( #41254 )
...
Signed-off-by: caitianchi <caitianchi@tc-mb.com >
Signed-off-by: tc-mb <157115220+tc-mb@users.noreply.github.com >
Co-authored-by: caitianchi <caitianchi@tc-mb.com >
2026-05-12 14:28:10 +00:00
d077622d60
[Build] Build bundled DeepGEMM _C per-Python so the wheel imports on every CPython ( #41516 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-12 10:27:29 -04:00
dd6b3a5ef5
[Perf] Use 2D-grid to eliminate divmod in W8W8 group quant ( #42153 )
...
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-12 10:01:30 -04:00
593d5a4033
[Bugfix] Fix mismatched kernel-per-logical blocks in NIXL HMA transfer ( #42097 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-12 15:53:30 +02:00
bnellnm and GitHub
6427603ae8
[MoE Refactor] Move remaining experts classes to experts directory ( #42334 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-12 09:19:46 -04:00
206eaed08d
[MoE Refactor] Move expert map related code into ExpertMapManager class ( #41046 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-05-12 09:18:27 -04:00
8f89381fc6
[Hybrid] Warmup Mamba2 SSD kernel ( #39822 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-12 12:46:22 +00:00
Dipika Sikka and GitHub
a7b801e26d
[MXFP4] Support for linear layers + compressed-tensors integration ( #41664 )
2026-05-12 07:49:33 -04:00
Kunshang Ji and GitHub
4df1be9547
[XPU] bump up vllm-xpu-kernels to v0.1.8 ( #42410 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-05-12 11:47:37 +00:00
bc03f280c8
[XPU] keep generator state of sycl kernel align with pytorch ( #41771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Qiming Zhang <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 11:44:47 +00:00
997132911e
[Doc] Fix typo in llm-d documentation link ( #42397 )
...
Signed-off-by: Florian Woerner <florian.woerner@onmyown.io >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-05-12 04:26:46 -07:00
haosdent and GitHub
fc8bf6eedb
[CI] De-flake Language Models Test (Extended Generation) test_models(False-False-5-32-bigcode/starcoder2-3b) ( #42392 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-12 10:46:48 +00:00
liuzhenwei and GitHub
07a40ede19
[UT][XPU] fix test_parallel_sampling due to global random state ( #42388 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-12 18:03:23 +08:00
Kevin H. Luu and GitHub
e1c8776e90
[CI] Move DockerHub and PyPI publish steps to end of release pipeline ( #42355 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 09:17:42 +00:00
Kevin H. Luu and GitHub
1ff9d33535
[CI] Migrate remaining B200 jobs to b200-k8s with test fixes ( #42387 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 02:00:37 -07:00
7f65f84428
[Bugfix] Fix empty channel/recipient in harmony for /v1/responses ( #35540 )
...
Signed-off-by: kg6-sleipnir <christopherhazen42@gmail.com >
Signed-off-by: chazen <45186108+kg6-sleipnir@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-12 08:45:51 +00:00
amitz-nv and GitHub
ef34592a1a
[Bugfix] Fix double reduce in flashinfer_nvlink_two_sided and flashinfer_nvlink_one_sided backends ( #41382 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-12 07:47:47 +00:00
Kevin H. Luu and GitHub
f69644caf8
[CI] Migrate more B200 jobs to b200-k8s queue ( #42356 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 00:38:31 -07:00
d37e25ffbe
[Frontend] Consolidate Speech to Text entrypoints. ( #42370 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 07:06:57 +00:00
8517cdaf90
[XPU] update dp rank w/o env-var isolation ( #39856 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 14:49:54 +08:00
Lucas Kabela and GitHub
4e498b5e5c
[Bugfix][Performance Improvement] Improve penalties triton kernel performance ( #40657 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-12 05:47:20 +00:00
28ee78af54
Implement custom dataset class for ASR benchmarking ( #41576 )
...
Signed-off-by: Yasmin Moslem <48152713+ymoslem@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 12:17:58 +08:00
ZiTian Zhao and GitHub
630492da30
[Fix] Gemma4 Mixed-Resolution Image Co-Batching Crash ( #42217 )
...
Signed-off-by: zitian.zhao <zitian.zhao@tencentmusic.com >
2026-05-12 03:13:03 +00:00
Chauncey and GitHub
920bf3ec84
[Bugifx] [Qwen3CoderTool] Restore supports_required_and_named for required tool_choice ( #42292 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-12 02:09:56 +00:00
pschlan-amd and GitHub
39dff5ff39
Add VLLM_USE_SPINLOOP_EXT to use more efficient busy polling ( #36517 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 16:11:49 -07:00
d7af6b34d8
[Model Runner V2] Bug fix: logprob dtype int64/int32 issue ( #41761 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 21:55:43 +00:00
bbee532988
[Perf][1/n] Eliminate various GPU<->CPU syncs ( #41429 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 20:36:03 +00:00
53181384e0
[Bugfix] Fix DSV4 swiglu_limit on marlin backend ( #42287 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 13:03:56 -07:00
wang.yuqi and GitHub
a0dc7a0f36
[CI] Consolidate Speech to Text tests ( #42274 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-11 19:50:17 +00:00
56e5810ff1
[BugFix] Prevent orphaned process on NCCL destroy ( #39846 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-11 15:25:26 -04:00
Flora Feng and GitHub
639cbfd274
[CI] Add tests/parser to CI coverage ( #41877 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-11 19:08:54 +00:00
a721315488
[ROCm][Perf] Fix RMSNorm+Quant fusion for gfx950 (non-fnuz) ( #41825 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Chuan Li <chuali@amd.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: Frida Andersson <frida-andersson@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-11 15:00:51 -04:00
6fdb49392e
[Bugfix] Fix int32 overflow in DeepGEMM SiLU/mul FP8 Triton kernel ( #42201 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Signed-off-by: Vensen <vensenmu@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 14:52:31 -04:00
cf0d279142
[Docs] Add Apple Silicon documentation for vLLM-Metal GPU support ( #41987 )
...
Signed-off-by: alexagriffith <agriffith96@gmail.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-05-11 11:34:25 -07:00
5497ffbf7c
Add documentation about vLLM FIPS compliance ( #42190 )
...
Signed-off-by: Vinay Damodaran <vrdn@hey.com >
Signed-off-by: Vinay R Damodaran <vrdn@hey.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
2026-05-11 18:17:02 +00:00
Nick Hill and GitHub
9af6a5ed75
[Model Runner V2] Fix seq_lens_cpu_upper_bound ( #42202 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-11 10:37:50 -07:00
Hexiang Wang and GitHub
7863fff6e5
[ROCm][DSv4] implement flash sparse mla with triton kernels ( #41812 )
...
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
2026-05-11 09:27:11 -07:00
Wentao Ye and GitHub
0d453e2336
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement ( #40408 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-11 12:20:58 -04:00
Wentao Ye and GitHub
3f9c0c25b3
[Bug] Fix kimi dtype issue with mm_projector_forward ( #42081 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 11:45:24 -04:00
Vadim Gimpelson and GitHub
a2e776d716
[Bugfix] Accept canonicalized modelopt_* quant_method in _extract_modelopt_quant_algo ( #42181 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-11 11:10:57 -04:00
4955990f1b
[kv_offload] Move FilterReusedOffloadingManager logic to CPUOffloadingManager ( #41727 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 18:09:29 +03:00
Wentao Ye and GitHub
4b64fc2cbf
[Refactor] Cleanup batch invariant dead code ( #41993 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 10:48:39 -04:00
pschlan-amd and GitHub
5f1b313900
[ROCm] Clean up a bit the AITER FA backend ( #41942 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 22:45:18 +08:00
724ed2fc35
[DSv4] Improved dequant gather K cache kernel ( #42236 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 10:41:12 -04:00
a51376b3f0
[Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA ( #40392 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
Co-authored-by: ElizaWszola <ewszola@redhat.com >
2026-05-11 14:10:50 +00:00
Martin Hickey and GitHub
8415bf2cdb
[kv_offload] Set offloading connector to prefer HND layout ( #41928 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-11 15:05:41 +03:00
Noa Neria and GitHub
ac062147fa
Avoid silent weights corruption when loading Nemotron Nano VL with reusable-buffer loaders like runai distributed streaming ( #42244 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-05-11 12:03:14 +00:00
Chauncey and GitHub
617239b70c
[Frontend]Responses API supports chat_template_kwargs ( #42272 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-11 11:59:39 +00:00
Kyungmin Lee and GitHub
27ae676364
Fix EXAONE-4.5 to align with Transformers update ( #42246 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
2026-05-11 10:25:31 +00:00
haosdent and GitHub
17ed5e61f5
[CI] Make Python-only Installation optional ( #42293 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-11 09:47:16 +00:00
Nicolò Lucchesi and GitHub
5672d100ed
[KV Connector][NIXL][Bugfix] Fix NIXL handshake failures not honoring kv_load_failure_policy ( #40364 )
...
When NIXL handshake fails (e.g., due to compatibility hash mismatch
between prefill and decode instances), requests fail with "engine dead"
error instead of gracefully falling back to local recomputation as configured
by kv_load_failure_policy='recompute'.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:37:21 +00:00
Nicolò Lucchesi and GitHub
770e9bd6b3
[Nixl][PD] Lease renewal TTL KV blocks on P ( #41383 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:27:30 +00:00
Cyrus Leung and GitHub
9efdddca28
[Model] Fix missing maybe_prefix ( #42280 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-11 09:04:06 +00:00
Qiu and GitHub
b1b59720b2
bugfix(flashinfer,dcp): remove kv_cache_layout for BatchDCPPrefillWrapper._new_tokens. ( #38895 )
...
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com >
2026-05-11 08:11:49 +00:00
f9f770ca0b
fix nixl side-channel host selection ( #41806 )
...
Signed-off-by: Shahar Mor <smor@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 07:40:37 +00:00
Haoqing Wang and GitHub
5cba6839e6
Document MolmoWeb hf_overrides ( #42163 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-10 23:58:22 -07:00
Jee Jee Li and GitHub
05d610e5cd
[CI/Build] Reduce LoRA model tests. ( #42266 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-11 14:49:08 +08:00
581b5e9afc
[Frontend] Return rendered prompt text in chat completion response ( #42052 )
...
Signed-off-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Cursor <cursor@cursor.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-11 13:53:39 +08:00
wangxiyuan and GitHub
5536fc0c01
[Misc] Replace mamba_type string literals with MambaAttentionBackendEnum ( #41188 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-05-11 03:59:36 +00:00
vllmellm and GitHub
7f95e66a11
[ROCm][Bugfix]: dynamically align BLOCK_DMODEL with Lv in MLA decode kernel ( #41119 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-11 11:14:19 +08:00
yzong-rh and GitHub
b1687527b8
[Bugfix] Gemma 4 chat template crash with missing tool name and tool id ( #42188 )
...
Signed-off-by: Yifan <yzong@redhat.com >
2026-05-11 03:07:45 +00:00
171019ab19
add fused mhc_post_pre kernel ( #41536 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-10 19:56:52 -07:00
Haoqing Wang and GitHub
879a8c3180
Fix Molmo2 image token metadata ( #42162 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-11 01:19:21 +00:00
1b57eb41f2
[MoE] Move various experts classes to fused_moe/experts/ ( #41979 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
2026-05-11 07:54:33 +08:00
Mohammad Miadh Angkad and GitHub
21943d4c25
[Performance] Make safetensors checkpoint prefetch settings configurable ( #41499 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 15:55:15 +00:00
f396bee56f
[DSV4] Add PP support for deepseek-v4 ( #41694 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: qizixi <22851944+zixi-qi@users.noreply.github.com >
2026-05-10 15:47:26 +00:00
215e2f7990
[Bugfix][Mamba] IMA in causal_conv1d kernel for long sequences ( #41617 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 12:38:28 +00:00
Ronen Schaffer and GitHub
e175192d33
[KV Offload] Pass ReqContext to touch(), complete_load(), and complete_store() ( #41366 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-10 15:09:25 +03:00
a54f0d1049
[CPU] Fix spec decode kernel signatures for synthetic mode compatibility ( #41932 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-05-10 12:07:15 +00:00
Isotr0py and GitHub
48698b1b9b
[Bugfix] Fuse Qwen3.5 in_qkvz_proj forwarding with LoRA enabled ( #37912 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-05-10 10:59:02 +00:00
Andreas Karatzas and GitHub
0a309b5ee9
[ROCm] Cap Triton paged attention block size to fix ROCm shared memory OOM ( #38502 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 10:03:00 +00:00
Jee Jee Li and GitHub
84f7a55340
[CI] Trigger LoRA test when changing MoE code. ( #42196 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-10 01:26:09 -07:00
Ethan Feng and GitHub
a2c9d548d7
[Docs] Fix broken local links ( #42160 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-10 01:15:38 -07:00
Yongye Zhu and GitHub
301305c093
Add @zyongye to CODEOWNERS ( #42200 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-10 16:07:32 +08:00
Mohammad Miadh Angkad and GitHub
efd0e7789d
Fix mypy failure on main ( #42197 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 07:55:57 +00:00
a5d0a5afba
[Frontend][Bugfix] Abort ASR engine requests on cancellation ( #41266 )
...
Signed-off-by: abdulrahman-cohere <abdulrahman.abdulrazzag@cohere.com >
Signed-off-by: <>
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-09 23:51:11 -07:00
Andreas Karatzas and GitHub
f2840120f6
[ROCm][CI] Fix NIXL spec-decode acceptance startup and diagnostics ( #41313 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 14:50:16 +08:00
Dao007forever and GitHub
3f5bd482f5
[Bugfix][KV Transfer][NIXL] Notify P node on pre-admission rejection to free stranded KV blocks ( #41269 )
2026-05-09 22:52:09 -07:00
Andreas Karatzas and GitHub
fb1ac806c5
[ROCm][CI] Stabilize ROCm shutdown and distributed compile CI ( #41573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 03:47:40 +00:00
Wei Zhao and GitHub
986edc858a
[Bugfix] Fix DeepSeek v4 topk numerical issue for unaligned max-model-len ( #42169 )
2026-05-09 20:30:08 -07:00
27d3bac272
docs: clarify Gemma 4 assistant speculative decoding ( #42180 )
...
Signed-off-by: AbhiOnGithub <abhiOnGithub@users.noreply.github.com >
Co-authored-by: AbhiOnGithub <abhiOnGithub@users.noreply.github.com >
2026-05-09 20:08:44 -07:00
00b0618a03
Use CU_MEMCPY_SRC_ACCESS_ORDER_ANY for batch KV cache swaps ( #39306 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <etelis2019@gmail.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Itay Etelis <etelis2019@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 05:57:09 +03:00
0d382ecde8
Handle optional bool-or-string CLI args in get_kwargs ( #40951 )
...
Signed-off-by: Christian Van <cvan20191@gmail.com >
Co-authored-by: Christian Van <cvan20191@gmail.com >
2026-05-09 19:47:21 -07:00
Isotr0py and GitHub
1029e5ef28
[CI/Build] Use modelscope's international site for regression test ( #42176 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-09 19:47:09 -07:00
0b272a6e01
[Bugfix] Fix SP pass for multimodal models and PP+SP residual handling ( #33322 )
...
Signed-off-by: Xingran Wang <wangxingran123456@outlook.com >
Signed-off-by: Hongjian Zhang <hirokenovo@gmail.com >
Co-authored-by: Hongjian Zhang <hirokenovo@gmail.com >
2026-05-09 19:44:16 -07:00
dcb3135af7
Fix: Nemotron 3 rescue whitespace-only final_content, not just None ( #41846 )
...
Signed-off-by: Nave Assaf <nassaf@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 02:07:58 +00:00
bc5fdc1e6a
Add NVFP4 all-gather GEMM fusion for AsyncTP ( #41882 )
...
Signed-off-by: roG0d <baonudesifeizhai@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 01:13:22 +00:00
aoshen02 and GitHub
006af4b956
[Bugfix] Skip routed-experts hot path when disabled ( #42148 )
2026-05-09 18:01:04 -07:00
Wentao Ye and GitHub
ea0e501bb1
[KV Connector] Remove compat support for pre-v0.12.0 constructor signatures without KVCacheConfig ( #39832 )
...
The v0.12.0 release contained initial support for HMA in KV Connectors. As part
of these changes, a KVCacheConfig argument was added to KV connector
constructors. Backwards compatibility support for out-of-tree connectors was
included in this change, with a very prominent warning. See #25712 and #27887 .
Since the warning has been around for over 5 months, we can safely remove
the support of it.
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-09 23:39:46 +00:00
Wentao Ye and GitHub
f80aa53c9d
[Refactor] Nixl util using lazy init ( #41392 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-09 17:46:52 -04:00
Juhi Mittal and GitHub
7a2b596982
[Quantization] Add ModelOpt NVFP4 W4A16 (4-bit weights, fp16/bf16 activations) support ( #41769 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
2026-05-09 21:15:50 +00:00
Jiangyun Zhu and GitHub
2ee8c2a56e
[SpecDecoding] extend mtp support for mimo 2.5 ( #41905 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-09 18:22:59 +00:00
SoluMilken and GitHub
cd74911d92
[Model] use AutoWeightsLoader for DeepSeekV2 ( #41706 )
...
Signed-off-by: SoluMilken <ypiheyn.imm02g@g2.nctu.edu.tw >
2026-05-10 01:55:25 +08:00
SoluMilken and GitHub
25abddc1a5
[BugFix] Fix Gemma4 'layers.0.moe.experts.0.down_proj_packed' KeyError issue ( #40708 )
...
Signed-off-by: SoluMilken <ypiheyn.imm02g@g2.nctu.edu.tw >
2026-05-09 17:20:44 +00:00
171d59ae8d
[Bugfix][PD] Fix DSv4 Disaggregated ( #41957 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: ZhanqiuHu <zhu@redhat.com >
2026-05-09 16:48:24 +00:00
3dda9aeb54
[Bugfix] Remove nested torch.compile in GDN rearrange_mixed_qkv causing CUDA graph capture failure ( #42070 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-05-09 08:30:55 -07:00
Kermit and GitHub
adb6d96516
[Bugfix] Fix GDN KKT precision loss on Hopper GPUs by aligning tl.dot operand layout with WGMMA ( #42076 )
...
Signed-off-by: kermit <ckeming@outlook.com >
2026-05-09 13:08:46 +00:00
Thien Tran and GitHub
530d371302
[DSv4] Improved fused Indexer Q quant kernel ( #41428 )
2026-05-09 01:20:32 -07:00
Micah Williamson and GitHub
34ab4f2565
[ROCm] Upgrade aiter to v0.1.13-rc5 ( #42113 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-09 08:13:45 +00:00
Jee Jee Li and GitHub
ecd0b60aad
[LoRA] Initial EP support for LoRA ( #40867 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-09 00:31:23 -07:00
d6563d693c
Require C++20 for compatibility with PyTorch ( #40380 )
...
Signed-off-by: Richard Barnes <rbarnes@meta.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-08 22:04:43 -07:00
Rishapveer Singh and GitHub
f6490a2841
[Bugfix] Preserve leading/trailing whitespace in GLM non-streaming tool parser ( #42026 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-05-08 21:49:15 -07:00
a2812becd6
[Models] Cohere Eagle + fix to Cohere MoE ( #42078 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-08 21:46:26 -07:00
e8f9038ebd
[ROCm][Bugfix] Re-tag AITER MoE weights as preshuffled after replace_parameter ( #42061 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-08 21:42:07 -07:00
df2636a9d8
[Bugfix] Fix LOGITPROC_SOURCE_ENTRYPOINT test to use spawn-compatible dist-info registration for XPU/ROCm ( #42040 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
Signed-off-by: David Zheng <153074367+dzhengAP@users.noreply.github.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 12:32:04 +08:00
Shengqi Chen and GitHub
97cc7685c4
Add @Harry-Chen in CODEOWNERS ( #42130 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-09 04:08:22 +00:00
haosdent and GitHub
e934e459e6
[CI][Bugfix] Make test_gpt2_cache_hit observable across V1 EngineCore ( #42037 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-09 11:53:15 +08:00
David Zheng and GitHub
845ca327ce
[Bugfix] Fix test_whisper distributed test process handling ( #42038 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-09 11:37:21 +08:00
4f6fa6341d
[XPU] update supported models on XPU ( #41911 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 10:44:03 +08:00
Ethan Feng and GitHub
a43bc34baf
[Docs] Update server entrypoint examples ( #42077 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-09 02:03:52 +00:00
Ethan Feng and GitHub
236bf9d152
[Docs] Fix RLHF example links ( #42073 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-09 02:03:42 +00:00
Lucas Wilkinson and GitHub
b1728c1e66
[Attention][Cleanup] Remove tree attention ( #42121 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-05-08 18:36:19 -07:00
be0dcc29dc
[XPU] remove q/k/v force contiguous for flash_attn ( #40356 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 01:19:05 +00:00
Sumanth R Hegde and GitHub
e3b65a5ba0
[feat] Add explicit /start_weight_update and /finish_weight_update APIs for weight transfer ( #39212 )
2026-05-08 18:03:33 -07:00
Harry Mellor and GitHub
30f519e947
Use pre-commit / pre-run-check to gate docs build too ( #42053 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-09 00:02:51 +00:00
Roy Wang and GitHub
60851b1d22
[Bugfix][KV Transfer] Reject NixlConnector + expandable_segments:True ( #41237 )
2026-05-08 16:47:33 -07:00
Michael Goin and GitHub
8bcd8a260c
[Bugfix] Fix FlashInfer CUTLASS MXFP4-MXFP8 MoE by restoring swizzled scale ( #42089 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-08 15:59:06 -07:00
John Calderon and GitHub
8a2fc80b84
[CUDA][CUTLASS] Enable cutlass scaled mm for non-compatible sizes ( #41868 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
2026-05-08 15:58:05 -07:00
6881c754e1
use HIP_VERSION variables to guard against duplicate atomicAdd definitions ( #41802 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-08 18:44:37 -04:00
Kevin H. Luu and GitHub
0c2e9d4892
[CI] Narrow misc.yaml source dependencies ( #42059 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 15:10:12 -07:00
Kevin H. Luu and GitHub
d2f22dfc9f
[CI] Narrow engine.yaml source dependencies ( #42055 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:55:33 -07:00
Kevin H. Luu and GitHub
f4dd5c116c
[CI] Narrow Platform Tests (CUDA) source dependencies ( #42054 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:54:06 -07:00
Kevin H. Luu and GitHub
f47ccc8b1c
[CI] Narrow pytorch.yaml compile job source dependencies ( #42057 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:43:17 -07:00
dbd86a67e3
[Bugfix][Gemma4] Fix infinite loop and array boundary issues in tool parser ( #41991 )
...
Signed-off-by: David Oy <david.oy@baseten.co >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-08 17:24:37 -04:00
2c6b59b807
[ROCm][Perf] Add Fused Shared Expert (FSE) support for Qwen3-Next ( #39280 )
...
Signed-off-by: nholmber <nholmber@users.noreply.github.com >
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Doug Lehr <douglehr@amd.com >
Co-authored-by: nholmber <nholmber@users.noreply.github.com >
Co-authored-by: Tres <tpopp@users.noreply.github.com >
Co-authored-by: Tres Popp <tres.popp@amd.com >
Co-authored-by: Doug Lehr <douglehr@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-05-08 15:38:00 -04:00
44e6b44a21
[CI][Elastic EP] Fix Elastic EP Scaling Test Failure ( #41792 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-08 15:17:44 -04:00
Hiroaki Mikami and GitHub
90f145aaf7
[Models][Gemma3/Gemma4] Support hidden_act variants in gated MLP ( #40588 )
...
Signed-off-by: Hiroaki Mikami <hiroaki8270+github@gmail.com >
2026-05-08 11:29:11 -07:00
Ethan Feng and GitHub
4140faa4a5
[Docs] Fix OpenAI batch model argument examples ( #42066 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-08 14:02:46 +00:00
liuzhenwei and GitHub
f2bbd575e2
[CI][XPU] Skip fork-dependent logits processor test ( #42013 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-08 06:10:19 -07:00
haosdent and GitHub
52458b60a8
[CI][Examples][RLHF] Disable async scheduling in rlhf_async_new_apis ( #42042 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 04:58:48 -07:00
Harry Mellor and GitHub
630820a59b
Make docs environment deterministic ( #41926 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-08 10:13:03 +00:00
Chaojun Zhang and GitHub
19df11f5d1
[CI][XPU]Ignore some lora tests from LoRA Intel CI pipeline ( #42010 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
2026-05-08 17:34:27 +08:00
haosdent and GitHub
36b2c79d4b
[CI][Bugfix] Drop duplicated examples/ prefix in tensorize_vllm_model command ( #42039 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 02:23:22 -07:00
haosdent and GitHub
160858cba4
[CI][Bugfix] Surface subprocess output in spawn_new_process_for_each_test ( #41943 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 10:39:37 +02:00
Simon Danielsson and GitHub
f9b9bf3bbb
[CI][ROCm] Ship RIXL with vllm/vllm-openai-rocm ( #41634 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-08 07:05:17 +00:00
445d747434
[Bugifx] Missing Renderer for fastokens mode ( #41984 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-07 23:45:14 -07:00
wang.yuqi and GitHub
77b13b9602
[Docs] Reorganize examples docs. ( #41082 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-07 23:23:44 -07:00
ed582b6a4c
[Aiter][ROCm] gdn_linear_attn kernel fusion ( #40711 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: hellozhuo <zhuo.su@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-07 23:11:37 -07:00
David Zheng and GitHub
1acd67a795
[Bugfix] Fix XPU/ROCm compatibility in spawn_new_process_for_each_test ( #41895 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-08 00:47:22 -04:00
0b99971352
[Kernel][Helion] Optimize Helion config parsing latency ( #40850 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-05-07 20:27:34 -07:00
baf068d8be
enable persistent mla for sparse mla backend ( #41990 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
2026-05-07 20:10:50 -07:00
SamareshSingh and GitHub
01b0f3adab
fix: default TILELANG_CLEANUP_TEMP_FILES=1 to avoid shared /tmp conflicts ( #41486 )
...
Signed-off-by: Samaresh Kumar Singh <ssam3003@gmail.com >
2026-05-07 19:59:00 -07:00
Nick Hill and GitHub
989c176c0a
[Perf][3/n] Eliminate GPU<->CPU syncs in attention impls ( #41434 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-07 19:44:24 -07:00
cd58e30872
[Perf] Use numpy zero-copy path for embedding float response serialization ( #41681 )
...
Signed-off-by: Shrinav Loka <lokashrinav@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 19:42:21 -07:00
1d694e78c9
[Examples][last/6] Resettle examples. ( #41084 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-07 19:42:12 -07:00
haosdent and GitHub
57c2f724c1
[CI][Bugfix] Fix CI failures for "PyTorch Compilation Unit Tests" ( #41940 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-07 19:42:00 -07:00
5f6a02812a
[CI][Bugfix] Fix failure CI step "PyTorch Fullgraph Smoke Test" ( #41953 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-07 19:41:56 -07:00
50f2db2555
add: LFM2/2.5 Tool Parser ( #39243 )
...
Signed-off-by: Jonathan Buchanan <jonathan.buchanan@liquid.ai >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-08 09:58:17 +08:00
09a7cc5ba9
[KV Connector] Opt DecodeBenchConnector into SupportsHMA ( #41770 )
...
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-07 16:10:00 -07:00
Nick Hill and GitHub
10ebb40d62
[Core] Avoid using extra thread in UniProcExecutor ( #40891 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-07 15:33:00 -07:00
54f548e9e5
[Bugfix] Restore moe_forward output shape invariant on TRTLLM MXFP4 path ( #41646 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-07 15:26:06 -07:00
Kyle Sayers and GitHub
c1819ca283
[Compressed Tensors] Allow configs with non-explicit ignores ( #41965 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-05-07 14:03:45 -07:00
969fbfb4a9
Laguna xs dflash support ( #41880 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-07 13:31:16 -07:00
akii96 and GitHub
3af561ec0a
[ROCm] Fix AITER AR+RMSNorm no-residual fusion ( #41972 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-07 13:14:58 -07:00
akii96 and GitHub
c936548ce6
[ROCm][DeepSeek] Enable V3.2 TP4 AITER MLA ( #41835 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-07 15:10:57 -05:00
TomerBN-Nvidia and GitHub
8189a15914
[Core] Replace routing replay with device cache and async D2H pipeline ( #39917 )
...
Signed-off-by: Tomer Barnatan <tbarnatan@nvidia.com >
2026-05-07 11:24:56 -07:00
Flora Feng and GitHub
8eb401134e
[Refactor] Consolidate required/named tool_choice streaming into DelegatingParser ( #41876 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-07 09:50:59 -07:00
Nicolò Lucchesi and GitHub
9d6500b89d
[Misc] Delay EPLB Nixl import until needed ( #41805 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-07 09:43:07 -07:00
zhrrr and GitHub
7a08b34fbf
[Model Runner V2] support qwen35 / mamba hybrid model ( #35520 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-05-07 09:31:05 -07:00
noobHappylife and GitHub
06a60d3dd0
Fix spec decode benchmark metrics ( #41916 )
...
Signed-off-by: noobhappylife <aratar1991@hotmail.com >
2026-05-07 09:23:21 -07:00
2a16ece2d3
tokenizer: Add fastokens support ( #41741 )
...
Signed-off-by: AlonKejzman <alonkeizman@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-07 22:49:42 +08:00
Andreas Karatzas and GitHub
003159d98b
[ROCm][CI] Avoid duplicate ROCm AITER norm-quant patterns ( #41534 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-07 06:33:30 -07:00
2a84da3b17
[XPU] Implement out-of-place all-reduce functionality ( #41808 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 05:58:05 -07:00
Chaojun Zhang and GitHub
805e9f7b77
[XPU] Fix lora bugs & enable UTs under tests/lora ( #38206 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
2026-05-07 05:58:00 -07:00
s-yanev and GitHub
75f0d516c4
[Bugfix] Fix GLM4-MoE weight loading for NVFP4 quantized checkpoints ( #41755 )
...
Signed-off-by: Stoyan Yanev <stoyan.yanev@cleverpine.com >
2026-05-07 05:55:52 -07:00
f650ace6de
[MM][Gemma4] Use video profiling hints in encoder budget ( #41837 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
2026-05-07 05:46:04 -07:00
Li, Jiang and GitHub
b3945cc316
[CPU] Bump up to the latest CPU kernels ( #41924 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-07 05:45:59 -07:00
ffee741626
[Model] Use AutoWeightsLoader for AXK1 ( #41901 )
...
Signed-off-by: liwenyi <lwy.lwy@163.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-07 05:40:29 -07:00
Jee Jee Li and GitHub
9c0812ffd0
[Bugfix] Fix FusedMoEWithLoRA has no attribute runner ( #41889 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-07 04:53:14 -07:00
Fadi Arafeh and GitHub
b20731d0ae
[CI][Arm] skip e2e model tests if HF_TOKEN is not set ( #41919 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-05-07 11:31:50 +00:00
d4b0048404
Eliminate redundant MoE buffer copies in AITER fused experts (without dependency on AITER changes) ( #41713 )
...
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 03:46:40 -07:00
6e6d182d18
[Bugfix] Fix OOM in tensorizer LoRA deserialization ( #41845 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-07 02:17:47 -07:00
tej and GitHub
8a4888be21
[ROCm] Profiler api support for ROCm MORI toy proxy server in PD Disaggregation ( #40264 )
...
Signed-off-by: Tej Kiran <kiran.tej@amd.com >
2026-05-07 16:58:38 +08:00
Yuwen Zhou and GitHub
713b28bd0b
[CPU] Add FP8 W8A16 MoE support ( #41314 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-05-06 23:17:07 -07:00
51f22dcfd0
[Feat][CPU] Enable Gated DeltaNet Attention (Qwen 3.5 / 3.6) ( #41025 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-07 12:57:09 +08:00
20cac26b19
[ROCm] Enable SimpleCPUOffloadConnector on ROCm ( #40549 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-06 20:52:02 -07:00
Russell Bryant and GitHub
5a0a8fc1ea
[Docs] add cache directory security guidance ( #38920 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
2026-05-06 16:54:29 -07:00
Micah Williamson and GitHub
7a576e2c72
[ROCm][CI] Remove TORCH_NCCL_BLOCKING_WAIT=1 After Bugfix In ROCm 7.2 ( #41840 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-06 16:37:11 -07:00
Yongye Zhu and GitHub
80d5e7d103
[Bugfix] Fix condition to clear persistent topk so that it can be captured regardless ( #41665 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 16:17:48 -07:00
95582868ef
[Bugfix] DeepSeekV32/v4: respect string='true|false' attribute andunwrap arguments/input wrapper ( #41801 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
2026-05-06 21:48:01 +00:00
50acdc5b5c
Fix Qwen3 streaming content routing ( #40820 )
...
Signed-off-by: xy3 <120182408@qq.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-06 17:22:01 -04:00
JackyLiu and GitHub
deb737e323
[Doc] Add ModernBertForSequenceClassification to scoring.md cross-en… ( #41832 )
...
Signed-off-by: JLiu4Coding <lzwgre@126.com >
2026-05-06 14:17:56 -07:00
Flora Feng and GitHub
f3f8efa73a
[CI] Enable gemma4 parser test on CI ( #41857 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-06 20:25:34 +00:00
Johnny Yang and GitHub
ca3e62d336
Upgrade tpu-inference to v0.19.0 ( #41844 )
...
Signed-off-by: Johnny Yang <johnnyyang@google.com >
2026-05-06 11:41:37 -07:00
Benjamin Chislett and GitHub
38e16678ba
[Bugfix] Align block table for TRTLLM MLA edge-case ( #39324 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-06 11:17:02 -07:00
27702f6d08
[Bugfix] Fix token loss in PP mode which causes degraded accuracy ( #41133 )
...
Signed-off-by: Jing Wang <jingwang96@qq.com >
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-06 14:07:32 -04:00
Divakar Verma and GitHub
22a3cbe152
[ROCm] aiter_unified_attn fp8 q scale refactor ( #38296 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-06 16:11:36 +00:00
Viktor Pus and GitHub
d5b31c954d
[Bugfix] Account for truncate_prompt_tokens when computing max_tokens ( #41800 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
2026-05-06 16:10:17 +00:00
David Zheng and GitHub
ee38750a75
[Bugfix] Fix spawn_new_process_for_each_test silently swallowing test failures ( #41423 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-06 11:17:15 -04:00
27e0057aed
[Spec Decode] Add Gemma4 MTP speculative decoding support ( #41745 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-06 22:39:29 +08:00
Ronen Schaffer and GitHub
f39bcf1e30
[KV Offload] Return None from lookup() for in-flight blocks ( #41795 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-06 17:31:21 +03:00
6467213a9f
fix(openai): tolerate empty content in forced tool choice ( #40148 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-06 07:16:03 -07:00
df8e63f4ed
nixl refactor: new transfer design ( #40731 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-06 06:16:25 -07:00
242afc6bf4
[MM][Gemma4] Respect max_soft_tokens in encoder budget ( #41799 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: gemini-code-assist <gemini-code-assist@google.com >
2026-05-06 05:54:42 -07:00
lyd1992 and GitHub
5d0fd87038
[CPU][RISC-V] Auto-bind OMP threads and harden nobind path ( #40569 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-05-06 11:38:08 +00:00
Harry Mellor and GitHub
d8deb5b7ad
Fix some legacy checkpoints with deprecated rope_type values ( #41734 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-06 11:13:12 +00:00
2e777d21a8
[Bugfix][Rocm]Aiter MoE re-uses existing tensor addresses after weight update. ( #40390 )
...
Signed-off-by: Yuankai Chen <yuankach@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-06 10:32:26 +00:00
Nicolò Lucchesi and GitHub
e43a791284
[Bugfix][CI] Fix Disaggregated test area path ( #41794 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-06 17:41:24 +08:00
66d1cc0c77
fix(rocm): remove workaround causing invalid argument on Qwen3.5 with TP=2 ( #40686 )
...
Co-authored-by: Test User <test@example.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-06 01:38:32 -07:00
1c58876618
[XPU] Disable CUDA graph memory estimate on XPU platform ( #41344 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-06 16:38:18 +08:00
51c1ee9b7c
[Examples] Resettle Disaggregated examples. ( #40759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-06 01:20:38 -07:00
Lucas Kabela and GitHub
213f10bfdd
[Bugfix] Fix codegen for unqualified names ( #40726 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-06 01:11:37 -07:00
e87e09a50a
[Feat] dnnl build for AVX2 W8A8 Int8 ( #41318 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-06 15:28:02 +08:00
Yuwen Zhou and GitHub
809b98e5b7
[CPU] Add FP8 W8A16 linear support ( #41186 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-05-06 07:05:27 +00:00
wi-adam and GitHub
b53c507bc9
[Bugfix] Skip PP sampled-token receive on last rank during async scheduling ( #40749 )
...
Signed-off-by: Adam Winstanley <adam@winstanley.industries >
2026-05-06 05:31:14 +00:00
2d7d6cf765
[Spec Decode] Allow multimodal models with a warning ( #41752 )
...
Signed-off-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Li Zhang <lzhanga@amazon.com >
2026-05-05 22:16:44 -07:00
Andreas Karatzas and GitHub
91740ca5ea
[ROCm][CI] Refine gating tests ( #37243 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-05 22:05:20 -07:00
e47c98ef7a
[Fix] Add missing stubs from cpu fp8 attention changes ( #41387 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-06 12:16:27 +08:00
aee190ac37
[Build] Fall back to system libgomp when torch has no vendored copy ( #40575 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-06 11:42:03 +08:00
16e336491e
[Mistral Tokenizer] allow more leniency in apply_chat_template ( #41658 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-05 19:56:15 -07:00