6773 Commits
Author SHA1 Message Date
d6247d7173 [Spec Decode][Perf] Replicate DSpark Markov head across TP ranks (#49731)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-07-29 11:43:51 -04:00
a0c092ee72 [BugFix] Fix num_output_placeholders preemption underflow (#48245)
Signed-off-by: Chris Eastwood <chris.eastwood@pwn4g3.dev>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Chris Eastwood <chris.eastwood@pwn4g3.dev>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
2026-07-29 06:41:54 -07:00
Taneem IbrahimandGitHub 43eaefba5a [ModelRunner V2] Enable sequence pooling for embedding and classification models (#48791)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
2026-07-29 06:30:45 -07:00
stefankoncarevicandGitHub 625871b52c [CI][Test] Fix pooling truncation test after VLLMError hierarchy change (#50241)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
2026-07-29 12:15:56 +00:00
f51193b9ae [Kernel][Mamba] Fused-kernel support for align-mode DS-conv state migration with num_accepted_tokens > 1 (#49291)
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
2026-07-29 18:54:52 +08:00
Guan-Ming ChiuandGitHub aeaa50a71c [Bugfix][Multimodal] Include media IO config in MM cache hash (#49975)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
2026-07-29 10:48:40 +00:00
542a8fad6d [KV Offload] Move CPUOffloadingSpec onto SharedOffloadRegion (#50094)
Signed-off-by: Change72 <changg@nvidia.com>
Co-authored-by: Cursor Agent <cursor-agent@cursor.com>
2026-07-29 13:05:50 +03:00
fxmarty-amdandGitHub 5b14019576 [CI] Fix MXFP8 MOE backend selection tests on gfx942 (#50222)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
2026-07-29 17:17:11 +08:00
df2735ea2e [Misc][Minimax-M3]add default video_processor (#50092)
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
2026-07-29 01:24:51 -07:00
Jared WenGitHubmergify[bot] <37929162+mergify[bot]@users.noreply.github.com>Cyrus Leung
32e657e689 [BugFix] eagle draft max position embeddings (#49343)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
2026-07-29 01:24:47 -07:00
ad5d29db70 [Model] Support Qwen3.5 text-only dense and MoE models (#50210)
Signed-off-by: Perkz Zheng <PerkzZheng@users.noreply.github.com>
Co-authored-by: Perkz Zheng <PerkzZheng@users.noreply.github.com>
2026-07-29 08:21:57 +00:00
Bugen ZhaoandGitHub f5a7cce9b6 [Model] Add Kimi K3 support: Python frontend [2/2] (#50093)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
2026-07-29 01:06:01 -07:00
Seiji EicherandGitHub 6370e53f24 [Frontend] Reuse prefill token ids on the decode chat path for disaggregated serving (#48145)
Signed-off-by: Seiji Eicher <seiji@anyscale.com>
2026-07-29 10:02:57 +02:00
Tianmu LiGitHubCodexLi, Jiang <jiang1.li@intel.com>
65a1a16594 [CPU] Fix FP8 attention scratchpad sizing (#50194)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
2026-07-29 15:34:37 +08:00
zofiaGitHubcopilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
7de49bab7e [XPU][UT][CI] add xpu config to run gpt-oss accuracy in ut and ci (#48703)
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-29 14:16:16 +08:00
+13 7c6729b769 [Model] Add Kimi K3 support: model files and kernels [1/N] (#50089)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Ziming Huang <zelda.huanghuang@gmail.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
Co-authored-by: Isotr0py <mozf@inferact.ai>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Summer Yang <girasoleyang@gmail.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
Co-authored-by: Bowen Wang <abmfy@icloud.com>
Co-authored-by: gnovack <novackgm@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: xiaozhoupy <peiyuanzhou1994@gmail.com>
Co-authored-by: Roy Wang <yasong.wang@inferact.ai>
Co-authored-by: Jeff (Junze) Ma <93145857+majunze2001@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-07-29 14:10:58 +08:00
6f00a1ae3b fused_moe: add VLLM_TRITON_USE_TD tensor-descriptor path (#42436)
Signed-off-by: Artur Fierka <artur.fierka@intel.com>
Signed-off-by: Lena Onyshchenko <162571002+oonyshch@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Lena Onyshchenko <162571002+oonyshch@users.noreply.github.com>
2026-07-29 13:15:10 +08:00
6f91edf96d [Test] dynamic_shapes_compilation (#49974)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
2026-07-29 04:54:57 +00:00
Philip PesicandGitHub dc1be79031 Add CachePolicyFactory for pluggable/external eviction policies (#49114)
Signed-off-by: Philip Pesic <philippesic06@gmail.com>
2026-07-29 07:34:55 +03:00
Andreas KaratzasandGitHub 0bb548b60e [CI][ROCm] Stabilize Qwen2-VL LoRA test (#50161)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
2026-07-29 04:27:45 +00:00
Zach ZhuandGitHub 58f9659397 [Frontend][Core] Standardize request error handling with VLLMError hierarchy (#49665)
Signed-off-by: Zach Zhu <zzqshu@126.com>
2026-07-29 04:20:40 +00:00
f37f03db4a [KV Connector] Support NIXL P/D for hybrid MLA+SSM models (#49762)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Jared Wen <w13431838023@gmail.com>
2026-07-29 11:50:30 +08:00
32a423ac0a Integrate CuTeDSL MoE for ReLU2 NVFP4 (#49580)
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
2026-07-28 19:46:38 -07:00
30c2718eaa [CompressedTensors] FP4 Qutlass Integration (#43229)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
Signed-off-by: Brian Dellabetta <bdellabe@redhat.com>
Signed-off-by: Brian Dellabetta <brian-dellabetta@users.noreply.github.com>
Co-authored-by: Brian Dellabetta <bdellabe@redhat.com>
Co-authored-by: Brian Dellabetta <brian-dellabetta@users.noreply.github.com>
Co-authored-by: Dipika Sikka <dipikasikka1@gmail.com>
2026-07-28 20:34:20 -06:00
Andreas KaratzasandGitHub 7398a30d79 [ROCm][CI] Stabilize ngram and suffix correctness test (#50190)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
2026-07-28 20:29:00 -06:00
17a74b745b [Model] Add Inkling compressed-tensors dynamic FP8 support (#48876)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-07-28 19:27:01 -07:00
6fbbcf2151 [BugFix] Stop dummy runs from writing mamba state through stale block-table rows (#49757)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Jeff Ma <jeffjma@umich.edu>
Co-authored-by: Jeff Ma <jeffjma@umich.edu>
2026-07-29 01:09:41 +00:00
Kevin GlynnandGitHub 56f31af62a [Bugfix] Fix /wake_up crash on hybrid models (Mamba/DeltaNet) (#41602)
Signed-off-by: Kevin Glynn <kevglynn@gmail.com>
2026-07-29 00:17:35 +00:00
Divakar VermaandGitHub fe65aa6a97 [CI][NIXL] Fix flaky DP+EP test port conflict (#50171)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
2026-07-29 00:12:50 +00:00
Andreas KaratzasandGitHub 176256b962 [ROCm][CI] Stabilize ROCm audio streaming test (#50163)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
2026-07-29 07:58:10 +08:00
fxmarty-amdandGitHub 5369f7b7b8 [MXFP8][ROCm] Fix MXFP8 MoE backend selection (#49747)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
2026-07-29 07:57:20 +08:00
liuzhenweiandGitHub e7f6a39db8 [Test] Make EPD correctness tests configurable for XPU (#50110)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
2026-07-29 07:47:48 +08:00
a07fac758f [Perf] Zero-copy torch.Tensor pickling in shm_broadcast MessageQueue (#48442)
Signed-off-by: Ruinan Ma <r7ma3088@gmail.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
2026-07-28 16:28:28 -07:00
Julien DebacheandGitHub bb3b61f2fd perf: dispatch non-grouped bias-less topk routing methods to fused path (#49618)
Signed-off-by: jdebache <jdebache@nvidia.com>
2026-07-28 14:57:22 -07:00
Shanshan ShenandGitHub 6c7e679f04 [ROCm][Bugfix] Sanitize AITER paged-MQA logits before sparse top-k for DeepSeek-V4 (#49714)
Signed-off-by: shen-shanshan <467638484@qq.com>
2026-07-28 11:08:50 -07:00
labAxiaomingandGitHub 1db989bbf1 [Bugfix][Multimodal] Fix video temporal padding estimates (#49030)
Signed-off-by: xiaoming <1259730330@qq.com>
2026-07-29 01:57:31 +08:00
Andreas KaratzasandGitHub 05a0814863 [ROCm] Fix and optimize GPT-J-style MRoPE (#49906)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
2026-07-28 10:50:10 -06:00
Bugen ZhaoandGitHub 01661cc57f [Rust][Benchmark] Make vllm bench serve Rust delegation opt-in (#50081)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
2026-07-28 16:21:12 +00:00
ba702e978e [Attention] Skip sparse indexer scoring for dense short prefills (#48407)
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-07-28 16:17:22 +00:00
6453fc0b8c [Bugfix] Don't reuse engine core payload buffer while zmq is sending it (#50053)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-07-28 16:09:28 +00:00
30217b0e80 [Bugfix][KV Offload][P2P] Scope serve state to fetch rounds (#49877)
Signed-off-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com>
2026-07-28 18:01:55 +03:00
Nick HillandGitHub 0d0504b54c [Core] Warm up runner-owned Triton kernels before the first request (#49903) 2026-07-28 07:58:02 -07:00
1e81853afc [Bugfix][KV Offload] Keep Mamba block span unscaled under DCP (#49964)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-07-28 17:41:55 +03:00
b6cbba8bc8 [Bugfix][Kernel] Fix batch invariance in RMSNorm kernels by pinning block size (#48391)
Signed-off-by: oops-oom <73481342@qq.com>
Signed-off-by: oops-oom <liubin8905@vip.qq.com>
Co-authored-by: oops-oom <73481342@qq.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
2026-07-28 22:24:02 +08:00
94100b5915 [CI] Wire untethered test files into CI jobs (#49340)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-28 14:02:31 +00:00
601fa9a74e [KV Connector] Support NIXL heterogeneous P/D block sizes for hybrid models (#49612)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 06:45:02 -07:00
948107acf7 [Bugfix] Enhance extra_config handling for layer name suffix matching (#48589)
Signed-off-by: Xin He <xin3.he@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-28 20:37:55 +08:00
Itay AlroyandGitHub 35efdf6b34 [Elastic EP] Async preparation (#47288)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
2026-07-28 05:19:13 -07:00
Jiangyun ZhuGitHubOpenAI Codexmergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
25ace8fe5d [CI] Increase Qwen3.5 MTP GSM8K generation length (#49881)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-28 18:04:36 +08:00
Liangliang MaGitHubmergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
88402a41c4 [Test] Skip ROCm AITER MLA prefill tests on non-ROCm platforms (#49945)
Signed-off-by: Liangliang-Ma <liangliang.ma@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-28 16:26:49 +08:00