Michael Goin and GitHub
5774aad9c5
[Perf][gpt-oss] Downgrade triton_kernels to v3.5.1 ( #43135 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-20 14:13:12 -07:00
363fc84407
Integrate flashinfer b12x MoE and FP4 GEMM kernels for SM120/121 ( #40082 )
...
Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-05-20 17:21:11 +00:00
f2d5e3d3ae
[CI] Lower granite-4.0-h-tiny gsm8k threshold for Hybrid SSM NixlConnector PD accuracy tests (4 GPUs) ( #43186 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-20 17:00:24 +00:00
2d6b3489b9
[R3] Add routed experts to openai entrypoint ( #38939 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-20 09:07:59 -07:00
Flora Feng and GitHub
a10d69116c
[Bugfix] Use shared coerce_to_schema_type in DeepSeekV32 tool parser ( #43019 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 10:21:00 -04:00
ded871201a
[Bug][Structured Outputs] Fix bug that leads to unconstrained generations with structural tags ( #42452 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-20 07:08:58 -07:00
Dipika Sikka and GitHub
df84fb07a6
Remove additional dead code as a follow-up to #42889 ( #43144 )
...
Signed-off-by: Dipika Sikka <dipikasikka1@gmail.com >
2026-05-20 10:01:45 -04:00
Kebe and GitHub
19cf334207
[Feature] Support manually enabling the cumem allocator ( #33648 )
...
Signed-off-by: Kebe <mail@kebe7jun.com >
2026-05-20 08:58:30 -04:00
Ronen Schaffer and GitHub
4f940896a3
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers ( #43076 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-20 03:32:08 +00:00
Michael Goin and GitHub
cd0ff26e7a
[CI] Add DSV4-Flash to gsm8k moe-refactor/config-b200.txt ( #42111 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-19 20:21:01 -07:00
fadf5d332c
add enqueue all option to throughput benchmark ( #42975 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Signed-off-by: pmaybank <113125070+pmaybank@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 20:16:02 -07:00
Benjamin Chislett and GitHub
c628a93a64
[Perf][Bugfix] Update dflash aux layer indexing ( #40727 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-19 20:15:57 -07:00
Terrence Zhao and GitHub
5774aaed0c
[Cohere] Enable Cohere MoE ( #43143 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-05-19 19:32:06 -07:00
Aaron Hao and GitHub
73dd2f33b7
[bug] fix WeightTransferConfig.backend to allow for all strings ( #43121 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
2026-05-19 21:01:29 -04:00
117afeea46
Fix error in Dynamic NTK scaling ( #41277 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
Signed-off-by: Max de Bayser <maxdebayser@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-19 17:27:54 -04:00
tomeras91 and GitHub
f54721bcc3
[Bugfix][MoE] FlashInfer one-sided: workspace union across heterogeneous layers ( #42976 )
...
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com >
2026-05-19 14:43:04 -04:00
Dom Brown and GitHub
d247a931cc
[feat] Add FP8 per-tensor Q scale support to Triton attention backend ( #42080 )
...
Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com >
2026-05-19 09:02:05 -07:00
Flora Feng and GitHub
42b4f1fdf7
[Refactor] Extract extract_types_from_schema utility from Minimax M2 tool parser ( #43025 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-19 11:21:12 -04:00
Wang Yiwen and GitHub
1c6158083a
[Model] Openvla support ( #42654 )
...
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com >
2026-05-19 08:17:42 -07:00
Nick Hill and GitHub
b82e908b4c
[Perf][4/n] Eliminate various GPU<->CPU syncs ( #42347 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-19 10:35:54 -04:00
Sage and GitHub
a78b842d0e
[Bugfix] Fix top logprobs token placeholders in /inference/v1/generate ( #42887 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-05-19 10:21:49 +00:00
129019f334
[CI] Add MTP + PD disagg test for Qwen3.5 ( #42677 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-19 11:44:33 +02:00
Woosuk Kwon and GitHub
07beaed842
[Model Refactoring] Rename deepseek_v4.py to model.py [4/N] ( #43077 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 01:12:46 -07:00
Yifan Qiao and GitHub
056bc2e166
[KVConnector][DSV4] HMA support for Mooncake store connector ( #42828 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-19 01:07:46 -07:00
Woosuk Kwon and GitHub
b14be81c1f
[Model Refactoring] Move deepseek_v4_ops to models/deepseek_v4 [3/N] ( #43073 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 00:52:54 -07:00
fab07e4d0f
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B ( #42289 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
Co-authored-by: gemini-code-assist <noreply@google.com >
2026-05-18 21:22:33 -07:00
3ca8db2ef8
add cutedsl dsv4 indexer fp8 kernel ( #42899 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-18 21:17:56 -07:00
fba010dd74
[Bugfix][MRV2] Fix KVCache tensor explicit kernel_block_size dim ( #42766 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-18 20:25:41 -07:00
Mohammad Miadh Angkad and GitHub
da03e549b3
[UX] Add a persistent cache for FlashInfer autotuning ( #42537 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-18 20:25:37 -07:00
Woosuk Kwon and GitHub
287471b994
[Model Refactoring] Migrate DeepSeek V4 to vllm/models/ [1/N] ( #43004 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-18 19:50:02 -07:00
Wentao Ye and GitHub
37ece593c1
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement ( #42774 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 16:38:12 -07:00
Flora Feng and GitHub
57fef4e0bf
[Refactor] Extract shared coerce_to_schema_type utility from Minimax M2 tool parser ( #43006 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-18 17:55:39 -04:00
Ronen Schaffer and GitHub
84747489de
Tier offload followup ( #42529 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-18 19:41:58 +00:00
Bowen Bao and GitHub
a2c8fc6657
[ROCm][Quantization][3/N] Refactor quark_moe w4a4 w/ oracle ( #41436 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-18 13:46:13 -04:00
Netanel Haber and GitHub
47829b1159
[Bugfix] mamba: run single-token extends as decodes ( #42430 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-18 15:26:00 +00:00
Blanc Swan and GitHub
4a39b4f553
[Model] Add Apertus Tool Parser ( #41154 )
...
Signed-off-by: Blanc <swan.blanc@infomaniak.com >
2026-05-18 11:20:04 -04:00
78e7a7b9b0
Refactor AWQ Marlin MoE onto modular WNA16 oracle ( #42483 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Signed-off-by: Siddharth Bedekar <104613085+bedeks@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 08:02:43 -07:00
e5417657e5
[KV Connector][Offloading] Flush all pending jobs on last step ( #42611 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 12:59:42 +00:00
Nicolò Lucchesi and GitHub
69c91d010a
[MRv2] Default to MRv1 when a connector is present ( #42955 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-18 20:34:16 +08:00
roikoren755 and GitHub
737bfa3a43
[Bugfix][Hybrid][NemotronH] Fix mamba_cache_mode=all + speculative decoding crash ( #41233 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-05-18 14:54:00 +03:00
Yuwen Zhou and GitHub
88a860d754
[CPU] Add MXFP4 W4A16 MoE support ( #41922 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-05-18 03:04:45 -07:00
7d5b033782
[LoRA] Support 2D and 3D MoE LoRA adapter at the same time ( #42242 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-18 15:22:26 +08:00
e3aeee5ff8
[Bugfix] moe lora align kernel grid ( #40131 )
...
Signed-off-by: TheDuyIT <nduy250299@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: dtnguyen <dtnguyen@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-18 00:17:53 -07:00
Andreas Karatzas and GitHub
b50646e5ef
[ROCm][CI] Stabilize ROCm pooling and multimodal CI ( #42909 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-18 03:57:59 +00:00
Soyaazz and GitHub
990f49bdcb
[MM][CG] Enable encoder Cudagraph for Step3VL ( #42224 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
2026-05-17 20:19:13 -07:00
107210442d
[CI] Add NIXL EP import canary ( #42567 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-05-17 19:11:46 -07:00
Taneem Ibrahim and GitHub
1c8e9c0399
Refactor: Pass num_labels explicitly to PoolerClassify instead of reading from global config ( #42851 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-17 14:40:21 +00:00
0867497368
[CI/Build] Bump flashinfer to v0.6.11.post2 ( #41711 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-16 14:55:12 -07:00
36e74c9ea4
[KV Connector] Support disk offloading in MooncakeStoreConnector ( #42689 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-16 13:34:15 -07:00
Taneem Ibrahim and GitHub
787bc0d031
Add unit tests for pooler activation functions ( #42824 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-16 14:58:16 -04:00