khluu and Claude Opus 4.6
f478d42cdb
[CI] Reorganize release pipeline: separate nightly vs release sections
...
Reorder the pipeline into clear sections for readability:
1. Build Python Wheels (always runs)
2. ROCm Wheel Pipeline (always runs)
3. Nightly Docker Images (NIGHTLY=1 only) - CUDA/Ubuntu builds,
multi-arch manifests, DockerHub publish, ROCm image + publish
4. Release (manual) - version input, PyPI upload, CPU image builds,
ROCm root index
Key changes:
- Extract CPU image builds (manual/blocked) from nightly-gated group
into their own "Build release CPU Docker images" group so they remain
available for actual releases without NIGHTLY=1
- Move ROCm wheel jobs (1-4) up next to CUDA wheel builds
- Remove redundant per-step NIGHTLY gates inside already-gated groups
- Rename groups: "Build release Docker images" -> "Build nightly Docker
images", "Publish release images" -> "Publish nightly images"
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-15 01:08:39 -07:00
khluu and Claude Opus 4.6
1a811d5747
[CI] Only build release Docker images when NIGHTLY=1
...
Gate the "Build release Docker images" group, "Publish release images"
group, and ROCm release image build behind NIGHTLY=1 to avoid expensive
image builds on every commit in the release pipeline.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-15 01:01:03 -07:00
zhanqiuhu and GitHub
799973af4e
[CI][NIXL] Fix PD CI breakage: pin nixl-cu{12,13} versions ( #39851 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-14 23:50:23 -07:00
bcc2306cef
[Bugfix] Respect VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY in prefetch offloader ( #37699 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-14 20:43:29 -07:00
wliao2 and GitHub
3abf858443
[Test] Refactor hard coded device string in test files under compile/quantization/models/model_executor folders ( #38901 )
...
Signed-off-by: Liao, Wei <wei.liao@intel.com >
2026-04-15 11:02:35 +08:00
f4b42df048
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity ( #38479 )
...
Signed-off-by: vibhavagarwal5 <vibhavagarwal5@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-14 19:57:13 -07:00
Giancarlo Delfin and GitHub
3bfe55a037
[Model Runner V2] Disable piecewise cudagraph mode fallback for eagle draft decodes ( #39773 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-14 17:47:57 -07:00
Andrey Talman and GitHub
b569620f72
[CI] Add PyTorch nightly build and test pipeline ( #37226 )
...
Signed-off-by: atalman <atalman@fb.com >
2026-04-14 17:13:24 -07:00
65b9808960
[Bugfix] Disable FlashInfer CUTLASS MoE on SM121 (DGX Spark) ( #39825 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-14 16:03:57 -07:00
Francesco Fusco and GitHub
507df79a29
[Hybrid] Simplify accepted token counting in spec decode for hybrid models ( #38372 )
2026-04-14 15:19:09 -07:00
1696c864b9
[Bugfix][Mooncake] Fix thread-local CUDA context for NVLink transfers in _send_blocks ( #39548 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-04-14 14:13:58 -07:00
Wentao Ye and GitHub
2ad1029233
[Bug] Fix batch invariance nvfp4 support ( #39820 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-14 17:08:17 -04:00
maobaolong and GitHub
b2f749dc97
fix(lmcache): correct store for cached requests while enable prefix cache ( #39719 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-04-14 20:51:27 +00:00
70ed01550c
[Reasoning][Frontend] Add model config to adjust_request in reasoning parser ( #37848 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com >
Signed-off-by: Aaron Pham <contact@aarnphm.xyz >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Aaron Pham <contact@aarnphm.xyz >
2026-04-14 20:29:06 +00:00
bnellnm and GitHub
19ec9a0a62
[MoE Refactor] Refactor ZeroExpertFusedMoE into new framework ( #35549 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-14 16:11:20 -04:00
1a9353bb02
[MoE] Move GPT OSS Triton kernel experts into fused_moe/experts/ ( #39007 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-14 19:27:39 +00:00
roikoren755 and GitHub
ecf5ff7ce3
[Mamba] Flashinfer selective_state_update ( #36162 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-14 15:10:58 -04:00
zhanqiuhu and GitHub
30679319e8
[CI][KVConnector][Metrics] Update multi KV connector edge case according to prefill stats changes ( #39808 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-14 18:59:15 +00:00
240f2636ca
[Kernel] Support TRTLLM GEN NVFP4 MoE for non-512-aligned hidden dims via weight padding ( #39510 )
...
Signed-off-by: root <root@lyris0017.lyris.clusters.nvidia.com >
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Co-authored-by: root <root@lyris0017.lyris.clusters.nvidia.com >
2026-04-14 11:49:56 -07:00
dc8df110bc
add warning when FP8 KV cache misses prefill query quantization ( #39752 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-lyris02.lyris.clusters.nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 14:43:05 -04:00
be0c855ebd
[KV Offload] Unified memory layout for offloading workers ( #37206 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-14 21:33:33 +03:00
Andrew Barnes and GitHub
e64b39ea71
[ROCm] Align AiterFlashAttentionImpl attn_type check with backend ( #39119 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-14 10:36:26 -07:00
Alessandro Sangiorgi and GitHub
2faad08362
[compile] Nest inductor cache under AOT compile dir ( #39718 )
...
Signed-off-by: Alessandro Sangiorgi <asangior@redhat.com >
2026-04-14 17:17:54 +00:00
Rohan Potdar and GitHub
23f3760217
[Bugfix][ROCm]: Allow gpt_oss_mxfp4 quantization method on rocm ( #39754 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-14 17:10:04 +00:00
Mark McLoughlin and GitHub
906a8c15d0
[Core][Metrics] Remove unused SchedulerStats.encoder_cache_usage ( #39693 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-14 12:53:57 -04:00
Micah Williamson and GitHub
4f4f8eaa78
[ROCm][CI] Fix condition for test_per_token_group_quant_fp8_packed ( #39730 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-14 16:14:31 +00:00
Netanel Haber and GitHub
b6890a120a
Bugfix: use_existing_torch.py: Glob recursive subdirs in requirements ( fixes #39024 ) ( #39793 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-14 23:11:46 +08:00
Lucas Kabela and GitHub
c08f3b2a62
Measure encoder compile time seperate from llm backbone ( #39240 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-14 10:52:49 -04:00
Hexiang Wang and GitHub
f02b3269e7
[PluggableLayer][3/N] Apply PluggableLayer to moe-related layers. ( #33556 )
...
Signed-off-by: whx-sjtu <2952154980@qq.com >
2026-04-14 09:55:00 -04:00
e1e318af01
[MoE Refactor] Remove MoE DP chunking ( #39107 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-14 09:48:05 -04:00
f7e62e3d66
[Bugfix] Fix mismatch between global and local attention heads in tensor-parallel mode for param2moe model ( #39707 )
...
Signed-off-by: bhargav-patel-29 <bhargav.patel@tihiitb.org >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 20:13:36 +08:00
18b1c77211
fix: handle ImportError in load_audio ( #39473 )
...
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-04-14 19:09:06 +08:00
Matthias Gehre and GitHub
1e4748c66a
[Bugfix] Fix vllm bench serve to count multimodal tokens in "total input tokens" ( #38654 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-14 11:00:40 +00:00
6f786f2c50
[Bugfix][Model] Fix Devstral Small 2 HF format weight loading ( #39293 )
...
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-14 10:11:18 +00:00
fxmarty-amd and GitHub
4eee77b877
[fix][MOE] Fix MOE experts intermediate_size dimension not being narrowed before weight loading ( #39688 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-04-14 09:35:28 +00:00
xiangdong and GitHub
a1993b96fd
[XPU][CI] Remove Arc in label-xpu ( #39776 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-14 02:27:38 -07:00
Julien Debache and GitHub
893b2affff
feat: add TxtSlicesDataset to allow sampling slices from txt file for benchmarking ( #30156 )
...
Signed-off-by: jdebache <jdebache@nvidia.com >
2026-04-14 09:20:03 +00:00
80118853f4
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3-VL video inference ( #38061 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-14 16:49:32 +08:00
c0ecaed950
[Frontend] Offload blocking preprocessing & postprocessing ops to thread pool for pooling entrypoints. ( #39763 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 08:29:25 +00:00
0008729abf
[Model] Use mm_features for Ernie-4.5 VL M-RoPE ( #39753 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-14 01:11:52 -07:00
d3af8c1831
[Core][Metrics][BugFix] Replace num_cached_tokens/num_external_computed_tokens with PrefillStats ( #37460 )
...
Related to `Counters can only be incremented by non-negative amounts`
error with the `vllm:prompt_tokens_by_source_total` metric.
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-04-14 09:00:45 +01:00
noobHappylife and GitHub
25b3242d8b
Fix Responses API streaming for multiple auto tool calls ( #39626 )
...
Signed-off-by: noobhappylife <aratar1991@hotmail.com >
2026-04-14 13:28:43 +08:00
b075604da1
[Bugfix] Fix Gemma4 tool parser converting bare null to string "null" ( #39679 )
...
Signed-off-by: KimuGenie <baby11686@naver.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-14 04:44:46 +00:00
Flora Feng and GitHub
db8a6d66bf
[Refactor][Parser] Migrate chat completion auto-tool/reasoning/plain streaming to parse_delta ( #39446 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-14 04:39:45 +00:00
Chauncey and GitHub
d2130a47bb
[Bugfix]: Fix MinimaxM2ToolParser missing tools parameter ( #39683 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-14 11:16:39 +08:00
c687bf226a
[LMCache][MP] optimize save when mla enabled ( #38810 )
...
Signed-off-by: idellzheng <idellzheng@tencent.com >
Co-authored-by: Yihua Cheng <yihua98@uchicago.edu >
2026-04-13 17:56:43 -07:00
Giancarlo Delfin and GitHub
ccf90ba784
[Model Runner V2] Add full cuda graph support for eagle prefill ( #37588 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-13 16:01:24 -07:00
Netanel Haber and GitHub
6adacfcb65
ParakeetExtractor performance and UX enhancements ( #39423 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-13 21:37:35 +00:00
Flora Feng and GitHub
14cb86c187
[Refactor][Parser] Simplify parse_delta ( #39728 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-13 21:02:13 +00:00
8213e8f880
Bug/test eagle dp v0 ( #38938 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-13 20:50:08 +00:00