![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1dfd64c1cc
|
[PluggableLayer][3/N] Apply PluggableLayer to llm_head and vocab embedding layer (#33465)
Signed-off-by: whx-sjtu <2952154980@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-10 16:12:59 +08:00 |
|
 
|
ad720aefe9
|
[Bugfix] Fix V1 dummy run writing NaN to KV cache null block (#39444)
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
|
2026-04-10 10:09:46 +02:00 |
|
 milesialandGitHub
|
270e8a4102
|
Nemotron Nano VL: Streamline pixel shuffle (#37580)
Signed-off-by: milesial <milesial@users.noreply.github.com>
|
2026-04-10 07:31:19 +00:00 |
|
 Richard ZouandGitHub
|
f44afef6d6
|
[compile] Allow strings in custom ops without regressing compilation times (#38123)
Signed-off-by: Richard Zou <zou3519@gmail.com>
|
2026-04-10 07:26:37 +00:00 |
|
 
|
447ce22212
|
[GGUF] Support non-standard quant types with prefix (e.g. UD-IQ1_S) (#39471)
Signed-off-by: Injae Ryou <injaeryou@gmail.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 07:22:53 +00:00 |
|
 Chendi.XueandGitHub
|
65e4e46f66
|
update CODEOWNERS file (#39439)
Signed-off-by: Chendi Xue <chendi.xue@intel.com>
|
2026-04-10 15:05:31 +08:00 |
|
 
|
49d20346e4
|
[Perf] Reduce H2D pageable memory copies (#38794)
Signed-off-by: jackcfwang <jackcfwang@tencent.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 15:03:26 +08:00 |
|
 Nick HillandGitHub
|
ef076c1b73
|
[Core] Change max_model_len in EngineCoreReadyResponse to be non-None (#39442)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-10 14:34:57 +08:00 |
|
 Yan MaandGitHub
|
ec68d53b2b
|
Add platform manual_seed_all API (#38468)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-10 13:43:50 +08:00 |
|
 ElhamandGitHub
|
13e6b1b908
|
[BugFix][CPU] Add CPU profiler summary file output (#38366)
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com>
|
2026-04-10 13:41:15 +08:00 |
|
 Isotr0pyandGitHub
|
58c0a928c9
|
[Bugfix] Fix broken explicit unquantized kv cache dtype support (#38922)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-09 22:27:53 -07:00 |
|
  
|
3dd60971de
|
[feat]: make DCP error msg clearer (#28443)
Signed-off-by: vnadathur <glvikramn@gmail.com>
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com>
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com>
Co-authored-by: vnadathur <glvikramn@gmail.com>
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com>
|
2026-04-10 13:27:22 +08:00 |
|
 Ronen SchafferandGitHub
|
a5b17fba8f
|
[KV Offload] Implement shutdown() in OffloadingConnector and related classes (#39182)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-04-10 08:06:22 +03:00 |
|
 Cyrus LeungandGitHub
|
c48b2b83bd
|
[Mergify] Update model vendor auto-label rules (#39312)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-10 04:25:37 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
e7a1387e73
|
Add EXAONE-4.5 (#39388)
Signed-off-by: lkm2835 <lkm2835@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-09 20:53:26 -07:00 |
|
 
|
f83de7196f
|
[BugFix] Fix OOB read in CUTLASS grouped GEMM with epilogue (#38571)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-09 23:52:52 -04:00 |
|
 Ganesh RandGitHub
|
445a2a4d1a
|
feat(cpu): add CPU support for draft model speculative decoding (#32662)
Signed-off-by: R <Ganesh.R@amd.com>
|
2026-04-10 11:49:52 +08:00 |
|
 Kunshang JiandGitHub
|
55d037e2e5
|
[CT][FP8][Marlin] refactor CompressedTensorsW8A16Fp8 to use kernel abstraction (#38244)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>
|
2026-04-10 09:58:35 +08:00 |
|
 ChaunceyandGitHub
|
ecbfbb8d61
|
[Feature] Add auto-detection for reasoning_config when only reasoning_parser is set (#38214)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-10 01:36:26 +00:00 |
|
 Chuan (Richard) LiandGitHub
|
e0613702ad
|
[ROCm] Fix AITER ops fake impl and minor bugs (#36092)
Signed-off-by: Li <chuali@amd.com>
|
2026-04-09 17:56:17 -07:00 |
|
 Ibrahim ArshadandGitHub
|
9853a3c159
|
fix(gdn): Align prefill warmup with real prefill path (#39169)
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com>
|
2026-04-10 00:49:50 +00:00 |
|
 Artem PerevedentsevandGitHub
|
bb6047db13
|
[Model][Perf] Enable checkpoints prefetching for Lustre FS by default (#39422)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-10 00:47:59 +00:00 |
|
 
|
467d3247c3
|
[LMCache] vLLM Block Allocation Event (#38856)
Signed-off-by: yuwei <yuwei@dev.local>
Co-authored-by: yuwei <yuwei@dev.local>
|
2026-04-09 17:30:29 -07:00 |
|
 Cyrus LeungandGitHub
|
e5de19ff9a
|
[CI/Build[ Don't auto-rebase PRs with CI failures (#39443)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 13:57:37 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
edee96519a
|
[Spec Decode] fix returning size mismatch on extract hidden states proposer (#38610)
Signed-off-by: Jaebok Lee <jaebok9541@naver.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-09 20:39:39 +00:00 |
|
 Rishi PuriandGitHub
|
adaabb8a55
|
Add nightly b200 test for spec decode eagle correctness (#38577)
Signed-off-by: Rishi Puri <riship@nvidia.com>
|
2026-04-09 20:09:09 +00:00 |
|
 Ekagra RanjanandGitHub
|
f7cad67412
|
[ASR] Fix spacing bw chunks in multi chunk audio transcription (#39116)
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com>
|
2026-04-09 12:46:33 -07:00 |
|
 Xinyu ChenandGitHub
|
a8134aef4e
|
[XPU] check is_xccl_available before oneccl warmup (#39302)
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com>
|
2026-04-09 12:42:17 -07:00 |
|
 Michael GoinandGitHub
|
2800706f06
|
[Refactor] Move NVFP4 GEMM management into NvFp4LinearKernel (#39129)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-04-09 15:05:36 -04:00 |
|
 Cyrus LeungandGitHub
|
0d310ffbeb
|
[CI/Build] Update auto-rebase rule (#39429)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 10:59:56 -07:00 |
|
 Micah WilliamsonandGitHub
|
d5f75fdf50
|
[ROCm] Correctly guard fused_silu_mul_block_quant on ROCm (#39387)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-04-09 17:59:03 +00:00 |
|
 PikaPikachuandGitHub
|
827268e98d
|
[Quantization] Support Quark W8A8 INT8 MoE inference (#36320)
Signed-off-by: kangletian <Letian.Kang@amd.com>
|
2026-04-09 17:24:43 +00:00 |
|
 Wentao YeandGitHub
|
56e19d7ee2
|
[Model Runner V2] Fix flex attention kv blocks calculation issue (#39353)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-09 13:07:43 -04:00 |
|
 Andreas KaratzasandGitHub
|
9036d4c464
|
[ROCm][CI] Resolved nvidia package deps issue (#39421)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-10 00:06:06 +08:00 |
|
 
|
a8c6ee9b78
|
[Performance Improvement] Update batched_count_greater_than to handle batch size 1 without recompile (#38933)
Signed-off-by: Lucas Kabela <lucaskabela@meta.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
|
2026-04-09 23:51:31 +08:00 |
|
 Cyrus LeungandGitHub
|
3b1d9c3156
|
[CI/Build] Fix memory cleanup in MM test (#39411)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 08:50:45 -07:00 |
|
 Cyrus LeungandGitHub
|
54d244f28f
|
[UX] Improve error message for MM input too long (#39409)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 13:20:19 +00:00 |
|
 Richard ZouandGitHub
|
6c749399b7
|
[BugFix] fix tests/kernels/moe/test_moe_layer.py (#39404)
Signed-off-by: Richard Zou <zou3519@gmail.com>
|
2026-04-09 08:48:59 -04:00 |
|
 
|
91eea72330
|
[Tests] Add Qwen3-VL multimodal memory leak check (#39268)
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com>
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 04:54:46 -07:00 |
|
 
|
df2503e125
|
nemotron-nano-vl: Allow use_audio_in_video to be passed at vllm serve time (#38538)
Signed-off-by: Andrii Skliar <askliar@nvidia.com>
Co-authored-by: Andrii Skliar <askliar@nvidia.com>
|
2026-04-09 11:44:39 +00:00 |
|
 Nick HillandGitHub
|
c8d98f81f6
|
[Core] Simplify API server handshake (#39364)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-09 18:56:15 +08:00 |
|
 Harry MellorandGitHub
|
d87fb264df
|
[Docs] Bring README updates into docs README (#39397)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-04-09 10:35:00 +00:00 |
|
 wang.yuqiandGitHub
|
66c079ae83
|
[Frontend][4/n] Improve pooling entrypoints | pooling. (#39153)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-09 10:09:45 +00:00 |
|
 Shengqi ChenandGitHub
|
b6c9be509e
|
[CI] fix possible user permission issues in nightly index generation (#39390)
Signed-off-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-04-09 08:14:07 +00:00 |
|
 
|
ed733802f0
|
Fix NUMA binding on non-CDMM Grace-Blackwell systems (#39361)
Signed-off-by: Qidong Su <soodoshll@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-09 07:36:51 +00:00 |
|
 Andrew BarnesandGitHub
|
8a34c5087a
|
[ROCm] Remove unnecessary fp8 roundtrip in gather cache NHD dequant (#39122)
Signed-off-by: Bortlesboat <bortstheboat@gmail.com>
|
2026-04-09 15:12:22 +08:00 |
|
 Wentao YeandGitHub
|
ed2f282bc8
|
[Perf] Optimize redundant sync for pooling model, 3.7% Throughput Improvement (#39113)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-08 23:12:23 -07:00 |
|
 Zhewen LiandGitHub
|
9e78555743
|
[Docker] Add fastsafetensors to NVIDIA Dockerfile (#38950)
|
2026-04-08 22:21:37 -07:00 |
|
 
|
e80e633927
|
[XPU] Skip VLLM_BATCH_INVARIANT for XPU in EAGLE DP test (#39164)
Signed-off-by: sihao.li <sihao.li@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-04-09 12:45:16 +08:00 |
|
 
|
490f17d0c7
|
[Multimodal] Fix nested_tensors_equal: add length check for lists and tuple support (#38388)
Signed-off-by: khairulkabir1661 <khairulkabir1661@users.noreply.github.com>
Co-authored-by: khairulkabir1661 <khairulkabir1661@users.noreply.github.com>
|
2026-04-09 04:40:37 +00:00 |
|