  
|
d4cb783c10
|
[Bugfix] Fix GDN FLA kernel crashes with NULL_BLOCK_ID=0 CUDA graph padding (#39064)
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com>
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-11 08:35:19 +00:00 |
|
 Li, JiangandGitHub
|
eb92ba740a
|
[CI/Build] Fix sentence-transformers version in CPU test (#39557)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-04-11 07:04:25 +00:00 |
|
 z1yingandGitHub
|
a3e750c0a5
|
[Misc] Update deprecation warning for --model flag (#39518)
Signed-off-by: Ziying Tao <tzzying@outlook.com>
|
2026-04-10 23:25:20 -07:00 |
|
 Lee YongjunandGitHub
|
da72daced2
|
[Bugfix] add SupportsMultiModal to Exaone4_5_MTP (#39526)
Signed-off-by: leeyongjun <jqueen.astro@gmail.com>
|
2026-04-10 22:57:26 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
8d0aabdde9
|
Fix the order of _free_encoder_inputs (#38907)
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-10 22:47:48 -07:00 |
|
 
|
0f3ce4c74b
|
[XPU] Fix spec-decode UTs under tests/v1/spec_decode (#38491)
Signed-off-by: Yan Ma <yan.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-04-11 01:31:00 +00:00 |
|
 Benjamin ChislettandGitHub
|
af661a182d
|
Revert "Add nightly b200 test for spec decode eagle correctness (#38577)" (#39512)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-04-10 20:07:32 -04:00 |
|
 Michael GoinandGitHub
|
7f0b8f2020
|
[Docs] Use --torch-backend=auto for editable install docs (#39511)
Signed-off-by: Michael Goin <mgoin64@gmail.com>
|
2026-04-10 15:27:02 -07:00 |
|
 Michael GoinandGitHub
|
11e2375fe2
|
[Refactor] Move MXFP8 GEMM management into MxFp8LinearKernel (#39205)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-04-10 14:02:03 -07:00 |
|
 
|
fc645f1acc
|
Add structure to requirements/ directory (#39024)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-10 13:46:41 -07:00 |
|
 Fynn Schmitt-UlmsandGitHub
|
2d80cf9d6e
|
Fix pre-commit labeled trigger system (#39523)
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com>
|
2026-04-10 13:54:49 -06:00 |
|
  
|
e7cfd7c5b9
|
Add Gemma4 Eagle3 support (#39450)
Signed-off-by: Rahul-Tuli <rtuli@redhat.com>
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com>
Co-authored-by: Rahul-Tuli <rtuli@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-04-10 12:35:35 -07:00 |
|
 yzong-rhandGitHub
|
e816a8811f
|
[Bugfix] Fix FlashInfer crash with kv_cache_dtype_skip_layers (#39002)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-04-10 18:50:47 +00:00 |
|
 zhanqiuhuandGitHub
|
e281cb721c
|
[CI] Add MultiConnector (Nixl+Offloading) e2e edge case tests (#39343)
Signed-off-by: ZhanqiuHu <zhu@redhat.com>
|
2026-04-10 17:35:03 +00:00 |
|
 ManuandGitHub
|
51cfc0e76c
|
perf(moe): add tuned fused_moe config for RTX PRO 6000 Blackwell Server Edition (#39183)
Signed-off-by: manu <fortin.emmanuel@gmail.com>
|
2026-04-10 11:32:42 -06:00 |
|
 
|
b87575d24b
|
feat: add logit_scale to PoolerConfig for affine score calibration (#39435)
Signed-off-by: Jesus Federico <jefp@amazon.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 17:21:14 +00:00 |
|
 TJianandGitHub
|
42c6bb4b75
|
[ROCm] [AITER] Revert AITER version to v0.1.10.post3 (#39509)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-04-10 16:25:52 +00:00 |
|
 Jee Jee LiandGitHub
|
ecd1ea1363
|
[Kernel] Porting the TRTLLM minimax_allreduce_rms kernels (#37045)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-04-11 00:20:20 +08:00 |
|
 zhrrrandGitHub
|
8f121f7879
|
[Model Runner V2] support auto resolve cudagraph mode/sizes based on attn backend (#32936)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
|
2026-04-10 08:27:15 -07:00 |
|
 wang.yuqiandGitHub
|
cb5f7501cb
|
[New Model]: jinaai/jina-reranker-v3 (#38800)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-10 15:20:40 +00:00 |
|
 Peter NguyenandGitHub
|
8d0f908b98
|
[Model] Implement LoRA support for Qwen3ASRForConditionalGeneration (#37247)
Signed-off-by: Peter Nguyen <petern0408@gmail.com>
|
2026-04-10 18:34:31 +04:00 |
|
 Nicolò LucchesiandGitHub
|
c9dddc144b
|
[CI] Add Nixl+OffloadingConnector e2e integration tests (#39200)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-04-10 21:40:40 +08:00 |
|
 
|
c1cc7344fb
|
[ROCm] Add RDNA 3.5/4 device IDs (gfx1150, gfx1151, gfx1201) (#38455)
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com>
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-04-10 11:35:07 +00:00 |
|
 xaguilar-amdandGitHub
|
f976e3b98b
|
[Performance] Remove unnecessary zero-fill of MLA decode output tensor in Aiter backend (#37539)
Signed-off-by: xaguilar-amd <xaguilar@amd.com>
|
2026-04-10 11:27:35 +00:00 |
|
  
|
d468322dc1
|
[Kernel][Hardware][AMD] Add TritonW4A16LinearKernel for ROCm (#37352)
Signed-off-by: jatseng-ai <jatseng@amd.com>
Signed-off-by: jatseng-ai <janet.tseng@amd.com>
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com>
|
2026-04-10 10:25:27 +00:00 |
|
 
|
967146e7bd
|
[model] support FireRedLID (#39290)
Signed-off-by: PatchouliTaisa <patchychen@tencent.com>
Co-authored-by: PatchouliTaisa <patchychen@tencent.com>
|
2026-04-10 08:43:58 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
8e8a3becd1
|
[ZenCPU] Make PT Backport Patch Accessible to vLLM (#38205)
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
|
2026-04-10 08:29:35 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1dfd64c1cc
|
[PluggableLayer][3/N] Apply PluggableLayer to llm_head and vocab embedding layer (#33465)
Signed-off-by: whx-sjtu <2952154980@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-10 16:12:59 +08:00 |
|
 
|
ad720aefe9
|
[Bugfix] Fix V1 dummy run writing NaN to KV cache null block (#39444)
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
|
2026-04-10 10:09:46 +02:00 |
|
 milesialandGitHub
|
270e8a4102
|
Nemotron Nano VL: Streamline pixel shuffle (#37580)
Signed-off-by: milesial <milesial@users.noreply.github.com>
|
2026-04-10 07:31:19 +00:00 |
|
 Richard ZouandGitHub
|
f44afef6d6
|
[compile] Allow strings in custom ops without regressing compilation times (#38123)
Signed-off-by: Richard Zou <zou3519@gmail.com>
|
2026-04-10 07:26:37 +00:00 |
|
 
|
447ce22212
|
[GGUF] Support non-standard quant types with prefix (e.g. UD-IQ1_S) (#39471)
Signed-off-by: Injae Ryou <injaeryou@gmail.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 07:22:53 +00:00 |
|
 Chendi.XueandGitHub
|
65e4e46f66
|
update CODEOWNERS file (#39439)
Signed-off-by: Chendi Xue <chendi.xue@intel.com>
|
2026-04-10 15:05:31 +08:00 |
|
 
|
49d20346e4
|
[Perf] Reduce H2D pageable memory copies (#38794)
Signed-off-by: jackcfwang <jackcfwang@tencent.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 15:03:26 +08:00 |
|
 Nick HillandGitHub
|
ef076c1b73
|
[Core] Change max_model_len in EngineCoreReadyResponse to be non-None (#39442)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-10 14:34:57 +08:00 |
|
 Yan MaandGitHub
|
ec68d53b2b
|
Add platform manual_seed_all API (#38468)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-10 13:43:50 +08:00 |
|
 ElhamandGitHub
|
13e6b1b908
|
[BugFix][CPU] Add CPU profiler summary file output (#38366)
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com>
|
2026-04-10 13:41:15 +08:00 |
|
 Isotr0pyandGitHub
|
58c0a928c9
|
[Bugfix] Fix broken explicit unquantized kv cache dtype support (#38922)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-09 22:27:53 -07:00 |
|
  
|
3dd60971de
|
[feat]: make DCP error msg clearer (#28443)
Signed-off-by: vnadathur <glvikramn@gmail.com>
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com>
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com>
Co-authored-by: vnadathur <glvikramn@gmail.com>
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com>
|
2026-04-10 13:27:22 +08:00 |
|
 Ronen SchafferandGitHub
|
a5b17fba8f
|
[KV Offload] Implement shutdown() in OffloadingConnector and related classes (#39182)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-04-10 08:06:22 +03:00 |
|
 Cyrus LeungandGitHub
|
c48b2b83bd
|
[Mergify] Update model vendor auto-label rules (#39312)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-10 04:25:37 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
e7a1387e73
|
Add EXAONE-4.5 (#39388)
Signed-off-by: lkm2835 <lkm2835@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-09 20:53:26 -07:00 |
|
 
|
f83de7196f
|
[BugFix] Fix OOB read in CUTLASS grouped GEMM with epilogue (#38571)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-09 23:52:52 -04:00 |
|
 Ganesh RandGitHub
|
445a2a4d1a
|
feat(cpu): add CPU support for draft model speculative decoding (#32662)
Signed-off-by: R <Ganesh.R@amd.com>
|
2026-04-10 11:49:52 +08:00 |
|
 Kunshang JiandGitHub
|
55d037e2e5
|
[CT][FP8][Marlin] refactor CompressedTensorsW8A16Fp8 to use kernel abstraction (#38244)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>
|
2026-04-10 09:58:35 +08:00 |
|
 ChaunceyandGitHub
|
ecbfbb8d61
|
[Feature] Add auto-detection for reasoning_config when only reasoning_parser is set (#38214)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-10 01:36:26 +00:00 |
|
 Chuan (Richard) LiandGitHub
|
e0613702ad
|
[ROCm] Fix AITER ops fake impl and minor bugs (#36092)
Signed-off-by: Li <chuali@amd.com>
|
2026-04-09 17:56:17 -07:00 |
|
 Ibrahim ArshadandGitHub
|
9853a3c159
|
fix(gdn): Align prefill warmup with real prefill path (#39169)
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com>
|
2026-04-10 00:49:50 +00:00 |
|
 Artem PerevedentsevandGitHub
|
bb6047db13
|
[Model][Perf] Enable checkpoints prefetching for Lustre FS by default (#39422)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-10 00:47:59 +00:00 |
|
 
|
467d3247c3
|
[LMCache] vLLM Block Allocation Event (#38856)
Signed-off-by: yuwei <yuwei@dev.local>
Co-authored-by: yuwei <yuwei@dev.local>
|
2026-04-09 17:30:29 -07:00 |
|