 
|
b87575d24b
|
feat: add logit_scale to PoolerConfig for affine score calibration (#39435)
Signed-off-by: Jesus Federico <jefp@amazon.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 17:21:14 +00:00 |
|
 TJianandGitHub
|
42c6bb4b75
|
[ROCm] [AITER] Revert AITER version to v0.1.10.post3 (#39509)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-04-10 16:25:52 +00:00 |
|
 Jee Jee LiandGitHub
|
ecd1ea1363
|
[Kernel] Porting the TRTLLM minimax_allreduce_rms kernels (#37045)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-04-11 00:20:20 +08:00 |
|
 zhrrrandGitHub
|
8f121f7879
|
[Model Runner V2] support auto resolve cudagraph mode/sizes based on attn backend (#32936)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
|
2026-04-10 08:27:15 -07:00 |
|
 wang.yuqiandGitHub
|
cb5f7501cb
|
[New Model]: jinaai/jina-reranker-v3 (#38800)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-10 15:20:40 +00:00 |
|
 Peter NguyenandGitHub
|
8d0f908b98
|
[Model] Implement LoRA support for Qwen3ASRForConditionalGeneration (#37247)
Signed-off-by: Peter Nguyen <petern0408@gmail.com>
|
2026-04-10 18:34:31 +04:00 |
|
 Nicolò LucchesiandGitHub
|
c9dddc144b
|
[CI] Add Nixl+OffloadingConnector e2e integration tests (#39200)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-04-10 21:40:40 +08:00 |
|
 
|
c1cc7344fb
|
[ROCm] Add RDNA 3.5/4 device IDs (gfx1150, gfx1151, gfx1201) (#38455)
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com>
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-04-10 11:35:07 +00:00 |
|
 xaguilar-amdandGitHub
|
f976e3b98b
|
[Performance] Remove unnecessary zero-fill of MLA decode output tensor in Aiter backend (#37539)
Signed-off-by: xaguilar-amd <xaguilar@amd.com>
|
2026-04-10 11:27:35 +00:00 |
|
  
|
d468322dc1
|
[Kernel][Hardware][AMD] Add TritonW4A16LinearKernel for ROCm (#37352)
Signed-off-by: jatseng-ai <jatseng@amd.com>
Signed-off-by: jatseng-ai <janet.tseng@amd.com>
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com>
|
2026-04-10 10:25:27 +00:00 |
|
 
|
967146e7bd
|
[model] support FireRedLID (#39290)
Signed-off-by: PatchouliTaisa <patchychen@tencent.com>
Co-authored-by: PatchouliTaisa <patchychen@tencent.com>
|
2026-04-10 08:43:58 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
8e8a3becd1
|
[ZenCPU] Make PT Backport Patch Accessible to vLLM (#38205)
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
|
2026-04-10 08:29:35 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1dfd64c1cc
|
[PluggableLayer][3/N] Apply PluggableLayer to llm_head and vocab embedding layer (#33465)
Signed-off-by: whx-sjtu <2952154980@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-10 16:12:59 +08:00 |
|
 
|
ad720aefe9
|
[Bugfix] Fix V1 dummy run writing NaN to KV cache null block (#39444)
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
|
2026-04-10 10:09:46 +02:00 |
|
 milesialandGitHub
|
270e8a4102
|
Nemotron Nano VL: Streamline pixel shuffle (#37580)
Signed-off-by: milesial <milesial@users.noreply.github.com>
|
2026-04-10 07:31:19 +00:00 |
|
 Richard ZouandGitHub
|
f44afef6d6
|
[compile] Allow strings in custom ops without regressing compilation times (#38123)
Signed-off-by: Richard Zou <zou3519@gmail.com>
|
2026-04-10 07:26:37 +00:00 |
|
 
|
447ce22212
|
[GGUF] Support non-standard quant types with prefix (e.g. UD-IQ1_S) (#39471)
Signed-off-by: Injae Ryou <injaeryou@gmail.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 07:22:53 +00:00 |
|
 Chendi.XueandGitHub
|
65e4e46f66
|
update CODEOWNERS file (#39439)
Signed-off-by: Chendi Xue <chendi.xue@intel.com>
|
2026-04-10 15:05:31 +08:00 |
|
 
|
49d20346e4
|
[Perf] Reduce H2D pageable memory copies (#38794)
Signed-off-by: jackcfwang <jackcfwang@tencent.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-10 15:03:26 +08:00 |
|
 Nick HillandGitHub
|
ef076c1b73
|
[Core] Change max_model_len in EngineCoreReadyResponse to be non-None (#39442)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-10 14:34:57 +08:00 |
|
 Yan MaandGitHub
|
ec68d53b2b
|
Add platform manual_seed_all API (#38468)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-10 13:43:50 +08:00 |
|
 ElhamandGitHub
|
13e6b1b908
|
[BugFix][CPU] Add CPU profiler summary file output (#38366)
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com>
|
2026-04-10 13:41:15 +08:00 |
|
 Isotr0pyandGitHub
|
58c0a928c9
|
[Bugfix] Fix broken explicit unquantized kv cache dtype support (#38922)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-04-09 22:27:53 -07:00 |
|
  
|
3dd60971de
|
[feat]: make DCP error msg clearer (#28443)
Signed-off-by: vnadathur <glvikramn@gmail.com>
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com>
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com>
Co-authored-by: vnadathur <glvikramn@gmail.com>
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com>
|
2026-04-10 13:27:22 +08:00 |
|
 Ronen SchafferandGitHub
|
a5b17fba8f
|
[KV Offload] Implement shutdown() in OffloadingConnector and related classes (#39182)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-04-10 08:06:22 +03:00 |
|
 Cyrus LeungandGitHub
|
c48b2b83bd
|
[Mergify] Update model vendor auto-label rules (#39312)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-10 04:25:37 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
e7a1387e73
|
Add EXAONE-4.5 (#39388)
Signed-off-by: lkm2835 <lkm2835@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-09 20:53:26 -07:00 |
|
 
|
f83de7196f
|
[BugFix] Fix OOB read in CUTLASS grouped GEMM with epilogue (#38571)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-09 23:52:52 -04:00 |
|
 Ganesh RandGitHub
|
445a2a4d1a
|
feat(cpu): add CPU support for draft model speculative decoding (#32662)
Signed-off-by: R <Ganesh.R@amd.com>
|
2026-04-10 11:49:52 +08:00 |
|
 Kunshang JiandGitHub
|
55d037e2e5
|
[CT][FP8][Marlin] refactor CompressedTensorsW8A16Fp8 to use kernel abstraction (#38244)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>
|
2026-04-10 09:58:35 +08:00 |
|
 ChaunceyandGitHub
|
ecbfbb8d61
|
[Feature] Add auto-detection for reasoning_config when only reasoning_parser is set (#38214)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-10 01:36:26 +00:00 |
|
 Chuan (Richard) LiandGitHub
|
e0613702ad
|
[ROCm] Fix AITER ops fake impl and minor bugs (#36092)
Signed-off-by: Li <chuali@amd.com>
|
2026-04-09 17:56:17 -07:00 |
|
 Ibrahim ArshadandGitHub
|
9853a3c159
|
fix(gdn): Align prefill warmup with real prefill path (#39169)
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com>
|
2026-04-10 00:49:50 +00:00 |
|
 Artem PerevedentsevandGitHub
|
bb6047db13
|
[Model][Perf] Enable checkpoints prefetching for Lustre FS by default (#39422)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-10 00:47:59 +00:00 |
|
 
|
467d3247c3
|
[LMCache] vLLM Block Allocation Event (#38856)
Signed-off-by: yuwei <yuwei@dev.local>
Co-authored-by: yuwei <yuwei@dev.local>
|
2026-04-09 17:30:29 -07:00 |
|
 Cyrus LeungandGitHub
|
e5de19ff9a
|
[CI/Build[ Don't auto-rebase PRs with CI failures (#39443)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 13:57:37 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
edee96519a
|
[Spec Decode] fix returning size mismatch on extract hidden states proposer (#38610)
Signed-off-by: Jaebok Lee <jaebok9541@naver.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-09 20:39:39 +00:00 |
|
 Rishi PuriandGitHub
|
adaabb8a55
|
Add nightly b200 test for spec decode eagle correctness (#38577)
Signed-off-by: Rishi Puri <riship@nvidia.com>
|
2026-04-09 20:09:09 +00:00 |
|
 Ekagra RanjanandGitHub
|
f7cad67412
|
[ASR] Fix spacing bw chunks in multi chunk audio transcription (#39116)
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com>
|
2026-04-09 12:46:33 -07:00 |
|
 Xinyu ChenandGitHub
|
a8134aef4e
|
[XPU] check is_xccl_available before oneccl warmup (#39302)
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com>
|
2026-04-09 12:42:17 -07:00 |
|
 Michael GoinandGitHub
|
2800706f06
|
[Refactor] Move NVFP4 GEMM management into NvFp4LinearKernel (#39129)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-04-09 15:05:36 -04:00 |
|
 Cyrus LeungandGitHub
|
0d310ffbeb
|
[CI/Build] Update auto-rebase rule (#39429)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 10:59:56 -07:00 |
|
 Micah WilliamsonandGitHub
|
d5f75fdf50
|
[ROCm] Correctly guard fused_silu_mul_block_quant on ROCm (#39387)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-04-09 17:59:03 +00:00 |
|
 PikaPikachuandGitHub
|
827268e98d
|
[Quantization] Support Quark W8A8 INT8 MoE inference (#36320)
Signed-off-by: kangletian <Letian.Kang@amd.com>
|
2026-04-09 17:24:43 +00:00 |
|
 Wentao YeandGitHub
|
56e19d7ee2
|
[Model Runner V2] Fix flex attention kv blocks calculation issue (#39353)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-09 13:07:43 -04:00 |
|
 Andreas KaratzasandGitHub
|
9036d4c464
|
[ROCm][CI] Resolved nvidia package deps issue (#39421)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-10 00:06:06 +08:00 |
|
 
|
a8c6ee9b78
|
[Performance Improvement] Update batched_count_greater_than to handle batch size 1 without recompile (#38933)
Signed-off-by: Lucas Kabela <lucaskabela@meta.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
|
2026-04-09 23:51:31 +08:00 |
|
 Cyrus LeungandGitHub
|
3b1d9c3156
|
[CI/Build] Fix memory cleanup in MM test (#39411)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 08:50:45 -07:00 |
|
 Cyrus LeungandGitHub
|
54d244f28f
|
[UX] Improve error message for MM input too long (#39409)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-04-09 13:20:19 +00:00 |
|
 Richard ZouandGitHub
|
6c749399b7
|
[BugFix] fix tests/kernels/moe/test_moe_layer.py (#39404)
Signed-off-by: Richard Zou <zou3519@gmail.com>
|
2026-04-09 08:48:59 -04:00 |
|