 
|
de2d76f352
|
[Build] Switch CUDA 12.9 wheel builds to PyTorch manylinux_2_28 base (#41668)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-15 13:46:16 -07:00 |
|
 
|
9a7a273dfe
|
Add HumanEval and GSM8K benchmarks to datasets (#42648)
Signed-off-by: southfreebird <yvorott@gmail.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
|
2026-05-15 13:01:21 -07:00 |
|
 
|
b2c58ee942
|
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants (#42685)
Signed-off-by: Lanze Liu <lanzetech@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-15 15:34:59 -04:00 |
|
 frida-anderssonandGitHub
|
4d67d3bde2
|
[ROCm] Restore fast top_k_per_row kernels for sparse MLA when topk_tokens=2048 (#42072)
Signed-off-by: Frida Andersson <fanderss@amd.com>
|
2026-05-15 19:02:57 +00:00 |
|
  
|
06d020bb6e
|
[Bugfix] Fix SM121 (DGX Spark) exclusion from Marlin/CUTLASS FP8 paths (#35568)
Signed-off-by: Blake Ledden <blake@secondnaturecomputing.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Pavani Majety <pmajety@nvidia.com>
|
2026-05-15 10:59:00 -07:00 |
|
 chunxiaozhengandGitHub
|
f45c210885
|
[LMCacheMPConnector] Prioritize importing the lmcache_mp_connector from lmcache (#42596)
Signed-off-by: idellzheng <idellzheng@tencent.com>
|
2026-05-15 17:46:31 +00:00 |
|
 akii96andGitHub
|
be7a03ea65
|
[ROCm] Widen AITER fused AR RMSNorm 1-stage gate (#42409)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
|
2026-05-15 17:44:38 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
6147c70224
|
[Model Runner v2] Support reload weights (sleep mode) (#42673)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-15 16:41:23 +00:00 |
|
 
|
0162596603
|
[Model Runner V2] FP32 gumbel sampling. (#41775)
Signed-off-by: PatchouliTaisa <patchychen@tencent.com>
Co-authored-by: PatchouliTaisa <patchychen@tencent.com>
|
2026-05-15 09:20:08 -07:00 |
|
   
|
46a95815d3
|
[ROCm][MLA] FP8 ASM prefill for AITER dense MLA backend on gfx950 (#42509)
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com>
Co-authored-by: clintg6 <clint.greene@amd.com>
Co-authored-by: frida-andersson <frida.andersson@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-15 23:56:58 +08:00 |
|
 BadrBasowidandGitHub
|
fb5bd03f51
|
[Perf] Set IR Op Priority Once at Worker Init (#42631)
Signed-off-by: BadrBasowid <badr.basowid@gmail.com>
|
2026-05-15 15:56:13 +00:00 |
|
 Mohammad Miadh AngkadandGitHub
|
ee58665aac
|
[Bugfix] Fix DeepGEMM context lens contiguity in MLA indexer (#42135)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-15 23:29:58 +08:00 |
|
 Wentao YeandGitHub
|
491e8d8539
|
[Perf] Optimize MLA attention _v_up_proj bmm by removing additional copy (#42561)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-15 08:14:26 -07:00 |
|
 Wentao YeandGitHub
|
af9616d845
|
[Model Runner V2] Fix kv_connector pre_forward order (#42676)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-15 08:13:59 -07:00 |
|
 
|
d792d993c1
|
[ROCm] Widen OAI Triton MoE capability range to include gfx12 (RDNA4) (#37826)
Signed-off-by: L.B.R. <lbr@mmonad.com>
Co-authored-by: L.B.R. <lbr@mmonad.com>
|
2026-05-15 07:59:57 -07:00 |
|
 Aaron HaoandGitHub
|
e0a45f1455
|
[Feat][RL] IPC weight sync optimizations: multigpu support and chunked packed tensors (#37476)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
Signed-off-by: hao-aaron <ahao@anyscale.com>
|
2026-05-15 22:53:06 +08:00 |
|
 Benjamin ChislettandGitHub
|
0fe7550254
|
[Bugfix] DFlash FP8 KV-Cache (#42692)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-05-15 08:29:45 -06:00 |
|
 Li, JiangandGitHub
|
95cfe102a5
|
[Bugfix] Ensure embeding model compilation on CPU (#42709)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-05-15 18:58:19 +08:00 |
|
 
|
1dc3fe08ea
|
gemma3 multi-gpu bug-fix (#42630)
Signed-off-by: Philip Maybank <pmaybank@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-05-15 02:32:05 -07:00 |
|
 
|
d26a28ab03
|
fix: propagate revision/code_revision pins to all artifact boundaries (#42616)
Signed-off-by: jperezde <jperezde@redhat.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-05-15 02:31:54 -07:00 |
|
 Andreas KaratzasandGitHub
|
d735968f6d
|
[ROCm][CI] Stage B gating (#42025)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
v0.21.1rc0
|
2026-05-15 01:49:27 -07:00 |
|
 
|
ccde9540be
|
DeepSeekV4-Pro enable cuda graph full and piecewise mode (#42604)
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-05-15 01:45:30 -07:00 |
|
 wang.yuqiandGitHub
|
75fd68c7a5
|
[Entrypoints] Split the pooling offline API into PoolingOfflineMixin. (#42267)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-15 08:05:57 +00:00 |
|
 Yifan QiaoandGitHub
|
4b364f810e
|
[Core][DSV4] Skip caching SWA blocks that can never serve a prefix-cache hit (#42258)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-05-15 15:59:18 +08:00 |
|
 
|
31fa757cf9
|
[Misc] Make it simpler to replace out-of-tree layer classes with related LoRA layers. (#42306)
Signed-off-by: paulyu12 <507435917@qq.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-05-15 15:20:42 +08:00 |
|
 Cyrus LeungandGitHub
|
2676ab1e0b
|
[Deprecation] Remove old locations of get_tokenizer and resolve_hf_chat_template (#35024)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-05-15 00:13:32 -07:00 |
|
 
|
27b85d2084
|
[Bugfix] Clarify CPU backend memory error messages reference shared flag (#42479)
Signed-off-by: daniel-devlab <282598346+daniel-devlab@users.noreply.github.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-05-15 06:35:05 +00:00 |
|
 Louie TsaiandGitHub
|
e30f39c4f1
|
Update Intel Xeon model list and vLLM Benchmark Suite BKMs (#42607)
Signed-off-by: louie-tsai <louie.tsai@intel.com>
|
2026-05-15 05:14:03 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
bf610c2f56
|
[Bugfix] Fix inverted condition causing thinking_token_budget to be silently ignored (#41674)
Signed-off-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-15 12:48:49 +08:00 |
|
 
|
faa4b76afa
|
[Model] Support InternS2 Preview (#42705)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: zxy <46674730+CUHKSZzxy@users.noreply.github.com>
|
2026-05-14 21:30:26 -07:00 |
|
 
|
f351455f0f
|
[CPU][RISC-V] Add RVV-optimized attention kernels for RISC-V Vector Extension (#40119)
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-15 12:08:23 +08:00 |
|
 Cyrus LeungandGitHub
|
56434e8651
|
[Bugfix] Fix incorrect chat template format for Qwen3.5 (#42660)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-05-14 20:52:52 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
0d4d334eaa
|
Bump llguidance to 1.7 (#42150)
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-14 20:35:27 -04:00 |
|
  
|
fa2a33b893
|
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py (#38288)
Signed-off-by: Chengyi Nie <cnie@roblox.com>
Co-authored-by: Chengyi Nie <cnie@roblox.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-15 08:25:52 +08:00 |
|
 Giancarlo DelfinandGitHub
|
3b6a204789
|
[Model Runner V2][Bug Fix][DSV4] Ensure lazy attention state initializations happen during cudagraph capture (#42444)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-05-14 16:16:17 -07:00 |
|
 
|
f8848b2f2d
|
[Bugfix] Add swiglu limits to deepgemm fp8 methods (#41986)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-14 15:43:13 -07:00 |
|
 Charlie FuandGitHub
|
4cfcc0866f
|
[CI][ROCm] Remove unsupported cases in test_fusion.py (#38680)
Signed-off-by: charlifu <charlifu@amd.com>
|
2026-05-14 17:37:18 -04:00 |
|
 
|
f887aa1a53
|
[Aiter][ROCm] RMSNormGated+GroupedQuantFP8 fusion (#40710)
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Tres Popp <trespopp@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-14 15:37:09 -04:00 |
|
 Matthew BonanniandGitHub
|
9898f94abe
|
[Attention] Remove deprecated MLA prefill arguments (#42555)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-14 10:34:06 -07:00 |
|
 
|
ae4f59f0ec
|
[Model Runner v2] Oracle for model runner v2 - qwen3 dense model by default [1/N] (#39337)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-14 10:02:33 -07:00 |
|
 RanranandGitHub
|
f3d5360591
|
[Bugfix][Multimodal] PyAV video backend returns keyframes labeled as targets (#42586)
Signed-off-by: Ranran <hzz5361@psu.edu>
|
2026-05-14 08:56:59 -07:00 |
|
 Baorun (Lauren) MuandGitHub
|
a7737cb4f3
|
[Fix] Misc Fixes in ViT CUDA Graph (#38040)
Signed-off-by: Baorun Mu <bmu@nvidia.com>
|
2026-05-14 23:49:06 +08:00 |
|
 Cyrus LeungandGitHub
|
b8a25d0e12
|
[Bugfix] Fix LM detection for Nemotron Parse (#42641)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-05-14 23:42:10 +08:00 |
|
 frida-anderssonandGitHub
|
f07b1da797
|
[ROCm] Enable gluon paged MQA logits on gfx950 (MI355X) (#42062)
Signed-off-by: Frida Andersson <fanderss@amd.com>
|
2026-05-14 15:39:26 +00:00 |
|
 
|
f60c6b33a5
|
[V1][DP][LB] Publish request counts at the start of each engine step (#41626)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-14 15:39:24 +00:00 |
|
 
|
24337fb860
|
PD disagg with NIXL Connector: GDN support (Qwen3.5) (#41869)
Signed-off-by: Zhanqiu Hu <zhu@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-14 16:33:01 +02:00 |
|
   
|
c7560af424
|
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support (#39568)
Signed-off-by: xhx1022 <1737006628@qq.com>
Signed-off-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com>
|
2026-05-14 14:12:30 +00:00 |
|
 Mohammad Miadh AngkadandGitHub
|
2317682f95
|
[Bugfix] Fix TRTLLM ragged MLA prefill workspace warmup (#42112)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-14 09:48:56 -04:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
5bd8c71e79
|
[kv_offload] Implement reset_cache() for the offloading connector (#41956)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
|
2026-05-14 16:00:10 +03:00 |
|
 Wentao YeandGitHub
|
6548560496
|
[Compile] Fix compile warning with topk softplus sqrt (#41261)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-14 05:12:50 -07:00 |
|