 
|
4ff865c38e
|
[Bugfix] Disable allreduce_rms_fusion when pipeline_parallel_size > 1 (#43616)
Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-29 22:57:43 +08:00 |
|
 
|
5502c3b52d
|
[Misc] added unit tests for the core pooling methods (#43818)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-29 14:40:31 +00:00 |
|
 Chunyang WenandGitHub
|
f191d5630e
|
docs: clarify ITL acronym in optimization docs (#43922)
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com>
|
2026-05-29 07:40:05 -07:00 |
|
 
|
11dfa3169d
|
Add vLLM library info to Hugging Face Hub requests (#43857)
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Lucain Pouget <lucain@huggingface.co>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-05-29 14:04:58 +00:00 |
|
 Li, JiangandGitHub
|
3f6f508e14
|
[Bugfix][CPU] Remove invalid extra deps (#43977)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-05-29 22:02:09 +08:00 |
|
 Harry MellorandGitHub
|
0585b5ba2e
|
Skip docs build if PR doesn't affect docs (#43972)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-29 12:09:52 +00:00 |
|
 Thien TranandGitHub
|
d2889722ff
|
[Bugfix] Corrupted MLA + linear attention (#43961)
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2026-05-29 05:00:51 -07:00 |
|
 
|
0b56815a24
|
[ROCm][Perf] DSv3.2 MI355X TP4 decode-step orchestration cleanup (3 micro-opts) (#42982)
Signed-off-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-29 04:26:57 -07:00 |
|
 
|
ab12aab127
|
[Bugfix] [ROCm] [DSV4] Fix AITER MXFP4 MoE weight loading and shuffle… (#42595)
Co-authored-by: MHYangAMD <MHYangAMD@users.noreply.github.com>
|
2026-05-29 04:08:33 -07:00 |
|
 JartXandGitHub
|
0cff0741ff
|
[Kernel][ROCm] Native W4A16 kernel for AMD RDNA3 (gfx1100) — fp16 + bf16 (#41394)
Signed-off-by: JartX <sagformas@epdcenter.es>
|
2026-05-29 11:04:40 +00:00 |
|
 
|
60a7a2214f
|
[Bugfix] Fix Step3 pipeline parallel KeyError for residual tensor (#37622)
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-05-29 03:04:02 -07:00 |
|
 Nicolò LucchesiandGitHub
|
7ebc0ec104
|
[CI] Nixl+SimpleCPUOffloadingConnector unit tests (#43871)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-29 02:40:42 -07:00 |
|
 
|
e8b5199973
|
[XPU] support MTP of gdn attention (#43565)
Signed-off-by: mayuyuace <qiming1.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-29 17:10:24 +08:00 |
|
 Simon DanielssonandGitHub
|
b7fb747d8d
|
[CI][ROCm] Don't skip MoRI-IO Connector tests (#43703)
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
2026-05-29 17:06:23 +08:00 |
|
 Kunshang JiandGitHub
|
30c6289b8e
|
[XPU] fix xpu install document triton-xpu version (#43947)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-29 02:05:12 -07:00 |
|
 Andreas KaratzasandGitHub
|
ff990d0d32
|
[ROCm][CI] Fix AITER unified attention for encoder-decoder cross-attention (#43945)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-29 16:43:39 +08:00 |
|
 ChaunceyandGitHub
|
87f12e5c7c
|
[Frontend]Responses API supports chat_template_kwargs (#43761)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-05-29 07:58:19 +00:00 |
|
 kliuaeandGitHub
|
ab7521d77c
|
[ROCm][DSv4] Remove device pipeline stall in sparse attention (#43898)
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
|
2026-05-29 15:42:40 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
94d3f4d205
|
[CPU Backend] CPU top-k and top-p sampling kernels using Triton (#43633)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-29 15:02:39 +08:00 |
|
 
|
04516eabc8
|
[XPU] add gelu_tanh to xpu moe backend supported activations (#42822)
Signed-off-by: yintong-lu <yintong.lu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-29 14:37:20 +08:00 |
|
 
|
648c3ebee6
|
[CI] Separate non-root smoke tests from image build step (#43712)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-28 23:34:16 -07:00 |
|
  
|
22a58640b4
|
[9/n] Migrate attention and cache kernels to torch stable ABI (continued) (#43717)
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-29 04:44:45 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
710f077617
|
[Refactor] Remove dead code (#43234)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-29 00:29:56 -04:00 |
|
 
|
d63108fb18
|
[kv_offload] Skip decode-phase blocks in CPU offload (#43797)
Signed-off-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
|
2026-05-29 06:39:43 +03:00 |
|
 
|
9636709372
|
[XPU] add scale transpose to prepare_fp8_moe_layer_for_xpu and bump up kernels (#43277)
Signed-off-by: mayuyuace <qiming1.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-29 03:22:51 +00:00 |
|
 Weida HongandGitHub
|
dfe8ba7c80
|
Adjust design around encoder_cudagraph_forward (#42288)
Signed-off-by: Weida Hong <wdhongtw@google.com>
|
2026-05-29 03:02:52 +00:00 |
|
  
|
212deff2ec
|
[feat] add GlmgaProcessor specific logits in glm4_1v.py (#43575)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
|
2026-05-29 02:56:02 +00:00 |
|
 Woosuk KwonandGitHub
|
7bd45da585
|
[DSv4] Move mHC tilelang kernels & Don't use CustomOP in dsv4/nvidia (#43905)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-29 10:25:02 +08:00 |
|
 
|
bf18d7e0b4
|
[Misc][NUMA] Auto-bind to PCT priority cores on DGX B300 + widen EngineCore across shard NUMA nodes (#43270)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: Cursor <noreply@cursor.com>
|
2026-05-29 10:07:44 +08:00 |
|
 Bugen ZhaoandGitHub
|
1521173c17
|
[Rust Frontend] Add /version endpoint using engine-reported value (#43854)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-29 00:32:27 +00:00 |
|
    
|
b690b2bb67
|
[Model]Support Step-3.7-Flash (#43859)
Signed-off-by: luotingdan <luotingdan@stepfun.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: luotingdan <luotingdan@stepfun.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Yu Huang <yuhuang@nvidia.com>
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-28 17:01:48 -07:00 |
|
 yzong-rhandGitHub
|
325a1ec4fb
|
[CI] Enable prefix caching in BFCL benchmark (#43925)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-05-28 23:36:31 +00:00 |
|
 
|
69c9f19957
|
fix(frontend): Add multimodal placeholders to Gemma4 tool message template (#41459)
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com>
Co-authored-by: Ben Browning <bbrownin@redhat.com>
|
2026-05-28 14:48:12 -07:00 |
|
 rasmithandGitHub
|
9769e2df2a
|
[AMD][CI][BugFix] Fix Distributed Compile Unit Tests (2xH100-2xMI300) group (#43120)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-05-28 14:39:01 -07:00 |
|
 Michael GoinandGitHub
|
03f03f9630
|
Refactor output filename handling in ci-fetch-log.sh (#43901)
Signed-off-by: Michael Goin <mgoin64@gmail.com>
|
2026-05-28 14:20:12 -07:00 |
|
 Benjamin ChislettandGitHub
|
9202ea6fda
|
[Spec Decode] Allow causal DFlash (#43445)
|
2026-05-28 21:18:44 +00:00 |
|
 Woosuk KwonandGitHub
|
69b8956dcd
|
[Model Refactoring] Remove unncessary torch op registration for DSv4 (#43891)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-28 14:04:55 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
a3ed5ab10c
|
[KV Offload] Add per-request offloading policy via on_new_request lifecycle hook (#43205)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-28 20:45:18 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
7e53283b1c
|
[Core] Cleanup KVConnector handling with PP + fix MRV2 (#43732)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-28 13:12:03 -07:00 |
|
 
|
9090368b65
|
[Feat] Add support for per GPU worker RDMA NIC selection (#42083)
Signed-off-by: Raj Joshi <rajjoshi@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-28 12:45:23 -07:00 |
|
 Harry MellorandGitHub
|
085ac221a3
|
Deprecate JAISLMHeadModel (#43784)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-28 18:29:12 +00:00 |
|
 Hua HuangandGitHub
|
9006204e90
|
[MM][CG] Avoid over-padding Qwen2.5-VL encoder cudagraph window metadata (#42796)
Signed-off-by: Hua Huang <huah@nvidia.com>
|
2026-05-28 11:22:56 -07:00 |
|
 
|
ed7fe831da
|
[ROCm] Enable the aiter top-k/top-p sampler by default (#43331)
Signed-off-by: John Qin <yanyuan.qin@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-05-28 13:19:59 -05:00 |
|
 Nicolò LucchesiandGitHub
|
5b115bb8a3
|
[Attention][AMD] Standardize kv layout to blocks first for AMD (#43660)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-28 12:28:50 -05:00 |
|
 
|
53a2088675
|
Allow native KV cache dtype in Triton cache update (#43330)
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com>
|
2026-05-28 16:51:40 +00:00 |
|
 Chao-Ju ChenandGitHub
|
099024762c
|
[Rust Frontend] Optimize multimodal prompt expansion (#43670)
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com>
|
2026-05-28 09:46:18 -07:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
9aa131f944
|
Add Cosmos3 Reasoner model (#43356)
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-28 09:43:55 -07:00 |
|
 Micah WilliamsonandGitHub
|
1b5437cec8
|
[ROCm] Bump ROCm to 7.2.3 (#43136)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-05-28 09:42:43 -07:00 |
|
  
|
3207e7680e
|
[XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) (#41426)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-28 16:30:48 +00:00 |
|
 Matthias GehreandGitHub
|
a9ec46d4b7
|
[ROCm][Perf] Support N=5 in wvSplitK skinny GEMM kernels for speculative decoding (#40687)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
|
2026-05-28 16:28:21 +00:00 |
|