    
|
ee5a89f4d7
|
[ROCm][MiniMax-M3] Add AITER sparse paged attention (#47287)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-07-12 19:27:29 -07:00 |
|
 
|
e26264f3ef
|
[Kernel] Implement CUDA kernel for ReLUSquaredActivation (relu^2) (#39058)
Signed-off-by: Tanish Malekar <tanishmalekar32@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-07-12 19:18:03 -07:00 |
|
    
|
51878e5b6e
|
[2/N][KV-Cache Layout Refactor] Pack K/V into the content dim across attention backends (#44455)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Signed-off-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-07-11 11:11:16 -04:00 |
|
   
|
19069bcbd5
|
FP32 router GEMV optimization (#48335)
Signed-off-by: peiyuanz <peiyuanz@inferact.ai>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: peiyuanz <peiyuanz@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: zhouzhou <zhouzhou@zhouzhoudeMacBook-Pro.local>
|
2026-07-11 13:07:48 +00:00 |
|
 gnovackandGitHub
|
f378f79b7c
|
handle topk_ids padding in align sum kernel (#47785)
Signed-off-by: gnovack <novackgm@gmail.com>
|
2026-07-10 13:33:28 -07:00 |
|
 Michael GoinandGitHub
|
08dfd68610
|
[Model] Add LongCat-Flash-Lite (n-gram embedding) (#47857)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-07-10 07:17:50 -07:00 |
|
  
|
c85d72076a
|
[HARDWARE][POWER] optimize math functions of VSX power (#47321)
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com>
Signed-off-by: Akash Kaothalkar <akash.kaothalkar@ibm.com>
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com>
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com>
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com>
Co-authored-by: Akash Kaothalkar <akash.kaothalkar@ibm.com>
|
2026-07-07 09:35:47 +00:00 |
|
 Jee Jee LiandGitHub
|
5d23ca47ab
|
[Kernel] Applies routed_scaling_factor internally (#47408)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-07-07 02:00:54 -07:00 |
|
 Li, JiangandGitHub
|
344609ab17
|
[CI/Build] Fix pre-commit check (#47695)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-07-06 08:24:24 +00:00 |
|
 velonica0andGitHub
|
990c2a0187
|
[RISC-V] Enable BF16 on VLEN=256 hardware (#45243)
Signed-off-by: velonica0 <like@mail.nankai.edu.cn>
|
2026-07-06 06:05:16 +00:00 |
|
 
|
e433634c78
|
[Performance][Hardware][RISC-V] Reduce LMUL pressure in INT4 LUT dequant (#47538)
Signed-off-by: liutong <liutong@iscas.ac.cn>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-06 05:58:56 +00:00 |
|
 
|
16f8110935
|
[Bugfix][CPU][RISC-V] Fix VLEN detection for RVV attention path (#47532)
Signed-off-by: liutong <liutong@iscas.ac.cn>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-06 05:58:03 +00:00 |
|
 Fadi ArafehandGitHub
|
f1073c050c
|
[CPU][BugFix] Multiple fixes to w4a8_int8 CPU MoE path (#46739)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
|
2026-07-06 05:39:20 +00:00 |
|
 gausah01andGitHub
|
26eb87204d
|
[Bugfix] Fix CPU split-KV scratchpad sizing (#45844)
Signed-off-by: Gauri Sahnan <gauri.sahnan@arm.com>
|
2026-07-04 06:47:23 +00:00 |
|
 Chris LeonardandGitHub
|
fbc9ba6d30
|
New stable abi cleanup (#46656)
Signed-off-by: Chris Leonard <chleonar@redhat.com>
|
2026-07-03 14:02:26 +08:00 |
|
 Michael GoinandGitHub
|
d715b3aa1e
|
Delete PagedAttention (#47361)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-07-02 12:31:26 -07:00 |
|
 TJianandGitHub
|
de2a8fc042
|
[ROCm] [PyTorch] Move to stable abi since ROCm upgraded to torch 2.11 (#47128)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-07-02 18:34:07 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
d0a2584773
|
[Misc] Use functions instead of PTX for the PDL instruction (#46984)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-01 19:38:35 -07:00 |
|
 
|
89e99202f2
|
[CPU][Perf]Added tanh AOR for faster gelu activations. (#44639)
Signed-off-by: Anna Mayne <anna.mayne@arm.com>
Signed-off-by: almayne <anna.mayne@arm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-06-30 23:24:40 -07:00 |
|
 Wentao YeandGitHub
|
9e84ec8648
|
[Refactor] Remove dead minimax allreduce rms kernel (#46842)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-30 08:29:21 -07:00 |
|
 
|
c8fb2963bd
|
[FS-Offloading] Batch Lookup in C (#46713)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-29 09:28:32 -07:00 |
|
 
|
c6554f321c
|
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build (#46769)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-26 14:32:21 -04:00 |
|
 MattandGitHub
|
1a4984520e
|
[Hardware][AMD][CI] Fix AMD CI image build (#46792)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 22:05:12 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
02a1f23711
|
[DFlash] Fuse precompute kv per-layer rmsnorms (#46761)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:32:07 -07:00 |
|
 Wentao YeandGitHub
|
cc7981599e
|
[Refactor] Remove dead kernel code (#46405)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-25 18:09:56 -07:00 |
|
 
|
e8c24a7695
|
[Kernel] Vectorized fp32 moe_sum reduction and support any topk (#46643)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-25 14:02:28 -07:00 |
|
 haoyangli0109andGitHub
|
1744adc256
|
[ROCM] [Communication] Add INT3 quantization method for quickreduce (#45666)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
|
2026-06-25 15:14:15 +00:00 |
|
  
|
638b1a99cc
|
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM (#45269)
Signed-off-by: wcy <233313160abc@gmail.com>
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-06-25 08:18:10 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
2396d91e93
|
[CPU][Spec Decode] Enable DFlash SD for CPU (#44029)
Signed-off-by: guybd <guy.boudoukh@intel.com>
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 15:32:48 +08:00 |
|
 Matthias GehreandGitHub
|
77c1d9fe9b
|
[ROCm][Perf] Tune wvSplitK on gfx1151 (#40784)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
|
2026-06-25 14:17:46 +08:00 |
|
  
|
1aad125815
|
[CPU] Enable chunked prefill and prefix caching for qwen3.5 (#46202)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-06-25 03:49:21 +00:00 |
|
 Jee Jee LiandGitHub
|
23aed9b0ee
|
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel (#46508)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-25 08:42:51 +08:00 |
|
 
|
b3a688cb9e
|
[ROCm] Fix OOB During Model Warmup With ROCM_ATTN and MRV2 (#46548)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-24 10:53:21 -05:00 |
|
 Cyrus LeungandGitHub
|
24d5186138
|
[Bugfix] Re-enable FP8 MoE on NVIDIA Thor (#46339)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-06-24 07:35:46 -07:00 |
|
 Fadi ArafehandGitHub
|
061043eaca
|
[CPU][Perf] Accelerate unquantized MoE for AArch64 (#46353)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
|
2026-06-24 14:14:35 +00:00 |
|
 Mohammad Miadh AngkadandGitHub
|
191826ec61
|
[CI/Build] Fix topk histogram build on SM75 (#46550)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-24 00:51:11 -07:00 |
|
 Jee Jee LiandGitHub
|
9d6fdc2901
|
[Kernel] GLM5 Router GEMM (#46385)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-23 22:54:50 -07:00 |
|
 Roberto L. CastroandGitHub
|
855cd4d787
|
[Perf][DSv4/DSv3.2] Add cluster-cooperative topK kernel for low-latency scenarios (#43008)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
|
2026-06-23 16:11:00 -07:00 |
|
   
|
68afd78897
|
[Bugfix][ROCm] Fix cumem sleep and teardown (#46203)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-24 02:45:31 +08:00 |
|
  
|
6691f087a6
|
[Minimax-M3] BF16/FP8 Indexer using MSA (#45892)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2026-06-23 10:28:49 -07:00 |
|
 Rukhaiya2004andGitHub
|
9f5117820f
|
[HARDWARE][POWER] Enable fp16 support for PowerPC (#46135)
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com>
|
2026-06-23 13:24:49 +00:00 |
|
  
|
7d47cff933
|
[Bugfix][KV Offload] Fix swap_blocks_batch on the default stream (#46379)
Signed-off-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com>
|
2026-06-23 05:45:27 -07:00 |
|
 Yifan QiaoandGitHub
|
aa4990a9a2
|
[Attention] Re-enable cross-layer KV cache layout for MLA via stride-aware kernels (#45111)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-06-22 06:57:02 -07:00 |
|
 
|
d2c671c29b
|
[CPU][RISC-V] Add RVV micro GEMM for WNA16 (#44324)
Signed-off-by: wcy <233313160abc@gmail.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-06-22 12:53:54 +00:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
ebfbcfe46a
|
Stop setting CUDA_VISIBLE_DEVICES internally in vLLM, add device_ids arg (#45026)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: kourosh hakhamaneshi <kouroshHakha@users.noreply.github.com>
|
2026-06-20 13:38:10 -07:00 |
|
 JasonLi314andGitHub
|
93bad11912
|
[Bugfix] Fix gridDim.y overflow for large row counts (#45255)
Signed-off-by: Jason Li <li.jason.cs@gmail.com>
|
2026-06-19 23:27:45 -04:00 |
|
 Chris LeonardandGitHub
|
b9a7cd464c
|
[12/n] final _C library kernel migration (#45415)
|
2026-06-19 06:57:26 -07:00 |
|
 Wentao YeandGitHub
|
225936a1dd
|
[CI Bug] Revert #42379 to fix CI Multi-Modal Models (Extended Generation 1) (#46070)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-18 12:37:39 -07:00 |
|
 HumphreyandGitHub
|
4583630b56
|
[Bugfix][Kernel] Check output alignment in vectorize_with_alignment (fixes misaligned-address crash for non-multiple-of-8 head sizes) (#45466)
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com>
|
2026-06-18 16:58:22 +00:00 |
|
 
|
021cdf72bc
|
Fix _riscv_supports_rvv_vlen128() to detect RVV on hardware without zvl flags (#43179)
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn>
Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn>
|
2026-06-18 21:22:35 +08:00 |
|