   
|
1053e248f0
|
[ROCm][Quantization][5/N] Refactor quark_moe w8a8-int8 w/ oracle (#46765)
Signed-off-by: amd-sourjya <amd-sourjya@users.noreply.github.com>
Co-authored-by: amd-sourjya <amd-sourjya@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 16:01:34 -05:00 |
|
 Wentao YeandGitHub
|
b5bcb3ce88
|
[Refactor] Remove dead code in multiple files (#49745)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-27 15:58:26 -04:00 |
|
 Wentao YeandGitHub
|
b2f9e4caa4
|
[DSv4 Perf] Adaptive topk width, 1.0% E2E throughput improvement (#50004)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-27 15:56:52 -04:00 |
|
  
|
831d3848f1
|
[Core] Fail fast when /dev/shm is too small for the shm ring buffer (#48879)
Signed-off-by: Dr Andrea Tassi <andrea@verticular.uk>
Co-authored-by: Dr Andrea Tassi <andrea@verticular.uk>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-27 19:48:35 +00:00 |
|
 
|
fd10e8946d
|
[Test] Regression test for hybrid-Mamba eagle cache-peek in Mooncake connector (#43559) (#48361)
Signed-off-by: Rishi Puri <riship@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-27 19:03:58 +00:00 |
|
  
|
ed13deb376
|
[Bugfix][CPU] Fall back to torch for unaligned swigluoai on NEON/vec MoE (#49985)
Signed-off-by: oops-oom <73481342@qq.com>
Co-authored-by: oops-oom <73481342@qq.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-27 18:35:57 +00:00 |
|
 
|
99de48e98f
|
Fix MLA padding and grouped topk routing in the Transformers modelling backend (#49982)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-07-27 18:32:32 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
bf2b45b5d6
|
[Attention] Integrate FlashAttention 4 SM100 headdim 256 support (#42669)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-07-27 18:25:50 +00:00 |
|
 
|
8112b6c997
|
[MRV2] Always build attn metadata at capture time (#49364) (#49995)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-07-27 17:25:17 +00:00 |
|
 TobyJBellandGitHub
|
15d65f8669
|
[Bugfix] Changed speech to text chunk timestamp to cumulative approach (#41131)
Signed-off-by: Toby Bell <toby.bell1702@hotmail.co.uk>
|
2026-07-27 17:06:46 +00:00 |
|
 
|
e3c2fc3b3c
|
[Rust Frontend][gRPC] Add server and model discovery (#49491)
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-27 09:53:27 -07:00 |
|
 Nicolò LucchesiandGitHub
|
2b465b2c42
|
[Misc][PD] Nixl cleanup get_backend_aware_kv_block_len and virtually_split_kv_in_blocks (#49988)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
|
2026-07-27 18:51:37 +02:00 |
|
 yzong-rhandGitHub
|
3f47a8384d
|
[Bugfix] Fix VLLM_ENFORCE_STRICT_TOOL_CALLING mutation in tests (#49846)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-07-27 16:12:14 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
04502deca2
|
[Perf] Hash videos by source bytes (#49607)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
Signed-off-by: Guan-Ming Chiu <105915352+guan404ming@users.noreply.github.com>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 23:57:55 +08:00 |
|
  
|
d2ca3002d9
|
[MRV2][Performance] Skip no-op FP32 logits materialization (#47711)
Signed-off-by: jesse <szxfml@gmail.com>
Signed-off-by: Song Zhixin <szxfml@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-07-27 15:31:36 +00:00 |
|
 Umut PolatandGitHub
|
27d7061ef6
|
[Bugfix] Restore truncate_prompt_tokens for Jina rerank/score online (#49963)
Signed-off-by: Umut Polat <52835619+umut-polat@users.noreply.github.com>
|
2026-07-27 22:52:59 +08:00 |
|
 Guan-Ming ChiuandGitHub
|
ef9975d021
|
[Bugfix] Reject pipeline parallelism for DiffusionGemma (#45828)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-27 14:37:54 +00:00 |
|
 Roberto L. CastroandGitHub
|
56c96b0d91
|
[Perf] Tune LL BF16 Router GEMM (#48774)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com>
|
2026-07-27 10:25:37 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
59a6b0411d
|
[Core] Fix internal LB load-balancing (#49204)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 14:00:46 +00:00 |
|
 Rui "Garry" GaoandGitHub
|
dbccc5ae32
|
[Model] Enable EVS for Qwen3.5 (#48912)
Signed-off-by: Rui "Garry" Gao <garrygaogg@gmail.com>
|
2026-07-27 13:42:35 +00:00 |
|
 liminfei-amdandGitHub
|
a89015c6df
|
[Perf] Make merge attention context count a runtime argument (#48739)
Signed-off-by: liminfei-amd <91481003+liminfei-amd@users.noreply.github.com>
|
2026-07-27 13:24:42 +00:00 |
|
 neweyesandGitHub
|
96fa3f42c9
|
[Perf] Skip ll_bf16 router GEMM warmup for non-MoE models (#49659)
Signed-off-by: neweyes <328719365@qq.com>
|
2026-07-27 05:16:42 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
81962bb699
|
[Bugfix]Reject invalid FlashInfer MNNVL workspaces (#49043)
Signed-off-by: lengrongfu <lenronfu@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 08:12:23 -04:00 |
|
 Harry MellorandGitHub
|
92e8518d37
|
Improve Transformers modelling backend fx tracer (#49957)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-27 12:51:16 +01:00 |
|
 Ronen SchafferandGitHub
|
77cba0259f
|
[KV Offloading] Per-request tier filtering with TierFilter/TierMatcher (#48123)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-07-27 13:29:57 +03:00 |
|
 Andreas KaratzasandGitHub
|
30fbd05537
|
[ROCm] Use backend-default dot precision for ReplaySSM (#49909)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 18:11:06 +08:00 |
|
 
|
0906123953
|
[ROCm] [Model] Enable TML inkling (#48841)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-27 10:05:46 +00:00 |
|
  
|
bc3629b1c4
|
[ROCm][CI] Skip three torchao tests of gfx950 until torchao==0.18 is released (#49732)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-27 09:36:38 +00:00 |
|
 
|
312ea82e75
|
[CI][ROCm] Make hf-xet reconstruction safe on shared NFS (#49837)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-27 17:23:12 +08:00 |
|
 
|
394beb633b
|
[Bugfix][ROCm] Use batch DMA for CPU KV cache loads (#49843)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-27 02:18:01 -07:00 |
|
  ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
7f599d7854
|
[communication] [bugfix] fix quickreduce acc error in cudagraph mode (#46913)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 01:38:45 -07:00 |
|
 
|
eb290ab673
|
[Bugfix][CPU] Zero-pad MoE intermediate size for grouped-gemm TP alignment (#49591)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
2026-07-27 16:32:23 +08:00 |
|
 Andreas KaratzasandGitHub
|
cbc3a87200
|
[Tokenizer] Use HF config for HF tokenizers (#49907)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 07:49:11 +00:00 |
|
 liuzhenweiandGitHub
|
afc94523c9
|
[XPU][CI] Use platform device in InputBatch V2 test (#49939)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-07-27 15:28:02 +08:00 |
|
 
|
8061dc26bd
|
[Bugfix] Normalize sparse MLA warmup compression ratios (#49392)
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
2026-07-27 07:02:48 +00:00 |
|
 
|
fd9d2ede6f
|
[Rust Frontend] Keep --max-model-len engine-owned (#49944)
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-07-27 14:50:08 +08:00 |
|
 Andreas KaratzasandGitHub
|
e09900436c
|
[CI][ROCm] Reduce kernel test runtime (#49915)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 06:44:21 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
5d07e268b1
|
[Quantization][INC]Add MXFP8 Linear Support (#47514)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
Co-authored-by: Yi Liu <yi4.liu@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 14:26:31 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
d742856610
|
[3/N][Core][KV Connector] Support reliable partial-tail KV offload for sub-block prompts (#49502)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-26 23:22:17 -07:00 |
|
  
|
544cb724c8
|
[CPU][Spec Decode] Optimize GDN conv path for speculative decoding (#48577)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-27 06:20:14 +00:00 |
|
 
|
c314af1abf
|
[CPU][Perf] INT8 Fused MoE Kernel for Arm CPUs (#48637)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-27 05:53:09 +00:00 |
|
 
|
f19ee27e39
|
[Hardware][Power] Add FAST_EXP for Power (#49571)
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com>
Co-authored-by: Akash kaothalkar <akash.kaothalkar@ibm.com>
|
2026-07-27 05:47:29 +00:00 |
|
 Andreas KaratzasandGitHub
|
5f89a03dcb
|
[CI] Explicitly tear down speculative decode runners (#49910)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 05:32:46 +00:00 |
|
 
|
53397fbfac
|
[Bugfix][KV Offload][P2P] Fix EngineCore crash reconnecting to a reaped peer (#49823)
Signed-off-by: Jason Yao <wsyjh8@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-27 08:07:00 +03:00 |
|
 Andreas KaratzasandGitHub
|
49f31d7cee
|
[ROCm] Make vllm_c RMSNorm output contiguous (#49913)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-26 23:59:14 -05:00 |
|
 Nick HillandGitHub
|
74d3b799e1
|
[Bugfix] Fix mHC block-M prenorm GEMM cross-row reduction carry-over (#49429)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-27 04:42:47 +00:00 |
|
 
|
29fdeab254
|
[XPU][CI] Add more test cases in Intel GPU CI (#49422)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-27 12:40:22 +08:00 |
|
  
|
ff6173997d
|
[CI] Add kimi and k3 auto-labeling rules (#49895)
Signed-off-by: Joe Cotant <joe@inferact.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-07-26 21:00:27 -07:00 |
|
 
|
8de50e46d4
|
[Docs] Document NVFP4 GEMM kernel selection and Marlin weight-only fallback (#49376)
Signed-off-by: harjoth <harjoth.khara@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-07-26 20:53:28 -07:00 |
|
 Andreas KaratzasandGitHub
|
da99ffcc13
|
[ROCm][CI] Keep native datasets cache off shared NFS (#49516)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
|
2026-07-26 22:31:24 -05:00 |
|