  
|
bc3629b1c4
|
[ROCm][CI] Skip three torchao tests of gfx950 until torchao==0.18 is released (#49732)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-27 09:36:38 +00:00 |
|
 
|
312ea82e75
|
[CI][ROCm] Make hf-xet reconstruction safe on shared NFS (#49837)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-27 17:23:12 +08:00 |
|
 
|
394beb633b
|
[Bugfix][ROCm] Use batch DMA for CPU KV cache loads (#49843)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-27 02:18:01 -07:00 |
|
  ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
7f599d7854
|
[communication] [bugfix] fix quickreduce acc error in cudagraph mode (#46913)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 01:38:45 -07:00 |
|
 
|
eb290ab673
|
[Bugfix][CPU] Zero-pad MoE intermediate size for grouped-gemm TP alignment (#49591)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
2026-07-27 16:32:23 +08:00 |
|
 Andreas KaratzasandGitHub
|
cbc3a87200
|
[Tokenizer] Use HF config for HF tokenizers (#49907)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 07:49:11 +00:00 |
|
 liuzhenweiandGitHub
|
afc94523c9
|
[XPU][CI] Use platform device in InputBatch V2 test (#49939)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-07-27 15:28:02 +08:00 |
|
 
|
8061dc26bd
|
[Bugfix] Normalize sparse MLA warmup compression ratios (#49392)
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
2026-07-27 07:02:48 +00:00 |
|
 
|
fd9d2ede6f
|
[Rust Frontend] Keep --max-model-len engine-owned (#49944)
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-07-27 14:50:08 +08:00 |
|
 Andreas KaratzasandGitHub
|
e09900436c
|
[CI][ROCm] Reduce kernel test runtime (#49915)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 06:44:21 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
5d07e268b1
|
[Quantization][INC]Add MXFP8 Linear Support (#47514)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
Co-authored-by: Yi Liu <yi4.liu@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 14:26:31 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
d742856610
|
[3/N][Core][KV Connector] Support reliable partial-tail KV offload for sub-block prompts (#49502)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-26 23:22:17 -07:00 |
|
  
|
544cb724c8
|
[CPU][Spec Decode] Optimize GDN conv path for speculative decoding (#48577)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: Codex <codex@openai.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-27 06:20:14 +00:00 |
|
 
|
c314af1abf
|
[CPU][Perf] INT8 Fused MoE Kernel for Arm CPUs (#48637)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-27 05:53:09 +00:00 |
|
 
|
f19ee27e39
|
[Hardware][Power] Add FAST_EXP for Power (#49571)
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com>
Co-authored-by: Akash kaothalkar <akash.kaothalkar@ibm.com>
|
2026-07-27 05:47:29 +00:00 |
|
 Andreas KaratzasandGitHub
|
5f89a03dcb
|
[CI] Explicitly tear down speculative decode runners (#49910)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 05:32:46 +00:00 |
|
 
|
53397fbfac
|
[Bugfix][KV Offload][P2P] Fix EngineCore crash reconnecting to a reaped peer (#49823)
Signed-off-by: Jason Yao <wsyjh8@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-27 08:07:00 +03:00 |
|
 Andreas KaratzasandGitHub
|
49f31d7cee
|
[ROCm] Make vllm_c RMSNorm output contiguous (#49913)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-26 23:59:14 -05:00 |
|
 Nick HillandGitHub
|
74d3b799e1
|
[Bugfix] Fix mHC block-M prenorm GEMM cross-row reduction carry-over (#49429)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-27 04:42:47 +00:00 |
|
 
|
29fdeab254
|
[XPU][CI] Add more test cases in Intel GPU CI (#49422)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-27 12:40:22 +08:00 |
|
  
|
ff6173997d
|
[CI] Add kimi and k3 auto-labeling rules (#49895)
Signed-off-by: Joe Cotant <joe@inferact.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-07-26 21:00:27 -07:00 |
|
 
|
8de50e46d4
|
[Docs] Document NVFP4 GEMM kernel selection and Marlin weight-only fallback (#49376)
Signed-off-by: harjoth <harjoth.khara@gmail.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-07-26 20:53:28 -07:00 |
|
 Andreas KaratzasandGitHub
|
da99ffcc13
|
[ROCm][CI] Keep native datasets cache off shared NFS (#49516)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
|
2026-07-26 22:31:24 -05:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
8040ef2426
|
[Frontend] expose stream_interval as req sampling param (#49754)
Signed-off-by: walterbm <walter.beller.morales@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-27 11:30:37 +08:00 |
|
 
|
bf4f633b4c
|
[XPU] Enable QK Norm + RoPE fusion pass on XPU (#49394)
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-27 03:22:29 +00:00 |
|
 Andreas KaratzasandGitHub
|
854c33f380
|
[CI][ROCm] Keep global GPU memory cleanup opt-in (#49911)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-26 22:19:55 -05:00 |
|
 Andreas KaratzasandGitHub
|
ac87549cbd
|
[CI][ROCm] Reduce V1 attention test runtime (#49916)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-27 11:12:45 +08:00 |
|
 
|
439f336212
|
[Core] Fix gpu<->cpu syncs in MRV2 mamba_hybrid.py (#49736)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-07-27 02:41:27 +00:00 |
|
 limewardandGitHub
|
ffc4f08c8e
|
[Core][KV-transfer] MoRIIO: heterogeneous TP<->DP prefill/decode read routing (#46116)
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io>
|
2026-07-27 02:10:00 +00:00 |
|
 Nick HillandGitHub
|
50aa830482
|
[BugFix][MRV2] Don't create dummy requests longer than max_model_len (#49751)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-27 02:04:20 +00:00 |
|
  
|
f0553889c0
|
[Bugfix] Prevent NaN poisoning in xpu_mla_sparse for fully-masked index chunks (#48366)
Signed-off-by: Nick Iusiumbeli <nickuspro@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-27 09:06:07 +08:00 |
|
 
|
fdaa0d9e59
|
[ModelRunner V2] Support encoder-only attention (#49331)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-27 00:38:44 +00:00 |
|
 
|
0934b26790
|
[CI/Build] Refresh tags before building macOS wheel (#49901)
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-26 13:07:32 -07:00 |
|
 
|
9e50e1037e
|
[Bugfix][CuMem] Make KV-cache wake cleanup tag-safe (#49857)
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
|
2026-07-26 12:09:13 -07:00 |
|
 Schwinn SaereesitthipitakandGitHub
|
b5b61c622c
|
[Core][Distributed] Add process-checkpoint lifecycle hooks for communicators (starting with Flashinfer) (#46877)
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
|
2026-07-26 14:47:50 -04:00 |
|
  
|
b68d7ef262
|
[Bugfix][KV Offload] Namespace auto cache dtype by effective dtype (#49438)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 20:59:09 +03:00 |
|
 
|
7154856f3d
|
[Bugfix] Fix handling 5D KV cache in kv_postprocess_layout_on_receive (#47791)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-26 22:22:42 +08:00 |
|
  
|
3f1d40960f
|
[KV Offload] Fix num_tokens_after_batch for different termination types (#49285)
Signed-off-by: Alex <jihuihuang@example.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
Signed-off-by: Alex <alex.tech.lab@outlook.com>
Signed-off-by: Alex <jihuihuang@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 16:22:58 +03:00 |
|
 Taneem IbrahimandGitHub
|
0da6e7f3d6
|
[Bugfix] Reject contradictory custom-op directives (#49134)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-26 08:42:14 -04:00 |
|
    
|
5559679229
|
[Bugfix][KV Offload] Bound unaligned SWA loads by physical GPU blocks (#49052)
Signed-off-by: Colton Ottley <colton@ottleyengineering.com>
Co-authored-by: Colton Ottley <colton@ottleyengineering.com>
Co-authored-by: jasl <jasl9187@hotmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 14:53:32 +03:00 |
|
  
|
da3a252fd1
|
[KVOffload][P2P] Generic P2P secondary tier: peer lookup and serving via ParentManager (#48021)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Signed-off-by: liranschour <liranschour@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 11:45:47 +03:00 |
|
 Guan-Ming ChiuandGitHub
|
21fd9e85a0
|
[Model] Support top_k and top_p sampling for DiffusionGemma (#45429)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-26 08:39:25 +00:00 |
|
 Guan-Ming ChiuandGitHub
|
8d28b48d01
|
[Perf] Isolate MM preprocessing on its own executor (#49524)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-26 08:04:17 +00:00 |
|
 Taneem IbrahimandGitHub
|
0164022c90
|
[CI] Fix speech correctness check rejecting improved WER (#49853)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-26 06:24:07 +00:00 |
|
 
|
30b0714031
|
[Perf] DeepSeek-OCR-2 TTFT Optimize (#49531)
Signed-off-by: RED <outofthewoods@qq.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
|
2026-07-26 05:53:06 +00:00 |
|
 Nils MattesonandGitHub
|
2e860de498
|
[Doc] Add compile cache volume example to the Docker deployment page (#49782)
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
|
2026-07-26 05:24:01 +00:00 |
|
 
|
7eca0e1a64
|
[KV Offload] Deduplicate replicated MLA KV in the shared CPU region (#48906)
Signed-off-by: Change72 <changg@nvidia.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-26 08:22:30 +03:00 |
|
 
|
7a29a3c54c
|
[Bugfix][KV Offload] Namespace persistent cache by model runner (#49440)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 08:21:52 +03:00 |
|
 Athrael SojuandGitHub
|
1240c74c0a
|
[Bugfix] Respect declared attention contract for ColQwen3.5 retrievers (#49372)
Signed-off-by: Athrael Soju <athrael.soju@gmail.com>
|
2026-07-26 04:08:07 +00:00 |
|
 
|
48ebd6f2f1
|
[Bugfix][KVConnector] Disable cross-layer KV blocks for per-token-head quant (#49226)
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 05:24:27 +03:00 |
|