 limewardandGitHub
|
ffc4f08c8e
|
[Core][KV-transfer] MoRIIO: heterogeneous TP<->DP prefill/decode read routing (#46116)
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io>
|
2026-07-27 02:10:00 +00:00 |
|
 Nick HillandGitHub
|
50aa830482
|
[BugFix][MRV2] Don't create dummy requests longer than max_model_len (#49751)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-27 02:04:20 +00:00 |
|
  
|
f0553889c0
|
[Bugfix] Prevent NaN poisoning in xpu_mla_sparse for fully-masked index chunks (#48366)
Signed-off-by: Nick Iusiumbeli <nickuspro@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-27 09:06:07 +08:00 |
|
 
|
fdaa0d9e59
|
[ModelRunner V2] Support encoder-only attention (#49331)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-27 00:38:44 +00:00 |
|
 Schwinn SaereesitthipitakandGitHub
|
b5b61c622c
|
[Core][Distributed] Add process-checkpoint lifecycle hooks for communicators (starting with Flashinfer) (#46877)
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
|
2026-07-26 14:47:50 -04:00 |
|
 Taneem IbrahimandGitHub
|
0da6e7f3d6
|
[Bugfix] Reject contradictory custom-op directives (#49134)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-26 08:42:14 -04:00 |
|
  
|
da3a252fd1
|
[KVOffload][P2P] Generic P2P secondary tier: peer lookup and serving via ParentManager (#48021)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Signed-off-by: liranschour <liranschour@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 11:45:47 +03:00 |
|
 Guan-Ming ChiuandGitHub
|
8d28b48d01
|
[Perf] Isolate MM preprocessing on its own executor (#49524)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-26 08:04:17 +00:00 |
|
 Taneem IbrahimandGitHub
|
0164022c90
|
[CI] Fix speech correctness check rejecting improved WER (#49853)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-07-26 06:24:07 +00:00 |
|
 
|
7eca0e1a64
|
[KV Offload] Deduplicate replicated MLA KV in the shared CPU region (#48906)
Signed-off-by: Change72 <changg@nvidia.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-26 08:22:30 +03:00 |
|
 
|
7a29a3c54c
|
[Bugfix][KV Offload] Namespace persistent cache by model runner (#49440)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-26 08:21:52 +03:00 |
|
 Athrael SojuandGitHub
|
1240c74c0a
|
[Bugfix] Respect declared attention contract for ColQwen3.5 retrievers (#49372)
Signed-off-by: Athrael Soju <athrael.soju@gmail.com>
|
2026-07-26 04:08:07 +00:00 |
|
 Chang GuoandGitHub
|
7a6a5b3667
|
[CI] Compute speech WER directly with jiwer (#49773)
|
2026-07-25 20:51:21 -04:00 |
|
 
|
0111002323
|
[Kernel] TD operand loads for batched MoE GEMM (moe_mmk) on XPU (#46340)
Signed-off-by: oonyshch <xonyshch@gmail.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-26 08:50:07 +08:00 |
|
 
|
d30b1ecd1b
|
[Bugfix][KV Offloading] Defer request finalization until final store (#49671)
Signed-off-by: Rui Yin <2260891073@qq.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-25 20:58:09 +00:00 |
|
 
|
70009fb934
|
[MM][CG] Support ViT CUDA Graph for Gemma-4 (#46837)
Signed-off-by: Anthony Su <xsuanthony@gmail.com>
Co-authored-by: Linkun Chen <github@lkchen.net>
|
2026-07-25 15:02:09 -05:00 |
|
 Taneem IbrahimandGitHub
|
6b0103d1c9
|
[CI] Stabilize Pooling Rerank Equivalence Test (#49822)
|
2026-07-25 14:36:50 -04:00 |
|
 
|
9321aff536
|
[Bugfix] Wait for the linear bias before layerwise online processing (#49805)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-25 17:56:32 +00:00 |
|
 Harry MellorandGitHub
|
26d725c334
|
[Model] Add VaultGemma via Transformers modeling backend (#49803)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-25 16:54:15 +00:00 |
|
+1        
|
0b0bd2b5f6
|
[Feature] Add fault tolerance framework (simplified) for DP+EP external LB deployments (#44428)
Signed-off-by: fangyuchu <fangyuchu@qq.com>
Signed-off-by: a798347923 <2645302020@qq.com>
Signed-off-by: TianZhuo <2770730562@qq.com>
Signed-off-by: a798347923 <39047817+a798347923@users.noreply.github.com>
Signed-off-by: 205150940 <112750056+205150940@users.noreply.github.com>
Signed-off-by: w00689259 <wangzhuo66@huawei.com>
Signed-off-by: zWaNg3 <37772915+zWaNg3@users.noreply.github.com>
Signed-off-by: zWaNg3 <389750525@qq.com>
Signed-off-by: yzchang-plus <1078477584@qq.com>
Signed-off-by: Jade Zheng <zheng.shoujian@outlook.com>
Co-authored-by: zWaNg3 <37772915+zWaNg3@users.noreply.github.com>
Co-authored-by: a798347923 <2645302020@qq.com>
Co-authored-by: TianZhuo <2770730562@qq.com>
Co-authored-by: 205150940 <112750056+205150940@users.noreply.github.com>
Co-authored-by: a798347923 <39047817+a798347923@users.noreply.github.com>
Co-authored-by: w00689259 <wangzhuo66@huawei.com>
Co-authored-by: zWaNg3 <389750525@qq.com>
Co-authored-by: yzchang-plus <1078477584@qq.com>
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com>
|
2026-07-25 11:49:09 -04:00 |
|
 
|
1423569ff5
|
[Bugfix][Tool Parser] Fix dropped streaming arguments in Jamba and InternLM2 parsers (#48852)
Signed-off-by: mosya415 <263250241+mosya415@users.noreply.github.com>
Co-authored-by: mosya415 <263250241+mosya415@users.noreply.github.com>
|
2026-07-25 09:34:43 -04:00 |
|
 Harry MellorandGitHub
|
9a50464698
|
[CI] Stop flaky test from downloading model every time (#49800)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-25 13:29:36 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
b9b6306ebe
|
feat[vLLM × v5]: Add audio support for the Transformers backend (#39330)
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-25 04:20:58 -07:00 |
|
 
|
ca0defa343
|
Make bare hugging_face imports forbidden (#49726)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-25 04:18:39 -07:00 |
|
 
|
0b1a8bb1f6
|
[Bugfix][CI] Fix stale Mooncake lookup expectation broken by a merge race (#49802)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-25 03:10:26 -07:00 |
|
 
|
fe5145765f
|
[Core] Keep attention backends eligible for text-only serving of prefix-LM models (#48796)
Signed-off-by: qtris123 <voquangtri2021@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-07-25 03:06:14 -07:00 |
|
 
|
dbcc1cdd0a
|
[Model] Remove Ouro (#49786)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-25 02:50:08 -07:00 |
|
 
|
94682b79f4
|
[multimodal] Make PyNvVideoCodec decoder concurrency configurable (#49753)
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-24 23:31:41 -07:00 |
|
   
|
0ba2aa35a8
|
Stabilize GPU memory teardown between ROCm CI tests (#49242)
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-25 04:21:36 +00:00 |
|
 Divakar VermaandGitHub
|
aaaeda98dc
|
[CI] fix compile test | refactor VLLM_DISABLE_COMPILE_CACHE for tests (#49770)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-07-24 23:15:19 -05:00 |
|
 
|
70052fb924
|
[Bugfix][KV Connector][Mooncake] Keep TP-sharded Mamba state out of the KV-head dedup (#49499)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-24 21:02:22 -07:00 |
|
 
|
6a1acac3fe
|
[BUGFIX] Fix log capture in KV test (#49655)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-25 00:36:53 +00:00 |
|
 Woosuk KwonandGitHub
|
213f681f81
|
Revert "[Perf][GLM-5.2] Blackwell decode optimizations" (#49768)
|
2026-07-24 16:04:59 -07:00 |
|
 Aarushi JainandGitHub
|
33c4f3551c
|
[ROCm][CI] Wait for ROCm VRAM to settle between compiled and eager LL… (#49739)
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
|
2026-07-24 17:19:46 -05:00 |
|
  
|
7513d071bd
|
[ROCm][CI] Fix XPASS(strict) on mixed audio embeds test (#49733)
Signed-off-by: Djordje Ramic <djoramic@amd.com>
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-24 16:09:18 -05:00 |
|
 Harry MellorandGitHub
|
89f6aa3a9e
|
[KV Offload][CI] Fall back to buffered I/O without O_DIRECT; fix flaky api-server test (#49734)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-24 13:56:57 -07:00 |
|
 Harry MellorandGitHub
|
84d26b9ee3
|
[Model] Remove Plamo2 (#49729)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-24 13:54:24 -07:00 |
|
 
|
9e6746b3c7
|
[CI] Stabilize memory-sensitive compile and structured output tests (#49749)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-24 16:45:09 -04:00 |
|
 
|
972848f276
|
[Bugfix] Support non-uniform page sizes in KVBlockZeroer (#49704)
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-07-24 13:38:59 -07:00 |
|
    
|
5d8e90a966
|
[WideEP] Update NCCL to 2.30.7 to enable DeepEPv2 in the vllm/vllm-openai image (#45321)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Signed-off-by: Tyler Michael Smith <tyler@vllm.ai>
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Ilya Markov <ilmarkov@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Codex <noreply@openai.com>
|
2026-07-24 13:00:02 -07:00 |
|
 
|
c064fa52b6
|
Fix GLM-4.1V video placeholder token ID handling. (#49484)
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-24 12:35:16 -07:00 |
|
 Andreas KaratzasandGitHub
|
9863102ed9
|
[CI] Reuse loaded config for cached tokenizer (#49509)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
|
2026-07-24 12:32:40 -07:00 |
|
 
|
8c13ee5735
|
Add sm_107 for Rubin (#49387)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
|
2026-07-24 11:59:49 -07:00 |
|
  ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
866fea2b99
|
[Kernel] ReplaySSM: cache SSM inputs for faster Mamba2 standard decode (#48018)
Signed-off-by: Johnny-Liou <a897111@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-07-24 09:39:49 -07:00 |
|
 
|
d02df748bf
|
[Bugfix] Accept RFC 2397 parameters in base64 data URLs (#48973)
Signed-off-by: Thomas Fahrner <thomas.fahrner@parasail.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 08:23:08 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
7b40fb9645
|
[UX] Reject incompatible nested runtime overrides (#49247)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-24 07:16:22 -07:00 |
|
 BadrBasowidandGitHub
|
8eac21a602
|
[ROCM] Fix AITER Fused AllReduce RMSNorm for Transformers Backend (#49673)
Signed-off-by: BadrBasowid <badr.basowid@gmail.com>
|
2026-07-24 07:16:17 -07:00 |
|
 
|
a454a1dd25
|
[Bugfix][Benchmarks] Restore --skip-tokenizer-init with custom dataset (#49180)
Signed-off-by: Michele Gazzetti <michele.gazzetti1@ibm.com>
Co-authored-by: Kevin H. Luu <khluu000@gmail.com>
|
2026-07-24 04:30:19 -07:00 |
|
  
|
833483f357
|
Encoder cache extension hooks (#48218)
Signed-off-by: hotTea <958436561@qq.com>
Signed-off-by: hanxi-java <634498162@qq.com>
Co-authored-by: hanxi-java <634498162@qq.com>
Co-authored-by: 韩熙 <63780107+hanxi-java@users.noreply.github.com>
|
2026-07-24 02:52:00 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
589a5b884b
|
[PD][NixlPush][Bugfix] Fix blocking handshake call on writer thread (#49221)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-24 01:41:37 -07:00 |
|