7b5d60cc37
[Bugfix][V1] Clean up compiled-model bytecode hooks on VllmRunner exit ( #45195 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 20:31:17 -07:00
556b063e45
[XPU] Fix test_spec_decode_logprobs: use FLASH_ATTN for XPU in GPU_DETERMINISM_KWARGS ( #44468 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 11:07:04 +08:00
4bf699d310
[Kernel] Support DS Mamba tail copy for MTP align mode ( #45473 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Co-authored-by: Thomas Parnell <tom.parnell@gmail.com >
2026-06-16 22:50:30 +00:00
Stan Wozniak and GitHub
520828789c
Apply LRU policy only to proper cache entries ( #42656 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-16 21:49:15 +00:00
9d4dc4ca2f
[Kernel] Support GLM-5 dimensions for TRT-LLM ragged MLA prefill ( #43525 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 20:49:47 +00:00
Nick Hill and GitHub
d8d95998dc
[Core] Add prefill step cadence for better non-PD DP balancing ( #44558 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 13:17:18 -07:00
8e27a9c215
[PERF] Fuse multi-group block table staged writes ( #44944 )
...
Signed-off-by: jesse <szxfml@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 10:53:27 -07:00
44b2512767
[KV Connector][Mooncake] Add cache_prefix to namespace store keys ( #45767 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 10:24:20 -07:00
188c68798e
[KVConnector][MoRIIO] Allow overriding the advertised host IP ( #45488 )
...
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 17:18:37 +00:00
Carl Y and GitHub
eb04c769d3
feat: MLA prefill enable FA4 fp8 output ( #43050 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-06-16 07:10:59 -07:00
Hank Han and GitHub
d53f4593ce
[KV Connector][Mooncake] Pipeline-parallel support for PD-disaggregated serving with Mooncake connector ( #44528 )
...
Signed-off-by: hanhan.hank <hanhan.hank@bytedance.com >
Signed-off-by: Hank Han <hanhan7630@outlook.com >
2026-06-16 04:35:38 -07:00
ad32608e24
[MM][Perf][CG] Support dual-path ViT full CUDA graph for DeepSeek-OCR ( #43586 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 04:35:20 -07:00
7ad894c86a
[Bugfix] Prevent cuMemcpyBatchAsync segfault with MTP and KV offloading ( #44784 )
...
Signed-off-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-16 07:58:39 +00:00
d467a2a7f2
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer ( #45357 )
...
Signed-off-by: llx-08 <2596671364@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-15 21:36:09 +00:00
7e612a0f06
[KV Offloading] Implement reset_cache for TieringOffloadingManager ( #44541 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 18:42:53 +00:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
Yejing Lai and GitHub
9872921c5f
[XPU] skip UT test_with_ngram_gpu_spec_decoding ( #44423 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-06-15 08:46:30 +00:00
4ef4492e9b
[V1][Spec Decode] Add Dynamic SD ( #32374 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-14 00:14:27 -07:00
521b88c29e
[Bugfix] Reject structured outputs for diffusion decoders with a clear error ( #45468 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-13 12:04:01 -07:00
Juan Pérez de Algaba and GitHub
470229c37e
[Security] Fix DoS via prompt_embeds on M-RoPE models ( #45252 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-13 10:17:38 +00:00
2b3006076c
[Security] Add timeout guard for regex compilation in structured outp… ( #45118 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 09:52:56 +00:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
aab639c705
[Core][AMD] Propagate shutdown timeout to MultiprocExecutor ( #43154 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-12 15:13:31 -05:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
4171ae406c
[V1][Metrics] Add MLA attention metrics for DeepSeek MFU estimation ( #39457 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-06-12 14:28:40 +01:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
88ed636218
[KV Connector]: Support KV push from Prefill to Decode node using Nixl KV Connector ( #35264 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-06-12 10:38:41 +00:00
+1
eb28452b10
[Model] Add DiffusionGemma Support ( #45163 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Martin Kukla <martin.kukla@cantab.net >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Dipika Sikka <dsikka@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
Co-authored-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Alec Kohlhoff <134344302+aleckohlhoff@users.noreply.github.com >
Co-authored-by: Porras Huang <20535584+porrashuang@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: scoootscooob <167050519+scoootscooob@users.noreply.github.com >
2026-06-11 22:17:35 -07:00
Dao007forever and GitHub
6fbfdd1831
[NIXL] Per-region KV transfer classification for mixed full-attn + MLA groups ( #44583 )
2026-06-11 21:42:41 -07:00
b927004c44
[Bugfix] Mamba CPU Offloading ( #44599 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-11 21:07:35 -07:00
Nick Hill and GitHub
2263f8a3de
[CI][BugFix] Fix broken test_mamba_prefix_cache.py due to stale mock ( #45345 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 03:26:17 +00:00
4bc83323f2
[Bugfix] OffloadingConnector: respect skip_reading_prefix_cache flag ( #44592 )
...
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: littlecircle0730 <43994952+littlecircle0730@users.noreply.github.com >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-06-12 02:20:39 +00:00
8a91228dbe
[Bugfix][KVConnector][Mooncake] Close MooncakeDistributedStore on connector teardown ( #45206 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 14:33:48 -07:00
4085ff7cb4
[Core] Add kvcache watermark to reduce preemptions ( #44594 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-11 08:27:31 -07:00
Nicolò Lucchesi and GitHub
750aab5b8e
[Bugfix] Fix CPU memory leak related to not cleaning up old remotes data ( #44424 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-11 07:54:52 -07:00
55911db580
[PD][Core] Fix Mamba prefix cache hit rate in PD disaggregation ( #44243 )
...
Co-authored-by: lHrHenry233 <2381623149@qq.com >
Co-authored-by: underfituu <hzhucong@163.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-06-11 14:10:25 +00:00
ebc6ef971a
Hidden states extraction improvements ( #43805 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 09:44:45 -04:00
c3662b36ea
[KV offload] Parallel-agnostic fs-tier cache for single full-attention group ( #44733 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-06-11 15:48:37 +03:00
Yifan Qiao and GitHub
f272dfdce1
[KV Connector] Mooncake store: prefix-cache retention interval for sparse attention ( #44774 )
2026-06-10 21:36:34 -07:00
Stan Wozniak and GitHub
dc66e01a70
[Hybrid] Marconi-style admission policy for hybrid cache ( #37898 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-10 10:03:13 -07:00
0bae1d3848
[MRV2][Spec Decode] DFlash ( #44586 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-10 08:47:46 -07:00
af65e08fc5
KV-Cache multi-tier offloading async batched lookup ( #44193 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 14:59:30 +00:00
9dfc313bdc
Feature/offloading manager stats ( #35669 )
...
Signed-off-by: Sriusa4414@gmail.com
Signed-off-by: srinivas_oo7 <Sriusa4414@gmail.com >
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 12:44:55 +00:00
Juan Pérez de Algaba and GitHub
8a5cf1ccd6
[Security] Fix remote DoS via invalid recovered token reinjection ( #44744 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 02:31:43 -07:00
Andreas Karatzas and GitHub
82a42234be
[ROCm][CI] Defer AITER sampler import and isolate server test PYTHONPATH ( #44823 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 08:56:11 +00:00
c1d754d681
[Mooncake] Use all HCAs on multi-NIC hosts instead of GPU-indexed RNIC selection ( #43799 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-06-09 11:05:36 -07:00
Nicolò Lucchesi and GitHub
6690a0c4de
[PD][Bugfix] Fix KV Cache sharing with HMA ( #44629 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 06:10:06 -07:00
Nicolò Lucchesi and GitHub
dab60fc658
[Bugfix][CI] Fix test_offloading_connector.py::test_fs_tiering_offloading ( #44903 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 00:57:34 -07:00
Andreas Karatzas and GitHub
05cb606cad
[ROCm][CI] Re-route NixlConnector jobs ( #44809 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 18:57:11 -05:00
Wentao Ye and GitHub
2c27c294c0
[Model Runner V2] Fix mrv2 mm lora issue ( #44450 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-08 14:30:09 -04:00