2e2e626b40
[Bugfix] Count per-group blocks in get_max_concurrency_for_kv_cache_config ( #48317 )
...
Signed-off-by: David Orman <ormandj@corenode.com >
Co-authored-by: Luke Alonso <lalonso@gmail.com >
Co-authored-by: Martin Vit <martin@voipmonitor.org >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-21 00:29:10 +00:00
yzong-rh and GitHub
a287eb163f
[Front-end] [Messages] Populate num_cache_creation_tokens ( #48535 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-07-18 13:04:35 -04:00
5a65ba5f17
[Refactor] Move iteration logging to the frontend ( #46647 )
...
Signed-off-by: maxyanghu <hyoung2991@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Shang Wang <shangw@nvidia.com >
2026-07-15 17:59:05 -07:00
32aef44388
[Bugfix] Include inline per-token-head scales in offloaded page transfer width ( #48411 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-07-14 16:07:26 +03:00
Nick Hill and GitHub
8ac8375270
[Core] Preserve Marconi caching with selective hybrid cache retention ( #47782 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-13 21:24:20 +01:00
56a357ed33
[Bugfix][KV Cache] Don't route uniform-page-size MLA+SWA models into DeepseekV4 packing ( #48256 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-13 08:16:24 +00:00
8df14cfc8c
[EC Connector] Add EC Transfer Params ( #42433 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-12 14:35:33 +03:00
481e481be7
[2/N][Core] support partial prefix cache hit for hybrid model ( #46384 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-07-12 05:37:51 +00:00
e5588e49bc
[Core][KV events] Report prefix-cache-reused blocks in full report mode ( #45261 )
...
Signed-off-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Lei Gong <gonglei25@huawei.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 22:46:54 -07:00
cd0de48d08
[Bugfix][V1] Free out-of-window blocks on the processed-token basis under async scheduling ( #47728 )
...
Signed-off-by: Saddss <28726669061@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Saddss <28726669061@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 13:34:19 +01:00
55da232db6
[Bugfix] Pad Mamba page size instead of scaling block_size in unify_kv_cache_spec_page_size ( #45207 )
...
Signed-off-by: Sahil170595 <147995121+Sahil170595@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 22:01:34 +00:00
93e2ab7111
Disable dynamic speculative decoding when DP is enabled ( #45963 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 12:52:41 +00:00
c5b66233b2
[Bugfix][Spec Decode] Skip uniform spec-decode padding for diffusion models ( #47464 )
...
Signed-off-by: kl527 <kl527@cornell.edu >
Signed-off-by: Kyung Sub Lee (Daniel) <66861800+kl527@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 09:25:12 +00:00
373eb314af
[Bugfix][Core] Fix num_output_placeholders underflow with async scheduling + spec decode ( #46066 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-06 13:50:38 +00:00
e7c9df9449
[Bugfix][Structured Output][Spec Decode] Constrain bitmask and trim grammar advance at the reasoning boundary ( #44297 )
...
Signed-off-by: Allen.Yu <yuyue0225sc@163.com >
Signed-off-by: yue.yu <yuyue0225sc@163.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-07-04 09:08:45 +00:00
Nick Hill and GitHub
e392bf7a68
[BugFix][MRV2] Ensure all req slots are accounted for when scheduling ( #46974 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:24 -07:00
c6741b2ad4
[Model] Support Unlimited OCR ( #46564 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-27 23:09:18 -07:00
Nick Hill and GitHub
658b54efe4
[ModelRunner V2] Update scheduler tests to cover MRV2 paths ( #46771 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 09:36:31 -07:00
d490b98162
[Core] Avoid mixed length specdec batches via padding ( #45237 )
...
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
2026-06-25 08:34:44 -07:00
Lucas Wilkinson and GitHub
e7df232288
[KV Offload] Gate packed HMA KV cache on cross-layer config ( #46252 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-06-24 11:55:30 -04:00
6af0559ddb
[Core][DP] Throttle prefills based on local prefill work ( #46532 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 02:27:12 +00:00
430a95ae3a
[v1][kvcache] Honor prefix-cache retention interval for Mamba/linear attention ( #45845 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 19:51:11 -07:00
3e6529cc0e
[Bugfix][Spec Decode] Fix EAGLE drafter multimodal encoder cache misses ( #46315 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 18:14:02 +00:00
6bc6f2d86d
[1/N][Core] add partial prefix cache primitives ( #45939 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 23:43:10 -07:00
2cac89f9da
[Spec Decode] Support mixed KV page sizes for DFlash ( #45181 )
...
Signed-off-by: Alex Steiner <asteiner@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 22:45:14 +08:00
6e919960af
[Perf] Skip/shrink all_token_ids copy in scheduler for non-async and V2 runner ( #45840 )
...
Signed-off-by: amanchugh89 <amanchugh.89@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 22:36:57 +00:00
cc22621b51
[KV Offload] Support packed HMA KV cache layout ( #46205 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:19:40 +00:00
01192139bf
[DSv4] Pack KV caches into contiguous per-block allocations for DeepSeek V4 ( #44577 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-19 12:55:42 -04:00
Stan Wozniak and GitHub
520828789c
Apply LRU policy only to proper cache entries ( #42656 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-16 21:49:15 +00:00
Nick Hill and GitHub
d8d95998dc
[Core] Add prefill step cadence for better non-PD DP balancing ( #44558 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 13:17:18 -07:00
d467a2a7f2
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer ( #45357 )
...
Signed-off-by: llx-08 <2596671364@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-15 21:36:09 +00:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
4085ff7cb4
[Core] Add kvcache watermark to reduce preemptions ( #44594 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-11 08:27:31 -07:00
Stan Wozniak and GitHub
dc66e01a70
[Hybrid] Marconi-style admission policy for hybrid cache ( #37898 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-10 10:03:13 -07:00
303916e93d
[Bugfix]: Fix assertion in MambaManager.allocate_slots() ( #39562 )
...
Signed-off-by: Holworth <kangqihan17@mails.ucas.ac.cn >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-08 00:34:37 -04:00
Nick Hill and GitHub
3b3d5287fa
[BugFix] Resolve multiple async kv load deadlock ( #44560 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-06 23:05:47 +00:00
a6183563b6
[Prefix Caching] DeepSeekv4 - Support selective prefix-cache retention for sliding-window KV cache ( #43447 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-04 00:48:31 -07:00
0c6631f02a
[KVCache] Support Pluggable KVCacheSpec ( #37505 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 09:05:16 -07:00
Andy Lo and GitHub
95b1615ec9
[Perf] Improve multimodal item handling from O(n) to O(log n) per step ( #44212 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-06-03 11:00:26 +00:00
Yifan Qiao and GitHub
e9e08c49b9
[Bugfix] Cache the EAGLE/MTP lookahead block in the SWA prefix-cache mask ( #44082 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 12:21:07 -07:00
d247a9dc13
[EC Connector] Non blocking EC Connector lookup ( #41627 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 08:48:25 +00:00
Yifan Qiao and GitHub
7c37096620
[Core][Refactor]: thread scheduler_block_size into KVCacheManager and KVCacheCoordinator ( #44165 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 01:14:44 -07:00
1fc2cee50a
[KVConnector][Mooncake] Wire reset_cache cascade end-to-end ( #42694 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Signed-off-by: Ao Shen <aoshen@inferact.ai >
Co-authored-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 20:52:35 -07:00
Gabriel Wu and GitHub
82536acc54
Keep scheduler alive for delayed KV connector frees ( #43433 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-05-23 06:23:32 +00:00
Yifan Qiao and GitHub
4b364f810e
[Core][DSV4] Skip caching SWA blocks that can never serve a prefix-cache hit ( #42258 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-15 15:59:18 +08:00
c7560af424
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support ( #39568 )
...
Signed-off-by: xhx1022 <1737006628@qq.com >
Signed-off-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com >
2026-05-14 14:12:30 +00:00
Yan Ru Pei and GitHub
bcb9c133ba
feat(kv-events): emit KV cache metadata ( #40984 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com >
2026-05-12 15:58:48 +00:00