 Or OzeriandGitHub
|
357fddf614
|
[kv_offload]: Add DSv4 support (#43142)
Signed-off-by: Or Ozeri <oro@il.ibm.com>
|
2026-05-24 11:10:12 +03:00 |
|
 
|
0902d8e62f
|
[KV Connector] Keep MooncakeStore full hits block-aligned (#43494)
Signed-off-by: Dao Le <daole@inferact.ai>
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-23 23:15:03 -07:00 |
|
 Dao007foreverandGitHub
|
819c610f9b
|
[Mooncake] Add metrics for MooncakeStoreConnector operations (#43392)
|
2026-05-23 13:34:40 -07:00 |
|
  
|
5bb8d2767a
|
[Kernel] Batch invariant NVFP4 linear using cutlass (#39912)
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-23 09:41:12 -04:00 |
|
 Gabriel WuandGitHub
|
82536acc54
|
Keep scheduler alive for delayed KV connector frees (#43433)
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com>
|
2026-05-23 06:23:32 +00:00 |
|
 
|
d19db10974
|
[Bugfix] Fix native Triton top-k/top-p kernel assumes contiguous logi… (#42739)
Signed-off-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com>
|
2026-05-22 22:56:16 -07:00 |
|
 
|
84e351555a
|
[Bugfix] Auto-raise max_num_batched_tokens for prefix-LM multimodal models (#43051)
Signed-off-by: Ashwin Giridharan <girida@amazon.com>
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com>
|
2026-05-22 21:23:50 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4e2eba28be
|
[Perf] Optimize hidden state extraction logic (#37374)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-22 18:23:08 -04:00 |
|
 
|
b21f3d56d4
|
[KV Connector] MooncakeStore: don't co-queue save with load to avoid double delayed-free (#43371)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-22 16:14:11 +00:00 |
|
 Isotr0pyandGitHub
|
f0feb15e7f
|
[Multimodal] Simplify ViT CUDA graph interfaces (#41234)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-22 22:31:00 +08:00 |
|
 Weida HongandGitHub
|
6bb8753db1
|
Correcting the mock classes for MM GC tests (#43321)
Signed-off-by: Weida Hong <wdhongtw@google.com>
|
2026-05-22 15:21:35 +08:00 |
|
 
|
5ea76fa89a
|
[CI] Fix test_lora_with_spec_decode on V2 model runner (#43314)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-22 14:24:18 +08:00 |
|
 Lanze LiuandGitHub
|
39d5fa96a7
|
[Bugfix] Zero stale is_prefilling in padded CUDA graph rows for Mamba (#41873)
Signed-off-by: Lanze Liu <lanzetech@gmail.com>
|
2026-05-21 15:42:42 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
b730c46352
|
[Perf] [Hybrid] Fused Triton kernel for GPU-side Mamba state postprocessing (#40172)
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-21 04:50:54 -07:00 |
|
 
|
f2d5e3d3ae
|
[CI] Lower granite-4.0-h-tiny gsm8k threshold for Hybrid SSM NixlConnector PD accuracy tests (4 GPUs) (#43186)
Signed-off-by: haosdent <haosdent@gmail.com>
Signed-off-by: NickLucche <nlucches@redhat.com>
Co-authored-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 17:00:24 +00:00 |
|
 
|
ded871201a
|
[Bug][Structured Outputs] Fix bug that leads to unconstrained generations with structural tags (#42452)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-20 07:08:58 -07:00 |
|
 KebeandGitHub
|
19cf334207
|
[Feature] Support manually enabling the cumem allocator (#33648)
Signed-off-by: Kebe <mail@kebe7jun.com>
|
2026-05-20 08:58:30 -04:00 |
|
 Ronen SchafferandGitHub
|
4f940896a3
|
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers (#43076)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-20 03:32:08 +00:00 |
|
 Benjamin ChislettandGitHub
|
c628a93a64
|
[Perf][Bugfix] Update dflash aux layer indexing (#40727)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-05-19 20:15:57 -07:00 |
|
 Nick HillandGitHub
|
b82e908b4c
|
[Perf][4/n] Eliminate various GPU<->CPU syncs (#42347)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-19 10:35:54 -04:00 |
|
 
|
129019f334
|
[CI] Add MTP + PD disagg test for Qwen3.5 (#42677)
Signed-off-by: ZhanqiuHu <zhu@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-19 11:44:33 +02:00 |
|
 Yifan QiaoandGitHub
|
056bc2e166
|
[KVConnector][DSV4] HMA support for Mooncake store connector (#42828)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-05-19 01:07:46 -07:00 |
|
  
|
fab07e4d0f
|
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B (#42289)
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist <noreply@google.com>
|
2026-05-18 21:22:33 -07:00 |
|
 Ronen SchafferandGitHub
|
84747489de
|
Tier offload followup (#42529)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-18 19:41:58 +00:00 |
|
 Netanel HaberandGitHub
|
47829b1159
|
[Bugfix] mamba: run single-token extends as decodes (#42430)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
|
2026-05-18 15:26:00 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
e5417657e5
|
[KV Connector][Offloading] Flush all pending jobs on last step (#42611)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Signed-off-by: liranschour <liranschour@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-18 12:59:42 +00:00 |
|
 roikoren755andGitHub
|
737bfa3a43
|
[Bugfix][Hybrid][NemotronH] Fix mamba_cache_mode=all + speculative decoding crash (#41233)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-05-18 14:54:00 +03:00 |
|
 
|
107210442d
|
[CI] Add NIXL EP import canary (#42567)
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-05-17 19:11:46 -07:00 |
|
  
|
36e74c9ea4
|
[KV Connector] Support disk offloading in MooncakeStoreConnector (#42689)
Signed-off-by: Zhewen Li <zhewenli@inferact.ai>
Co-authored-by: Zhewen Li <zhewenli@inferact.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-16 13:34:15 -07:00 |
|
 Jiangyun ZhuandGitHub
|
8a56da3845
|
[Experimental] Breakable CUDA graph (#42304)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
|
2026-05-16 22:04:12 +08:00 |
|
 Yifan QiaoandGitHub
|
4b364f810e
|
[Core][DSV4] Skip caching SWA blocks that can never serve a prefix-cache hit (#42258)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-05-15 15:59:18 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
bf610c2f56
|
[Bugfix] Fix inverted condition causing thinking_token_budget to be silently ignored (#41674)
Signed-off-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-15 12:48:49 +08:00 |
|
 Matthew BonanniandGitHub
|
9898f94abe
|
[Attention] Remove deprecated MLA prefill arguments (#42555)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-14 10:34:06 -07:00 |
|
 Baorun (Lauren) MuandGitHub
|
a7737cb4f3
|
[Fix] Misc Fixes in ViT CUDA Graph (#38040)
Signed-off-by: Baorun Mu <bmu@nvidia.com>
|
2026-05-14 23:49:06 +08:00 |
|
 
|
24337fb860
|
PD disagg with NIXL Connector: GDN support (Qwen3.5) (#41869)
Signed-off-by: Zhanqiu Hu <zhu@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-14 16:33:01 +02:00 |
|
   
|
c7560af424
|
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support (#39568)
Signed-off-by: xhx1022 <1737006628@qq.com>
Signed-off-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com>
|
2026-05-14 14:12:30 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
5bd8c71e79
|
[kv_offload] Implement reset_cache() for the offloading connector (#41956)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
|
2026-05-14 16:00:10 +03:00 |
|
 
|
0d2732dd91
|
[MLA Attention Backend] Add TOKENSPEED_MLA backend for DSR1/Kimi K25 prefill + decode on Blackwell (#41778)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Roger Wang <hey@rogerw.io>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-13 23:48:02 -07:00 |
|
 Siddharth BedekarandGitHub
|
f51f6844f9
|
[Bugfix][Spec Decode] Wire draft_probs into probabilistic draft_model rejection (#40269)
|
2026-05-13 21:04:03 -04:00 |
|
 liangel-02andGitHub
|
6b5c389ee3
|
expose flex block size for batch invariant mode (#41252)
Signed-off-by: Angel Li <liangel@meta.com>
|
2026-05-13 14:11:57 -07:00 |
|
 
|
2f821faeae
|
[Spec Decode] Support hybrid attention models in extract_hidden_states (#39949)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-13 10:45:53 -07:00 |
|
 Wentao YeandGitHub
|
e35c0d4c63
|
[Feature] Support compile mode for batch invariance on SM80 (#42456)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-13 11:02:39 -04:00 |
|
 Ronen SchafferandGitHub
|
11f6b545d4
|
[kv_offload] Add multi-tier KV cache offloading framework (#40020)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-13 17:21:43 +03:00 |
|
 Ronen SchafferandGitHub
|
79fd1bc7ed
|
[kv_offload] Add req_id to ReqContext for per-request tracking (#42507)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-13 11:11:10 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
13bf242100
|
[Feat][KVConnector] Add bind_gpu_block_pool() to KVConnectorBase_V1 (#39654)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-13 02:10:29 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9ce74042d3
|
[Bugfix][SimpleCPUOffloadBackend] Dedup in-flight CPU offload stores across scheduler steps (#41289)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-13 01:53:32 -07:00 |
|
 Nicolò LucchesiandGitHub
|
71bcd02ef3
|
[Bugfix][PD] Fix multi-node TP (TP>8) (#39907)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-12 22:20:57 -07:00 |
|
+4        
|
ebeb09d822
|
[KV Transfer] Add MooncakeStoreConnector for KV cache offloading via Mooncake distributed store (#40900)
Signed-off-by: leichao.lc <leichao.lc@antgroup.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: leichao.lc <leichao.lc@antgroup.com>
Co-authored-by: ivanium <yifanqiao@inferact.ai>
Co-authored-by: aoshen524 <aoshen@inferact.ai>
Co-authored-by: Dao007forever <daole@inferact.ai>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
Co-authored-by: foraxe <1055696449@qq.com>
Co-authored-by: Skywalker-EP <173423846@qq.com>
Co-authored-by: fems14 <1804143737@qq.com>
Co-authored-by: jianzs <zheng.shoujian@outlook.com>
Co-authored-by: baxingpiaochong <771405853@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-12 16:09:10 -07:00 |
|
 Nick HillandGitHub
|
fe8b42e80c
|
[CI] Fix test_async_scheduling.py flakiness (#42455)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-12 21:38:32 +00:00 |
|
 Giancarlo DelfinandGitHub
|
fe5b4e0fe7
|
[Model Runner V2] Apply synthetic mode to probabilistic rejection sampler (#41035)
|
2026-05-12 13:37:03 -07:00 |
|