 Artem PerevedentsevandGitHub
|
1cb224430b
|
[GDN] Enable FI Blackwell GDN prefill kernel (#40717)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-05-20 01:46:55 -07:00 |
|
 Harry MellorandGitHub
|
9b343dd4f5
|
Enable mermaid diagrams in the docs (#43192)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-20 08:10:00 +00:00 |
|
  
|
07aeaf9d4d
|
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) (#42663)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-20 00:18:12 -07:00 |
|
 Nicolò LucchesiandGitHub
|
40651c0207
|
[Docs][PD][NIXL] Bidirectional kv-cache transfer (#43097)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 09:02:36 +02:00 |
|
 Nicolò LucchesiandGitHub
|
7e4bc2cecb
|
[Docs][PD][NIXL] Lease extension mechanism for blocks on P (#43099)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 08:58:25 +02:00 |
|
 Kevin H. LuuandGitHub
|
85959567c3
|
[ci] Revert model executor test back to L4 (#43188)
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
|
2026-05-19 23:01:41 -07:00 |
|
 Ronen SchafferandGitHub
|
4f940896a3
|
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers (#43076)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-20 03:32:08 +00:00 |
|
 Michael GoinandGitHub
|
cd0ff26e7a
|
[CI] Add DSV4-Flash to gsm8k moe-refactor/config-b200.txt (#42111)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-05-19 20:21:01 -07:00 |
|
 Izik GolanandGitHub
|
2ae910ed88
|
[Perf] Avoid forward scan for async output placeholders (#42938)
|
2026-05-19 20:16:07 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
fadf5d332c
|
add enqueue all option to throughput benchmark (#42975)
Signed-off-by: Philip Maybank <pmaybank@amd.com>
Signed-off-by: pmaybank <113125070+pmaybank@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-19 20:16:02 -07:00 |
|
 Benjamin ChislettandGitHub
|
c628a93a64
|
[Perf][Bugfix] Update dflash aux layer indexing (#40727)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-05-19 20:15:57 -07:00 |
|
 Terrence ZhaoandGitHub
|
5774aaed0c
|
[Cohere] Enable Cohere MoE (#43143)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
|
2026-05-19 19:32:06 -07:00 |
|
 Nick HillandGitHub
|
39bba710be
|
[MRV2][BugFix] Fix default-stream CG capture in P/W LoRA case (#43160)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-19 19:19:05 -07:00 |
|
 Aaron HaoandGitHub
|
73dd2f33b7
|
[bug] fix WeightTransferConfig.backend to allow for all strings (#43121)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
|
2026-05-19 21:01:29 -04:00 |
|
 Fadi ArafehandGitHub
|
be16785998
|
[CPU][DOC] Fix installation commands for Arm CPUs (#43115)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
|
2026-05-19 23:31:15 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
117afeea46
|
Fix error in Dynamic NTK scaling (#41277)
Signed-off-by: Max de Bayser <mbayser@br.ibm.com>
Signed-off-by: Max de Bayser <maxdebayser@gmail.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-19 17:27:54 -04:00 |
|
 Doğaç EldenkandGitHub
|
1242196295
|
[Model] Support post-norm architecture for EAGLE-3 supeculators (#42764)
Signed-off-by: Doğaç Eldenk <dogacel@gmail.com>
|
2026-05-19 13:39:00 -07:00 |
|
 Kevin H. LuuandGitHub
|
a65093c1a3
|
[ci] Move language models tests (hybrid) back to L4 (#43129)
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
|
2026-05-19 11:51:34 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9aaf83ef50
|
[CI failure] Temporarily disable using persistent cache for flashinfer autotune (#43119)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-19 11:44:32 -07:00 |
|
 tomeras91andGitHub
|
f54721bcc3
|
[Bugfix][MoE] FlashInfer one-sided: workspace union across heterogeneous layers (#42976)
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com>
|
2026-05-19 14:43:04 -04:00 |
|
 
|
aed2eb355a
|
[Docs] Fix MooncakeStoreConnector role in disaggregated example (#42994)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-19 11:14:43 -07:00 |
|
 Dom BrownandGitHub
|
d247a931cc
|
[feat] Add FP8 per-tensor Q scale support to Triton attention backend (#42080)
Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com>
|
2026-05-19 09:02:05 -07:00 |
|
 Jinzhen LinandGitHub
|
8200fbe1ac
|
[Misc] add humming to dependencies (#42540)
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com>
|
2026-05-19 08:36:47 -07:00 |
|
 Flora FengandGitHub
|
42b4f1fdf7
|
[Refactor] Extract extract_types_from_schema utility from Minimax M2 tool parser (#43025)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-19 11:21:12 -04:00 |
|
 Wang YiwenandGitHub
|
1c6158083a
|
[Model] Openvla support (#42654)
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com>
|
2026-05-19 08:17:42 -07:00 |
|
 Xinyu ChenandGitHub
|
d740e2c029
|
[XPU] update xpu graph usage (#43043)
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com>
|
2026-05-19 23:09:07 +08:00 |
|
 Nick HillandGitHub
|
b82e908b4c
|
[Perf][4/n] Eliminate various GPU<->CPU syncs (#42347)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-19 10:35:54 -04:00 |
|
 SageandGitHub
|
a78b842d0e
|
[Bugfix] Fix top logprobs token placeholders in /inference/v1/generate (#42887)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
|
2026-05-19 10:21:49 +00:00 |
|
 
|
129019f334
|
[CI] Add MTP + PD disagg test for Qwen3.5 (#42677)
Signed-off-by: ZhanqiuHu <zhu@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-19 11:44:33 +02:00 |
|
 Shanshan ShenandGitHub
|
ef54a4d604
|
[Misc][MM] Remove redundant code in CLIPAttention (#43046)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-05-19 08:43:16 +00:00 |
|
 Woosuk KwonandGitHub
|
07beaed842
|
[Model Refactoring] Rename deepseek_v4.py to model.py [4/N] (#43077)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-19 01:12:46 -07:00 |
|
 Yifan QiaoandGitHub
|
056bc2e166
|
[KVConnector][DSV4] HMA support for Mooncake store connector (#42828)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-05-19 01:07:46 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
f34623bf3c
|
[bug] AsyncScheduler drops first post-resume token after pause_generation + clear_cache (#42117)
Signed-off-by: hao-aaron <ahao@anyscale.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-19 01:06:21 -07:00 |
|
 Woosuk KwonandGitHub
|
b14be81c1f
|
[Model Refactoring] Move deepseek_v4_ops to models/deepseek_v4 [3/N] (#43073)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-19 00:52:54 -07:00 |
|
 wang.yuqiandGitHub
|
301d986473
|
[Frontend] Consolidate beam search by BeamSearchMixin. (#42946)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-19 07:37:40 +00:00 |
|
  ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
257af77bc2
|
[Docs] Reorganize online serving docs. (#41907)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-19 14:43:18 +08:00 |
|
 Taneem IbrahimandGitHub
|
4a4fdabe28
|
[Misc] Aligning tokwise pooler heads for consistency (#43041)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-05-19 06:16:42 +00:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
f1e3f0e6d6
|
[XPU] Use custom op collective behavior (#41354)
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-19 14:14:59 +08:00 |
|
  
|
9fd8487d2f
|
[Docs] Add SVG images for pooling models. (#42626)
Signed-off-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-18 22:50:38 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
27f4ba9481
|
fix: use keyword arguments for shard_id and expert_id in weight_loade… (#42671)
Signed-off-by: junyanxu <junyanxu5513@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-19 05:29:04 +00:00 |
|
 
|
6e889b582b
|
[ci] Route 28 gpu_1_queue tests to h200_35gb queue (#43030)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-18 21:58:36 -07:00 |
|
  
|
fab07e4d0f
|
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B (#42289)
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist <noreply@google.com>
|
2026-05-18 21:22:33 -07:00 |
|
 
|
3ca8db2ef8
|
add cutedsl dsv4 indexer fp8 kernel (#42899)
Signed-off-by: george <george@inferact.ai>
Co-authored-by: george <george@inferact.ai>
|
2026-05-18 21:17:56 -07:00 |
|
 Woosuk KwonandGitHub
|
87b08c5f64
|
[Model Refactoring] Move DeepSeek V4 layers to models/deepseek_v4/ [2/N] (#43039)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-18 21:00:58 -07:00 |
|
 
|
fba010dd74
|
[Bugfix][MRV2] Fix KVCache tensor explicit kernel_block_size dim (#42766)
Signed-off-by: NickLucche <nlucches@redhat.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-18 20:25:41 -07:00 |
|
 Mohammad Miadh AngkadandGitHub
|
da03e549b3
|
[UX] Add a persistent cache for FlashInfer autotuning (#42537)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-18 20:25:37 -07:00 |
|
 Kunshang JiandGitHub
|
36dcaf25d8
|
[XPU] add gptq(int4) support (#37844)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-19 11:17:09 +08:00 |
|
 Ofir ZafrirandGitHub
|
8f16c4a5c0
|
[BugFix][CPU][Spec Decode] Fix Eagle implementation on CPU backend (#42468)
Signed-off-by: Ofir Zafrir <ofir.zafrir@intel.com>
|
2026-05-19 03:16:07 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
afd7b1dce9
|
[Bugfix] Use platform-agnostic device in example_connector load (#42926)
Signed-off-by: Revital Sur <eres@il.ibm.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-19 03:12:04 +00:00 |
|
 Woosuk KwonandGitHub
|
287471b994
|
[Model Refactoring] Migrate DeepSeek V4 to vllm/models/ [1/N] (#43004)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-18 19:50:02 -07:00 |
|