 Ronen SchafferandGitHub
|
4bfa0f2b14
|
[KV Offload] Rename SecondaryTierManager.get_finished() to get_finished_jobs() (#43870)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-28 16:00:18 +00:00 |
|
 Vadim GimpelsonandGitHub
|
5d126dd155
|
[Bugfix] Exclude Ray DP from #42585's deferred port allocation (#43864)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
|
2026-05-28 15:55:14 +00:00 |
|
 
|
c08ebebf30
|
[Perf] Add do_not_specialize to Mamba SSD chunk kernels (#43803)
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com>
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com>
|
2026-05-28 15:40:02 +00:00 |
|
 Wentao YeandGitHub
|
be4062fd6c
|
[Bug] Fix tests/distributed/test_elastic_ep.py - assert False (#43813)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-28 11:00:56 -04:00 |
|
 
|
577d693838
|
[rust] fix: aggregate is_sleeping and reset_prefix_cache across DP engines (#43429)
Signed-off-by: Will.hou <1205157517@qq.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-28 07:56:56 -07:00 |
|
 Bugen ZhaoandGitHub
|
61a1e30473
|
[Rust Frontend] Reduce Gemma4 tool parser args scan complexity (#43850)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-28 14:52:29 +00:00 |
|
 Bugen ZhaoandGitHub
|
3a282230ee
|
[Rust Frontend] Add hy_v3 tool parser (#43872)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-28 14:42:47 +00:00 |
|
 Li, JiangandGitHub
|
20d69d100a
|
[CPU] Migrate cpu_awq into awq_marlin (#43841)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-05-28 22:36:31 +08:00 |
|
 Simon DanielssonandGitHub
|
552eb81918
|
[Bugfix][ROCm] Resolve MoRI connector hangs at high concurrency (#40344)
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
2026-05-28 14:30:21 +00:00 |
|
 Woosuk KwonandGitHub
|
9957e4d240
|
[Model Refactoring] Remove torch compile dependency in DSv4 (#43746)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-28 14:26:25 +00:00 |
|
 
|
864990e8d9
|
Add token-offset based selective offload in OffloadConnector (#39983)
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com>
Co-authored-by: Or Ozeri <or@ozery.com>
|
2026-05-28 14:11:02 +00:00 |
|
 
|
f3b2a819f7
|
[Perf][KDA] Fuse gate softplus, chunk-local cumsum, and RCP_LN2 scaling (#43667)
Signed-off-by: haojiangzheng <justineric096@gmail.com>
Co-authored-by: haojiangzheng <justineric096@gmail.com>
|
2026-05-28 13:47:08 +00:00 |
|
 Wentao YeandGitHub
|
64e1218673
|
[Perf] Optimize moe permute by pre-allocate buffer, 9~14% kernel performance improvement (#43014)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-28 06:18:26 -07:00 |
|
 Julien DenizeandGitHub
|
02606b0b09
|
[BUGFIX] Multimodal benchmark with MistralTokenizer (#42965)
Signed-off-by: juliendenize <julien.denize@mistral.ai>
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com>
|
2026-05-28 05:36:24 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
19af4e6dd4
|
Fix OlmoHybridForCausalLM not initialising (#43846)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-28 05:33:31 -07:00 |
|
 omerpaz95andGitHub
|
811d805195
|
[EC Connector] Add shutdown API to EC Connector. (#42423)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
|
2026-05-28 12:28:01 +00:00 |
|
 Vadim GimpelsonandGitHub
|
c1c4db8b4b
|
Log dummy DP step in iteration details (#41406)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>
|
2026-05-28 12:18:39 +00:00 |
|
 ChaunceyandGitHub
|
d692b89c2c
|
[Feature] Add structured output and effort support to Anthropic Messages API (#42396)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-05-28 12:06:48 +00:00 |
|
 Bugen ZhaoandGitHub
|
8e0580f4ee
|
[CI] Auto-apply rust label to relevant PRs (#43866)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-28 11:57:22 +00:00 |
|
 
|
61288b5458
|
[Bugfix] Fix HyperCLOVAX CI failure after upstream removed remote code (#43860)
Signed-off-by: Kevin Luu <kevin@inferact.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-28 03:37:36 -07:00 |
|
  
|
a583c84e2b
|
[Bugfix][ROCm] Fix Accuracy Drop in Sparse Indexer on gfx950 (#43781)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
|
2026-05-28 03:37:15 -07:00 |
|
  
|
4ec2817313
|
[Model][Bugfix] Rename weight_mapper to hf_to_vllm_mapper in LlamaNemotronVL pooling models (#43581)
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com>
Co-authored-by: opencode <noreply@opencode.ai>
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>
|
2026-05-28 03:32:22 -07:00 |
|
 Wei ZhaoandGitHub
|
f2caefe226
|
[UX] Increase DP Coordinator startup timeout from 30s to 120s (#42343)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-05-28 03:31:45 -07:00 |
|
 Animesh TrivediandGitHub
|
bfb9ebc211
|
[Feature] Add support for timed trace replay in vllm bench serve to replay Moonshot and Alibaba workload traces (#39795)
Signed-off-by: Animesh Trivedi <Animesh.Trivedi@ibm.com>
|
2026-05-28 03:31:34 -07:00 |
|
 Andreas KaratzasandGitHub
|
a9bc0ad8e4
|
[ROCm][CI] Move workload from MI300 to MI325 (#43824)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-28 03:31:29 -07:00 |
|
  
|
b372ad3e90
|
[Bugfix] Stream DeepSeek DSML tool-call argument deltas incrementally (#42879)
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
|
2026-05-28 17:50:23 +08:00 |
|
 Harry MellorandGitHub
|
2a781756a1
|
Restore Literal for WeightTransferConfig.backend (#43183)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-28 09:39:41 +00:00 |
|
 Woosuk KwonandGitHub
|
a04afd76aa
|
[DSV4] Remove AMD/XPU path in deepseek_v4/nvidia (#43829)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-05-28 08:00:52 +00:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
6cc8577421
|
[Kernel] Marlin MoE: include SM 12.x in default arch list (#40923)
Signed-off-by: Tony Liu <tonyliu0512@gmail.com>
Co-authored-by: Tony Liu <tonyliu0512@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-28 15:30:26 +08:00 |
|
 
|
d6b48f928f
|
[BugFix] Fix hard-coded timeout for multi-API-server startup (#43768)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-28 00:09:13 -07:00 |
|
 Rotem ShavittandGitHub
|
1b16f2ddc9
|
change name of fs_python secondary tier to fs. (#43600)
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com>
|
2026-05-28 07:05:48 +00:00 |
|
 TJianandGitHub
|
0ba46d4b11
|
[ROCm][DSV4] Enable Tilelang MHC replacing torch/triton mhc (#43679)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-05-28 07:05:28 +00:00 |
|
 JINO ROHITandGitHub
|
e1814f822d
|
minor docs: fix incorrect example path (#43830)
Signed-off-by: JINO-ROHIT <find.jinorohit@gmail.com>
|
2026-05-27 22:58:43 -07:00 |
|
 
|
7909f82a45
|
[Bugfix][Frontend] streaming tool-call serializer drops first args chunk when name and args share a DeltaMessage (#42683)
Signed-off-by: ignaciosica <mignacio.sica@gmail.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-28 05:20:55 +00:00 |
|
 Nick HillandGitHub
|
626fa9bba5
|
[BugFix] Fix blocked reasoning parsing with MRV2 (#43808)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-28 04:59:34 +00:00 |
|
 Thien TranandGitHub
|
e54eff769d
|
[Bugfix] Pass routed_scaling_factor to FlashInfer TRTLLM BF16 MoE (#43769)
|
2026-05-27 21:29:14 -07:00 |
|
 
|
05ac829629
|
fix: parse Qwen3 XML JSON arguments first (#43243)
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-05-28 03:35:59 +00:00 |
|
 Andreas KaratzasandGitHub
|
33e94fc3ad
|
[ROCm][CI] Stabilize Cargo cache and pre-test image checks (#43815)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-28 11:24:44 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
413ac5c070
|
[Misc][Rocm] Remove redundant AiterUnifiedAttentionBackend block size log (#43664)
Signed-off-by: NickLucche <nlucches@redhat.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-27 22:19:11 -05:00 |
|
 Yongye ZhuandGitHub
|
2d2c660104
|
[MoE] Remove inplace fused experts mechanism (#43727)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-27 20:00:19 -07:00 |
|
 Benjamin BartelsandGitHub
|
05eec7120e
|
Fix RunAI streamer tensor buffer reuse during weight loading (#43464)
Signed-off-by: bbartels <benjamin@bartels.dev>
|
2026-05-27 19:16:52 -07:00 |
|
 Bugen ZhaoandGitHub
|
c87f62ccf8
|
[Rust Frontend] Introduce mock engine for benchmark baseline (#43469)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-28 01:40:35 +00:00 |
|
 
|
1223732dda
|
[ModelRunnerV2][Hybrid model] Support kernel block size in hybrid model (#38831)
Signed-off-by: MengqingCao <cmq0113@163.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Mengqing Cao <cmq0113@163.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-28 00:55:55 +00:00 |
|
 amitz-nvandGitHub
|
381edde1b9
|
[Bugfix][Kernel] TRTLLM NVFP4 MoE chunking (#43599)
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com>
|
2026-05-28 00:36:21 +00:00 |
|
 Andreas KaratzasandGitHub
|
094124af15
|
Add @AndreasKaratzas to CODEOWNERS (#43740)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-27 16:14:50 -07:00 |
|
 Dakai AnandGitHub
|
5963c19478
|
Fix Qwen3-VL and Qwen3-omni-thinker accuracy degradation from deepstack inputs under torch.compile (#43617)
Signed-off-by: Dakai An <dakaian108@gmail.com>
|
2026-05-27 15:34:08 -07:00 |
|
 
|
7fb9c0197a
|
[Bugfix][DFlash]allocate the proper number of lookahead slots (#43733)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com>
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@gmail.com>
|
2026-05-27 21:45:34 +00:00 |
|
 Harry MellorandGitHub
|
2c2c966669
|
Validate against some config fields being set to 0 (#43794)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-27 21:14:49 +00:00 |
|
 Harry MellorandGitHub
|
2616f67faa
|
Remove Transformers forward/backward compatibility tests (#43785)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-27 12:46:36 -07:00 |
|
 
|
206b72c982
|
[Quantization] Fix Humming RoutedExperts import (#43540)
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-27 10:51:56 -07:00 |
|