 dependabot[bot]andGitHub
|
e898f259a3
|
Bump fsspec from 2024.12.0 to 2026.6.0
Bumps [fsspec](https://github.com/fsspec/filesystem_spec) from 2024.12.0 to 2026.6.0.
- [Commits](https://github.com/fsspec/filesystem_spec/compare/2024.12.0...2026.6.0)
---
updated-dependencies:
- dependency-name: fsspec
dependency-version: 2026.6.0
dependency-type: direct:production
update-type: version-update:semver-major
...
Signed-off-by: dependabot[bot] <support@github.com>
|
2026-07-14 04:58:47 +00:00 |
|
    
|
95aab66e95
|
[ROCm][MiniMax-M3][Spec Decode] Support speculative decode with AITER sparse PA (#47984)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-07-14 04:12:53 +00:00 |
|
 nemanjaudovicandGitHub
|
dcf4072da9
|
[Perf][ROCm] Fix GDN KKT warmup regression on RDNA by avoiding fp32 tl.dot (#45000)
Signed-off-by: Saeid Rostami <srostami@amd.com>
Signed-off-by: nemanjaudovic <nudovic@amd.com>
|
2026-07-13 20:48:54 -07:00 |
|
 
|
382bbd5144
|
[ROCm][Kernel] Add HybridW4A16LinearKernel: Triton prefill + HIP skinny decode (#40977)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-13 20:22:00 -07:00 |
|
 
|
b50ef9c6ed
|
[ROCm][MiniMax-M2] Dispatch fused QK-norm + AllReduce via AITER (#44849)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Pawel Kowalski <pawel.kowalski@amd.com>
|
2026-07-14 03:10:16 +00:00 |
|
 Dan BlanaruandGitHub
|
9e289c553c
|
up FI fp8 moe topk to 32 (#44462)
|
2026-07-14 02:58:16 +00:00 |
|
 
|
c4f5cd60da
|
[1/N] Add dense MHA path for sparse MLA short sequences (#47327)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-14 00:29:56 +00:00 |
|
 
|
0b0ef8d7eb
|
[Quantization][INC][ARK] Support INT2 XPU WOQ Linear (#47521)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-14 08:29:45 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
21472f32ea
|
add pad-aware swiglu limit kernel (#48287)
Signed-off-by: gnovack <novackgm@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-13 16:48:19 -07:00 |
|
 
|
fec64fea75
|
[BugFix] Correct OTEL span start time for Dynamo compilation (#40698)
Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai>
Co-authored-by: Simon Mo <simon.mo@hey.com>
|
2026-07-13 16:25:03 -07:00 |
|
 
|
8b8af2caf7
|
[Frontend] Expose logprob_token_ids on Python OpenAI endpoints (#43463)
Signed-off-by: Lang Zhao <lang.zhao@galileo.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-13 14:40:21 -07:00 |
|
 SnehlataandGitHub
|
7738ef35b8
|
[Feat] Add Support for BertForMaskedLM to vLLM (#48463)
Signed-off-by: atalhens <sneh.lata@nutanix.com>
|
2026-07-13 20:56:25 +00:00 |
|
 
|
9a21f0d1a3
|
[BugFix] Initialize model_config for Qwen3-VL MoE (#44863)
Signed-off-by: wenpengw-nv <wenpengw@nvidia.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-07-13 13:43:53 -07:00 |
|
 Nick HillandGitHub
|
8ac8375270
|
[Core] Preserve Marconi caching with selective hybrid cache retention (#47782)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-13 21:24:20 +01:00 |
|
 shanjiazandGitHub
|
7dc447dda7
|
Added sliding window attention support for qwen-eagle3 architecture (#47568)
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com>
|
2026-07-13 20:20:44 +00:00 |
|
 
|
7fc97042c3
|
Add DCP + Eagle support for Tokenspeed MLA backends (#48180)
Signed-off-by: Pavani Majety <pmajety@nvidia.com>
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com>
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com>
|
2026-07-13 11:46:02 -07:00 |
|
 Micah WilliamsonandGitHub
|
18c4067a54
|
[ROCm][CI] Unblock AMD: Language Models Test (Extended Pooling) (#48513)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-07-13 18:38:10 +00:00 |
|
 
|
550218b136
|
[Bugfix][Frontend] Flush engine reasoning parser at engine-reasoning → tool streaming boundary (#47606)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>
|
2026-07-13 14:06:10 -04:00 |
|
 Gavin MorrisandGitHub
|
5c342876a6
|
[Doc] Add DeepseekV32ForCausalLM to supported_models.md (#48293)
Signed-off-by: Gavin Morris <gmorriscs@gmail.com>
|
2026-07-13 17:43:59 +00:00 |
|
 
|
9427c45386
|
[ROCm][CI] Transformers: pass only one of input_ids/inputs_embeds (#48258)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-13 17:28:50 +00:00 |
|
 
|
43c8cbf79b
|
[EC Connector] CPU Offloading EC Connector (#47423)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
Signed-off-by: Or Ozeri <oro@il.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-13 20:09:41 +03:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
62286308c9
|
[Misc] Improve Matryoshka pooling dimensions validation (#48057)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-13 12:57:36 -04:00 |
|
 Nick HillandGitHub
|
26587f9519
|
[BugFix][ModelRunner V2] Fix stale attn metadata in speculator prefill cudagraph capture (#48261)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-13 09:39:15 -07:00 |
|
 
|
93e3bc8f30
|
[XPU][CI]Adjust timeout_in_minutes in Intel GPU CI (#48418)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-13 23:11:16 +08:00 |
|
 Yan MaandGitHub
|
c2c9f7c5e2
|
remove force channels_last in Idefics3MultiModalProcessor (#48467)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-07-13 14:18:57 +00:00 |
|
 Omer Ullman ArgovandGitHub
|
1be6e937b2
|
lower memory required for capturing cudagraphs for large cudagraph sizes (#48483)
Signed-off-by: Omer Ullman Argov <118735753+omera-nv@users.noreply.github.com>
|
2026-07-13 10:14:25 -04:00 |
|
 Wentao YeandGitHub
|
b3cfca996c
|
[Mypy Fix] Split mypy work (#48490)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-13 12:42:42 +00:00 |
|
 Bugen ZhaoandGitHub
|
487dfb3418
|
[CI] Add SPDX license header to Rust/Protobuf sources (#48472)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-07-13 10:22:47 +01:00 |
|
  
|
107a03ba63
|
[Core] Support fp32 lm_head for generation models via head_dtype (RFC #48305 §3.6) (#48390)
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-07-13 16:43:34 +08:00 |
|
 
|
56a357ed33
|
[Bugfix][KV Cache] Don't route uniform-page-size MLA+SWA models into DeepseekV4 packing (#48256)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-13 08:16:24 +00:00 |
|
  
|
bea70c7cfc
|
[Attention] Make sliding-window support an explicit backend capability (#48011)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-13 01:07:56 -07:00 |
|
 Mohammad Miadh AngkadandGitHub
|
75fe92a316
|
[Distributed][Perf] Enable FlashInfer MNNVL allreduce RMS quant fusion (#48064)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-13 15:02:59 +08:00 |
|
 
|
b7b58d1eba
|
[ROCm][CI] Cache Rust builds by source inputs (#46527)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-07-13 01:14:08 -05:00 |
|
 Canlin GuoandGitHub
|
36484e464a
|
[BugFix] Restore full tokens for Qwen MTP When MoE SP (#48429)
Signed-off-by: gcanlin <canlinguosdu@gmail.com>
|
2026-07-13 13:29:41 +08:00 |
|
 
|
9e57de7197
|
[CPU] Create Proper Numa topology for s390x (#40714)
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-13 12:58:43 +08:00 |
|
 Yejing LaiandGitHub
|
8c5dafcd09
|
[Bugfix][UT]Fix EagleMiniCPMForCausalLM meet TypeError (#48452)
Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
|
2026-07-13 04:37:23 +00:00 |
|
  
|
05fa8183a6
|
[CPU][Spec Decode] Support DFlash speculative decoding for GDN models on CPU (#46090)
Signed-off-by: guybd <guy.boudoukh@intel.com>
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-07-13 04:16:18 +00:00 |
|
 
|
d973cce3ca
|
Re-disable CUDA graph memory profiling on ROCm (#48440)
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-13 03:59:20 +00:00 |
|
 
|
775c1589ea
|
[Bugfix][ROCm] Keep TP all_gather on base-class collective (#48446)
Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-07-13 03:53:53 +00:00 |
|
 zztandGitHub
|
2595d5cebc
|
[Model] Optimize Qwen3.5 on H20 (#48350)
Signed-off-by: zzt <zengzetang.zzt@antgroup.com>
|
2026-07-13 03:30:48 +00:00 |
|
    
|
ee5a89f4d7
|
[ROCm][MiniMax-M3] Add AITER sparse paged attention (#47287)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-07-12 19:27:29 -07:00 |
|
 
|
e26264f3ef
|
[Kernel] Implement CUDA kernel for ReLUSquaredActivation (relu^2) (#39058)
Signed-off-by: Tanish Malekar <tanishmalekar32@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-07-12 19:18:03 -07:00 |
|
 AlexHuangandGitHub
|
4c81772e8b
|
[Bugfix][KV Offloading] Fix stale transfer_jobs after reset_cache + harden job completion (#48102)
Signed-off-by: Alex <alex.tech.lab@outlook.com>
|
2026-07-12 20:00:04 +03:00 |
|
 Bugen ZhaoandGitHub
|
27c3e579f0
|
[CI][Rust Frontend] Pin cargo tool versions (#48222)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-07-12 16:34:26 +01:00 |
|
 
|
8df14cfc8c
|
[EC Connector] Add EC Transfer Params (#42433)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-12 14:35:33 +03:00 |
|
 Jiangyun ZhuandGitHub
|
370b678a02
|
[CI][2/N] reduce CI time (#48394)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
|
2026-07-12 04:16:55 -07:00 |
|
 
|
5c0c987c03
|
Make tiering offload region DP-replica aware (#47987)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-12 13:10:21 +03:00 |
|
 Hugo CentenoandGitHub
|
5f8e73cb8b
|
[Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions (#48330)
Signed-off-by: hcenteno <hugo.centeno@estudiantat.upc.edu>
|
2026-07-12 09:39:27 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
83762b77b0
|
[Frontend] Add /abort_requests to the RLHF dev API router (#47173)
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-12 14:21:02 +08:00 |
|
 
|
a02984ed47
|
[Perf][Qwen] Replace MOE all-reduce with reduce-scatter (#47006)
Signed-off-by: gcanlin <canlinguosdu@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Co-authored-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-12 06:14:49 +00:00 |
|