Commit Graph
18709 Commits
Author SHA1 Message Date
dependabot[bot]andGitHub e898f259a3 Bump fsspec from 2024.12.0 to 2026.6.0
Bumps [fsspec](https://github.com/fsspec/filesystem_spec) from 2024.12.0 to 2026.6.0.
- [Commits](https://github.com/fsspec/filesystem_spec/compare/2024.12.0...2026.6.0)

---
updated-dependencies:
- dependency-name: fsspec
  dependency-version: 2026.6.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-14 04:58:47 +00:00
95aab66e95 [ROCm][MiniMax-M3][Spec Decode] Support speculative decode with AITER sparse PA (#47984)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
2026-07-14 04:12:53 +00:00
nemanjaudovicandGitHub dcf4072da9 [Perf][ROCm] Fix GDN KKT warmup regression on RDNA by avoiding fp32 tl.dot (#45000)
Signed-off-by: Saeid Rostami <srostami@amd.com>
Signed-off-by: nemanjaudovic <nudovic@amd.com>
2026-07-13 20:48:54 -07:00
382bbd5144 [ROCm][Kernel] Add HybridW4A16LinearKernel: Triton prefill + HIP skinny decode (#40977)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-07-13 20:22:00 -07:00
b50ef9c6ed [ROCm][MiniMax-M2] Dispatch fused QK-norm + AllReduce via AITER (#44849)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Pawel Kowalski <pawel.kowalski@amd.com>
2026-07-14 03:10:16 +00:00
Dan BlanaruandGitHub 9e289c553c up FI fp8 moe topk to 32 (#44462) 2026-07-14 02:58:16 +00:00
c4f5cd60da [1/N] Add dense MHA path for sparse MLA short sequences (#47327)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 00:29:56 +00:00
0b0ef8d7eb [Quantization][INC][ARK] Support INT2 XPU WOQ Linear (#47521)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-14 08:29:45 +08:00
gnovackGitHubmergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
21472f32ea add pad-aware swiglu limit kernel (#48287)
Signed-off-by: gnovack <novackgm@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 16:48:19 -07:00
fec64fea75 [BugFix] Correct OTEL span start time for Dynamo compilation (#40698)
Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai>
Co-authored-by: Simon Mo <simon.mo@hey.com>
2026-07-13 16:25:03 -07:00
8b8af2caf7 [Frontend] Expose logprob_token_ids on Python OpenAI endpoints (#43463)
Signed-off-by: Lang Zhao <lang.zhao@galileo.ai>
Co-authored-by: Claude <noreply@anthropic.com>
2026-07-13 14:40:21 -07:00
SnehlataandGitHub 7738ef35b8 [Feat] Add Support for BertForMaskedLM to vLLM (#48463)
Signed-off-by: atalhens <sneh.lata@nutanix.com>
2026-07-13 20:56:25 +00:00
9a21f0d1a3 [BugFix] Initialize model_config for Qwen3-VL MoE (#44863)
Signed-off-by: wenpengw-nv <wenpengw@nvidia.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
2026-07-13 13:43:53 -07:00
Nick HillandGitHub 8ac8375270 [Core] Preserve Marconi caching with selective hybrid cache retention (#47782)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
2026-07-13 21:24:20 +01:00
shanjiazandGitHub 7dc447dda7 Added sliding window attention support for qwen-eagle3 architecture (#47568)
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com>
2026-07-13 20:20:44 +00:00
7fc97042c3 Add DCP + Eagle support for Tokenspeed MLA backends (#48180)
Signed-off-by: Pavani Majety <pmajety@nvidia.com>
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com>
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com>
2026-07-13 11:46:02 -07:00
Micah WilliamsonandGitHub 18c4067a54 [ROCm][CI] Unblock AMD: Language Models Test (Extended Pooling) (#48513)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
2026-07-13 18:38:10 +00:00
550218b136 [Bugfix][Frontend] Flush engine reasoning parser at engine-reasoning → tool streaming boundary (#47606)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>
2026-07-13 14:06:10 -04:00
Gavin MorrisandGitHub 5c342876a6 [Doc] Add DeepseekV32ForCausalLM to supported_models.md (#48293)
Signed-off-by: Gavin Morris <gmorriscs@gmail.com>
2026-07-13 17:43:59 +00:00
9427c45386 [ROCm][CI] Transformers: pass only one of input_ids/inputs_embeds (#48258)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-13 17:28:50 +00:00
43c8cbf79b [EC Connector] CPU Offloading EC Connector (#47423)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
Signed-off-by: Or Ozeri <oro@il.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-13 20:09:41 +03:00
Taneem IbrahimGitHubmergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
62286308c9 [Misc] Improve Matryoshka pooling dimensions validation (#48057)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-13 12:57:36 -04:00
Nick HillandGitHub 26587f9519 [BugFix][ModelRunner V2] Fix stale attn metadata in speculator prefill cudagraph capture (#48261)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
2026-07-13 09:39:15 -07:00
93e3bc8f30 [XPU][CI]Adjust timeout_in_minutes in Intel GPU CI (#48418)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-13 23:11:16 +08:00
Yan MaandGitHub c2c9f7c5e2 remove force channels_last in Idefics3MultiModalProcessor (#48467)
Signed-off-by: Yan Ma <yan.ma@intel.com>
2026-07-13 14:18:57 +00:00
Omer Ullman ArgovandGitHub 1be6e937b2 lower memory required for capturing cudagraphs for large cudagraph sizes (#48483)
Signed-off-by: Omer Ullman Argov <118735753+omera-nv@users.noreply.github.com>
2026-07-13 10:14:25 -04:00
Wentao YeandGitHub b3cfca996c [Mypy Fix] Split mypy work (#48490)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
2026-07-13 12:42:42 +00:00
Bugen ZhaoandGitHub 487dfb3418 [CI] Add SPDX license header to Rust/Protobuf sources (#48472)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
2026-07-13 10:22:47 +01:00
107a03ba63 [Core] Support fp32 lm_head for generation models via head_dtype (RFC #48305 §3.6) (#48390)
Signed-off-by: Karthik Kothuri <karthikkothuri2009@gmail.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
2026-07-13 16:43:34 +08:00
56a357ed33 [Bugfix][KV Cache] Don't route uniform-page-size MLA+SWA models into DeepseekV4 packing (#48256)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude <noreply@anthropic.com>
2026-07-13 08:16:24 +00:00
bea70c7cfc [Attention] Make sliding-window support an explicit backend capability (#48011)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-07-13 01:07:56 -07:00
Mohammad Miadh AngkadandGitHub 75fe92a316 [Distributed][Perf] Enable FlashInfer MNNVL allreduce RMS quant fusion (#48064)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-07-13 15:02:59 +08:00
b7b58d1eba [ROCm][CI] Cache Rust builds by source inputs (#46527)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
2026-07-13 01:14:08 -05:00
Canlin GuoandGitHub 36484e464a [BugFix] Restore full tokens for Qwen MTP When MoE SP (#48429)
Signed-off-by: gcanlin <canlinguosdu@gmail.com>
2026-07-13 13:29:41 +08:00
Rehan KhanGitHubLi, Jiang <jiang1.li@intel.com>
9e57de7197 [CPU] Create Proper Numa topology for s390x (#40714)
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
2026-07-13 12:58:43 +08:00
Yejing LaiandGitHub 8c5dafcd09 [Bugfix][UT]Fix EagleMiniCPMForCausalLM meet TypeError (#48452)
Signed-off-by: Lai, Yejing <yejing.lai@intel.com>
2026-07-13 04:37:23 +00:00
guybdGitHubCursorLi, Jiang <jiang1.li@intel.com>
05fa8183a6 [CPU][Spec Decode] Support DFlash speculative decoding for GDN models on CPU (#46090)
Signed-off-by: guybd <guy.boudoukh@intel.com>
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
2026-07-13 04:16:18 +00:00
d973cce3ca Re-disable CUDA graph memory profiling on ROCm (#48440)
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 03:59:20 +00:00
775c1589ea [Bugfix][ROCm] Keep TP all_gather on base-class collective (#48446)
Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-13 03:53:53 +00:00
zztandGitHub 2595d5cebc [Model] Optimize Qwen3.5 on H20 (#48350)
Signed-off-by: zzt <zengzetang.zzt@antgroup.com>
2026-07-13 03:30:48 +00:00
ee5a89f4d7 [ROCm][MiniMax-M3] Add AITER sparse paged attention (#47287)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
2026-07-12 19:27:29 -07:00
e26264f3ef [Kernel] Implement CUDA kernel for ReLUSquaredActivation (relu^2) (#39058)
Signed-off-by: Tanish Malekar <tanishmalekar32@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-12 19:18:03 -07:00
AlexHuangandGitHub 4c81772e8b [Bugfix][KV Offloading] Fix stale transfer_jobs after reset_cache + harden job completion (#48102)
Signed-off-by: Alex <alex.tech.lab@outlook.com>
2026-07-12 20:00:04 +03:00
Bugen ZhaoandGitHub 27c3e579f0 [CI][Rust Frontend] Pin cargo tool versions (#48222)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
2026-07-12 16:34:26 +01:00
8df14cfc8c [EC Connector] Add EC Transfer Params (#42433)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-12 14:35:33 +03:00
Jiangyun ZhuandGitHub 370b678a02 [CI][2/N] reduce CI time (#48394)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
2026-07-12 04:16:55 -07:00
5c0c987c03 Make tiering offload region DP-replica aware (#47987)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-12 13:10:21 +03:00
Hugo CentenoandGitHub 5f8e73cb8b [Bugfix] Guard mixed-dtype allreduce RMSNorm quant fusions (#48330)
Signed-off-by: hcenteno <hugo.centeno@estudiantat.upc.edu>
2026-07-12 09:39:27 +00:00
aoshen02GitHubClaude Opus 4.8mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
83762b77b0 [Frontend] Add /abort_requests to the RLHF dev API router (#47173)
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-12 14:21:02 +08:00
a02984ed47 [Perf][Qwen] Replace MOE all-reduce with reduce-scatter (#47006)
Signed-off-by: gcanlin <canlinguosdu@gmail.com>
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Co-authored-by: yewentao256 <zhyanwentao@126.com>
2026-07-12 06:14:49 +00:00