Commit Graph
19137 Commits
Author SHA1 Message Date
limewardandGitHub ffc4f08c8e [Core][KV-transfer] MoRIIO: heterogeneous TP<->DP prefill/decode read routing (#46116)
Signed-off-by: Edwin Lim <edwin.lim@mangoboost.io>
2026-07-27 02:10:00 +00:00
Nick HillandGitHub 50aa830482 [BugFix][MRV2] Don't create dummy requests longer than max_model_len (#49751)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
2026-07-27 02:04:20 +00:00
f0553889c0 [Bugfix] Prevent NaN poisoning in xpu_mla_sparse for fully-masked index chunks (#48366)
Signed-off-by: Nick Iusiumbeli <nickuspro@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-27 09:06:07 +08:00
fdaa0d9e59 [ModelRunner V2] Support encoder-only attention (#49331)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
2026-07-27 00:38:44 +00:00
0934b26790 [CI/Build] Refresh tags before building macOS wheel (#49901)
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-07-26 13:07:32 -07:00
9e50e1037e [Bugfix][CuMem] Make KV-cache wake cleanup tag-safe (#49857)
Signed-off-by: aoshen02 <aoshen02@users.noreply.github.com>
Co-authored-by: aoshen02 <aoshen02@users.noreply.github.com>
2026-07-26 12:09:13 -07:00
Schwinn SaereesitthipitakandGitHub b5b61c622c [Core][Distributed] Add process-checkpoint lifecycle hooks for communicators (starting with Flashinfer) (#46877)
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com>
2026-07-26 14:47:50 -04:00
b68d7ef262 [Bugfix][KV Offload] Namespace auto cache dtype by effective dtype (#49438)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 20:59:09 +03:00
7154856f3d [Bugfix] Fix handling 5D KV cache in kv_postprocess_layout_on_receive (#47791)
Signed-off-by: Daniel Socek <daniel.socek@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-26 22:22:42 +08:00
3f1d40960f [KV Offload] Fix num_tokens_after_batch for different termination types (#49285)
Signed-off-by: Alex <jihuihuang@example.com>
Signed-off-by: Alex <jihui.huang@daocloud.io>
Signed-off-by: Alex <alex.tech.lab@outlook.com>
Signed-off-by: Alex <jihuihuang@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 16:22:58 +03:00
Taneem IbrahimandGitHub 0da6e7f3d6 [Bugfix] Reject contradictory custom-op directives (#49134)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
2026-07-26 08:42:14 -04:00
5559679229 [Bugfix][KV Offload] Bound unaligned SWA loads by physical GPU blocks (#49052)
Signed-off-by: Colton Ottley <colton@ottleyengineering.com>
Co-authored-by: Colton Ottley <colton@ottleyengineering.com>
Co-authored-by: jasl <jasl9187@hotmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 14:53:32 +03:00
da3a252fd1 [KVOffload][P2P] Generic P2P secondary tier: peer lookup and serving via ParentManager (#48021)
Signed-off-by: Liran Schour <lirans@il.ibm.com>
Signed-off-by: liranschour <liranschour@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 11:45:47 +03:00
Guan-Ming ChiuandGitHub 21fd9e85a0 [Model] Support top_k and top_p sampling for DiffusionGemma (#45429)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
2026-07-26 08:39:25 +00:00
Guan-Ming ChiuandGitHub 8d28b48d01 [Perf] Isolate MM preprocessing on its own executor (#49524)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
2026-07-26 08:04:17 +00:00
Taneem IbrahimandGitHub 0164022c90 [CI] Fix speech correctness check rejecting improved WER (#49853)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
2026-07-26 06:24:07 +00:00
30b0714031 [Perf] DeepSeek-OCR-2 TTFT Optimize (#49531)
Signed-off-by: RED <outofthewoods@qq.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
2026-07-26 05:53:06 +00:00
Nils MattesonandGitHub 2e860de498 [Doc] Add compile cache volume example to the Docker deployment page (#49782)
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
2026-07-26 05:24:01 +00:00
7eca0e1a64 [KV Offload] Deduplicate replicated MLA KV in the shared CPU region (#48906)
Signed-off-by: Change72 <changg@nvidia.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-07-26 08:22:30 +03:00
7a29a3c54c [Bugfix][KV Offload] Namespace persistent cache by model runner (#49440)
Signed-off-by: Jonguk Cheong <jdal3031@snu.ac.kr>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 08:21:52 +03:00
Athrael SojuandGitHub 1240c74c0a [Bugfix] Respect declared attention contract for ColQwen3.5 retrievers (#49372)
Signed-off-by: Athrael Soju <athrael.soju@gmail.com>
2026-07-26 04:08:07 +00:00
48ebd6f2f1 [Bugfix][KVConnector] Disable cross-layer KV blocks for per-token-head quant (#49226)
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-26 05:24:27 +03:00
liuzhenweiandGitHub b153ae6089 [XPU][CI] add heterogeneous TP UT (#49651)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
2026-07-26 02:01:15 +00:00
Chang GuoandGitHub 7a6a5b3667 [CI] Compute speech WER directly with jiwer (#49773) 2026-07-25 20:51:21 -04:00
0111002323 [Kernel] TD operand loads for batched MoE GEMM (moe_mmk) on XPU (#46340)
Signed-off-by: oonyshch <xonyshch@gmail.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
2026-07-26 08:50:07 +08:00
d30b1ecd1b [Bugfix][KV Offloading] Defer request finalization until final store (#49671)
Signed-off-by: Rui Yin <2260891073@qq.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
2026-07-25 20:58:09 +00:00
Taneem IbrahimandGitHub dbd80cc031 [UX] DCP Topology Validation (#49777)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
2026-07-25 16:53:57 -04:00
70009fb934 [MM][CG] Support ViT CUDA Graph for Gemma-4 (#46837)
Signed-off-by: Anthony Su <xsuanthony@gmail.com>
Co-authored-by: Linkun Chen <github@lkchen.net>
2026-07-25 15:02:09 -05:00
Tyler Michael SmithandGitHub ee1d996367 [Build] Fix for DeepEP manylinux pidfd sycall usage (#49814) 2026-07-25 15:29:38 -04:00
Taneem IbrahimandGitHub 6b0103d1c9 [CI] Stabilize Pooling Rerank Equivalence Test (#49822) 2026-07-25 14:36:50 -04:00
9321aff536 [Bugfix] Wait for the linear bias before layerwise online processing (#49805)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 17:56:32 +00:00
Harry MellorandGitHub 26d725c334 [Model] Add VaultGemma via Transformers modeling backend (#49803)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-25 16:54:15 +00:00
Wentao YeandGitHub 7fe6d3c76b [Perf] Fix moe reduce_scatter perf regression by removing additional comm, 5% E2E throughput gain back. (#48763)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
2026-07-25 16:36:19 +00:00
Harry MellorandGitHub 2e0da24150 Mergify message not on cancelled (#45117)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-25 16:20:24 +00:00
+1 0b0bd2b5f6 [Feature] Add fault tolerance framework (simplified) for DP+EP external LB deployments (#44428)
Signed-off-by: fangyuchu <fangyuchu@qq.com>
Signed-off-by: a798347923 <2645302020@qq.com>
Signed-off-by: TianZhuo <2770730562@qq.com>
Signed-off-by: a798347923 <39047817+a798347923@users.noreply.github.com>
Signed-off-by: 205150940 <112750056+205150940@users.noreply.github.com>
Signed-off-by: w00689259 <wangzhuo66@huawei.com>
Signed-off-by: zWaNg3 <37772915+zWaNg3@users.noreply.github.com>
Signed-off-by: zWaNg3 <389750525@qq.com>
Signed-off-by: yzchang-plus <1078477584@qq.com>
Signed-off-by: Jade Zheng <zheng.shoujian@outlook.com>
Co-authored-by: zWaNg3 <37772915+zWaNg3@users.noreply.github.com>
Co-authored-by: a798347923 <2645302020@qq.com>
Co-authored-by: TianZhuo <2770730562@qq.com>
Co-authored-by: 205150940 <112750056+205150940@users.noreply.github.com>
Co-authored-by: a798347923 <39047817+a798347923@users.noreply.github.com>
Co-authored-by: w00689259 <wangzhuo66@huawei.com>
Co-authored-by: zWaNg3 <389750525@qq.com>
Co-authored-by: yzchang-plus <1078477584@qq.com>
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com>
2026-07-25 11:49:09 -04:00
Canlin GuoandGitHub 33ef67e9fb [BugFix] Increase the max supported duration for MOSS-TD (#49403)
Signed-off-by: Canlin Guo <canlinguosdu@gmail.com>
2026-07-25 14:48:12 +00:00
Harry MellorandGitHub 3e74c60b9c [Docs] Use gen-files for generated docs content (#49587)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-25 14:43:21 +00:00
rongfu.lengandGitHub d1a8ba63d9 [Bugfix][MiniMax-M3] Fix token-major top-k buffer handling in Triton … (#49149)
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
2026-07-25 06:45:27 -07:00
1423569ff5 [Bugfix][Tool Parser] Fix dropped streaming arguments in Jamba and InternLM2 parsers (#48852)
Signed-off-by: mosya415 <263250241+mosya415@users.noreply.github.com>
Co-authored-by: mosya415 <263250241+mosya415@users.noreply.github.com>
2026-07-25 09:34:43 -04:00
Harry MellorandGitHub 9a50464698 [CI] Stop flaky test from downloading model every time (#49800)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-25 13:29:36 +00:00
Harshal JanjaniGitHubHarry Mellormergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
b9b6306ebe feat[vLLM × v5]: Add audio support for the Transformers backend (#39330)
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-25 04:20:58 -07:00
ca0defa343 Make bare hugging_face imports forbidden (#49726)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-07-25 04:18:39 -07:00
0b1a8bb1f6 [Bugfix][CI] Fix stale Mooncake lookup expectation broken by a merge race (#49802)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 03:10:26 -07:00
fe5145765f [Core] Keep attention backends eligible for text-only serving of prefix-LM models (#48796)
Signed-off-by: qtris123 <voquangtri2021@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
2026-07-25 03:06:14 -07:00
dbcc1cdd0a [Model] Remove Ouro (#49786)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 02:50:08 -07:00
a82f1b388f [Perf][V1] Skip LRU hash-split in free_blocks when prefix caching is off (#48017)
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com>
Signed-off-by: Agata Dobrzyniewicz <160237065+adobrzyn@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 09:06:20 +00:00
Johnny-LiouandGitHub 190be7dad2 [Docs] Fix confusing docstring indentation in nemotron_h.py (#49781)
Signed-off-by: Johnny-Liou <a897111@gmail.com>
2026-07-25 07:33:43 +00:00
94682b79f4 [multimodal] Make PyNvVideoCodec decoder concurrency configurable (#49753)
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com>
Co-authored-by: OpenAI Codex <noreply@openai.com>
2026-07-24 23:31:41 -07:00
0ba2aa35a8 Stabilize GPU memory teardown between ROCm CI tests (#49242)
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-07-25 04:21:36 +00:00
Divakar VermaandGitHub aaaeda98dc [CI] fix compile test | refactor VLLM_DISABLE_COMPILE_CACHE for tests (#49770)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
2026-07-24 23:15:19 -05:00