2ded1b24e7
[KV Connector][Mooncake] Apply SWA lookup mask before hashing/key build ( #47317 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-09 19:51:23 +00:00
weishu and GitHub
2285cfca46
[KVConnector] MultiConnector: give every sub-connector the request's real blocks in update_state_after_alloc ( #46865 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-07-09 11:10:39 -07:00
85b3a7264b
[Bugfix][Model Runner V2] Order uniform decodes first so spec decodes aren't misclassified as prefills ( #47381 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-09 14:26:27 +01:00
412414d8e0
Remove PersimmonForCausalLM and FuyuForCausalLM model architectures ( #48096 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-07-09 04:59:08 -07:00
e87521626f
Sanitize server file paths from validation error responses ( #46415 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-09 17:46:29 +08:00
1cd75b3dd4
[Bugfix] Fix race condition in KVBlockZeroer ( #48085 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-07-09 09:18:19 +00:00
ab7961a14a
Remove TeleChatForCausalLM ( #47989 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-09 00:34:54 -07:00
a07765c6bd
[Bugfix] Fix Qwen3-ASR transcription streaming postprocessing ( #42478 )
...
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
Signed-off-by: JooHo Lee <jooho414@gmail.com >
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
2026-07-09 00:33:27 -07:00
7802c20c4e
[KVConnector][NIXL] Support pipeline-parallel prefill in push mode ( #45880 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 16:49:23 -07:00
95d6d6f4bb
[Bugfix] Use int8 workspace for FlashInfer MLA decode ( #48046 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-08 23:39:40 +00:00
Roberto L. Castro and GitHub
5f85975624
[Feat] Add runtime monitor for post-warmup TileLang compilation ( #46718 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
2026-07-08 22:11:28 +00:00
dcdd756d75
[CI] GSM8K eval integration test for KV offloading ( #46893 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 17:59:48 -04:00
Kaihang Jiang and GitHub
089e412878
[Perf] Integrate TRTLLM BF16 MoE Modular Kernel ( #45182 )
...
Signed-off-by: Kaihang Jiang <kaihangj@login-lyris02.lyris.clusters.nvidia.com >
2026-07-09 01:36:14 +04:00
Nick Hill and GitHub
a5d19cbb95
[Core] Move MRV1 late_interaction_runner.py out of MRV2 subtree ( #48014 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 18:30:11 +00:00
b2cf70ea3a
[CI] BugFix Eval Small Models Distributed test for DiffusionGemma ( #47980 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-07-08 17:00:56 +00:00
almayne and GitHub
d1f1d86797
[Bugfix] Re-enable benchmarking of librispeech dataset. ( #47033 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
2026-07-08 16:19:26 +00:00
Tyler Michael Smith and GitHub
68b4a1d582
Fix NVML capability lookup for visible devices ( #47892 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-07-08 09:07:44 -04:00
cd0de48d08
[Bugfix][V1] Free out-of-window blocks on the processed-token basis under async scheduling ( #47728 )
...
Signed-off-by: Saddss <28726669061@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Saddss <28726669061@qq.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-08 13:34:19 +01:00
Canlin Guo and GitHub
285c08c036
[Model] Support MOSS-Transcribe-Diarize ( #47729 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-07-08 04:05:45 -07:00
04a703e397
[Frontend] Support bad_words in the /v1/completions endpoint ( #46793 )
...
Signed-off-by: sungbin1015 <sbin@solbox.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:51:17 +00:00
Nicolò Lucchesi and GitHub
bd3bb4eb26
[Misc][Docs] Add human-readable integer support for more cli-args ( #47608 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-08 09:43:34 +00:00
99a85617bf
[Test] Skip DeepEP MoE layer tests without P2P access ( #47946 )
...
Signed-off-by: Tyler Michael Smith <tyler@tylermsmith.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-08 09:46:03 +01:00
51e5372f3d
[Model][HunyuanVL] Use native transformers processor and adapt to transformers 5.13 ( #47872 )
...
Co-authored-by: manayang <manayang@tencent.com >
2026-07-08 07:58:23 +00:00
Hongxia Yang and GitHub
2c64b4c1cc
[ROCm] fixed aiter master flag and expert parallelism compatibility on minimax-m3-mxfp8 ( #47158 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-07-08 15:26:17 +08:00
d35eba302f
[Bugfix] Avoid leaking Pydantic repr in tool_choice error message ( #47028 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-07-08 15:00:59 +08:00
Zach Zhu and GitHub
5d5fab0061
[Bugfix][Frontend] Fix http_requests_total metric recording some 4xx errors as 5xx ( #44303 )
...
Signed-off-by: Zach Zhu <zzqshu@126.com >
2026-07-08 05:33:21 +00:00
d9e57ea82e
[ROCm][Perf] MXFP8 dense-linear + grouped-MoE GEMM optimizations for MiniMax-M3 ( #46117 )
...
Signed-off-by: amd-ethany <amd-ethany@users.noreply.github.com >
Co-authored-by: amd-ethany <amd-ethany@users.noreply.github.com >
2026-07-08 04:03:34 +00:00
9021589498
[Minimax-M3] Using tok_sparse_select from MSA instead of triton kernels ( #47502 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-07 21:01:12 -07:00
Ting SUN and GitHub
0303f37a54
[Bugfix][Pooling] Align CrossEncoder token type ids after truncation ( #47772 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-07-08 03:59:22 +00:00
Walter Beller-Morales and GitHub
dd127d82ed
[Core][Engine] only materialize tokens when thinking budget is in req ( #47053 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-07-07 21:02:38 -06:00
0ca6eee743
[Core] Pass request context to CPU offload cache policy touch ( #47744 )
...
Signed-off-by: jacklin78911-collab <jacklin78911@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-08 05:56:25 +03:00
Martin Hickey and GitHub
f7fc0ca993
[Frontend] Add endpoint plugins framework ( #47454 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-07-08 10:00:41 +08:00
Rahul Vishwakarma and GitHub
f7efab58ec
[CPU][Bugfix] Fix flaky ShortConv prefill test on ARM (uninitialized weights) ( #47848 )
...
Signed-off-by: Rahul Vishwakarma <Rahul.Vishwakarma2@ibm.com >
2026-07-07 18:20:09 -07:00
e97c3cb303
[Core] Persist and reuse the memory-profiling result across boots (opt-in) ( #47388 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-08 00:53:02 +00:00
Juan Pérez de Algaba and GitHub
675f4295cd
fix(security): bound completion prompt list to prevent unbounded engine fan-out ( #47845 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-07-07 22:48:20 +00:00
Jason Li and GitHub
d99adcebdc
[BugFix] Fix ModelOpt quantization inference for fused siblings ( #47445 )
...
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
2026-07-08 03:19:42 +05:00
55da232db6
[Bugfix] Pad Mamba page size instead of scaling block_size in unify_kv_cache_spec_page_size ( #45207 )
...
Signed-off-by: Sahil170595 <147995121+Sahil170595@users.noreply.github.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-07 22:01:34 +00:00
Rishabh Saini and GitHub
2f3f441f84
fix: include topic frame in KV events replay response ( #45177 )
...
Signed-off-by: RishabhSaini <rishabhsaini01@gmail.com >
2026-07-07 14:48:23 -04:00
d6875196ad
[Bugfix] Exclude kv_cache_memory_bytes from CacheConfig.compute_hash ( #47356 )
...
Signed-off-by: Nils Matteson <nils@thaw.sh >
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com >
Co-authored-by: Nils Matteson <nils@thaw.sh >
2026-07-07 10:46:51 -07:00
bdc6f3bfa1
[Bug] Fix tmp directory for lm_eval ( #47755 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 16:40:12 +00:00
392d1b4d2e
[BugFix][LoRA] Refresh punica metadata when LoRA slots are reassigned under an unchanged mapping ( #47725 )
...
Signed-off-by: AmeenP <ameenp360@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-07 08:53:55 -07:00
c46ced1ee3
[kv_offload] Establish tier-owned KV event handling ( #46544 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Signed-off-by: Chang Guo <changg@nvidia.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:55:20 +03:00
65dcde1695
[Bugfix] Fix PD disagg + MTP correctness for Qwen3.5(GDN) ( #47466 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 13:51:22 +00:00
65a7b46284
[KV-Offloading] Support workload identity for objectstore secondary tier ( #47063 )
...
Signed-off-by: Pierangelo Di Pilato <pierdipi@redhat.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-07-07 16:30:16 +03:00
93e2ab7111
Disable dynamic speculative decoding when DP is enabled ( #45963 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-07 12:52:41 +00:00
8b91cd5b20
[Bugfix][Core] Close underlying iterator in merge_async_iterators single-iterator fast path ( #44726 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-07 05:13:41 -07:00
Harry Mellor and GitHub
dd94484577
Bump Transformers version to 5.10.4 ( #41359 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 05:13:28 -07:00
danielafrimi and GitHub
0a2965b1b3
[BugFix] Fix ModelOpt mixed-precision quantization for sparse quantized_layers configs. ( #47318 )
...
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Signed-off-by: <dafrimi@nvidia.com >
2026-07-07 11:45:13 +00:00
Harry Mellor and GitHub
0ed05b6f82
[CI] Fix Transformers modeling backend LoRA test ( #47832 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-07 11:40:00 +00:00
Guan-Ming Chiu and GitHub
ed051fab54
[Bugfix] Reject sampling params unsupported by diffusion models ( #45418 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-07-07 11:25:36 +00:00