Chauncey and GitHub
d2130a47bb
[Bugfix]: Fix MinimaxM2ToolParser missing tools parameter ( #39683 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-14 11:16:39 +08:00
c687bf226a
[LMCache][MP] optimize save when mla enabled ( #38810 )
...
Signed-off-by: idellzheng <idellzheng@tencent.com >
Co-authored-by: Yihua Cheng <yihua98@uchicago.edu >
2026-04-13 17:56:43 -07:00
Giancarlo Delfin and GitHub
ccf90ba784
[Model Runner V2] Add full cuda graph support for eagle prefill ( #37588 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-13 16:01:24 -07:00
Netanel Haber and GitHub
6adacfcb65
ParakeetExtractor performance and UX enhancements ( #39423 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-13 21:37:35 +00:00
Flora Feng and GitHub
14cb86c187
[Refactor][Parser] Simplify parse_delta ( #39728 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-13 21:02:13 +00:00
8213e8f880
Bug/test eagle dp v0 ( #38938 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-13 20:50:08 +00:00
Pedram Razavi and GitHub
3693f922ff
[Bugfix][Pooling] Fix silent weight corruption with buffer-reusing iterators ( #39650 )
...
Signed-off-by: Pedram Razavi <pedram.razavi@gmail.com >
2026-04-13 19:37:27 +00:00
5c18b961d6
[Core][Metrics] expose waiting request breakdown via labeled metric (capacity/deferred) ( #38435 )
...
Signed-off-by: Mukesh Baphna <mukesh@hippocraticai.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-04-13 15:30:55 -04:00
f72b20976c
[Bugfix] Reject non-nvfp4 dtypes when using the flashinfer_nvlink_one_sided all2all backend ( #39717 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-13 19:13:51 +00:00
610a3efcaf
[Doc] Fix Python-only build 404 fallback guidance ( #38052 )
...
Signed-off-by: George-ao <yuyiao772@gmail.com >
Signed-off-by: Yuyi Ao <yuyiao772@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-13 12:09:31 -07:00
JartX and GitHub
f414f90601
[Bugfix][Kernel][ROCm] Fix triton_w4a16 scales mismatch when BLOCK_K > group_size ( #39705 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-04-13 14:29:45 -04:00
Nicolò Lucchesi and GitHub
8625ec267b
[Misc] Multi-turn benchmark output performance json ( #39572 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-13 18:15:23 +00:00
995e9a209e
[Bugfix] Use is_integrated to detect UMA GPUs for memory reporting ( #35356 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Co-authored-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-13 11:07:40 -07:00
Yongye Zhu and GitHub
739e5945dc
[Quantization] [Refactor] Create special "GptOssMxfp4MoeMethod" ( #39604 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-13 12:53:58 -04:00
Santino Ramos and GitHub
4d042ed85f
[Bugfix] Fix tensor shape mismatch in sparse attention with speculative decoding ( #39542 )
...
Signed-off-by: Santino Ramos <santinor@inferact.ai >
2026-04-13 08:57:38 -07:00
zhanqiuhu and GitHub
10d9872d3a
[CI][Metrics] Fix local_cache_hit assertion after prompt tokens metrics updates ( #39709 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-13 15:16:56 +00:00
ccd0d1d906
[Bug] Fix rocm sparse attn indexer issue ( #39225 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-13 07:53:45 -07:00
Yi Liu and GitHub
d8ddb31644
[Bugfix][CT] Fix KV cache scale handling ( #39418 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
2026-04-13 10:50:16 -04:00
Ekagra Ranjan and GitHub
1ce0318c68
[Bugfix] stream failure when model name not in audio endpoints ( #36679 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-04-13 14:20:07 +00:00
Tihomir Elek and GitHub
8d825b87d6
[Bug] Fix TypeError when hf_config.architectures is None during model loading ( #38849 )
...
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com >
2026-04-13 12:13:21 +01:00
zofia and GitHub
1b19bd7589
[MXFP8] [XPU] add a new compressed tensor schema and add a xpu mxfp8 gemm kernel ( #38707 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-04-13 16:59:20 +08:00
200a727e94
[Bugfix] Fix Responses API instructions leaking through previous_response_id ( #37727 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-13 08:46:33 +00:00
edbc1abd1c
feat: add max_tokens_per_doc in rerank request. ( #38827 )
...
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-13 01:24:09 -07:00
Flora Feng and GitHub
0e39202ca9
[Bugfix] Fix GLM tool parser streaming with MTP or stream interval ( #39253 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-13 05:10:30 +00:00
9dd5ee0117
[XPU]Enhance environment collection for Intel XPU and optimize layout ( #35698 )
...
Signed-off-by: sihao.li <sihao.li@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-13 12:51:46 +08:00
fa6ae31177
feat: rename logit_bias/logit_scale to logit_mean/logit_sigma for affine score calibration ( #39530 )
...
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-13 04:43:44 +00:00
maobaolong and GitHub
2a3c32ce67
fix(lmcache): correct store for cached requests and num_scheduled_tokens in lmcache_mp_connector.py ( #39655 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-04-13 03:29:19 +00:00
4beeb0689c
fused qknorm+rope kernel optimization for SM9.0 ( #37376 )
...
Signed-off-by: EricccYang <yangyang4991@gmail.com >
Signed-off-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-04-12 19:58:37 -07:00
cae984060f
[compile] Enable AOT compile with batch invariance mode. ( #39201 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-12 19:58:33 -07:00
Jee Jee Li and GitHub
715681c127
[LoRA] Support dual CUDA streams-Linear Layer ( #35721 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-13 10:57:07 +08:00
Kunshang Ji and GitHub
dc02271d76
[XPU] revert torch-xpu to 2.10 ( #39656 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-13 10:50:29 +08:00
Andreas Karatzas and GitHub
4e4ad41d11
[ROCm][CI] Removed stale tests and extended acceptance test ( #39651 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-13 10:40:26 +08:00
Yongye Zhu and GitHub
620e8924d9
[Bugfix] [Tests] Enforce out tensor device in kernel/moe/test_cutedsl_moe.py ( #39644 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-12 17:08:08 -07:00
Animesh Jain and GitHub
f00c5539d7
[compile] Bug fix for _decompose_size_nodes ( #38360 )
...
Signed-off-by: Animesh Jain <anijain@umich.edu >
2026-04-12 20:20:24 +00:00
Le Yang and GitHub
21fab0a3db
fix(moe): fix RoutedExpertsCapturer assertion failure with DP>1 and MK path ( #37879 )
2026-04-12 10:28:17 -04:00
Nicolò Lucchesi and GitHub
3244a2ebf2
[KVConnector][NIXL] Organize NIXL connector into its own directory ( #39354 )
...
The number of features supported by the connector has grown substantially
and the `nixl_connector.py` file has accumulated a lot of code. Creates a separate
directory and isolates connector/scheduler code in the hope of improving clarity
and maintainability.
Further refactor of components aimed at improving clarity and simplifying code
will follow soon.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-12 13:10:50 +00:00
Mark McLoughlin and GitHub
72ff142c37
[Core][Metrics] Remove vllm:prompt_tokens_recomputed metric ( #38709 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-12 12:22:01 +03:00
Nick Hill and GitHub
ee3c0c83db
[Pooling] Disable async scheduling by default for pooling models ( #39592 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-12 07:23:42 +00:00
cc07dad789
[HMA] [KVEvent] Enable GPU-side KV events for HMA ( #37688 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-04-12 10:01:02 +03:00
17e787a779
fix(kimi_k25): resolve media_placeholder_token_id from tokenizer ( #39344 )
...
Signed-off-by: r266-tech <r266.tech@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-11 21:10:24 -07:00
639402f5a2
Support FP8 KVCache on XPU ( #37731 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-12 03:53:32 +00:00
Andreas Karatzas and GitHub
0f7be0f2f7
[ROCm][CI/Build] Fix memory cleanup in MM test ( #39555 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-12 11:13:34 +08:00
Yan Ma and GitHub
394ff86965
[XPU][CT] support per-channel quantization in xpu fp8 linear method ( #38316 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-12 02:46:28 +00:00
EdalatiAli and GitHub
df1e30e74b
[Quant] add CompressedTensorsW8A8Mxfp8 for linear and MoE layers ( #38815 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
2026-04-11 17:21:36 -06:00
bd8bd52308
[Bugfix] Runtime driver check for cuMemcpyBatchAsync in swap_blocks_batch ( #38919 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-11 11:02:34 -06:00
Wei Zhao and GitHub
59b2f7b640
[Perf] Fuse Zero Initializer for FP8 DeepGemm Block Quant Kernel ( #39547 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-04-11 07:16:51 -07:00
ShubyM and GitHub
92feb9991d
[Gemma4][Bugfix]: Enable Gemma4ForCasualLM to load lora adapters correctly ( #38844 )
...
Signed-off-by: ShubyM <shubymishra20@gmail.com >
2026-04-11 09:06:49 +00:00
d4cb783c10
[Bugfix] Fix GDN FLA kernel crashes with NULL_BLOCK_ID=0 CUDA graph padding ( #39064 )
...
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-11 08:35:19 +00:00
Li, Jiang and GitHub
eb92ba740a
[CI/Build] Fix sentence-transformers version in CPU test ( #39557 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-11 07:04:25 +00:00
z1ying and GitHub
a3e750c0a5
[Misc] Update deprecation warning for --model flag ( #39518 )
...
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-10 23:25:20 -07:00