f2145efcb6
[BugFix] KeyError on scope["method"] for realtime api websocket in AuthenticationMiddleware ( #36934 )
...
Signed-off-by: daniebrill <50454544+daniebrill@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-15 16:15:01 +00:00
Roy Huang and GitHub
ed33310552
[KVConnector][LMCache] Propagate cache_salt through MP connector for per-user cache isolation ( #39837 )
...
Signed-off-by: royyhuang <royyhuang@gmail.com >
Signed-off-by: royyhuang <roy.y.huang@gmail.com >
2026-04-15 09:10:49 -07:00
3cc328a4be
[SpecDecode][Benchmark] Add SPEED-bench support to benchmarking CLI ( #36029 )
...
Signed-off-by: talora <talora@nvidia.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-15 12:00:07 -04:00
3beb57a238
[XPU] properly handle q_descale on XPU as quant query input not supported ( #39676 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-15 21:52:58 +08:00
8b5531933a
FIX: support language_model.backbone naming in NemotronH Nano VL quantization config ( #39901 )
...
Signed-off-by: <>
Co-authored-by: root <root@lyris0144.lyris.clusters.nvidia.com >
2026-04-15 13:49:48 +00:00
Chauncey and GitHub
db8d4a4a06
[BugFix][Graph] fix: handle empty sym_shape_indices in PiecewiseBackend. ( #39395 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-15 09:28:09 -04:00
zofia and GitHub
fc701c8058
[XPU][MXFP4] add mxfp4 quant op for XPU ( #39857 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-04-15 12:28:19 +00:00
Csrayz and GitHub
68be0f853e
[Metrics] Add request_id to FinishedRequestStats to enable correlation between metrics and requests ( #39710 )
...
Enables external `StatLogger` plugins to correlate per-request metrics
with request-level context. Also, this is a pre-requisite for Prometheus
exemplars in #30972 .
Signed-off-by: Csrayz <33659823+Csrayz@users.noreply.github.com >
2026-04-15 11:24:17 +00:00
Zhenzhong Xu and GitHub
60995c05b4
[Quantization][Autoround][CPU] Add W4A16 Support ( #38192 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
2026-04-15 18:38:31 +08:00
Yan Ma and GitHub
29e5d10205
fix online fp8 for MiniCPM models ( #39862 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-15 09:09:20 +00:00
Or Ozeri and GitHub
235e1f930a
[kv_offload+HMA][3/N]: Remove block_size from KVEvents ( #36644 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-15 11:53:19 +03:00
+86
431cea3eea
[Bugfix] Fix tool_calls Iterable consumed when debug logging is enabled ( #34844 )
...
Signed-off-by: Wojciech Wais <wojciech.wais@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Jaebok Lee <jaebok9541@naver.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Signed-off-by: yuwei <yuwei@dev.local >
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Signed-off-by: Li <chuali@amd.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: lkm2835 <lkm2835@gmail.com >
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Signed-off-by: vnadathur <glvikramn@gmail.com >
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com >
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com >
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: jackcfwang <jackcfwang@tencent.com >
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
Signed-off-by: Richard Zou <zou3519@gmail.com >
Signed-off-by: milesial <milesial@users.noreply.github.com >
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Signed-off-by: whx-sjtu <2952154980@qq.com >
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Signed-off-by: jatseng-ai <jatseng@amd.com >
Signed-off-by: jatseng-ai <janet.tseng@amd.com >
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Signed-off-by: xaguilar-amd <xaguilar@amd.com >
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com >
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Peter Nguyen <petern0408@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: manu <fortin.emmanuel@gmail.com >
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Rahul-Tuli <rtuli@redhat.com >
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: leeyongjun <jqueen.astro@gmail.com >
Signed-off-by: Ziying Tao <tzzying@outlook.com >
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Signed-off-by: ShubyM <shubymishra20@gmail.com >
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: r266-tech <r266.tech@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Signed-off-by: Animesh Jain <anijain@umich.edu >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
Signed-off-by: EricccYang <yangyang4991@gmail.com >
Signed-off-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Signed-off-by: sihao.li <sihao.li@intel.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com >
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Santino Ramos <santinor@inferact.ai >
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: George-ao <yuyiao772@gmail.com >
Signed-off-by: Yuyi Ao <yuyiao772@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Mukesh Baphna <mukesh@hippocraticai.com >
Signed-off-by: Pedram Razavi <pedram.razavi@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: Rishi Puri <riship@nvidia.com >
Co-authored-by: zzaebok <44357534+zzaebok@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: yuwei <yuwei@dev.local >
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Co-authored-by: Chuan (Richard) Li <chuali@amd.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Ganesh R <ganesh.r@amd.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Kyungmin Lee <30465912+lkm2835@users.noreply.github.com >
Co-authored-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Co-authored-by: vnadathur <glvikramn@gmail.com >
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Elham <elham.harirpoush@arm.com >
Co-authored-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Chaofan Wang <jackcfwang@tencent.com >
Co-authored-by: Chendi.Xue <chendi.xue@intel.com >
Co-authored-by: Injae Ryou <injaeryou@gmail.com >
Co-authored-by: Richard Zou <zou3519@users.noreply.github.com >
Co-authored-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: Elvir Crnčević <elvircrn@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
Co-authored-by: Hexiang Wang <56632993+whx-sjtu@users.noreply.github.com >
Co-authored-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: PatchyTIS <58251192+PatchouliTIS@users.noreply.github.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: jatseng-ai <janet.tseng@amd.com >
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: xaguilar-amd <xavier.aguilarfruto@amd.com >
Co-authored-by: Ravitez Dondeti <dondetir@users.noreply.github.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Peter Nguyen <petern0408@gmail.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: zhrrr <43847754+izhuhaoran@users.noreply.github.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Jesus Federico <14651+jefp@users.noreply.github.com >
Co-authored-by: Manu <efortin@users.noreply.github.com >
Co-authored-by: zhanqiuhu <49648934+ZhanqiuHu@users.noreply.github.com >
Co-authored-by: yzong-rh <yzong@redhat.com >
Co-authored-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Rahul-Tuli <rtuli@redhat.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: Lee Yongjun <35302114+elwhyjay@users.noreply.github.com >
Co-authored-by: z1ying <55220715+z1ying@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
Co-authored-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com >
Co-authored-by: ShubyM <shubymishra20@gmail.com >
Co-authored-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: r266-tech <r2668940489@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Le Yang <562593859@qq.com >
Co-authored-by: Animesh Jain <anijain@umich.edu >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Zhengxu Chen <zhxchen17@fb.com >
Co-authored-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Co-authored-by: maobaolong <baoloongmao@tencent.com >
Co-authored-by: sihao_li <165983188+1643661061leo@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: Tihomir Elek <tiho.elek@gmail.com >
Co-authored-by: Yi Liu <yi4.liu@intel.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Santino Ramos <51103228+santiramos27@users.noreply.github.com >
Co-authored-by: haosdent <haosdent@gmail.com >
Co-authored-by: JartX <sagformas@epdcenter.es >
Co-authored-by: Yuyi Ao <yuyiao772@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: mukesh-hai <mukesh@hippocraticai.com >
Co-authored-by: Pedram Razavi <pedram@sierra.ai >
2026-04-15 01:32:47 -07:00
zhanqiuhu and GitHub
799973af4e
[CI][NIXL] Fix PD CI breakage: pin nixl-cu{12,13} versions ( #39851 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-04-14 23:50:23 -07:00
bcc2306cef
[Bugfix] Respect VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY in prefetch offloader ( #37699 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-14 20:43:29 -07:00
wliao2 and GitHub
3abf858443
[Test] Refactor hard coded device string in test files under compile/quantization/models/model_executor folders ( #38901 )
...
Signed-off-by: Liao, Wei <wei.liao@intel.com >
2026-04-15 11:02:35 +08:00
f4b42df048
[Attention Backend] TurboQuant: 2-bit KV cache compression with 4x capacity ( #38479 )
...
Signed-off-by: vibhavagarwal5 <vibhavagarwal5@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-04-14 19:57:13 -07:00
Giancarlo Delfin and GitHub
3bfe55a037
[Model Runner V2] Disable piecewise cudagraph mode fallback for eagle draft decodes ( #39773 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-14 17:47:57 -07:00
Andrey Talman and GitHub
b569620f72
[CI] Add PyTorch nightly build and test pipeline ( #37226 )
...
Signed-off-by: atalman <atalman@fb.com >
2026-04-14 17:13:24 -07:00
65b9808960
[Bugfix] Disable FlashInfer CUTLASS MoE on SM121 (DGX Spark) ( #39825 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-14 16:03:57 -07:00
Francesco Fusco and GitHub
507df79a29
[Hybrid] Simplify accepted token counting in spec decode for hybrid models ( #38372 )
2026-04-14 15:19:09 -07:00
1696c864b9
[Bugfix][Mooncake] Fix thread-local CUDA context for NVLink transfers in _send_blocks ( #39548 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-04-14 14:13:58 -07:00
Wentao Ye and GitHub
2ad1029233
[Bug] Fix batch invariance nvfp4 support ( #39820 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-14 17:08:17 -04:00
maobaolong and GitHub
b2f749dc97
fix(lmcache): correct store for cached requests while enable prefix cache ( #39719 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-04-14 20:51:27 +00:00
70ed01550c
[Reasoning][Frontend] Add model config to adjust_request in reasoning parser ( #37848 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com >
Signed-off-by: Aaron Pham <contact@aarnphm.xyz >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Aaron Pham <contact@aarnphm.xyz >
2026-04-14 20:29:06 +00:00
bnellnm and GitHub
19ec9a0a62
[MoE Refactor] Refactor ZeroExpertFusedMoE into new framework ( #35549 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-14 16:11:20 -04:00
1a9353bb02
[MoE] Move GPT OSS Triton kernel experts into fused_moe/experts/ ( #39007 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-14 19:27:39 +00:00
roikoren755 and GitHub
ecf5ff7ce3
[Mamba] Flashinfer selective_state_update ( #36162 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-14 15:10:58 -04:00
zhanqiuhu and GitHub
30679319e8
[CI][KVConnector][Metrics] Update multi KV connector edge case according to prefill stats changes ( #39808 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-14 18:59:15 +00:00
240f2636ca
[Kernel] Support TRTLLM GEN NVFP4 MoE for non-512-aligned hidden dims via weight padding ( #39510 )
...
Signed-off-by: root <root@lyris0017.lyris.clusters.nvidia.com >
Signed-off-by: Daniel Afrimi <dafrimi@nvidia.com >
Co-authored-by: root <root@lyris0017.lyris.clusters.nvidia.com >
2026-04-14 11:49:56 -07:00
dc8df110bc
add warning when FP8 KV cache misses prefill query quantization ( #39752 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Albert Cheng (Engrg-Hardware 1) <albecheng@login-lyris02.lyris.clusters.nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 14:43:05 -04:00
be0c855ebd
[KV Offload] Unified memory layout for offloading workers ( #37206 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-14 21:33:33 +03:00
Andrew Barnes and GitHub
e64b39ea71
[ROCm] Align AiterFlashAttentionImpl attn_type check with backend ( #39119 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-14 10:36:26 -07:00
Alessandro Sangiorgi and GitHub
2faad08362
[compile] Nest inductor cache under AOT compile dir ( #39718 )
...
Signed-off-by: Alessandro Sangiorgi <asangior@redhat.com >
2026-04-14 17:17:54 +00:00
Rohan Potdar and GitHub
23f3760217
[Bugfix][ROCm]: Allow gpt_oss_mxfp4 quantization method on rocm ( #39754 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-14 17:10:04 +00:00
Mark McLoughlin and GitHub
906a8c15d0
[Core][Metrics] Remove unused SchedulerStats.encoder_cache_usage ( #39693 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-14 12:53:57 -04:00
Micah Williamson and GitHub
4f4f8eaa78
[ROCm][CI] Fix condition for test_per_token_group_quant_fp8_packed ( #39730 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-14 16:14:31 +00:00
Netanel Haber and GitHub
b6890a120a
Bugfix: use_existing_torch.py: Glob recursive subdirs in requirements ( fixes #39024 ) ( #39793 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-14 23:11:46 +08:00
Lucas Kabela and GitHub
c08f3b2a62
Measure encoder compile time seperate from llm backbone ( #39240 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-14 10:52:49 -04:00
Hexiang Wang and GitHub
f02b3269e7
[PluggableLayer][3/N] Apply PluggableLayer to moe-related layers. ( #33556 )
...
Signed-off-by: whx-sjtu <2952154980@qq.com >
2026-04-14 09:55:00 -04:00
e1e318af01
[MoE Refactor] Remove MoE DP chunking ( #39107 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-14 09:48:05 -04:00
f7e62e3d66
[Bugfix] Fix mismatch between global and local attention heads in tensor-parallel mode for param2moe model ( #39707 )
...
Signed-off-by: bhargav-patel-29 <bhargav.patel@tihiitb.org >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 20:13:36 +08:00
18b1c77211
fix: handle ImportError in load_audio ( #39473 )
...
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-04-14 19:09:06 +08:00
Matthias Gehre and GitHub
1e4748c66a
[Bugfix] Fix vllm bench serve to count multimodal tokens in "total input tokens" ( #38654 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-14 11:00:40 +00:00
6f786f2c50
[Bugfix][Model] Fix Devstral Small 2 HF format weight loading ( #39293 )
...
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-14 10:11:18 +00:00
fxmarty-amd and GitHub
4eee77b877
[fix][MOE] Fix MOE experts intermediate_size dimension not being narrowed before weight loading ( #39688 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-04-14 09:35:28 +00:00
xiangdong and GitHub
a1993b96fd
[XPU][CI] Remove Arc in label-xpu ( #39776 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-14 02:27:38 -07:00
Julien Debache and GitHub
893b2affff
feat: add TxtSlicesDataset to allow sampling slices from txt file for benchmarking ( #30156 )
...
Signed-off-by: jdebache <jdebache@nvidia.com >
2026-04-14 09:20:03 +00:00
80118853f4
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3-VL video inference ( #38061 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Shanshan Shen <87969357+shen-shanshan@users.noreply.github.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-14 16:49:32 +08:00
c0ecaed950
[Frontend] Offload blocking preprocessing & postprocessing ops to thread pool for pooling entrypoints. ( #39763 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 08:29:25 +00:00
0008729abf
[Model] Use mm_features for Ernie-4.5 VL M-RoPE ( #39753 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-14 01:11:52 -07:00