 Dakai AnandGitHub
|
5963c19478
|
Fix Qwen3-VL and Qwen3-omni-thinker accuracy degradation from deepstack inputs under torch.compile (#43617)
Signed-off-by: Dakai An <dakaian108@gmail.com>
|
2026-05-27 15:34:08 -07:00 |
|
 
|
206b72c982
|
[Quantization] Fix Humming RoutedExperts import (#43540)
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-27 10:51:56 -07:00 |
|
 jatseng-aiandGitHub
|
05c50c721e
|
[ROCm] mori: add InterNodeV1LL inter-node kernel selection via VLLM_MORI_INTERNODE_KERNEL (#41751)
Signed-off-by: jatseng-ai <jatseng@amd.com>
|
2026-05-28 00:33:32 +08:00 |
|
 
|
2272062471
|
[Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes (#43731)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
|
2026-05-27 18:36:27 +08:00 |
|
 akii96andGitHub
|
de12f5ca0b
|
[ROCm][GPT-OSS] Avoid repeated compile-time cos_sin_cache.to(bf16) casts in rotary path (#42833)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
|
2026-05-27 16:22:27 +08:00 |
|
 Nico HolmbergandGitHub
|
7b54690244
|
[ROCm][Perf] Expose AITER MoE sorting dispatch policy via env var (#39177)
Signed-off-by: nholmber <nholmber@users.noreply.github.com>
|
2026-05-27 13:11:02 +08:00 |
|
 Xin YangandGitHub
|
d8eebe6d97
|
[Perf] Optimize Fp8BlockScaledMMLinearKernel input_scale tensor using new_empty() (#43677)
Signed-off-by: Xin Yang <xyangx@amazon.com>
|
2026-05-26 19:55:52 -07:00 |
|
 Andreas KaratzasandGitHub
|
5bdb181df5
|
[ROCm][CI] Fix ROCm multimodal Qwen2.5-VL activation compile and Phi4MM ragged image mask handling (#43647)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-26 19:53:34 -07:00 |
|
 Jee Jee LiandGitHub
|
6e503868ca
|
[Kernel] Porting fuse_minimax_qk_norm to manual fusion (#43410)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-26 13:16:03 -07:00 |
|
 Wei-Ming ChenandGitHub
|
6f5b533241
|
Add LM head quantization support for ModelOpt (#42124)
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
2026-05-26 09:21:05 -07:00 |
|
 
|
f51bbc694d
|
[MoE Refactor] W4a8 int8 oracle (#42789)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-05-26 11:15:42 -04:00 |
|
 
|
b226ddacfd
|
[MoE Refactor] Migrate ModelOptMxFp8FusedMoE to oracle (#42768)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-05-26 11:14:14 -04:00 |
|
 Mohammad Miadh AngkadandGitHub
|
a970fb5a1a
|
Fix CuPy runtime deps and restore humming (#43530)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-26 05:59:40 -07:00 |
|
 
|
ebd0692f80
|
[Model] Use AutoWeightsLoader for InternLM2 (#38278)
Signed-off-by: Jesus De Jesus <dejesus.9297@gmail.com>
Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-05-26 03:39:26 -07:00 |
|
 
|
97e4022c6c
|
[Bugfix] Apply fc_norm in Eagle3DeepseekV2 combine_hidden_states (#43482)
Signed-off-by: Yubo Wang <yubowang2019@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-26 00:46:10 -07:00 |
|
 Hank_andGitHub
|
b3269454b1
|
[chores][log] change registry log from warning to debug (#43045)
Signed-off-by: Hank <hcc.mayday@gmail.com>
|
2026-05-26 00:13:46 -07:00 |
|
 Thien TranandGitHub
|
d56612c621
|
[GDN] GDN Prefill kernel for SM100 (#43273)
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2026-05-26 14:02:11 +08:00 |
|
 
|
6f955986e1
|
[Bugfix][Model] Fix GPT2ForSequenceClassification sub-module prefix (#43579)
Signed-off-by: QingZhou-YangHY <3868850350@qq.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-05-25 22:43:19 -07:00 |
|
 Yan MaandGitHub
|
f815c99954
|
[Bugfix] fix device mismatch in MiniCPM-o-4_5 resampler (#43194)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-05-26 13:12:50 +08:00 |
|
 Jee Jee LiandGitHub
|
ec5de7fa7d
|
[LoRA] Add one shot triton kernel For MoE LoRA (#42290)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-05-25 19:47:04 -07:00 |
|
 Jee Jee LiandGitHub
|
d4004455d2
|
[Kernel] Remove NormGateLinear (#43554)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-25 09:49:19 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
5c1aec3dc0
|
Reduce memory usage for granite_speech. (#42933)
Signed-off-by: Yihuki <wangbovbvb@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-25 14:12:57 +08:00 |
|
 weizhoublueandGitHub
|
6cbe448eed
|
fix: MoE model using shared routed experts crashes on AMD GPUs (#42373)
Signed-off-by: weizhou.lan@daocloud.io <weizhou.lan@daocloud.io>
|
2026-05-25 12:03:05 +08:00 |
|
 Jee Jee LiandGitHub
|
b06813e872
|
[Kernel] Add mhc_pre_big_fuse_with_norm_tilelang (#43474)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-25 01:19:45 +00:00 |
|
 
|
d56285c747
|
Tuning script and configs for Triton Mamba SSU kernel (#43083)
Signed-off-by: Banani Ghosh <bg2502@nyu.edu>
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
Co-authored-by: Banani Ghosh <bg2502@nyu.edu>
|
2026-05-24 20:12:44 +03:00 |
|
 TJianandGitHub
|
1806d1adfc
|
[ROCm] [DSv4] [Perf] Support DeepSeek v4 MTP (#43385)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-05-24 18:43:08 +08:00 |
|
 Michael GoinandGitHub
|
10d264a2b9
|
Revert "[Misc] add humming to dependencies" (#43492)
|
2026-05-23 14:21:13 -07:00 |
|
 
|
4438b6e7dc
|
[MoE] Migrate W4A8 CT to oracle kernel setup (#42680)
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-05-23 13:56:01 -04:00 |
|
 
|
a0be71ee47
|
[MM] Enable FlashInfer metadata support for Qwen2.5-VL vision attention (#42787)
Signed-off-by: Hua Huang <huah@nvidia.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-23 16:08:40 +00:00 |
|
  
|
5bb8d2767a
|
[Kernel] Batch invariant NVFP4 linear using cutlass (#39912)
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-23 09:41:12 -04:00 |
|
 GuangYaoZhengandGitHub
|
3f3e862681
|
fix(eagle3): read norm_before_fc from eagle_config for NVIDIA checkpoint (#42143)
Signed-off-by: FERRARIZHENG <popkart06@gmail.com>
|
2026-05-23 08:21:34 +00:00 |
|
 Wei-Ming ChenandGitHub
|
09a219c075
|
[ModelOpt] Support Qwen3.5/3.6 VLM quantized prefix mapping (#42546)
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
2026-05-23 06:23:31 +00:00 |
|
 Taneem IbrahimandGitHub
|
3a1c062151
|
[Misc] Added missing return type annotations to improve mypy and IDE tooling (#43383)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-05-23 13:28:22 +08:00 |
|
 
|
54d153637b
|
[XPU] reudce host overhead of XPU MOE (#42915)
Signed-off-by: mayuyuace <qiming1.zhang@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-23 13:09:34 +08:00 |
|
 
|
a5bbd81e2e
|
[XPU]feat: enable FP8 block-scaled quantization on XPU (#42952)
Signed-off-by: Ma Jian <jian1.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-23 12:33:18 +08:00 |
|
 Andreas KaratzasandGitHub
|
d28bdf9344
|
[ROCm][CI] Fix ROCm LoRA Transformers fallback with full CUDA graphs (#41577)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-23 04:31:32 +00:00 |
|
 Itay AlroyandGitHub
|
6d30655b13
|
elastic_ep: stage/commit MoE quant method on reconfigure (#40881)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-05-22 18:57:26 -04:00 |
|
 
|
8de5cabeb7
|
[XPU]fix: add XPU platform guards to DeepSeek-V4 ops (#42950)
Signed-off-by: Ma Jian <jian1.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-23 06:29:45 +08:00 |
|
 Juhi MittalandGitHub
|
e203006a8b
|
[Quantization][ModelOpt] W4A16 NVFP4 fused MoE + mixed-precision dispatch (#42566)
Signed-off-by: Juhi Mittal <juhim@nvidia.com>
|
2026-05-22 20:51:49 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4e597b7491
|
[Bugfix] Clear error message for FP8 torchao quantization on unsupported GPUs (#36854)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-22 20:09:17 +00:00 |
|
 Artem PerevedentsevandGitHub
|
23f7b11bf4
|
[Bugfix] Detect wrong libcute_dsl_runtime.so variant in FlashInfer GDN (#43427)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-05-22 19:33:33 +00:00 |
|
 Isotr0pyandGitHub
|
f0feb15e7f
|
[Multimodal] Simplify ViT CUDA graph interfaces (#41234)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-22 22:31:00 +08:00 |
|
 sychen52andGitHub
|
fb21d8b4f9
|
Add NVFP4 MOE support for Deepseek V4. (#42209)
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
2026-05-22 07:21:51 -07:00 |
|
 
|
79ff0ffa98
|
[BugFix] wire make_empty_intermediate_tensors on AyaVision and Voxtral (#43118)
Signed-off-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: Keyi Li <likey6688@gmail.com>
|
2026-05-22 05:26:41 -07:00 |
|
  
|
b3c7ffcab8
|
[Misc] Replace assert with proper exceptions for security and validation in pooling (#43286)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-22 18:43:33 +08:00 |
|
 
|
d3d1cf6972
|
[XPU]feat: add XPU fallback for MoE topk routing and MXFP4 backend (#42951)
Signed-off-by: Ma Jian <jian1.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-22 10:22:45 +00:00 |
|
 wangxiyuanandGitHub
|
7e1b45a092
|
[Attention] Mamba attention module refactor (#41126)
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
|
2026-05-22 17:13:12 +08:00 |
|
 haosdentandGitHub
|
025d4f5cd2
|
[CI] Fix "test_awq_load[gemma4-moe-*]" failure (#43296)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-22 07:13:59 +00:00 |
|
 tc-mbandGitHub
|
fa1ff88b31
|
[Model] Fix MiniCPM-V 4.6 vit_merger qkv weight loading (#43213)
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
|
2026-05-21 22:44:06 -07:00 |
|
 Furkan FandGitHub
|
e746a2eebf
|
[Model] Use AutoWeightsLoader for Voyage (#42972)
Signed-off-by: Furkan Fidan <dev@yufufi.com>
|
2026-05-22 05:28:23 +00:00 |
|