 
|
67eb6083e3
|
Revert "[Misc] Move pyav and soundfile to common requirements" (#40276)
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-04-21 09:08:06 -07:00 |
|
 Harry MellorandGitHub
|
6ee081d1d0
|
Add new tp plan styles to the Transformers modelling backend (#40467)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-04-21 08:51:30 -07:00 |
|
 
|
66cc3fa559
|
[Model Runner V2] Multiple prompt logprobs support (#39937)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-21 15:49:05 +00:00 |
|
 Vadim GimpelsonandGitHub
|
6d85b36a9f
|
Revert #38730 and #38791 (#40032)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>
|
2026-04-21 11:44:11 -04:00 |
|
 Matthew BonanniandGitHub
|
ab5666eb7c
|
[UX] Bump version in CG memory profiling log message (#40465)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-04-21 15:26:06 +00:00 |
|
 roikoren755andGitHub
|
f819265a4a
|
Default to 'align' mamba cache mode for Mamba-based models when speculative decoding is enabled (#40454)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-04-21 14:51:43 +00:00 |
|
 Shanshan ShenandGitHub
|
936e0b79aa
|
[MM][CG] Optimize default max_frames_per_batch auto-infer for ViT CUDA graph video inference (#40445)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-21 14:47:53 +00:00 |
|
 
|
b2a5518679
|
[XPU][CI] Add misc, engine and lora cases on Intel GPU in CI (#39887)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-04-21 22:30:46 +08:00 |
|
 ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟andGitHub
|
908a713488
|
[Bugfix] LoRA: extend expert base_layer loading to Qwen3.5 and Step3.x (#37114)
Signed-off-by: Hollow Man <hollowman@opensuse.org>
|
2026-04-21 14:17:03 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
ec5ef0ac73
|
[Doc] Add Qwen3 AWQ models to documentation (#40034)
Signed-off-by: Yusuf <yusufmohammad@live.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-21 09:37:41 -04:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
7b1e0b07d0
|
[Bugfix] Fix dataset name and path argument validation bug in vllm bench serve (#40288)
Signed-off-by: talora <talora@nvidia.com>
Signed-off-by: Talor Abramovich <talor19@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-21 06:14:28 -07:00 |
|
  
|
d249a9e90e
|
Add Granite 4.1 Vision as built-in multimodal model (#40282)
Signed-off-by: Artem Spector <artems@il.ibm.com>
Signed-off-by: artemspector <artems@il.ibm.com>
Co-authored-by: artemspector <artems@il.ibm.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-21 05:43:39 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
d2e2e856ad
|
[Frontend] Remove frontend pooling multi task support. (#37861)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-21 12:27:44 +00:00 |
|
 
|
766cb65d00
|
feat(multimodal): support externally processed mm_kwargs with cache injection (#39502)
Signed-off-by: Krish Hung <krishung5@gmail.com>
Signed-off-by: krishung5 <krish@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 11:31:09 +00:00 |
|
 Jhao-Ting ChenandGitHub
|
28c222157b
|
fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs (#39391)
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
|
2026-04-21 15:04:41 +04:00 |
|
 wang.yuqiandGitHub
|
3975eb6de6
|
Revert "[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm" (#40438)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-21 08:47:18 +00:00 |
|
 Zeyu ZhangandGitHub
|
5a94a19824
|
[Bugfix] Normalize malformed dict prompts that carry token IDs in prompt (#40339)
Signed-off-by: Alchuang22-dev <2584829494@qq.com>
|
2026-04-21 07:44:36 +00:00 |
|
 hangy-amdandGitHub
|
f95c11a848
|
[Feat] dflash support for ROCm (#39703)
Signed-off-by: Hang Yang <hangy@amd.com>
|
2026-04-21 14:58:20 +08:00 |
|
 milesialandGitHub
|
257015d5e5
|
[MoE] Triton MoE Perf regression - restore low latency path (#39016)
|
2026-04-21 02:37:11 -04:00 |
|
 Shanshan ShenandGitHub
|
b47840019e
|
[MM][Misc] Support image+video mixed inputs (per prompt) for VLM examples (#40335)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-21 03:43:25 +00:00 |
|
 SeongJun LeeandGitHub
|
989cc12d88
|
[Fix] Add missing space in IP fallback warning (#40359)
Signed-off-by: lesj0610 <lesj0610@gmail.com>
|
2026-04-20 20:26:06 -07:00 |
|
 Wentao YeandGitHub
|
301024aa9c
|
[Deprecation] Deprecate cprofile and cprofile_context (#39100)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-21 11:25:22 +08:00 |
|
 Simon MoandGitHub
|
8256833fe6
|
[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm (#40331)
Signed-off-by: simon-mo <simon@inferact.ai>
|
2026-04-21 10:49:32 +08:00 |
|
 Shanshan ShenandGitHub
|
8097591286
|
[Doc] Update ViT CUDA graph doc for mixed (image+video) inputs (#40355)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-21 02:31:09 +00:00 |
|
 
|
20d3743491
|
[Bugfix] Gemma4: fix multimodal embedder norm order to match HF reference (#40411)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
|
2026-04-21 02:28:26 +00:00 |
|
 ChaunceyandGitHub
|
18563f2072
|
[Misc] Reduce attention logging levels (#40086)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-21 02:09:25 +00:00 |
|
 
|
0e884fe638
|
[Bugfix] Fix _CONFIG_REGISTRY types getting wrong config class when on-disk model_type differs (#39554)
Signed-off-by: Misa <misaAle@users.noreply.github.com>
Signed-off-by: Misael Casarez <misacasa@amazon.com>
Co-authored-by: Misael Casarez <misacasa@amazon.com>
|
2026-04-20 19:04:48 -07:00 |
|
  
|
fe5c115ee4
|
[vLLM IR] Add IR op testing and benchmarking infrastructure (#40167)
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com>
Co-authored-by: Theresa Shan <Theresa.Shan@amd.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 00:23:03 +00:00 |
|
 
|
6867bcd076
|
[Bugfix] Replace code that disabled shared expert overlap (#39222)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-20 19:36:16 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
c075702eae
|
[Misc][UX] Suppress confusing num_gpu_blocks log lines (#40402)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-20 22:32:46 +00:00 |
|
 Rita BrugarolasandGitHub
|
21b086d0aa
|
[ROCm] Hotfix: guard MLA dual RMS norm fusion against older AITer versions (#40386)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
|
2026-04-20 16:20:05 -05:00 |
|
 Sage MooreandGitHub
|
3173441b0f
|
[EPLB] Consolidate is_unchanged/is_received_locally into TransferMetadata (#37341)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
|
2026-04-20 21:12:42 +00:00 |
|
 Cao QianandGitHub
|
8b1f3bebca
|
[LMCache MP Connector] Add num_lmcache_extra_cached_token in KVTransferParams (#39843)
Signed-off-by: aeon-x <talexcao@gmail.com>
|
2026-04-20 20:42:49 +00:00 |
|
  
|
2390caf157
|
Enable building MoRI with AMD AINIC stack (#38371)
Signed-off-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com>
Signed-off-by: Theresa Shan <theresa.shan@amd.com>
Co-authored-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-04-20 11:17:59 -07:00 |
|
 Frederik GossenandGitHub
|
87805fa11e
|
[Core] Cache InductorPass.hash_source with functools.cache (#39328)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 14:06:15 -04:00 |
|
 Nicolò LucchesiandGitHub
|
304d5ba1a0
|
[Bugfix][CI] Fix tests/distributed/test_torchrun_example_moe.py (#40349)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-04-20 11:05:44 -07:00 |
|
 Tyler Michael SmithandGitHub
|
81d954f454
|
[WideEP] Remove naive all2all. Use allgather_reducescatter instead (#33728)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-04-20 17:53:55 +00:00 |
|
 Frederik GossenandGitHub
|
47fcb8ca68
|
[Core] Pass donate_graph_module=True to standalone_compile (#39733)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 17:40:52 +00:00 |
|
 baiandGitHub
|
191e3fdaa1
|
Update flashinfer to 0.6.8 (#39959)
Signed-off-by: bai <v@gor.io>
|
2026-04-20 10:37:23 -07:00 |
|
 Frederik GossenandGitHub
|
b9cf629bd0
|
[Core] Label torch trace logging overhead with dynamo_timed (#39329)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 17:31:03 +00:00 |
|
 
|
3461c8b027
|
[EPLB] Refactor Async EPLB synchronization logic (#37601)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2026-04-20 17:05:41 +00:00 |
|
 
|
726efe177b
|
[MoE Refactor] Move the shared/fused expert output sum into MoERunnerBase (#35949)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-20 12:28:46 -04:00 |
|
 Yan MaandGitHub
|
595562651a
|
[XPU] fix MoE triton backend in online fp8 quantization (#40109)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-20 11:31:39 -04:00 |
|
 Hashem HashemiandGitHub
|
3a30eaa1d7
|
Properly enable wvSplitK fp8 path for RDNA (#37712)
Signed-off-by: Hashem Hashemi <hashem.hashemi@amd.com>
|
2026-04-20 10:09:24 -05:00 |
|
 Rita BrugarolasandGitHub
|
fb5635d3f9
|
[ROCm] Add MLA dual RMS norm fusion (Q, KV) pass for DeepSeek/Kimi-K2 (#39242)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
|
2026-04-20 14:56:27 +00:00 |
|
 Wentao YeandGitHub
|
b42e878ec0
|
[Bug] Fix dcp error message (#40053)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-20 10:52:32 -04:00 |
|
 
|
7243e02aa1
|
[ROCm][Feature] Enable AITER MLA attention backend to work with Eagle3 speculative decoding on ROCm (#39616)
Signed-off-by: larryli2-amd <larryli2@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-04-20 09:44:43 -05:00 |
|
 Sage MooreandGitHub
|
def8f52200
|
[CI][EPLB] Add Async EPLB end-to-end integration test to CI (#40168)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
|
2026-04-20 10:22:54 -04:00 |
|
 Vasiliy KuznetsovandGitHub
|
38fa87caca
|
mxfp8 online quant move to new frontend (#40152)
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com>
|
2026-04-20 06:26:12 -07:00 |
|
 
|
a023edfa5b
|
[bugfix] Use only onlines CPUs in lscpu (#40161)
Signed-off-by: kse <kevin.sejourne@cloud-temple.com>
Co-authored-by: kse <kevin.sejourne@cloud-temple.com>
|
2026-04-20 13:19:57 +00:00 |
|