![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
d2e2e856ad
|
[Frontend] Remove frontend pooling multi task support. (#37861)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-21 12:27:44 +00:00 |
|
 
|
766cb65d00
|
feat(multimodal): support externally processed mm_kwargs with cache injection (#39502)
Signed-off-by: Krish Hung <krishung5@gmail.com>
Signed-off-by: krishung5 <krish@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 11:31:09 +00:00 |
|
 Jhao-Ting ChenandGitHub
|
28c222157b
|
fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs (#39391)
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
|
2026-04-21 15:04:41 +04:00 |
|
 wang.yuqiandGitHub
|
3975eb6de6
|
Revert "[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm" (#40438)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-21 08:47:18 +00:00 |
|
 Zeyu ZhangandGitHub
|
5a94a19824
|
[Bugfix] Normalize malformed dict prompts that carry token IDs in prompt (#40339)
Signed-off-by: Alchuang22-dev <2584829494@qq.com>
|
2026-04-21 07:44:36 +00:00 |
|
 hangy-amdandGitHub
|
f95c11a848
|
[Feat] dflash support for ROCm (#39703)
Signed-off-by: Hang Yang <hangy@amd.com>
|
2026-04-21 14:58:20 +08:00 |
|
 milesialandGitHub
|
257015d5e5
|
[MoE] Triton MoE Perf regression - restore low latency path (#39016)
|
2026-04-21 02:37:11 -04:00 |
|
 Shanshan ShenandGitHub
|
b47840019e
|
[MM][Misc] Support image+video mixed inputs (per prompt) for VLM examples (#40335)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-21 03:43:25 +00:00 |
|
 SeongJun LeeandGitHub
|
989cc12d88
|
[Fix] Add missing space in IP fallback warning (#40359)
Signed-off-by: lesj0610 <lesj0610@gmail.com>
|
2026-04-20 20:26:06 -07:00 |
|
 Wentao YeandGitHub
|
301024aa9c
|
[Deprecation] Deprecate cprofile and cprofile_context (#39100)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-21 11:25:22 +08:00 |
|
 Simon MoandGitHub
|
8256833fe6
|
[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm (#40331)
Signed-off-by: simon-mo <simon@inferact.ai>
|
2026-04-21 10:49:32 +08:00 |
|
 Shanshan ShenandGitHub
|
8097591286
|
[Doc] Update ViT CUDA graph doc for mixed (image+video) inputs (#40355)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-04-21 02:31:09 +00:00 |
|
 
|
20d3743491
|
[Bugfix] Gemma4: fix multimodal embedder norm order to match HF reference (#40411)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
|
2026-04-21 02:28:26 +00:00 |
|
 ChaunceyandGitHub
|
18563f2072
|
[Misc] Reduce attention logging levels (#40086)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-21 02:09:25 +00:00 |
|
 
|
0e884fe638
|
[Bugfix] Fix _CONFIG_REGISTRY types getting wrong config class when on-disk model_type differs (#39554)
Signed-off-by: Misa <misaAle@users.noreply.github.com>
Signed-off-by: Misael Casarez <misacasa@amazon.com>
Co-authored-by: Misael Casarez <misacasa@amazon.com>
|
2026-04-20 19:04:48 -07:00 |
|
  
|
fe5c115ee4
|
[vLLM IR] Add IR op testing and benchmarking infrastructure (#40167)
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com>
Co-authored-by: Theresa Shan <Theresa.Shan@amd.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 00:23:03 +00:00 |
|
 
|
6867bcd076
|
[Bugfix] Replace code that disabled shared expert overlap (#39222)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-20 19:36:16 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
c075702eae
|
[Misc][UX] Suppress confusing num_gpu_blocks log lines (#40402)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-20 22:32:46 +00:00 |
|
 Rita BrugarolasandGitHub
|
21b086d0aa
|
[ROCm] Hotfix: guard MLA dual RMS norm fusion against older AITer versions (#40386)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
|
2026-04-20 16:20:05 -05:00 |
|
 Sage MooreandGitHub
|
3173441b0f
|
[EPLB] Consolidate is_unchanged/is_received_locally into TransferMetadata (#37341)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
|
2026-04-20 21:12:42 +00:00 |
|
 Cao QianandGitHub
|
8b1f3bebca
|
[LMCache MP Connector] Add num_lmcache_extra_cached_token in KVTransferParams (#39843)
Signed-off-by: aeon-x <talexcao@gmail.com>
|
2026-04-20 20:42:49 +00:00 |
|
  
|
2390caf157
|
Enable building MoRI with AMD AINIC stack (#38371)
Signed-off-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com>
Signed-off-by: Theresa Shan <theresa.shan@amd.com>
Co-authored-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-04-20 11:17:59 -07:00 |
|
 Frederik GossenandGitHub
|
87805fa11e
|
[Core] Cache InductorPass.hash_source with functools.cache (#39328)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 14:06:15 -04:00 |
|
 Nicolò LucchesiandGitHub
|
304d5ba1a0
|
[Bugfix][CI] Fix tests/distributed/test_torchrun_example_moe.py (#40349)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-04-20 11:05:44 -07:00 |
|
 Tyler Michael SmithandGitHub
|
81d954f454
|
[WideEP] Remove naive all2all. Use allgather_reducescatter instead (#33728)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-04-20 17:53:55 +00:00 |
|
 Frederik GossenandGitHub
|
47fcb8ca68
|
[Core] Pass donate_graph_module=True to standalone_compile (#39733)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 17:40:52 +00:00 |
|
 baiandGitHub
|
191e3fdaa1
|
Update flashinfer to 0.6.8 (#39959)
Signed-off-by: bai <v@gor.io>
|
2026-04-20 10:37:23 -07:00 |
|
 Frederik GossenandGitHub
|
b9cf629bd0
|
[Core] Label torch trace logging overhead with dynamo_timed (#39329)
Signed-off-by: Frederik Gossen <frgossen@meta.com>
|
2026-04-20 17:31:03 +00:00 |
|
 
|
3461c8b027
|
[EPLB] Refactor Async EPLB synchronization logic (#37601)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2026-04-20 17:05:41 +00:00 |
|
 
|
726efe177b
|
[MoE Refactor] Move the shared/fused expert output sum into MoERunnerBase (#35949)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-04-20 12:28:46 -04:00 |
|
 Yan MaandGitHub
|
595562651a
|
[XPU] fix MoE triton backend in online fp8 quantization (#40109)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-20 11:31:39 -04:00 |
|
 Hashem HashemiandGitHub
|
3a30eaa1d7
|
Properly enable wvSplitK fp8 path for RDNA (#37712)
Signed-off-by: Hashem Hashemi <hashem.hashemi@amd.com>
|
2026-04-20 10:09:24 -05:00 |
|
 Rita BrugarolasandGitHub
|
fb5635d3f9
|
[ROCm] Add MLA dual RMS norm fusion (Q, KV) pass for DeepSeek/Kimi-K2 (#39242)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
|
2026-04-20 14:56:27 +00:00 |
|
 Wentao YeandGitHub
|
b42e878ec0
|
[Bug] Fix dcp error message (#40053)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-20 10:52:32 -04:00 |
|
 
|
7243e02aa1
|
[ROCm][Feature] Enable AITER MLA attention backend to work with Eagle3 speculative decoding on ROCm (#39616)
Signed-off-by: larryli2-amd <larryli2@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-04-20 09:44:43 -05:00 |
|
 Sage MooreandGitHub
|
def8f52200
|
[CI][EPLB] Add Async EPLB end-to-end integration test to CI (#40168)
Signed-off-by: Sage Moore <sage@neuralmagic.com>
|
2026-04-20 10:22:54 -04:00 |
|
 Vasiliy KuznetsovandGitHub
|
38fa87caca
|
mxfp8 online quant move to new frontend (#40152)
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com>
|
2026-04-20 06:26:12 -07:00 |
|
 
|
a023edfa5b
|
[bugfix] Use only onlines CPUs in lscpu (#40161)
Signed-off-by: kse <kevin.sejourne@cloud-temple.com>
Co-authored-by: kse <kevin.sejourne@cloud-temple.com>
|
2026-04-20 13:19:57 +00:00 |
|
 
|
b82fc1364d
|
[Anthropic][Frontend] Added chat_template_kwargs to /v1/messages (#40125)
Signed-off-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com>
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com>
|
2026-04-20 06:10:45 -07:00 |
|
 Yan MaandGitHub
|
e06de7f005
|
[XPU] enable triton attention test on XPU by removing cuda device binding (#39627)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-04-20 20:57:11 +08:00 |
|
 zhanqiuhuandGitHub
|
cc3993b05d
|
nixl refactor [2/N]: unify TpKVTopology + HeteroTPTransferConfig into TransferTopology (#39529)
Signed-off-by: Zhanqiu Hu <zhu@redhat.com>
|
2026-04-20 12:39:08 +02:00 |
|
 Ilya MarkovandGitHub
|
50dd4cb427
|
[EPLB] Add nixl-based eplb communicator (#36276)
Signed-off-by: ilmarkov <markovilya197@gmail.com>
Signed-off-by: Markov Ilya <markovilya19@gmail.com>
|
2026-04-20 10:24:23 +00:00 |
|
 
|
f774ba028a
|
[kv_offload+HMA][4/N]: Support sliding window lookup (#36645)
Signed-off-by: Or Ozeri <oro@il.ibm.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-04-20 12:53:51 +03:00 |
|
 Fadi ArafehandGitHub
|
2aab9acf48
|
[CPU][BugFix] Fix inter-node pipeline parallel (#40150)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
|
2026-04-20 17:21:12 +08:00 |
|
 nemanjaudovicandGitHub
|
58631d7c3f
|
[Bugfix] Fix scaled_mm output narrowing for 3D input tensors (#38093)
Signed-off-by: nemanjaudovic <nudovic@amd.com>
|
2026-04-20 16:58:39 +08:00 |
|
 Andreas KaratzasandGitHub
|
a943839e9a
|
[ROCm][CI] Introducing new MI300 nodes (#39531)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-20 16:09:58 +08:00 |
|
 milesialandGitHub
|
6d8b80802b
|
[Docs] Fix thinking_token_budget docs (#40316)
Signed-off-by: milesial <milesial@users.noreply.github.com>
|
2026-04-20 08:09:44 +00:00 |
|
 wuyingjunandGitHub
|
77fd2c8631
|
[Bugfix] Forward mm_processor_kwargs in offline generate APIs (#40251)
Signed-off-by: wuyingjun <wuyingjun_yewu@cmss.chinamobile.com>
|
2026-04-20 00:56:56 -07:00 |
|
 San-NguyenandGitHub
|
e729cc823d
|
[Fix] Add Spacing when Requesting Output Token > max_model_len (#40324)
Signed-off-by: San-Nguyen <san.nguyen@ibm.com>
|
2026-04-20 00:25:06 -07:00 |
|
 velonica0andGitHub
|
ec7aafc02a
|
[CPU][RISC-V] Support multiple RVV VLEN targets via compile-time dispatch (#39478)
Signed-off-by: velonica0 <like@mail.nankai.edu.cn>
|
2026-04-20 14:36:59 +08:00 |
|