 Andreas KaratzasandGitHub
|
76ea1d5d2f
|
[ROCm][CI] Stabilize Granite tool-use and test URL construction (#43017)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-23 12:21:11 +08:00 |
|
 Andreas KaratzasandGitHub
|
6a4723a2e0
|
[ROCm][CI] Stabilize runner teardown between sampler tests (#43023)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-23 12:19:54 +08:00 |
|
 Yongye ZhuandGitHub
|
367cb81966
|
[DSV4] More multi-stream enablement for c4a (#42925)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-23 09:22:27 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
3cb83c9592
|
Add model to WeightTransferEngine.__init__ (#42922)
Signed-off-by: SumanthRH <sumanthrh99@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-22 17:52:15 -07:00 |
|
 Duncan MossandGitHub
|
552bbe6f4e
|
[Attention] Add head_dim=512 support for FlashInfer trtllm attention backend (#38822)
|
2026-05-22 20:27:35 -04:00 |
|
 Itay AlroyandGitHub
|
6d30655b13
|
elastic_ep: stage/commit MoE quant method on reconfigure (#40881)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-05-22 18:57:26 -04:00 |
|
 
|
8de5cabeb7
|
[XPU]fix: add XPU platform guards to DeepSeek-V4 ops (#42950)
Signed-off-by: Ma Jian <jian1.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-23 06:29:45 +08:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4e2eba28be
|
[Perf] Optimize hidden state extraction logic (#37374)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-22 18:23:08 -04:00 |
|
 gnovackandGitHub
|
f743254143
|
DSv4 fused Q-norm kernel grid refactor (#42353)
|
2026-05-22 15:21:33 -07:00 |
|
 Nick HillandGitHub
|
47d4407d7c
|
[Model Runner V2] Support sharing kv cache layers (#35045)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-22 22:18:23 +00:00 |
|
 Juhi MittalandGitHub
|
e203006a8b
|
[Quantization][ModelOpt] W4A16 NVFP4 fused MoE + mixed-precision dispatch (#42566)
Signed-off-by: Juhi Mittal <juhim@nvidia.com>
|
2026-05-22 20:51:49 +00:00 |
|
 
|
08cb46789d
|
mhc_post - remove sts & add vectorized copies (#43437)
Signed-off-by: george <george@inferact.ai>
Co-authored-by: george <george@inferact.ai>
|
2026-05-22 13:44:29 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
4e597b7491
|
[Bugfix] Clear error message for FP8 torchao quantization on unsupported GPUs (#36854)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-22 20:09:17 +00:00 |
|
 Artem PerevedentsevandGitHub
|
23f7b11bf4
|
[Bugfix] Detect wrong libcute_dsl_runtime.so variant in FlashInfer GDN (#43427)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-05-22 19:33:33 +00:00 |
|
 
|
977703aa94
|
[RFC][EPLB][#32028] Remove dead torch.accelerator.synchronize() from sync path (#40733)
Signed-off-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com>
Co-authored-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com>
|
2026-05-22 15:19:24 -04:00 |
|
 
|
2b94d1c0ca
|
[Frontend] Simplify AuthenticationMiddleware path extraction (#43426)
Signed-off-by: Russell Bryant <rbryant@redhat.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-22 11:59:14 -07:00 |
|
 Yongye ZhuandGitHub
|
843715739b
|
[Refactor] Extract DeepSeek V4 sparse MLA impl into model folder (#43149)
|
2026-05-22 10:06:31 -07:00 |
|
 
|
b21f3d56d4
|
[KV Connector] MooncakeStore: don't co-queue save with load to avoid double delayed-free (#43371)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-22 16:14:11 +00:00 |
|
 
|
c7624bea5e
|
[Bugfix] Source num_qo_heads from Attention layers in Flashinfer/Triton metadata builders (#42650)
Signed-off-by: zhanda <zhandazhu@gmail.com>
Co-authored-by: Shang Wang <shangw@nvidia.com>
|
2026-05-22 16:10:03 +00:00 |
|
 Bugen ZhaoandGitHub
|
91f5b92438
|
[Rust Frontend] [Refactor] Extract a newtype for utility call ID (#43405)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-05-22 08:22:11 -07:00 |
|
 Isotr0pyandGitHub
|
f0feb15e7f
|
[Multimodal] Simplify ViT CUDA graph interfaces (#41234)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-22 22:31:00 +08:00 |
|
 sychen52andGitHub
|
fb21d8b4f9
|
Add NVFP4 MOE support for Deepseek V4. (#42209)
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
2026-05-22 07:21:51 -07:00 |
|
 haosdentandGitHub
|
a377631d21
|
[CI] Fix AMD docker build tests (#43329)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-22 14:06:24 +00:00 |
|
 
|
d3a563501b
|
[EPLB] Change default EPLB communicator (#43110)
Signed-off-by: Markov Ilya <markovilya19@gmail.com>
Co-authored-by: Markov Ilya <markovilya19@gmail.com>
|
2026-05-22 09:43:27 -04:00 |
|
 Jee Jee LiandGitHub
|
15f7cd33dc
|
[LoRA] Reduce memory of 2D weights when EP is set (#42737)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-22 06:41:56 -07:00 |
|
 
|
79ff0ffa98
|
[BugFix] wire make_empty_intermediate_tensors on AyaVision and Voxtral (#43118)
Signed-off-by: Keyi Li <likey6688@gmail.com>
Co-authored-by: Keyi Li <likey6688@gmail.com>
|
2026-05-22 05:26:41 -07:00 |
|
 Tobias WasnerandGitHub
|
4658bf882b
|
[Bugfix] Clear P0 mm sender cache on sleep/pause to fix mm_hash desync (#43001)
Signed-off-by: Tobias Wasner <wasnertobias@gmail.com>
|
2026-05-22 03:54:29 -07:00 |
|
  
|
b3c7ffcab8
|
[Misc] Replace assert with proper exceptions for security and validation in pooling (#43286)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-22 18:43:33 +08:00 |
|
 
|
d3d1cf6972
|
[XPU]feat: add XPU fallback for MoE topk routing and MXFP4 backend (#42951)
Signed-off-by: Ma Jian <jian1.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-22 10:22:45 +00:00 |
|
 wangxiyuanandGitHub
|
7e1b45a092
|
[Attention] Mamba attention module refactor (#41126)
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
|
2026-05-22 17:13:12 +08:00 |
|
 Li, JiangandGitHub
|
65b7a812a2
|
[CPU] Experimentally enable Triton and MRV2 (#43225)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-05-22 01:48:17 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
2380bfc210
|
[Docs] Note image preprocessing difference between qwen_vl_utils and vllm. (#43393)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-22 01:43:14 -07:00 |
|
 mrjunwan-langandGitHub
|
a761697717
|
Fix the docker build failure in tpu-inference (#43360)
Signed-off-by: mrjunwan-lang <mrjunwan@google.com>
|
2026-05-22 01:36:17 -07:00 |
|
 Nick HillandGitHub
|
694d9a81bb
|
[BugFix] Fix setuptools-rust dep in requirements files (#43377)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-22 15:25:10 +08:00 |
|
 Weida HongandGitHub
|
6bb8753db1
|
Correcting the mock classes for MM GC tests (#43321)
Signed-off-by: Weida Hong <wdhongtw@google.com>
|
2026-05-22 15:21:35 +08:00 |
|
 haosdentandGitHub
|
025d4f5cd2
|
[CI] Fix "test_awq_load[gemma4-moe-*]" failure (#43296)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-22 07:13:59 +00:00 |
|
 
|
5ea76fa89a
|
[CI] Fix test_lora_with_spec_decode on V2 model runner (#43314)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-22 14:24:18 +08:00 |
|
 tc-mbandGitHub
|
fa1ff88b31
|
[Model] Fix MiniCPM-V 4.6 vit_merger qkv weight loading (#43213)
Signed-off-by: tc-mb <tianchi_cai@icloud.com>
|
2026-05-21 22:44:06 -07:00 |
|
 Furkan FandGitHub
|
e746a2eebf
|
[Model] Use AutoWeightsLoader for Voyage (#42972)
Signed-off-by: Furkan Fidan <dev@yufufi.com>
|
2026-05-22 05:28:23 +00:00 |
|
 haosdentandGitHub
|
1fe3303983
|
[CI] De-flake renderers/test_hf.py::test_resolve_content_format_fallbacks[Qwen/Qwen-VL-string] (#43064)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-22 12:15:22 +08:00 |
|
 
|
8c8b1825eb
|
[XPU] Enable multiple key kernels for sparse attention (#37888)
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com>
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-22 12:02:51 +08:00 |
|
 
|
18a27cc9a3
|
[Bugfix] Make CuMemAllocator free callback stream-aware (#43020)
Signed-off-by: zixi-qi <zixi@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-22 03:36:22 +00:00 |
|
   
|
0ddd7dd656
|
[Frontend] DP Supervisor (#40841)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: robertgshaw2-redhat <robertgshaw2@gmail.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-21 20:33:16 -07:00 |
|
 
|
60af5c16ee
|
[Frontend] Add truncation side to OpenAI endpoints (#43260)
Signed-off-by: Rui Zhang <rza21.bc@gmail.com>
Signed-off-by: Rui Zhang <rui.zhang@globalrelay.net>
Co-authored-by: Rui Zhang <rui.zhang@globalrelay.net>
|
2026-05-21 20:32:31 -07:00 |
|
 Divakar VermaandGitHub
|
35d0141a0b
|
[ROCm][CI] add warmup to mem_util test before measurement (#43236)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-05-22 03:17:54 +00:00 |
|
 Simon DanielssonandGitHub
|
86ccef7d44
|
[ROCm] Add XGMI backend for MoRI Connector (#41753)
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
2026-05-22 03:06:40 +00:00 |
|
 
|
2998a047aa
|
[Bugfix] Fix DSV4 Base model swiglu limit issue in FP8 path (#42855)
Signed-off-by: Chengze Fan <chengze@meta.com>
Signed-off-by: Chengze Fan <fancz2002@gmail.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
|
2026-05-21 19:43:01 -07:00 |
|
 Isotr0pyandGitHub
|
ba369b7eb5
|
[CI] Fix dockerfile dependency graph failure for pre-commit (#43378)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-22 10:26:05 +08:00 |
|
     
|
39910f2b25
|
[Rust Frontend] Move code from vllm-frontend-rs (#43283)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
Signed-off-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com>
Signed-off-by: Will.hou <1205157517@qq.com>
Signed-off-by: Will.hou <willamhou@ceresman.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Eric Curtin <eric.curtin@docker.com>
Co-authored-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com>
Co-authored-by: Will.hou <1205157517@qq.com>
Co-authored-by: Will.hou <willamhou@ceresman.com>
Please see https://github.com/Inferact/vllm-frontend-rs for full original commit history.
|
2026-05-21 17:21:48 -07:00 |
|
 Lanze LiuandGitHub
|
39d5fa96a7
|
[Bugfix] Zero stale is_prefilling in padded CUDA graph rows for Mamba (#41873)
Signed-off-by: Lanze Liu <lanzetech@gmail.com>
|
2026-05-21 15:42:42 -07:00 |
|