 Nick HillandGitHub
|
1d41009e81
|
[ModelRunner V2] Fix cross-attention block table sizing (#46753)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-26 16:34:21 -07:00 |
|
 Nick HillandGitHub
|
b94f212e37
|
[ModelRunner V2] Deduplicate ModelState init logic (#46776)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-26 16:32:45 -07:00 |
|
 Harry MellorandGitHub
|
d8eb734d94
|
Fix Transformers backend FP8 MoE and remove some boilerplate (#46820)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-27 00:16:05 +01:00 |
|
 
|
2ff76a5e85
|
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 (#46760)
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-26 21:58:40 +00:00 |
|
 Yifan QiaoandGitHub
|
75fdcc82a5
|
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas (#46873)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-06-26 21:48:53 +00:00 |
|
 yzong-rhandGitHub
|
77f8796d16
|
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. (#46437)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-06-26 17:18:47 -04:00 |
|
 
|
c40d307731
|
[Core] Remove FlashAttention block size restriction for hybrid models (#36701)
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-06-26 21:16:39 +00:00 |
|
 Woosuk KwonandGitHub
|
65e655d295
|
[GLM-5] Add DSV3.2/GLM5 to vllm/models/ (#46808)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-26 14:09:05 -07:00 |
|
 Charlie FuandGitHub
|
6e2fb02fe5
|
[ROCm][CI] Fix rlhf_nccl.py on ROCm (#46851)
Signed-off-by: charlifu <charlifu@amd.com>
|
2026-06-26 15:41:49 -05:00 |
|
 Micah WilliamsonandGitHub
|
274325dd43
|
[ROCm][CI] Remove V1 Sample + Logits from mi250 Queue (#46867)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-06-26 15:38:38 -05:00 |
|
 MattandGitHub
|
95e6442a6b
|
[Hardware][AMD][CI] Fix Kernels Quantization test timeout (#46859)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-26 15:19:16 -05:00 |
|
  
|
701a23d99f
|
[Bugfix][Model] Support tensor parallelism for DiffusionGemma (#45719) (#46177)
Signed-off-by: Carlos Alvarado <carlos-alvarado@outlook.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
|
2026-06-26 20:05:04 +00:00 |
|
 Ben BrowningandGitHub
|
dccb412e2c
|
[Bugfix][Parser] Pass token IDs to parser.parse() in Responses API and batch serving (#46843)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-26 19:29:52 +00:00 |
|
 
|
c6554f321c
|
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build (#46769)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-26 14:32:21 -04:00 |
|
 Julien DenizeandGitHub
|
3d3b96488f
|
Migrate Voxtral to mistral-common 1.11.5 audio API (#46705)
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com>
|
2026-06-26 11:06:31 -07:00 |
|
 Nick HillandGitHub
|
658b54efe4
|
[ModelRunner V2] Update scheduler tests to cover MRV2 paths (#46771)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-26 09:36:31 -07:00 |
|
 Li, JiangandGitHub
|
abc71548ef
|
[CI/Build][CPU] Add test image cache clean-up (#46831)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-06-26 23:28:49 +08:00 |
|
 Nick HillandGitHub
|
4e07ca2c92
|
[Core] Add VLLM_GPU_SYNC_CHECK env var (#44800)
|
2026-06-26 08:24:33 -07:00 |
|
 Bugen ZhaoandGitHub
|
e71bc6da85
|
[Rust Frontend] Use oss-harmony for Harmony output processing (#46799)
|
2026-06-26 08:24:13 -07:00 |
|
 fxmarty-amdandGitHub
|
37ce34922f
|
[CI] Fix failing CUDA graph capture in Triton MOE (#46735)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
|
2026-06-26 07:21:20 -07:00 |
|
 
|
c2507fb293
|
[ROCm] [MoE] [Perf] Shared-expert fusion for bias-routed MoE; enable on MiniMax-M3 mxfp8 model (#46545)
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-26 07:05:20 -07:00 |
|
 TJianandGitHub
|
8921c4be88
|
[ROCm] [Performance] Optimize aiter moe for DeepSeekV4 (#46122)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-26 06:43:27 -07:00 |
|
 
|
8e394244a5
|
[ROCm]Enable AITER MoE backend for MiniMax-M3-MXFP4 (#46419)
Signed-off-by: Qiang Li <qiang.li2@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-26 06:35:35 -07:00 |
|
 TJianandGitHub
|
302954e5f6
|
[ROCm] [CI] fix transcription flakiness AMD: Entrypoints Integration (API Server OpenAI - Part 1) (mi325_1) (#46823)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-26 21:33:35 +08:00 |
|
 Hyunkyun MoonandGitHub
|
950ee4c2e4
|
[API] Add token offsets to render endpoints (/v1/.../render) (#44226)
Signed-off-by: HyunKyun Moon <mhg5303@gmail.com>
|
2026-06-26 05:02:52 -07:00 |
|
 
|
d980a3cc6e
|
[ROCm] Fix AITER_UNIFIED_ATTN Dispatching After AITER Bump (#46780)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com>
|
2026-06-26 02:09:56 -07:00 |
|
 
|
bf292b5f6b
|
[Docs] Remove BambaForCausalLM from supported hybrid models list (#46071)
Signed-off-by: liejiang <jianglie2023@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
2026-06-26 08:02:50 +00:00 |
|
 wang.yuqiandGitHub
|
5e3dad04b1
|
[Misc] Move the legacy api_server.py to the examples directory. (#46783)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-26 07:43:29 +00:00 |
|
 Joe RowellandGitHub
|
63e161f296
|
[Bugfix][Tool Parser] PoolsideV1: fix string whitespace and required named tool choice (#46486)
Signed-off-by: Joe Rowell <joerowell4@gmail.com>
|
2026-06-26 06:05:16 +00:00 |
|
 Tiezhen WANGandGitHub
|
c7645bce04
|
Remove grok model arch from vllm (#46706)
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com>
|
2026-06-25 23:02:10 -07:00 |
|
 
|
35a49fcfc2
|
[CI][Bugfix] Spawn engine in mm cache sleep test to fix ROCm HIP error (#46749)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-26 00:38:26 -05:00 |
|
 peizhang56andGitHub
|
915e99ec67
|
[ROCm][Bugfix] Fix HIP fork re-init in multimodal offline examples (#46741)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
|
2026-06-26 00:37:47 -05:00 |
|
 Nick HillandGitHub
|
5b33041746
|
[ModelRunner V2] Fix whisper test (#46773)
|
2026-06-25 22:10:36 -07:00 |
|
 MattandGitHub
|
1a4984520e
|
[Hardware][AMD][CI] Fix AMD CI image build (#46792)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 22:05:12 -07:00 |
|
 ReidandGitHub
|
e312c5cb25
|
[Rust Frontend] Make Granite4 string argument scanning incremental (#46507)
Signed-off-by: reidliu41 <reid201711@gmail.com>
|
2026-06-26 03:54:03 +00:00 |
|
 Matti4andGitHub
|
1502cf6274
|
Fix relative allowed local media paths (#45263)
|
2026-06-25 20:45:20 -07:00 |
|
 
|
d350fa8ddd
|
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens (#46733)
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-26 03:41:33 +00:00 |
|
 
|
dbc49b6b99
|
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 (#45166)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
Signed-off-by: ovidiusm <ovidium@nvidia.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-06-25 19:33:42 -07:00 |
|
 fxmarty-amdandGitHub
|
552a9dbe59
|
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears (#44667)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
|
2026-06-25 19:33:00 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
02a1f23711
|
[DFlash] Fuse precompute kv per-layer rmsnorms (#46761)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:32:07 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
652d962bc9
|
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation (#46448)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:30:07 -07:00 |
|
 Giancarlo DelfinandGitHub
|
5314665bad
|
[Model Runner V2][DFlash] Enable dflash attention backend selection (#46770)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-06-25 19:29:25 -07:00 |
|
 Michael GoinandGitHub
|
3daea7ceb9
|
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models (#46759)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-06-25 19:03:09 -07:00 |
|
 Wentao YeandGitHub
|
cc7981599e
|
[Refactor] Remove dead kernel code (#46405)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-25 18:09:56 -07:00 |
|
 Nick HillandGitHub
|
32bb3195f0
|
[ModelRunner V2] Bound memory for large logprobs requests (#46746)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-25 18:04:06 -07:00 |
|
 
|
ad28d605e6
|
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) (#45544)
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
|
2026-06-25 17:46:28 -07:00 |
|
 Bugen ZhaoandGitHub
|
ae7c8ec223
|
[Rust Frontend] Switch rustls to native-tls/OpenSSL (#46696)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 17:19:44 -07:00 |
|
 Bugen ZhaoandGitHub
|
1d3f4cb3a4
|
[Rust Frontend] Extract renderer fixture test utilities (#46719)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 17:12:38 -07:00 |
|
 Bugen ZhaoandGitHub
|
f9e684499f
|
[Rust Frontend] Migrate gemma4 to unified parser (#46602)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 16:59:57 -07:00 |
|
 Giancarlo DelfinandGitHub
|
c53994e134
|
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling (#46665)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-06-25 23:46:10 +00:00 |
|