 
|
a2abce646f
|
[EPLB] Mask padding in EPLB load recording (#38128)
Signed-off-by: ilmarkov <markovilya197@gmail.com>
Signed-off-by: Markov Ilya <markovilya19@gmail.com>
Co-authored-by: Markov Ilya <markovilya19@gmail.com>
|
2026-06-28 19:43:58 -07:00 |
|
 Harry MellorandGitHub
|
311ad689ad
|
Remove boilerplate missed by #46820 (#46956)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 08:11:17 +08:00 |
|
 Woosuk KwonandGitHub
|
0472436541
|
[Spec Decode] Avoid redundant hidden-states gather in draft prefill (#46968)
|
2026-06-28 17:04:01 -07:00 |
|
  
|
4dfbf1503b
|
[Model] Add support for openai/privacy-filter (#41026)
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-06-28 16:18:22 -07:00 |
|
 Wei ZhaoandGitHub
|
95528527ea
|
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 (#46855)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-06-28 14:36:23 -07:00 |
|
  
|
c2127a25c7
|
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm (#46895)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:50:30 -05:00 |
|
 
|
03c6d01c30
|
[OCP MX ] Add back emulation to available OCP MX backends list (#46629)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:43:19 -05:00 |
|
 Woosuk KwonandGitHub
|
4b643c463e
|
[GLM5] Fix minor typo (#46961)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-28 08:37:00 -07:00 |
|
 
|
7544286b04
|
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE (#46552)
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-28 15:19:28 +00:00 |
|
 Woosuk KwonandGitHub
|
89876b0c54
|
[GLM5] Implement op fusion for GLM5/DSV3.2 (#46876)
|
2026-06-28 08:17:39 -07:00 |
|
 Wentao YeandGitHub
|
5c91039c41
|
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement (#46635)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-28 14:55:54 +00:00 |
|
 
|
5ecae3266c
|
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) (#45033)
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-28 07:52:00 -07:00 |
|
 
|
6eb63a1da6
|
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers (#46600)
Signed-off-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 01:37:44 -07:00 |
|
 
|
09841ae705
|
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation (#46846)
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>
|
2026-06-28 00:07:33 -07:00 |
|
 MattandGitHub
|
a2a92cbbaa
|
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection (#46930)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-28 00:07:14 -07:00 |
|
 
|
35e6c86caa
|
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL (#46034)
Signed-off-by: shen-shanshan <467638484@qq.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-28 00:06:43 -07:00 |
|
  
|
c7ca0bccae
|
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 (#44313)
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com>
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com>
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com>
|
2026-06-28 00:04:08 -07:00 |
|
  
|
c6741b2ad4
|
[Model] Support Unlimited OCR (#46564)
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-06-27 23:09:18 -07:00 |
|
 
|
a65f93fb2e
|
[ROCm][CI] Add ci_base metadata for external cache orchestration (#46886)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Codex <codex@example.invalid>
Co-authored-by: Codex <codex@example.invalid>
|
2026-06-28 12:51:19 +08:00 |
|
 ChaunceyandGitHub
|
11a12305c0
|
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models (#46786)
|
2026-06-27 18:38:07 -07:00 |
|
 
|
798185d438
|
[KV-Offloading] Fix tensors_per_block stride (#46888)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-27 21:01:45 -04:00 |
|
 MattandGitHub
|
9036c89ee4
|
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now (#46928)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-27 17:30:49 -05:00 |
|
 Giancarlo DelfinandGitHub
|
b6caeb5a09
|
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance (#46878)
|
2026-06-27 14:09:25 -07:00 |
|
 Taneem IbrahimandGitHub
|
8bf064f8d3
|
Fixed chunked embedding aggregation with request-id metadata (#46782)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-06-27 20:57:47 +00:00 |
|
 
|
ea2ead1db3
|
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod (#46818)
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com>
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com>
|
2026-06-27 20:23:59 +00:00 |
|
 Wentao YeandGitHub
|
56aa067bf0
|
[CI Bug] Fix h100 AssertionError: Cold-start child failed (#46927)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-27 20:17:33 +00:00 |
|
 xiaolinchenandGitHub
|
35e3850fa9
|
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) (#46915)
Signed-off-by: wentian-byte <2990624738@qq.com>
|
2026-06-27 14:30:10 -04:00 |
|
  
|
51a99565c3
|
[ROCm][Perf] Fused shared expert for Minimax M3 (#46474)
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-27 12:34:17 +00:00 |
|
    
|
867fd5e8ed
|
[ROCm][Perf] Use flydsl moe with Minimax-M3 mxfp8 weights on gfx950 and implemented moe-backend selection (#46184)
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com>
|
2026-06-27 10:22:57 +00:00 |
|
 
|
9fd00ee006
|
[ROCm][CI] Move remaining mi250_2 tests out of the MI250 queue (#46905)
Signed-off-by: Codex <codex@example.invalid>
Co-authored-by: Codex <codex@example.invalid>
|
2026-06-27 17:08:54 +08:00 |
|
 
|
091d13976c
|
[ROCm][CI] Add TRITON_ATTN score absolute tolerance floor (#46891)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-27 06:35:50 +00:00 |
|
 Wentao YeandGitHub
|
b588f66dc2
|
[GLM5.2 Perf] fused_indexer_q_rope_quant triton kernel, 1.9% ~ 3.3% E2E Throughput improvement. (#46862)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-26 22:16:20 -07:00 |
|
 Benjamin ChislettandGitHub
|
455f25aa13
|
[CLI] Add flag to print TTFT and TPS in vllm chat (#46775)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-06-26 22:15:10 -07:00 |
|
     
|
d706dec904
|
fix: Correct reasoning-end detection for prompt history (#44551)
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn>
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com>
Signed-off-by: Jason Ozuzu <jasonozuzu@cohere.com>
Signed-off-by: walterbm <walter.beller.morales@gmail.com>
Co-authored-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: walterbm <walter.beller.morales@gmail.com>
Co-authored-by: Walter Beller-Morales <walterbm@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-06-26 22:15:06 -07:00 |
|
 Divakar VermaandGitHub
|
68ee8300a0
|
[ROCm][CI]Fix test_concat_and_cache_mla_rope_fused on ROCm (#46409)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-06-27 12:38:13 +08:00 |
|
  
|
ddd3855a28
|
[MoE Backend] add HPC-Ops MoE backend (#45924)
Signed-off-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: youkaichao <youkaichao@gmail.com>
|
2026-06-27 11:18:07 +08:00 |
|
 Divakar VermaandGitHub
|
00e045b7c7
|
[ROCm][CI TG] refactor and fix deepep_moe test group (#46758)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-06-27 10:45:23 +08:00 |
|
 Divakar VermaandGitHub
|
17a71d8702
|
[ROCm][CI] Relax fused layernorm quant test tolerances for one-ULP outliers (#46658)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-06-27 10:44:29 +08:00 |
|
 weizhoublueandGitHub
|
2e058851d3
|
fix(docker): eliminate race conditions in shared buildkit cache mounts (#44984)
|
2026-06-26 19:43:17 -07:00 |
|
 DāvisandGitHub
|
1a92dfcce4
|
[Build] Show error message when using ROCm with LTO and different compilers (#35232)
|
2026-06-26 19:43:00 -07:00 |
|
 Chris LeonardandGitHub
|
d0f800811b
|
[Build] Update vllm to point to vllm-project/flash-attention commit that builds FA3 with torch stable API. (#46644)
|
2026-06-26 19:42:46 -07:00 |
|
 Nick HillandGitHub
|
c6dd32a810
|
[ModelRunner V2] Support realtime embeddings (#46762)
|
2026-06-26 19:42:27 -07:00 |
|
 
|
af16446bf3
|
Vram semaphore infra (#44465)
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-06-26 17:32:51 -07:00 |
|
 Harry MellorandGitHub
|
3f67477497
|
[CI] Don't try and download files that we already know don't exist (#46854)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-26 23:56:39 +00:00 |
|
 Nick HillandGitHub
|
1d41009e81
|
[ModelRunner V2] Fix cross-attention block table sizing (#46753)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-26 16:34:21 -07:00 |
|
 Nick HillandGitHub
|
b94f212e37
|
[ModelRunner V2] Deduplicate ModelState init logic (#46776)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-26 16:32:45 -07:00 |
|
 Harry MellorandGitHub
|
d8eb734d94
|
Fix Transformers backend FP8 MoE and remove some boilerplate (#46820)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-27 00:16:05 +01:00 |
|
 
|
2ff76a5e85
|
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 (#46760)
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-26 21:58:40 +00:00 |
|
 Yifan QiaoandGitHub
|
75fdcc82a5
|
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas (#46873)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
|
2026-06-26 21:48:53 +00:00 |
|
 yzong-rhandGitHub
|
77f8796d16
|
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. (#46437)
Signed-off-by: Yifan Zong <yzong@redhat.com>
|
2026-06-26 17:18:47 -04:00 |
|