  
|
75ccdf3145
|
[Core] Update PyTorch to 2.13.0, torchvision to 0.28.0, triton to 3.7.1 (#48155)
Signed-off-by: Andrey Talman <atalman@users.noreply.github.com>
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-23 11:09:36 -07:00 |
|
 Andrey TalmanandGitHub
|
c6fe94b4d5
|
[CI] Bump PyTorch Compilation Unit Tests timeout to 150 min (#49606)
|
2026-07-23 11:08:45 -07:00 |
|
 
|
46f01a50ac
|
[CI][Bugfix] Fix test isolation in block_int8/ptpc_fp8 MoE kernel tests (#49609)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-23 12:26:02 -05:00 |
|
 Andreas KaratzasandGitHub
|
f00efc5265
|
[CI] Isolate cudagraph tests in child processes (#49510)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
|
2026-07-23 17:16:19 +00:00 |
|
 Wentao YeandGitHub
|
b0cb1da1bd
|
[DSv4 Perf] Skip topk and router when not needed, 3.4% E2E TTFT improvement for Decode case (#49486)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-23 13:08:08 -04:00 |
|
 Andreas KaratzasandGitHub
|
0e36e3bbd1
|
[CI] Use explicit devices in quantization tests (#49512)
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
|
2026-07-23 10:54:26 -06:00 |
|
 
|
494845e79f
|
Revert "[MRV2] Always build attn metadata at capture time" (#49364) (#49451)
Co-authored-by: vllm-agent CI bot <ci-bot@vllm-agent.local>
|
2026-07-23 09:51:33 -07:00 |
|
 yue.yuandGitHub
|
0416dab275
|
[Bugfix][Structured Output][Spec Decode] Advance grammar across reasoning boundary (#44993)
Signed-off-by: Allen.Yu <yuyue0225sc@163.com>
|
2026-07-23 09:14:15 -07:00 |
|
  
|
c8db00b16c
|
Fix GPTQ quantized Qwen3.5 MTP weight loading with spec decode (#48816)
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
Co-authored-by: noobHappylife <64898326+noobHappylife@users.noreply.github.com>
|
2026-07-23 06:59:48 -07:00 |
|
 Guan-Ming ChiuandGitHub
|
80c7683923
|
[Perf] Defer MM embeds loading off the event loop (#49477)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-23 13:58:50 +00:00 |
|
  
|
638d6e9757
|
[Bugfix][CI/Build] Fix Plamo2 HF runner crash on transformers v5 (_tied_weights_keys list→dict) (#44239)
Signed-off-by: Nikhil Kulkarni <nikhilkulkarni1755@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-23 12:25:14 +00:00 |
|
 Junpu YuandGitHub
|
1ad84fea86
|
[Bugfix][Spec Decode] Select earliest-completing stop string in check_stop_strings (#49391)
Signed-off-by: Junpu Yu <davidyu@nvidia.com>
|
2026-07-23 18:14:06 +08:00 |
|
 
|
12213c6795
|
[Bugfix] handle grammar compilation failures to avoid engine crash (#47312)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-23 18:13:49 +08:00 |
|
   
|
10c75477b0
|
[Bugfix][Core] shm_broadcast: bound idle reader waits and release read slots (#45224)
Signed-off-by: Chaemin Lim <chaemin.lim@mangoboost.io>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Edwin Lim <edwin.lim@mangoboost.io>
Co-authored-by: Jaeyoun Kim <jaeyoun.kim@mangoboost.io>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-23 18:13:29 +08:00 |
|
 
|
ac36a7a1e7
|
[MRV2][Spec Decode] Avoid rejection sampler OOM by chunking (#48630)
Signed-off-by: mgoin <mgoin64@gmail.com>
Signed-off-by: Michael Goin <mgoin64@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-23 18:13:13 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
521aa80f71
|
[Core] Simplify KVBlockZeroer index tensor handling (#48399)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-23 18:12:46 +08:00 |
|
  
|
a76df87db8
|
[MooncakeStore] Re-derive full external hits on stored boundaries (#49481)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-23 18:12:28 +08:00 |
|
   
|
a4904ba903
|
[Perf][KVConnector][Mooncake] Vectorize prepare_value on the KV load path (#48531)
Signed-off-by: girasoley <girasoley@inferact.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: girasoley <girasoley@inferact.ai>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-23 18:12:00 +08:00 |
|
 Rehan KhanandGitHub
|
f83de6d44c
|
[CPU][Docs] Update docs and dockerfile for s390x (#49523)
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com>
|
2026-07-23 07:17:24 +00:00 |
|
 Umut PolatandGitHub
|
239fc73553
|
[Misc] Use VLLMValidationError in chat_utils content-part validation (#49217)
Signed-off-by: Umut Polat <52835619+umut-polat@users.noreply.github.com>
|
2026-07-23 05:48:37 +00:00 |
|
 Mike GandGitHub
|
76bf55240c
|
[Bugfix] Fix DeepSeek-V4 DSpark draft shared-expert padding for TP > 8 (#49415)
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com>
|
2026-07-23 05:06:21 +00:00 |
|
  
|
9a698f3255
|
[Performance][Model] Avoid transient Inkling result allocations (performance, and OOM prevention on smaller memory configurations) (#49487)
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-23 05:03:34 +00:00 |
|
  
|
4080263bb2
|
[Bugfix][Model] Remove SciPy dependency from Inkling scale planning (#49485)
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-23 04:34:57 +00:00 |
|
 jcotant-inferactandGitHub
|
fc5fda105f
|
[Docs] Re-add Reo.dev analytics beacon (#49474)
|
2026-07-23 03:03:32 +00:00 |
|
 Matej SirovatkaandGitHub
|
b07ec92faa
|
[Bugfix] Make shared NVFP4 MoE scales writable (#49489)
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com>
|
2026-07-22 19:17:53 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
27ffbfde8d
|
Fused Shared Expert Support for AMD Quark DeepSeek-V4 Model Checkpoints (#48044)
Signed-off-by: Colin Zeng <Colin.Zeng@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-23 00:34:17 +00:00 |
|
 Nick HillandGitHub
|
229e01e9e1
|
[BugFix] Handle per-group prefix-hit divergence for hybrid models with KV connector (#48425)
|
2026-07-22 17:19:11 -07:00 |
|
 
|
191146dba5
|
Add quantization label automation (#49492)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-22 19:59:15 -04:00 |
|
 Summer YangandGitHub
|
f3a920a076
|
[Core][DSV4] Compact MXFP4 indexer KV cache and packed group overlays (#48993)
|
2026-07-22 16:58:44 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
149daf0d72
|
[Bugfix] Exclude location-derived path vars from torch.compile cache factors (#47573)
Signed-off-by: Nils Matteson <nilsmatteson@icloud.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-22 16:56:09 -07:00 |
|
 Michael GoinandGitHub
|
917fdb5bf7
|
[Bugfix] Fix DeepGEMM warmup when using FlashInferFp8DeepGEMMDynamicBlockScaledKernel (#49467)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-07-22 16:28:49 -07:00 |
|
 stefankoncarevicandGitHub
|
4b594b4aa1
|
[Bugfix][CI] Fix topk_softplus_sqrt no-op on non-XPU platforms (#49452)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
|
2026-07-22 15:36:17 -07:00 |
|
 
|
7d10a4cfce
|
[Bugfix] Retry config read to survive concurrent HF cache refresh (#49001)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-22 15:35:38 -07:00 |
|
 Nick HillandGitHub
|
910cc8543a
|
[Bugfix] Restore gather_and_maybe_dequant_cache OOB guard (#49427)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-22 13:47:37 -07:00 |
|
 Nick HillandGitHub
|
431934522b
|
[CI] Fix stale/fragile untethered kernels-root tests (#49423)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-22 14:43:07 -06:00 |
|
  
|
61a09532f2
|
Bump Flashinfer version to 0.6.15 (#48914)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Wei Zhao <weizha@oci-aga-slurm-1-vscode-02.cm.cluster>
Co-authored-by: Wei Zhao <weizha@oci-aga-slurm-1-vscode-02.cm.cluster>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-07-22 13:32:05 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
3de4b2bf3c
|
[Bugfix][Parser] Fix special tokens (EOS/BOS) leaking into reasoning content (#48748)
Signed-off-by: Ben Browning <56071+bbrowning@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-22 16:10:55 -04:00 |
|
 
|
b44311b6ef
|
[CI] stabilize GDN prefill CuTeDSL test (#49388)
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
Co-authored-by: Codex <noreply@openai.com>
|
2026-07-22 09:03:46 -07:00 |
|
 Nick HillandGitHub
|
b0d7875180
|
[CI] Increase timeout of pytorch-compilation-unit-tests (#49450)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-22 15:39:09 +00:00 |
|
 Divakar VermaandGitHub
|
53c2f20dd9
|
[ROCm][CI] skip moe weight padding for eplb (#49350)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-07-22 10:18:01 -05:00 |
|
 Wentao YeandGitHub
|
37e370fe93
|
[DSv4 Perf] Skip empty c128 kernel launch, around 2x kernel performance improvement. (#48957)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-22 10:55:20 -04:00 |
|
 Guan-Ming ChiuandGitHub
|
2dc5a72e7e
|
[Bugfix][Renderer] Rebuild vision chunk UUIDs in async render path (#49400)
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
|
2026-07-22 14:11:24 +00:00 |
|
 Andrey TalmanandGitHub
|
c79ff5f918
|
[Build] Bump vllm-flash-attn to C++20-compatible commit for torch-nightly (#49326)
Signed-off-by: Andrey Talman <atalman@fb.com>
|
2026-07-22 13:51:59 +00:00 |
|
 Teresa ChenandGitHub
|
1a659a0c37
|
Upgrade tpu-inference to v0.25.0 (#49431)
|
2026-07-22 11:52:56 +00:00 |
|
 SageandGitHub
|
0f6cf7f628
|
[Rust Frontend] Extract request preparation from the inference path (#49045)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
|
2026-07-22 11:31:36 +00:00 |
|
 
|
c79ad3ae21
|
[Rust Frontend][gRPC] Add abort control RPC (#49255)
Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Connor Carpenter <connorc@nvidia.com>
|
2026-07-22 11:31:01 +00:00 |
|
 wang.yuqiandGitHub
|
61c9ef986a
|
[Frontend] Parallelize preprocessing within the same request for pooling models online serving. (#49153)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-07-22 10:56:23 +00:00 |
|
 LiangqiusongandGitHub
|
d6dbdb9b0d
|
[XPU] WA of topk_softplus_sqrt arg mismatch on XPU (#49408)
Signed-off-by: xiaolong <xiaolong.guo@intel.com>
|
2026-07-22 16:16:13 +08:00 |
|
 liuzhenweiandGitHub
|
06da482fb4
|
[XPU] WA of topk_softmax arg mismatch on XPU (#49395)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-07-22 01:05:01 -07:00 |
|
 
|
2f75e7f712
|
[CI] Increase timeouts for jobs exceeding current limits (#49374)
Signed-off-by: khluu <khluu000@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-07-22 00:16:16 -07:00 |
|