Bugen Zhao
c716d13203
add json passthrough support
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-07 11:28:29 +08:00
Bugen Zhao
29d9da5829
extract convert_params
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-06 21:44:09 +08:00
Bugen Zhao
7acfc161ec
add tool schema resolve
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-06 16:59:48 +08:00
Bugen Zhao
0578b70ea9
rename wrapper to section
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 21:47:54 +08:00
Bugen Zhao
9ca4fb8bdb
introduce config-driven parser
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 19:05:26 +08:00
wang.yuqi and GitHub
a14f57a3ac
[Frontend] Refine the entrypoint class's inheritance hierarchy. ( #47498 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-07-03 10:50:06 +00:00
18f658bb31
[Bugfix][Frontend] Fix batch chat endpoint corrupting logprobs when return_token_ids is set ( #47384 )
...
Signed-off-by: David Feng <fenghourun@meta.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-03 03:01:34 -07:00
Isotr0py and GitHub
400a9c386d
[Rust Frontend] Bump llm-multimodal version ( #47530 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-07-03 09:48:36 +00:00
Max de Bayser and GitHub
bbdcbe4686
Move Roberta remaining nn.Embedding to VocabParallelEmbedding ( #47452 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
2026-07-03 09:47:50 +00:00
Kalyanam Dewri and GitHub
4875b4456b
[Doc] Fix VLM2Vec benchmark chat template path ( #47517 )
...
Signed-off-by: kalyanamdewri <kalyanampriyam@gmail.com >
2026-07-03 08:24:45 +00:00
Dakai An and GitHub
1f486d96a1
Add Triton Backend for Unlimited-OCR R-SWA ( #47102 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-07-03 00:11:50 -07:00
Bugen Zhao and GitHub
b790c84cde
[CI] Enable sccache for Rust build under CUDA/ROCm ( #45246 )
2026-07-02 23:45:41 -07:00
6429d5f527
[Rust Frontend] add repetition_detection support to sampling params ( #46684 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 14:06:23 +08:00
Chris Leonard and GitHub
fbc9ba6d30
New stable abi cleanup ( #46656 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-07-03 14:02:26 +08:00
xiangdong and GitHub
2dfaae752b
[XPU][CI]Fix dependency typo in Intel GPU CI ( #47510 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 04:11:47 +00:00
Evgeny Parshutin and GitHub
bd8d9021ce
[CPU][Build] Enable oneDNN ITT task collection by default for CPU primitive-level profiling ( #47467 )
...
Signed-off-by: Evgeny Parshutin <eugeny.parshutin@intel.com >
2026-07-03 04:00:19 +00:00
xiangdong and GitHub
3f0b773b30
[XPU][CI]Mv huggingface cache to larger disk in Intel GPU CI ( #47405 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 11:56:17 +08:00
Reid and GitHub
9b8e76589d
[Rust Frontend] Recover buffered text from incomplete tool calls at EOS ( #47289 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-07-03 03:45:03 +00:00
1aeabec355
[Bugfix][Rust Frontend] Tolerate out-of-vocab prompt ids in detokenizer ( #44682 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-03 03:41:53 +00:00
979f5511d7
[Bugfix][Gemma4] Keep image bidirectional attention within the sliding window ( #47217 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-07-02 19:57:41 -07:00
41de1380c2
[BugFix] Derive FlashInfer Q dtype from resolved per-group builder state ( #47485 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 19:33:28 -07:00
Nick Hill and GitHub
d85601c20f
[CI] Pin modelscope version to fix test breakage ( #47465 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:33:08 -07:00
Nick Hill and GitHub
276b837dc4
[ModelRunner V2][BugFix] Free all model refs on shutdown ( #47483 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 19:32:48 -07:00
34bf7b45a0
[CI] intel CI: add quantization and awq case for xpu ( #46456 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
2026-07-03 09:51:56 +08:00
adamkbaranowski and GitHub
4c3c64fcf7
Add Laguna XS.2.1 DFlash drafter support ( #46853 )
...
Signed-off-by: Adam Baranowski <adam.baranowski@poolside.ai >
2026-07-02 18:09:27 -07:00
Andreas Karatzas and GitHub
442ccc6098
[ROCm][CI] Adding extract hs 2gpu ( #47482 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:59:38 -07:00
Andreas Karatzas and GitHub
6768fbc76f
[ROCm][CI] Adding qwen3 dp4 eplb ( #47480 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:58:56 -07:00
Andreas Karatzas and GitHub
407f406300
[ROCm][CI] Adding metadata ( #47477 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-07-02 17:45:23 -07:00
Harry Mellor and GitHub
e24d1b24fe
Fix Transformers modeling backend usage stats ( #47472 )
2026-07-02 12:51:23 -07:00
d29125c085
Xqa decode kernels ( #43232 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-07-02 12:32:05 -07:00
Michael Goin and GitHub
d715b3aa1e
Delete PagedAttention ( #47361 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 12:31:26 -07:00
Joe Rowell and GitHub
258f8de91f
[Bugfix][Tool Parser] poolside_v1: accept tool calls without newline after function name ( #47311 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-07-02 12:08:37 -07:00
Nick Hill and GitHub
e392bf7a68
[BugFix][MRV2] Ensure all req slots are accounted for when scheduling ( #46974 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:24 -07:00
Nick Hill and GitHub
443e68cfa6
[Bugfix] Fix pooled Whisper encoder sliding-window kernel size ( #47437 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 10:19:11 -07:00
Chauncey and GitHub
320ee285c9
[Model Runner V2][Perf] Warm up GLM-5.2 DSA indexer prefill metadata kernel ( #47285 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-02 16:31:31 +00:00
Bugen Zhao and GitHub
ec0ffaacc8
[Rust Frontend] Improve scheduler stats logging parity ( #47435 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 15:55:25 +01:00
Yuxuan Zhang and GitHub
178fd56094
support GLM-5.2 gate use FP32 ( #47410 )
...
Signed-off-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk >
2026-07-02 22:45:39 +08:00
a47f38f825
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in block verification kernels ( #47383 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-07-02 07:32:38 -07:00
Nick Hill and GitHub
3e158ae62d
[ModelRunner V2] Fix Mamba2 crash on non-spec-decode ( #47428 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 07:05:16 -07:00
a2f713002d
[ModelRunner V2] Enable by default for all dense models ( #44443 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-02 18:48:57 +08:00
TJian and GitHub
de2a8fc042
[ROCm] [PyTorch] Move to stable abi since ROCm upgraded to torch 2.11 ( #47128 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-07-02 18:34:07 +08:00
Michael Goin and GitHub
84b9c2762f
Update DeepGEMM tag to point to latest nv-dev branch for sm120 support ( #47304 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-02 18:33:44 +08:00
Bugen Zhao and GitHub
25fcb65d51
[Rust Frontend] Use enum-backed domain types for engine outputs and structured outputs ( #47283 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 10:41:46 +01:00
08a8a4af3f
feat(rust): expose profiler control routes in Rust frontend ( #46306 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 08:46:07 +00:00
b0b8a286dd
[Model] Add LLaVA-OneVision-2 (LlavaOnevision2ForConditionalGeneration) ( #44785 )
...
Signed-off-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: chengzheng345 <209475443+chengzheng345@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-02 16:41:49 +08:00
3af8789559
[Feature] Universal speculative decoding for heterogeneous vocabularies (TLI) ( #38174 )
...
Signed-off-by: wan-danfeng <wandanfeng0802@gmail.com >
Signed-off-by: Wonderful <wandanfeng0802@gmail.com >
Co-authored-by: Wan_DF <wonderful199082@126.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-07-02 01:34:20 -07:00
Chaojun Zhang and GitHub
8357226f4f
[XPU][CI] Split test_punica_ops into separate pytest invocations for stability ( #47376 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-02 07:50:55 +00:00
Hiki and GitHub
2665ed704b
[Bugfix][Kernel] Correct FlashInfer CUTLASS MoE tuning token bound ( #46838 )
...
Signed-off-by: Haobin Guo <haobing@nvidia.com >
2026-07-02 05:11:00 +00:00
xaguilar-amd and GitHub
09663abde0
[ROCm][MLA] Fuse MLA q/kv RMSNorm + FP8 per-token quant in the FP8 attention path ( #44977 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
2026-07-02 13:00:41 +08:00
Giancarlo Delfin and GitHub
d63c8e9444
[BugFix][Spec Decode] Compact shared topk indices buffer after first MTP draft step ( #47238 )
2026-07-01 21:38:51 -07:00
Jee Jee Li and GitHub
1360c42fe6
[UX] Include NVTX in cuda.txt ( #47319 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 19:38:50 -07:00
d0a2584773
[Misc] Use functions instead of PTX for the PDL instruction ( #46984 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-07-01 19:38:35 -07:00
7fe7fa9cda
[CI][Bugfix] Rerun test_engine_log_metrics_ray on Ray GCS startup timeout ( #47208 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 21:32:09 -05:00
Michael Goin and GitHub
2b753ad200
[Spec Decode] DSpark speculators checkpoint support ( #47093 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-07-01 17:32:27 -07:00
e196268bad
[Docker] Remove unused Dockerfile.nightly_torch ( #47338 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:42 -07:00
e91f5f8439
[CI] Remove torch_nightly mirror tags (superseded by TORCH_NIGHTLY full-nightly build) ( #47342 )
...
Co-authored-by: Andrey Talman <atalman@users.noreply.github.com >
2026-07-01 16:19:06 -07:00
fa248139a0
[MoE] Plumb gemm1_alpha/beta/clamp_limit into TRT-LLM FP8 MoE ( #45723 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 14:34:05 -07:00
Yongye Zhu and GitHub
d3229431f9
[DSV4] Better MXFP8 quantization kernel ( #47229 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-07-01 14:33:51 -07:00
Nick Hill and GitHub
4787f2dd1b
[Bugfix] Don't read KV cache past seq_len in triton paged attn kernels ( #47305 )
2026-07-01 12:43:00 -07:00
Nick Hill and GitHub
8cfeb84dba
[ModelRunner V2] Warmup cross-attn properly in encoder-decoder case ( #47308 )
2026-07-01 12:36:48 -07:00
Chaitanya Sri Krishna Lolla and GitHub
5fd442187c
[ROCm][P/D] MoRIIO toy proxy: support JSON Content-Type for OpenAI clients. ( #46482 )
...
Signed-off-by: lcskrishna <lollachaitanya@gmail.com >
2026-07-01 19:17:05 +00:00
00eb7cefa3
[Bugfix] Prevent padding placeholders from reaching embeddings ( #47029 )
...
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 09:26:03 -07:00
Michał Ganczarenko and GitHub
c8bdcc0116
[Bench][BugFix] Fix empty decoder prompt for Cohere ASR in throughput benchmark ( #47135 )
...
Signed-off-by: Michal Ganczarenko <michal.ganczarenko@intel.com >
2026-07-01 15:42:27 +00:00
f5a8d73377
[Spec Decode] DSpark ( #46995 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-07-01 08:30:24 -07:00
63fcce4de1
[Bugfix] Fix GraniteMoeShared weight loading broken by #41184 ( #47031 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 22:39:12 +08:00
Bugen Zhao and GitHub
c638f9216a
[Rust Frontend] Split engine core DTOs into separate modules ( #47265 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-01 15:28:21 +01:00
Chaojun Zhang and GitHub
13c49f9845
[xpu][lora]: Align LoRA implementation with Punica GPU: fix _apply_expand rank mismatch, add_inputs hardcode, and MoE EP ( #45368 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-07-01 22:14:04 +08:00
Nick Hill and GitHub
f1cf6b0086
[CI] Fix segfault in tracing test ( #47299 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-07-01 14:00:37 +00:00
Harry Mellor and GitHub
a78c15616f
Migrate GPTBigCode and Starcoder2 to the Transformers modeling backend ( #30966 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 13:41:36 +00:00
5c4db60f01
docs(security): document gRPC interface as insecure for private use only ( #45903 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <rbryant@redhat.com >
2026-07-01 12:39:57 +00:00
4e5ca89cfe
[ROCm][MiniMax-M3] Cross-layer lightning-indexer top-k sharing ( #47269 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 10:50:09 +00:00
Harry Mellor and GitHub
a22e0dfc69
[Model] Remove AyaVision, MusicFlamingo ( #47263 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:39:33 +00:00
stevenkuang and GitHub
cc56379e28
[Model] Support Hy3 token suffix and JSON Schema array types ( #47192 )
...
Signed-off-by: stevenkuang-tencent <stevenkuang@tencent.com >
2026-07-01 10:16:07 +00:00
024b06b0dc
[Bugfix] Expose usage field in GenerateResponse for disaggregated serving ( #42748 )
...
Signed-off-by: AIvashov <ivashov.aleksey@proton.me >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-01 10:00:19 +00:00
Harry Mellor and GitHub
e7d0fcbc09
[CI] Fix various failures on main ( #47197 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:35:34 +01:00
akii96 and GitHub
aa8bb5562e
[ROCm][Perf][Bugfix] DSv4 indexer: use platform FP8 dtype (fnuz) for Q-quant on gfx942 ( #46730 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-07-01 17:33:55 +08:00
Andy Lo and GitHub
fa4bec9056
[Bugfix] Fix pooled Whisper sliding-window KV sizing ( #47071 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-07-01 11:33:19 +02:00
dee5da1dec
[Test] Run SageMaker handler-override tests in-process via TestClient ( #47250 )
...
Signed-off-by: Jyothirmai Kottu <jkottu@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 09:14:00 +00:00
ed41aa270a
[ROCm][DSV4] Use aiter mHC pre/post as the default ROCm path ( #43950 )
...
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com >
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 16:27:42 +08:00
77a9c5ae28
Weight sync refactor + move sparse nccl engine ( #44353 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Signed-off-by: haoaaron <ahao@anyscale.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-01 01:25:19 -07:00
f651a8a9a4
[XPU][UT]Enable ut qk_norm_rope_fusion ( #42486 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 07:38:03 +00:00
Jee Jee Li and GitHub
8f82be5705
[CI/Build] Fix LoRA testing ( #47242 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 15:36:13 +08:00
Nils Matteson and GitHub
a461070d1c
[Core] Make sleep-mode backend capability flags communicator-agnostic ( #47243 )
2026-07-01 07:17:44 +00:00
4470ae84de
Remove mantis ( #46806 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 07:13:58 +00:00
Chauncey and GitHub
697c34b97b
[Bugfix] Fix beam search candidate indexing when logprobs count varies ( #47126 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-01 07:07:06 +00:00
Blas Rodriguez Irizar and GitHub
5b431b905c
[Rust Frontend] Coerce completion max_tokens: null to default ( #47166 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-07-01 06:41:33 +00:00
89e99202f2
[CPU][Perf]Added tanh AOR for faster gelu activations. ( #44639 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Signed-off-by: almayne <anna.mayne@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-30 23:24:40 -07:00
Micah Williamson and GitHub
b446792306
[ROCm][Bugfix] Fix Triton "out of resource: shared memory" Error In One-Shot LoRA MoE ( #47209 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:36 -07:00
Micah Williamson and GitHub
c3b1f9e827
[ROCm][CI] Enable LoRA TP Distributed Test Group In AMD CI ( #47193 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:32 -07:00
Jonathan Mamou and GitHub
df802a87b7
[CPU] Remove speculative decoding stream overrides from CPUModelRunner ( #47162 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
2026-07-01 06:12:49 +00:00
Nils Matteson and GitHub
93d8f834dd
[Core] Pluggable sleep-mode backend abstraction (RFC #34303 ) ( #44074 )
2026-06-30 22:00:53 -07:00
Maria Guevara and GitHub
aeb35b90f0
[Rust Frontend] Add error context in tool parser failures ( #46512 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-07-01 12:48:55 +08:00
Gabriel Wu and GitHub
9a08a5118e
fix: skip cooperative top-K on SM120 ( #47164 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-06-30 21:32:54 -07:00
c5200d3565
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE ( #46076 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: GirasoleY <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 00:32:20 -04:00
Matt and GitHub
3c1396bab6
[Hardware][AMD][CI] Toggle test coredumps on ROCm debug agent ( #47222 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 23:30:10 -05:00
Benjamin Chislett and GitHub
9969466a59
[Spec Decode] Support SWA + DFlash for MiMo ( #46104 )
2026-06-30 20:34:47 -07:00
achyuthan.s and GitHub
3406e8f83d
[Bugfix][Frontend][gpt-oss] Return raw output when Harmony parser ends non-terminal ( #47062 )
...
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
2026-07-01 01:46:01 +00:00
a264e41975
[Distributed] Default FlashInfer allreduce to mnnvl on single node ( #47219 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-30 18:35:56 -07:00
Woosuk Kwon and GitHub
f098ee70c7
[GLM5] Support FlashMLA FP8 KV cache (Hopper & Blackwell) ( #47090 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-30 18:13:21 -07:00
9294dd27eb
fix(reasoning): guard rfind in ernie45 streaming </response> branch ( #46255 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-07-01 01:01:14 +00:00
yzong-rh and GitHub
b1190d03cc
[Refactor][GPT-OSS] Harmony Responses API Refactor to use HarmonyParser ( #47185 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-30 19:23:20 -04:00
92c7fac640
[Perf] Restore zero-init of swizzled NVFP4 scale buffer to recover Blackwell decode throughput ( #45739 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-30 22:56:56 +00:00
Ting SUN and GitHub
ac521f6237
[Bugfix][Structured Outputs] Reject degenerate structured_outputs that crash EngineCore ( #45346 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-30 22:41:33 +00:00
28242824e0
[Bugfix][Frontend] Normalize constrained Harmony recipients ( #45657 )
...
Signed-off-by: shaojunjie <626650687@qq.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-30 17:33:10 -04:00
VectorPeak and GitHub
68294739d1
[Bugfix] Align OpenCV video metadata timeline ( #47099 )
...
Signed-off-by: VectorPeak <73048950+VectorPeak@users.noreply.github.com >
2026-06-30 20:43:42 +00:00
c8d2f3cb14
[Bugfix] compressed-tensors: allow int8 grouped WNA16 MoE on Marlin ( #47154 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 12:50:46 -07:00
Matt and GitHub
345b28ff2f
[Hardware][AMD][CI] Bump timeouts of various test groups on AMD CI ( #47195 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 14:30:53 -05:00
248d1fbb71
[Feat][1/N] CuTeDSL warmup infrastructure, FA4 MLA ( #46182 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-30 12:17:34 -07:00
11b26c5528
[Bugfix][Tool Parser] PoolsideV1: fix logprobs AttributeError on Responses API ( #47138 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 19:14:09 +00:00
Roberto L. Castro and GitHub
20434c472e
[Feat] Improve Triton JIT diagnostics ( #46621 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-30 18:50:15 +00:00
Andreas Karatzas and GitHub
c8f9c156a5
[ROCm][V1][MLA] Clone prefill backend state per metadata builder ( #46993 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 11:43:54 -07:00
953bba488d
[PERF] Extend NCCL symmetric memory to AllGather and ReduceScatter ( #46703 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: snordmann <snordmann@nvidia.com >
2026-06-30 11:38:18 -07:00
Wentao Ye and GitHub
3a9784b82c
[Feature] DP supervisor using rust frontend ( #47076 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 14:34:05 -04:00
Giancarlo Delfin and GitHub
3cecee40f3
[Model Runner V2][Spec Decode] Fix stale values in idx_mapping from CG num reqs padding ( #47066 )
2026-06-30 11:25:32 -07:00
a7732537f4
[Bugfix] Restore part of bugfix #42650 after accidental deletion in #43241 ( #47039 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
Co-authored-by: Zhanda Zhu <49645678+zhandaz@users.noreply.github.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-30 11:07:59 -07:00
727971f1c1
Add Medusa speculative decoding e2e test ( #41396 )
...
Signed-off-by: Anshika Ojha <anshikao@nvidia.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Anshika Ojha <215760622+ojhaanshika@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-06-30 18:02:22 +00:00
25671cb520
[Parser][Bugfix] Ensure tool call or other special tokens don't leak in non-streaming tool parsing ( #46875 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-30 13:46:53 -04:00
27d5f78b63
[CI] Move distributed small LM eval to B200 ( #47048 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 13:34:25 -04:00
liuzhenwei and GitHub
7a341fa109
[XPU] Support ZE_AFFINITY_MASK passthrough in xpu_disagg_acc_test ( #47105 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-30 17:06:12 +00:00
Charlie Fu and GitHub
f41e8ddc97
[ROCm][CI] Move PyTorch Compilation Unit Tests to MI300(gfx942) ( #47065 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-30 11:32:58 -05:00
245888ff77
[Feature] Detect all2all peer fault with fault tolerance backend and prevent corrupted output ( #43637 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 09:00:25 -07:00
e840f0d3f5
[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 08:39:59 -07:00
fcaa84efa7
[BugFix] Gate MRV2 mixed sparse-MLA warmup on max_num_seqs > 1 ( #47050 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: ziminghuang <ziminghuang@inferact.ai >
2026-06-30 16:31:27 +01:00
Wentao Ye and GitHub
9e84ec8648
[Refactor] Remove dead minimax allreduce rms kernel ( #46842 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 08:29:21 -07:00
d8f483dc30
[Spec Decode] Fix hidden-state extraction block size for hybrid verifiers ( #46301 )
...
Signed-off-by: Igor Margulis <igor.margulis@intel.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-06-30 08:19:51 -07:00
Nicolò Lucchesi and GitHub
dc148dc4d7
[CI][Bugfix] Fix Hybrid SSM NixlConnector PD prefix cache test (2 GPUs) ( #47157 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 23:14:13 +08:00
tc-mb and GitHub
7cf7cbcd95
[Bugfix] MiniCPM-V 4.6: fix grid rows/cols swap in placeholder generation ( #45918 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-30 08:12:44 -07:00
c231d1f290
fix(security): bound tokenizer work when explicit truncation_side is set ( #47007 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 23:08:51 +08:00
Giancarlo Delfin and GitHub
db808b3961
[Model Runner V2][Spec Decode] Implement block verification for rejection sampling ( #46781 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-30 08:07:24 -07:00
Arsalan Shakil and GitHub
00ebf19cca
[Bugfix][Quant] Raise actionable error instead of bare assert for group-size/TP mismatch ( #46230 ) ( #46236 )
...
Signed-off-by: Arsalan Shakil <shakil.arsalan@yahoo.com >
2026-06-30 14:57:14 +00:00
ded6676458
[Bugfix] Seed RayExecutorV2 TCPStore port by DP rank to avoid collisions ( #45960 )
...
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:37:34 -07:00
Bugen Zhao and GitHub
7a327f0b4f
[Rust Frontend] Simplify unit tests with shared TestTokenizer ( #47125 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 15:34:43 +01:00
Harry Mellor and GitHub
1ab9522935
Remove more unnecessary load_weights methods ( #47058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 15:22:16 +01:00
0fc2512094
[KV Offload] Pass ScheduleEndContext to on_schedule_end hook ( #46450 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 17:07:12 +03:00
Harry Mellor and GitHub
62c7d8009f
Forward fix nightly errors from #44589 ( #47151 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 14:02:34 +00:00
Isotr0py and GitHub
ab80b3dff4
[CI/Build] Bump PyNvVideoCodec version ( #47139 )
2026-06-30 06:38:46 -07:00
Qiming Zhang and GitHub
91055efd36
[XPU] C++ implementation for get_memory_info ( #47134 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-30 21:34:47 +08:00
Bugen Zhao and GitHub
3675bcff67
[Rust Frontend] Refactor TLS serve path with unified MaybeTlsListener ( #47101 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:31:58 +01:00
Bugen Zhao and GitHub
bdbd7278b6
[Rust Frontend] Extend renderer/parser roundtrip tests to support token ids ( #47110 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:27:45 +01:00
Harry Mellor and GitHub
5dc36a4fa5
[Model] Remove Tarsier, Tarsier2 ( #47143 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:20:33 +00:00
aab7af0bcb
[Bugfix][ROCm][MLA] Pass q/kv dtypes to get_mla_metadata_v1 in FP8 decode ( #46997 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-30 05:31:16 -07:00
536047755e
Bump actions/checkout from 6.0.1 to 7.0.0 ( #33057 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 13:16:20 +01:00
1907d3854a
[Bugfix] Reject negative values for max_logprobs and long_prefill_token_threshold ( #44002 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:01:03 +01:00
Chaojun Zhang and GitHub
ea9ddf59fc
[XPU][CI] Enable shared loader test ( #45977 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 11:20:33 +00:00
8cf7c4d8ad
[Attention Backend] add HPC-Ops Attention backend ( #46020 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 18:17:43 +08:00
8e9d70fdd5
[Kernel][XPU] Adjust kernel unit tests for XPU ( #45140 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-30 09:57:27 +00:00
Juan Pérez de Algaba and GitHub
364ee36af1
fix(security): prevent image decompression bomb OOM denial of service ( #47010 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-30 09:39:22 +00:00
Nicolò Lucchesi and GitHub
06fae69114
[Misc] Mistral label alert ( #47132 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 09:02:07 +00:00
14f8660a18
[CI/Build] Add CPU test dependency pre-commit hooks ( #47032 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 07:59:13 +00:00
aed541def4
[Bugfix][Responses] Set completed status for Harmony function calls ( #46945 )
...
Signed-off-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-30 07:55:14 +00:00
2bc20e8aba
[Frontend] Add Streaming Parser Engine and new Kimi k2.5/k2.6/k2.7 Parser ( #46610 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:53:17 +00:00
Chaojun Zhang and GitHub
8cc242335d
[XPU] Optimize XPU worker shutdown logic to prevent resource leak ( #46433 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 15:27:21 +08:00
Andreas Karatzas and GitHub
ba22cb6765
[ROCm][Ray][CI] Keep assigned GPU visible for weight transfer ( #47000 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:59:18 +08:00
Uros Markovic and GitHub
81bcced482
[Bugfix][ROCm] Preserve MoE weight padding for unquantized Triton path ( #46381 )
...
Signed-off-by: Uros Markovic <umarkovi@amd.com >
2026-06-30 14:47:57 +08:00
Kunshang Ji and GitHub
fb42e5219e
[Platform] Replace torch.cuda.mem_get_info with torch.accelerator.get_memory_info ( #44825 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-30 14:39:52 +08:00
Dakai An and GitHub
0feca7ffa8
PD disagg with Mooncake Connector: GDN support (Qwen3.5) and MLA support (Deepseek-V4-Flash) ( #46807 )
2026-06-29 23:29:04 -07:00
97b5ce5c39
[Bugfix] Raise VLLMValidationError for non-integer logit_bias keys ( #46612 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-30 06:18:59 +00:00
Andreas Karatzas and GitHub
4236514098
[ROCm][CI][Multimodal] Use ROCm-aware FA availability check for Unlimited-OCR ( #47004 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:03:13 +08:00
Blas Rodriguez Irizar and GitHub
e45c8a9f4b
[Rust Frontend] Start current wave for a stale DP FirstRequest ( #46833 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-30 05:13:09 +00:00
Wei Zhao and GitHub
b153dd3f28
[Bugfix] Use larger workspace size for Flashinfer MLA LSE ( #47074 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-29 22:11:03 -07:00
Reid and GitHub
930f8dc0a1
[Bugfix][Rust Frontend] Reject prompt_logprobs for streaming generate ( #46839 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 05:10:07 +00:00
Reid and GitHub
a16dbd5b85
[Rust Frontend] Avoid LoRA registry scans without active LoRA requests ( #47040 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 04:58:19 +00:00
bec232a914
Secondary tier implementation for PD disaggregation ( #42285 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 07:51:44 +03:00
b5c9e1ac33
[LoRA] Add language-backbone LoRA support for MiniCPM-V 4.6 ( #46740 )
...
Signed-off-by: linitra24 <Joy25810@foxmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-30 04:19:31 +00:00
ae2c4f3db7
[XPU][UT]Fix xpu pass_config.fuse_norm_quant assert issue ( #46804 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 21:13:44 -07:00
ganesh and GitHub
fca432e60a
[Bugfix] Propagate default stop_token_ids to per-request SamplingParams ( #35076 )
...
Signed-off-by: sriganesh123 <arjulasriganesh@gmail.com >
2026-06-30 12:10:09 +08:00
af1ee8c475
fix(config): reject negative max_logprobs (except -1) and long_prefill_token_threshold ( #44070 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 04:02:36 +00:00
5b4cb69523
[Bugfix][MLA] Fix LSE log-base mismatch in DCP + FlashInfer MLA decode ( #47079 )
...
Signed-off-by: girasoley <girasoleyang@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 19:15:02 -07:00
9fc0c08026
[ROCm][CI] Make tests/v1/shutdown an importable package ( #47085 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 21:01:27 -05:00
f2b5fabb23
[ROCm][CI] Move LM Eval Large Models (8 GPUs) to mi300 pool ( #47094 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 20:59:08 -05:00
b8cb75b149
[Rust Frontend] Add static HTTPS and mTLS support for HTTP and gRPC ( #45890 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 01:45:59 +00:00
Thien Tran and GitHub
43916891b2
[GDN] Improve kkt kernel of CuteDSL prefill backend ( #46346 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-29 18:34:18 -07:00
cda05ee8c4
[Bugfix][Reasoning] Fix thinking_token_budget not enforced on re-entry after forced end ( #43757 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-06-30 01:04:25 +00:00
weishu and GitHub
77654d080c
[KVTransfer] MultiConnector: merge kv_transfer_params dicts across connectors ( #46777 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-06-30 00:25:05 +00:00
Wentao Ye and GitHub
75698e60b3
[Bug] Fix sparse attention issue for GLM5.2 non-torch compile path ( #47083 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-29 15:45:53 -07:00
Andreas Karatzas and GitHub
8632c884dc
[ROCm][CI] Use spawn around the threaded OTLP test ( #47003 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 16:34:05 -05:00
c3734e8334
[CI][Bugfix] Add cohere_melody to ROCm test requirements ( #47072 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 16:29:47 -05:00
53f7553f09
[ROCm][DeepEP] Stabilize high-throughput DBO for DP+EP ( #46990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-29 14:28:02 -07:00
4eb227992a
[ROCm][CI] Make memory sampling less racy in tests and sleep mode ( #45490 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-29 14:26:41 -07:00
Micah Williamson and GitHub
ebcf511ec3
[ROCm][CI] Soft Fail Spec Decode Ngram + Suffix and Entrypoints Integration (LLM) AMD Mirrors ( #47067 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-29 16:24:08 -05:00
Matthew Bonanni and GitHub
8fc1b2d046
Fix FA4 dynamic_causal for full attention layers ( #46659 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-29 14:23:34 -07:00
Harry Mellor and GitHub
5316638a5e
Fix transient dependency issues caused by requirements/common.txt ( #47015 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 14:20:33 -07:00
zhrrr and GitHub
61ab70ec3b
[Model Runner V2] support mamba hybrid models align prefix cache ( #42406 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-29 14:09:16 -07:00
Woosuk Kwon and GitHub
a309d4fe60
Support DCP with FlashInfer MLA ( #43729 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-29 13:24:29 -07:00
72f639927f
[XPU] [RMSNorm] revert weightless change on xpu ( #46987 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 19:03:06 +00:00
Nick Hill and GitHub
8ad4a01825
[ModelRunner V2] Simplify recent UnlimitedOCR-related changes ( #46975 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-29 09:56:17 -07:00
Jee Jee Li and GitHub
7be582697b
[Bugfix] Fix DeepseekV2Model hidden_size ( #46986 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-29 16:44:05 +00:00
030c9523bd
[Perf][1/N] Expand Triton kernel warmup coverage, DSv4 ( #46634 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-29 16:40:34 +00:00
4708292d48
Bump flashinfer version to 0.6.13 ( #46683 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-29 09:30:57 -07:00
debec6440b
Add MiniMax-M3 modelopt nvfp4 support ( #46756 )
...
Signed-off-by: Xin Li <xinli@nvidia.com >
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
Co-authored-by: Xin Li <xinli@nvidia.com >
2026-06-29 09:29:39 -07:00
c8fb2963bd
[FS-Offloading] Batch Lookup in C ( #46713 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 09:28:32 -07:00
HDCharles and GitHub
379acd4e4f
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression ( #46860 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-06-29 15:55:42 +00:00
Martin Hickey and GitHub
07d33e575b
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager ( #44657 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 16:42:35 +01:00
36bbecd643
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" ( #46958 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 07:54:34 -07:00
Nicolò Lucchesi and GitHub
6149187a4c
[Kernel] Triton MLA logits workspace ( #46819 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-29 07:54:29 -07:00
Xiaohong (Sean) Chen and GitHub
49e28e8e91
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope ( #44010 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-06-29 22:54:15 +08:00
0ca39c4f1f
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone ( #46973 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-29 10:00:31 -04:00
Blas Rodriguez Irizar and GitHub
6185d73882
[Rust Frontend] Keep literal "null" string for string-typed tool params ( #46827 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-29 13:46:33 +00:00
bc8481af09
[MoE Refactor] Standardize Humming MoE experts + utilities ( #43373 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-29 06:19:29 -07:00
59575da46d
[XPU] exclude unsupported models for test_tensor_sechma.py ( #47008 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 12:30:28 +00:00
wang.yuqi and GitHub
3483240b7e
[Frontend] Consolidate scale out entrypoints ( #44512 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 03:18:53 -07:00
Roberto L. Castro and GitHub
eddfd4cf21
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen ( #46750 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-29 10:10:07 +00:00
Martin Hickey and GitHub
a4e3cb40d0
[mypy] Enable mypy for tests directory ( #47018 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 09:29:09 +00:00
soaringk and GitHub
ab132ee98b
Fix model info cache for package models ( #46567 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-29 09:17:54 +00:00
e186107870
[Bugfix] Use native SiLU activation in CPU fused MoE ( #45961 )
...
Signed-off-by: Alden Lobo <alden.lobo@arm.com >
Co-authored-by: Alden Lobo <alden.lobo@arm.com >
2026-06-29 09:12:20 +00:00
0e207dac78
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models ( #46835 )
...
Signed-off-by: John Langford <jl@hunch.net >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:59:15 +00:00
wang.yuqi and GitHub
9e86352c60
[CI Failure] Add transformers version check for openai/privacy-filter ( #47011 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 08:57:26 +00:00
Harry Mellor and GitHub
5051698e41
Remove unnecessary load_weights methods ( #44589 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 01:52:23 -07:00
Andreas Karatzas and GitHub
db28ae2d07
[ROCm][CI] Explicitly tear down multimodal offline LLMs ( #46999 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 07:59:24 +00:00
Harry Mellor and GitHub
f6bb8682ee
Fix docs on main ( #47009 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 15:50:57 +08:00
4559c43a95
[MM][CG] Gemma3 Encoder CUDA Graph ( #43591 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-29 04:52:00 +00:00
Bugen Zhao and GitHub
5274c1181d
[Rust Frontend] Add Harmony Renderer for GPT-OSS ( #46800 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-29 03:39:04 +00:00
Yuwen Zhou and GitHub
58d6a6e60a
[CPU] Support cpu compressed-tensor w8a8 int8 moe ( #42920 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-06-29 03:04:05 +00:00
a2abce646f
[EPLB] Mask padding in EPLB load recording ( #38128 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-28 19:43:58 -07:00
Harry Mellor and GitHub
311ad689ad
Remove boilerplate missed by #46820 ( #46956 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:11:17 +08:00
Woosuk Kwon and GitHub
0472436541
[Spec Decode] Avoid redundant hidden-states gather in draft prefill ( #46968 )
2026-06-28 17:04:01 -07:00
4dfbf1503b
[Model] Add support for openai/privacy-filter ( #41026 )
...
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-28 16:18:22 -07:00
Wei Zhao and GitHub
95528527ea
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 ( #46855 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-28 14:36:23 -07:00
c2127a25c7
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm ( #46895 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:50:30 -05:00
03c6d01c30
[OCP MX ] Add back emulation to available OCP MX backends list ( #46629 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:43:19 -05:00
Woosuk Kwon and GitHub
4b643c463e
[GLM5] Fix minor typo ( #46961 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-28 08:37:00 -07:00
7544286b04
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE ( #46552 )
...
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-28 15:19:28 +00:00
Woosuk Kwon and GitHub
89876b0c54
[GLM5] Implement op fusion for GLM5/DSV3.2 ( #46876 )
2026-06-28 08:17:39 -07:00
Wentao Ye and GitHub
5c91039c41
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement ( #46635 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-28 14:55:54 +00:00
5ecae3266c
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) ( #45033 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-28 07:52:00 -07:00
6eb63a1da6
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers ( #46600 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 01:37:44 -07:00
09841ae705
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation ( #46846 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-28 00:07:33 -07:00
Matt and GitHub
a2a92cbbaa
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection ( #46930 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-28 00:07:14 -07:00
35e6c86caa
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL ( #46034 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-28 00:06:43 -07:00
c7ca0bccae
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 ( #44313 )
...
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com >
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com >
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
2026-06-28 00:04:08 -07:00
c6741b2ad4
[Model] Support Unlimited OCR ( #46564 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-27 23:09:18 -07:00
a65f93fb2e
[ROCm][CI] Add ci_base metadata for external cache orchestration ( #46886 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-28 12:51:19 +08:00
Chauncey and GitHub
11a12305c0
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models ( #46786 )
2026-06-27 18:38:07 -07:00
798185d438
[KV-Offloading] Fix tensors_per_block stride ( #46888 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-27 21:01:45 -04:00
Matt and GitHub
9036c89ee4
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now ( #46928 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-27 17:30:49 -05:00
Giancarlo Delfin and GitHub
b6caeb5a09
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance ( #46878 )
2026-06-27 14:09:25 -07:00
Taneem Ibrahim and GitHub
8bf064f8d3
Fixed chunked embedding aggregation with request-id metadata ( #46782 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-27 20:57:47 +00:00
ea2ead1db3
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod ( #46818 )
...
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
2026-06-27 20:23:59 +00:00
Wentao Ye and GitHub
56aa067bf0
[CI Bug] Fix h100 AssertionError: Cold-start child failed ( #46927 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-27 20:17:33 +00:00
xiaolinchen and GitHub
35e3850fa9
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) ( #46915 )
...
Signed-off-by: wentian-byte <2990624738@qq.com >
2026-06-27 14:30:10 -04:00
51a99565c3
[ROCm][Perf] Fused shared expert for Minimax M3 ( #46474 )
...
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-27 12:34:17 +00:00
867fd5e8ed
[ROCm][Perf] Use flydsl moe with Minimax-M3 mxfp8 weights on gfx950 and implemented moe-backend selection ( #46184 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
2026-06-27 10:22:57 +00:00
9fd00ee006
[ROCm][CI] Move remaining mi250_2 tests out of the MI250 queue ( #46905 )
...
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-27 17:08:54 +08:00
091d13976c
[ROCm][CI] Add TRITON_ATTN score absolute tolerance floor ( #46891 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-27 06:35:50 +00:00
Wentao Ye and GitHub
b588f66dc2
[GLM5.2 Perf] fused_indexer_q_rope_quant triton kernel, 1.9% ~ 3.3% E2E Throughput improvement. ( #46862 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-26 22:16:20 -07:00
Benjamin Chislett and GitHub
455f25aa13
[CLI] Add flag to print TTFT and TPS in vllm chat ( #46775 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-26 22:15:10 -07:00
d706dec904
fix: Correct reasoning-end detection for prompt history ( #44551 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Signed-off-by: Jason Ozuzu <jasonozuzu@cohere.com >
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: Walter Beller-Morales <walterbm@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-26 22:15:06 -07:00
Divakar Verma and GitHub
68ee8300a0
[ROCm][CI]Fix test_concat_and_cache_mla_rope_fused on ROCm ( #46409 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 12:38:13 +08:00
ddd3855a28
[MoE Backend] add HPC-Ops MoE backend ( #45924 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: youkaichao <youkaichao@gmail.com >
2026-06-27 11:18:07 +08:00
Divakar Verma and GitHub
00e045b7c7
[ROCm][CI TG] refactor and fix deepep_moe test group ( #46758 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:45:23 +08:00
Divakar Verma and GitHub
17a71d8702
[ROCm][CI] Relax fused layernorm quant test tolerances for one-ULP outliers ( #46658 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:44:29 +08:00
weizhoublue and GitHub
2e058851d3
fix(docker): eliminate race conditions in shared buildkit cache mounts ( #44984 )
2026-06-26 19:43:17 -07:00
Dāvis and GitHub
1a92dfcce4
[Build] Show error message when using ROCm with LTO and different compilers ( #35232 )
2026-06-26 19:43:00 -07:00
Chris Leonard and GitHub
d0f800811b
[Build] Update vllm to point to vllm-project/flash-attention commit that builds FA3 with torch stable API. ( #46644 )
2026-06-26 19:42:46 -07:00
Nick Hill and GitHub
c6dd32a810
[ModelRunner V2] Support realtime embeddings ( #46762 )
2026-06-26 19:42:27 -07:00
af16446bf3
Vram semaphore infra ( #44465 )
...
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-26 17:32:51 -07:00
Harry Mellor and GitHub
3f67477497
[CI] Don't try and download files that we already know don't exist ( #46854 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-26 23:56:39 +00:00
Nick Hill and GitHub
1d41009e81
[ModelRunner V2] Fix cross-attention block table sizing ( #46753 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:34:21 -07:00
Nick Hill and GitHub
b94f212e37
[ModelRunner V2] Deduplicate ModelState init logic ( #46776 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:32:45 -07:00
Harry Mellor and GitHub
d8eb734d94
Fix Transformers backend FP8 MoE and remove some boilerplate ( #46820 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-27 00:16:05 +01:00
2ff76a5e85
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 ( #46760 )
...
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-26 21:58:40 +00:00
Yifan Qiao and GitHub
75fdcc82a5
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas ( #46873 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-26 21:48:53 +00:00
yzong-rh and GitHub
77f8796d16
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. ( #46437 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-26 17:18:47 -04:00
c40d307731
[Core] Remove FlashAttention block size restriction for hybrid models ( #36701 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-26 21:16:39 +00:00
Woosuk Kwon and GitHub
65e655d295
[GLM-5] Add DSV3.2/GLM5 to vllm/models/ ( #46808 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-26 14:09:05 -07:00
Charlie Fu and GitHub
6e2fb02fe5
[ROCm][CI] Fix rlhf_nccl.py on ROCm ( #46851 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-26 15:41:49 -05:00
Micah Williamson and GitHub
274325dd43
[ROCm][CI] Remove V1 Sample + Logits from mi250 Queue ( #46867 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-26 15:38:38 -05:00
Matt and GitHub
95e6442a6b
[Hardware][AMD][CI] Fix Kernels Quantization test timeout ( #46859 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-26 15:19:16 -05:00
701a23d99f
[Bugfix][Model] Support tensor parallelism for DiffusionGemma ( #45719 ) ( #46177 )
...
Signed-off-by: Carlos Alvarado <carlos-alvarado@outlook.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-26 20:05:04 +00:00
Ben Browning and GitHub
dccb412e2c
[Bugfix][Parser] Pass token IDs to parser.parse() in Responses API and batch serving ( #46843 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-26 19:29:52 +00:00
c6554f321c
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build ( #46769 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 14:32:21 -04:00
Julien Denize and GitHub
3d3b96488f
Migrate Voxtral to mistral-common 1.11.5 audio API ( #46705 )
...
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-06-26 11:06:31 -07:00
Nick Hill and GitHub
658b54efe4
[ModelRunner V2] Update scheduler tests to cover MRV2 paths ( #46771 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 09:36:31 -07:00
Li, Jiang and GitHub
abc71548ef
[CI/Build][CPU] Add test image cache clean-up ( #46831 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-26 23:28:49 +08:00
Nick Hill and GitHub
4e07ca2c92
[Core] Add VLLM_GPU_SYNC_CHECK env var ( #44800 )
2026-06-26 08:24:33 -07:00
Bugen Zhao and GitHub
e71bc6da85
[Rust Frontend] Use oss-harmony for Harmony output processing ( #46799 )
2026-06-26 08:24:13 -07:00
fxmarty-amd and GitHub
37ce34922f
[CI] Fix failing CUDA graph capture in Triton MOE ( #46735 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-26 07:21:20 -07:00
c2507fb293
[ROCm] [MoE] [Perf] Shared-expert fusion for bias-routed MoE; enable on MiniMax-M3 mxfp8 model ( #46545 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 07:05:20 -07:00
TJian and GitHub
8921c4be88
[ROCm] [Performance] Optimize aiter moe for DeepSeekV4 ( #46122 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 06:43:27 -07:00
8e394244a5
[ROCm]Enable AITER MoE backend for MiniMax-M3-MXFP4 ( #46419 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-26 06:35:35 -07:00
TJian and GitHub
302954e5f6
[ROCm] [CI] fix transcription flakiness AMD: Entrypoints Integration (API Server OpenAI - Part 1) (mi325_1) ( #46823 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 21:33:35 +08:00
Hyunkyun Moon and GitHub
950ee4c2e4
[API] Add token offsets to render endpoints (/v1/.../render) ( #44226 )
...
Signed-off-by: HyunKyun Moon <mhg5303@gmail.com >
2026-06-26 05:02:52 -07:00
d980a3cc6e
[ROCm] Fix AITER_UNIFIED_ATTN Dispatching After AITER Bump ( #46780 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-26 02:09:56 -07:00
bf292b5f6b
[Docs] Remove BambaForCausalLM from supported hybrid models list ( #46071 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-26 08:02:50 +00:00
wang.yuqi and GitHub
5e3dad04b1
[Misc] Move the legacy api_server.py to the examples directory. ( #46783 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-26 07:43:29 +00:00
Joe Rowell and GitHub
63e161f296
[Bugfix][Tool Parser] PoolsideV1: fix string whitespace and required named tool choice ( #46486 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-06-26 06:05:16 +00:00
Tiezhen WANG and GitHub
c7645bce04
Remove grok model arch from vllm ( #46706 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
2026-06-25 23:02:10 -07:00
35a49fcfc2
[CI][Bugfix] Spawn engine in mm cache sleep test to fix ROCm HIP error ( #46749 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-26 00:38:26 -05:00
peizhang56 and GitHub
915e99ec67
[ROCm][Bugfix] Fix HIP fork re-init in multimodal offline examples ( #46741 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
2026-06-26 00:37:47 -05:00
Nick Hill and GitHub
5b33041746
[ModelRunner V2] Fix whisper test ( #46773 )
2026-06-25 22:10:36 -07:00
Matt and GitHub
1a4984520e
[Hardware][AMD][CI] Fix AMD CI image build ( #46792 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 22:05:12 -07:00
Reid and GitHub
e312c5cb25
[Rust Frontend] Make Granite4 string argument scanning incremental ( #46507 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-26 03:54:03 +00:00
Matti4 and GitHub
1502cf6274
Fix relative allowed local media paths ( #45263 )
2026-06-25 20:45:20 -07:00
d350fa8ddd
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens ( #46733 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-26 03:41:33 +00:00
dbc49b6b99
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 ( #45166 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
Signed-off-by: ovidiusm <ovidium@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-25 19:33:42 -07:00
fxmarty-amd and GitHub
552a9dbe59
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears ( #44667 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-25 19:33:00 -07:00
02a1f23711
[DFlash] Fuse precompute kv per-layer rmsnorms ( #46761 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:32:07 -07:00
652d962bc9
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation ( #46448 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:30:07 -07:00
Giancarlo Delfin and GitHub
5314665bad
[Model Runner V2][DFlash] Enable dflash attention backend selection ( #46770 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 19:29:25 -07:00
Michael Goin and GitHub
3daea7ceb9
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models ( #46759 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 19:03:09 -07:00
Wentao Ye and GitHub
cc7981599e
[Refactor] Remove dead kernel code ( #46405 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 18:09:56 -07:00
Nick Hill and GitHub
32bb3195f0
[ModelRunner V2] Bound memory for large logprobs requests ( #46746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-25 18:04:06 -07:00
ad28d605e6
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) ( #45544 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-25 17:46:28 -07:00
Bugen Zhao and GitHub
ae7c8ec223
[Rust Frontend] Switch rustls to native-tls/OpenSSL ( #46696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:19:44 -07:00
Bugen Zhao and GitHub
1d3f4cb3a4
[Rust Frontend] Extract renderer fixture test utilities ( #46719 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:12:38 -07:00
Bugen Zhao and GitHub
f9e684499f
[Rust Frontend] Migrate gemma4 to unified parser ( #46602 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 16:59:57 -07:00
Giancarlo Delfin and GitHub
c53994e134
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling ( #46665 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 23:46:10 +00:00
Matt and GitHub
27da2a2ac4
[Hardware][AMD][CI] Use Triton-based AITER MHA for LM Eval Qwen-3.5 Models Tests ( #46691 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 17:08:04 -05:00
Michael Goin and GitHub
a2e8ec3d52
[CI] Depend GPQA Eval DGX Spark job on arm64 image build ( #46736 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 17:07:04 -04:00
e8c24a7695
[Kernel] Vectorized fp32 moe_sum reduction and support any topk ( #46643 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 14:02:28 -07:00
Andreas Karatzas and GitHub
2a6f8f0c05
[ROCm][CI] Fine-tuning queues and test names ( #39238 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-25 13:24:09 -07:00
Robert Shaw and GitHub
c5e3c40877
Fix P/D with DP Supervisor ( #46628 )
...
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-25 13:13:08 -07:00
Wentao Ye and GitHub
8b4d93ba2b
[Perf] Remove redundant clone for GLM, Deepseek etc ( #46651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 13:09:00 -07:00
Michael Goin and GitHub
e8e7b592d1
[Kernel][MoE] Tune block-FP8 fused MoE for low-batch decode ( #46642 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 12:38:28 -07:00
Rohan Potdar and GitHub
e53a17232c
[ROCm]: Bump aiter to 0.1.16.post2 ( #46692 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-25 11:53:26 -07:00
Flora Feng and GitHub
96eb8ddc41
[CI] Re-enable skipped glm and seedoss parser tests ( #46671 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-25 13:11:41 -04:00
Gabriel Wu and GitHub
8fa36fbbeb
[Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 ( #46506 )
2026-06-25 09:12:00 -07:00
Ranran and GitHub
e45b279928
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3Next ( #46316 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
2026-06-25 09:05:04 -07:00
d490b98162
[Core] Avoid mixed length specdec batches via padding ( #45237 )
...
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
2026-06-25 08:34:44 -07:00
haoyangli0109 and GitHub
1744adc256
[ROCM] [Communication] Add INT3 quantization method for quickreduce ( #45666 )
...
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com >
2026-06-25 15:14:15 +00:00
Divakar Verma and GitHub
cdfa2fd7e9
[ROCm][CI] rm duplicate Distributed Torchrun ci test ( #46729 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-25 09:58:17 -05:00
6f3da461d1
[Pooling] Fix Cohere embed billed image token accounting for mixed-content inputs ( #46093 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 10:44:29 -04:00
Russell Bryant and GitHub
d3130d878c
[CI] Pin GitHub Actions to commit hashes in macos-smoke-test.yml ( #38290 )
2026-06-25 13:48:44 +00:00
9bfd878a48
[MoE] [MoE Refactor] Add moe kernel oracle abc 37753 ( #43461 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Signed-off-by: qyYue1389 <yueqiuyang1389@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 09:34:03 -04:00
Matt and GitHub
2365b7a8e7
[Hardware][AMD][CI] Mirror Basic Models (Others) and Weight Loading Multiple GPU test groups ( #46668 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 08:25:09 -05:00
15be78732b
[NIXL][Mamba] Add Mamba1 support to NIXL P/D disaggregation ( #45019 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 05:50:41 -07:00
92221485aa
[CPU][CI/Build] Allow more CPU CI agents ( #46702 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:33:39 +08:00
xiangdong and GitHub
a6f41ab678
[XPU][CI]Refine .buildkite/ci_config_intel.yaml for Intel GPU CI ( #46674 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-25 08:58:26 +00:00
c63cd4906c
[ROCm][ [Perf] sparse attention optimization on minimax-m3 ( #46546 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: yueliu14 <yue.liu4@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-25 16:56:00 +08:00
638b1a99cc
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM ( #45269 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-25 08:18:10 +00:00
72adb20a6a
[Model] Remove AquilaForCausalLM, AquilaModel ( #46605 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 08:08:26 +00:00
2396d91e93
[CPU][Spec Decode] Enable DFlash SD for CPU ( #44029 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 15:32:48 +08:00
9b215ae60b
[Rust Frontend] Forward VLLM_ENGINE_READY_TIMEOUT_S via --args-json ( #44610 )
...
Signed-off-by: kai <kai@example.com >
Co-authored-by: 图灵 <tuling.wk@alibaba-inc.com >
2026-06-25 07:25:08 +00:00
Bugen Zhao and GitHub
4d3b4b9b01
[Rust Frontend] Make ToolParserOutput a seq of ToolParserEvent to preserve order ( #46584 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 06:27:07 +00:00
Matthias Gehre and GitHub
77c1d9fe9b
[ROCm][Perf] Tune wvSplitK on gfx1151 ( #40784 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-06-25 14:17:46 +08:00
Jeff (Junze) Ma and GitHub
36fd7e8b86
[SimpleCPUOffloadConnector] Fix remaining global→block conversions under PCP/DCP ( #46394 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
2026-06-24 23:05:24 -07:00
fc61c6fc26
[Perf] Enable + tune FlashInfer fused allreduce at world_size=16 on SM 10.3 (GB300) ( #46392 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:04:17 -07:00
Matt and GitHub
e2af449c39
[Hardware][AMD][CI] Move Metrics, Tracing (2 GPUs) & make optional ( #46686 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 05:49:33 +00:00
3f5a1e1733
[ROCm][CI] Expand basic correctness target suites ( #46573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matt <156021403+mawong-amd@users.noreply.github.com >
2026-06-25 12:18:57 +08:00
710ebaa189
[ROCm][Bugfix] Fix chunk alignment when using context parallelism with TRITON_MLA ( #46114 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:07:28 -04:00
1aad125815
[CPU] Enable chunked prefill and prefix caching for qwen3.5 ( #46202 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-25 03:49:21 +00:00
dc55936f64
[AMD][CI] Fix Pipeline + Context Parallelism test group ( #46650 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 22:23:42 -05:00
Bugen Zhao and GitHub
76c3c4ff63
[Rust Frontend] Introduce unified parser interface & combined parser ( #46583 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 03:17:31 +00:00
efb5acffd5
[Bugfix] fix: stream Mimimax m2 tool call string arguments ( #46382 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-25 03:12:45 +00:00
6e3a983cf3
[ROCm] Remove erroneous inclusion of gptq_marlin as supported quant scheme on ROCm ( #46655 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 21:27:19 -05:00
Xin Yang and GitHub
1273a8f05a
[Kernel] Add swap AB optimization to fused_moe_kernel ( #36559 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-06-25 01:44:30 +00:00
9e88e969c0
[Perf][KVConnector][Mooncake] Parallelize KV load with a receive-thread pool ( #45971 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 18:25:12 -07:00
dda3aca47f
[Speculative Decoding] Propagate norm_output and fc_norm config for Eagle3 speculators ( #46488 )
...
Signed-off-by: Orestis Zambounis <orestis.zambounis@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:51:33 +00:00
Jee Jee Li and GitHub
23aed9b0ee
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel ( #46508 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-25 08:42:51 +08:00
Maxwill Lin and GitHub
cd347298e8
[Frontend] Port seed_oss to the streaming parser engine as a Qwen3 subclass ( #46314 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 20:08:42 -04:00
Yifan Qiao and GitHub
b69816043a
[Bugfix][MooncakeStore] track resumed requests via scheduler's resumed_req_ids ( #46595 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 23:50:56 +00:00
Kaihang Jiang and GitHub
fc7fc421e9
[Kernel][MoE] Allow FlashInfer MXINT4 MoE for gated SiLU ( #46518 )
...
Signed-off-by: Kaihang Jiang <kaihangj@nvidia.com >
2026-06-24 18:32:50 -05:00
cyq and GitHub
e06a83445c
[Bugfix] Normalize slashes in Helion GPU names ( #46101 )
...
Signed-off-by: cyq <15000851237@163.com >
2026-06-24 18:22:49 -05:00
d7ab9be775
[Bugfix] Support -1 (invalid/non-local) slots in topk_ids for Triton MoE ( #46408 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:59:42 -07:00
6a1570711c
[Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing ( #46406 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:52:09 -07:00
Micah Williamson and GitHub
d6696e2385
[ROCm] Begin Deprecation Window for CUDA_VISIBLE_DEVICES on ROCm ( #46636 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-24 20:40:28 +00:00
Chauncey and GitHub
84c2f9f0fb
[Frontend] Fix Kimi K2 tool call IDs for required tool choice ( #46344 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-24 19:59:40 +00:00
49f2104c53
[Feature] Support DCP with FP8 KV cache in MLA decode path ( #44044 )
...
Signed-off-by: shivampr <shivampr.dev@gmail.com >
Signed-off-by: Shivam <shivamprasad91@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 19:28:17 +00:00
d511b5bae9
Chore: Fix minor doc sentence, grammar, quote errors ( #40469 )
...
Signed-off-by: Ashwin Phadke <23502062+ashwin-phadke@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-24 18:58:24 +00:00
3c43237233
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels ( #46560 )
...
Signed-off-by: xiaojun.wei <jessiewei747@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-24 11:00:57 -07:00
56ca5997ea
Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference ( #46389 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-24 13:53:54 -04:00
Aarushi Jain and GitHub
cf57311187
Run DeepSeek-V2-Lite prefetch-offload eval eager on ROCm ( #46386 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-24 12:25:57 -05:00
Lucas Wilkinson and GitHub
e7df232288
[KV Offload] Gate packed HMA KV cache on cross-layer config ( #46252 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-06-24 11:55:30 -04:00
b3a688cb9e
[ROCm] Fix OOB During Model Warmup With ROCM_ATTN and MRV2 ( #46548 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-24 10:53:21 -05:00
Wentao Ye and GitHub
1cd3e0e945
[Bug] Fix IndentationError: expected an indented block after 'with' statement ( #46627 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-24 23:14:17 +08:00
Yiwei Hu and GitHub
f889325c51
[KV Offload] Use background thread for mmap / cpu_tensors pinning ( #45850 )
...
Signed-off-by: Sorryhorizon <arikara6666@gmail.com >
2026-06-24 18:13:27 +03:00
bb61177e49
[KV Offloading] Replace bool|None lookup return with LookupResult enum ( #46363 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 18:06:08 +03:00
7f99e80c3b
[Perf][ThinkingBudget] reduce search space for thinking tokens ( #46425 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:02:25 +08:00
2801b11156
[Test] Pin block_size in auto-fit max_model_len test ( #45914 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:56:21 +08:00
007b5a52ed
[Log] Update to log once ( #46511 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-24 14:45:16 +00:00
Cyrus Leung and GitHub
24d5186138
[Bugfix] Re-enable FP8 MoE on NVIDIA Thor ( #46339 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-24 07:35:46 -07:00
Nemani Harsha Vardhan and GitHub
7dc036058b
[Doc] Document Qwen3.6 (dense + MoE) ViT CUDA graph support ( #44720 )
...
Signed-off-by: harsha20032020 <nhvardhan2020@gmail.com >
2026-06-24 14:35:08 +00:00
61ee183d28
[ROCm] Fix AITER FP8 quantization schema tests ( #46414 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:29:19 +08:00
84c62e1cbd
[Model Runner V2][MM] Support EVS ( #46535 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:18:56 -04:00
Fadi Arafeh and GitHub
061043eaca
[CPU][Perf] Accelerate unquantized MoE for AArch64 ( #46353 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-24 14:14:35 +00:00
93ec645878
[Bugfix] Fix illegal memory access from a forward during a partial wake_up ( #44483 )
...
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 22:12:23 +08:00
Kunshang Ji and GitHub
563c628968
[XPU] bump up vllm_xpu_kernels to v0.1.10.1 ( #46607 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 10:05:31 -04:00
0bc479e6eb
[Perf][LoRA] Replace O(n) list.index() with a dict in convert_mapping ( #46542 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 21:41:46 +08:00
Tae Jeong and GitHub
62890e204c
Fix duplicated logging when loading a corrupt or partial video ( #46467 )
...
Signed-off-by: hhhhhhhhhhhhhhhhho <man2719@naver.com >
2026-06-24 06:14:13 -07:00
Nicolò Lucchesi and GitHub
a2cb08b3d5
[Misc][PD] Disable bidirectional xfer mode for NixlPushConnector ( #46473 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-24 21:14:05 +08:00
cf9fd6457e
Fix KV offload request-finished lifecycle contract ( #46284 )
...
Signed-off-by: test test <2260891073@qq.com >
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 15:42:40 +03:00
Kunshang Ji and GitHub
d4448b511d
[XPU][Docker] switch to ubuntu 24.04 as base image ( #45973 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 20:39:20 +08:00
f1a6703edd
[Bugfix][Config] Keep pydantic validation for fields with a TYPE_CHECKING Literal alias ( #46220 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 12:25:50 +00:00
Roy Wang and GitHub
160c80a34c
[Rust Frontend] Raise frontend JSON body limit ( #46582 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-06-24 12:15:31 +00:00
Martin Hickey and GitHub
f237e16b41
[KV Offload] Replace OffloadingHandler with OffloadingWorker ( #45053 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-24 14:44:24 +03:00
70749fdcca
[Feature] Triton INT4 per-token-head KV cache quantization ( #40835 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:21:25 +00:00
d20dbf921b
[Mooncake] Only check and store new KV cache range ( #46412 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 03:10:50 -07:00
ede54b926e
set AttentionCGSupport.UNIFORM_BATCH for fa2 on xpu ( #46555 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 18:05:02 +08:00
52fbe12283
[Perf][Multimodal] Avoid building a full timestamps list in video frame sampling ( #46543 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 09:38:27 +00:00
Dakai An and GitHub
dc0d318177
[Attention] Add FLASH_ATTN_MLA_SPARSE backend for Hopper sparse MLA ( #46189 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-06-24 09:33:10 +00:00
soaringk and GitHub
d7c1821b5a
[Model][MiniMax-M3] Add pipeline parallelism support ( #45810 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-24 08:23:03 +00:00
4cd1a84c88
[Model] Remove BaiChuanForCausalLM and BaichuanForCausalLM ( #46362 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 16:13:57 +08:00
Mohammad Miadh Angkad and GitHub
191826ec61
[CI/Build] Fix topk histogram build on SM75 ( #46550 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-24 00:51:11 -07:00
Andreas Karatzas and GitHub
549c7074cd
[ROCm][CI] Skip the MoE Marlin tile-padding helper assertion ( #46580 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 07:31:33 +00:00
489abadfb8
feat: support to OpenMOSS-Team ( #44124 )
...
Signed-off-by: nagisa-kun <1434936049@qq.com >
Signed-off-by: nagisa19 <1434936049@qq.com >
Signed-off-by: nagisa <1434936049@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 00:08:13 -07:00
Woosuk Kwon and GitHub
96de8bb389
[MoE] Free unused MXFP4 scales in OAI Triton Backend ( #46549 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-24 00:06:41 -07:00
Jee Jee Li and GitHub
9d6fdc2901
[Kernel] GLM5 Router GEMM ( #46385 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:54:50 -07:00
Benjamin Chislett and GitHub
4c5bc41ba6
[Bugfix][Spec Decode] Fix probabilistic sampling for parallel drafting ( #45956 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-24 05:36:23 +00:00
Michał Ganczarenko and GitHub
ac1fa74616
[Bugfix] Fix NemotronLayerNorm1P hardcoded cuda device type ( #46495 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
2026-06-24 13:21:02 +08:00
Sting Lin and GitHub
556bc4e3a0
Upgrade tpu-inference to v0.23.0 ( #46568 )
2026-06-23 21:15:14 -07:00
Wei Zhao and GitHub
05a0caba91
[Mooncake] Optimize lookup pool key string construction ( #46188 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-24 11:51:53 +08:00
Nick Hill and GitHub
7ee4d22009
[Spec Decode] Reject placeholder (-1) draft tokens in rejection sampler ( #46533 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 03:32:32 +00:00
ce9f64020b
[Rust Frontend] Pass effective reasoning_parser_kwargs for structured output ( #46360 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 03:13:44 +00:00
4ed8eaafb0
[Rust Frontend] Integrate xgrammar-structural-tag for strict and required tool calling ( #46057 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 10:46:49 +08:00
6af0559ddb
[Core][DP] Throttle prefills based on local prefill work ( #46532 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 02:27:12 +00:00
e2bdc24612
[ROCm][Bugfix] Fix use_v2_model_runner inside Ray driver thread ( #45998 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 08:41:56 +08:00
Andreas Karatzas and GitHub
bcbeaac786
[ROCm][CI] Stage C-II of gating additional test groups ( #46537 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 17:36:40 -07:00
Maxwill Lin and GitHub
e48f2aa4ca
[Bugfix][Frontend] Emit a content block for empty Anthropic completions ( #46525 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 00:04:26 +00:00
Roberto L. Castro and GitHub
d86c66c981
[Feat] Add runtime monitor for post-warmup CuTeDSL compilation ( #46167 )
2026-06-23 23:33:17 +00:00
Nico Holmberg and GitHub
80e511772f
[ROCm][Bugfix][Perf] enable shared expert fusion for Qwen3.5 ( #44434 )
...
Signed-off-by: Nico Holmberg <nico.holmberg@amd.com >
2026-06-23 23:19:51 +00:00
Roberto L. Castro and GitHub
855cd4d787
[Perf][DSv4/DSv3.2] Add cluster-cooperative topK kernel for low-latency scenarios ( #43008 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-23 16:11:00 -07:00
3cc871aaf1
[Perf] Skip detokenization in online beam search ( #46422 )
...
Signed-off-by: Guy Stone <guys@spotify.com >
Signed-off-by: Guy Stone <guystone3@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 15:46:09 -07:00
0a3e2dbc09
[Optimization] Skip DP padding tokens in MoE ( #46428 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 14:54:46 -07:00
84f13374b3
[CI] Fix test_auto_gptq on ROCm CI ( #46164 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 16:38:06 -05:00
Micah Williamson and GitHub
b28103e1ca
[ROCm][CI] Shard LM Eval Qwen3-5 Models (B200-MI355) in AMD CI ( #46520 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 16:32:05 -05:00
Wentao Ye and GitHub
abc33134fa
[CI Test] Mark batch invariance test flaky ( #46530 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 21:01:34 +00:00
6617db1bfb
[Bugfix][Frontend] Emit non-ASCII tool-call arguments without \uXXXX escapes ( #46308 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-23 20:43:11 +00:00
899d72a58c
[Bugfix][ToolParser] Handle braces in required tool streaming strings ( #45389 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-23 20:29:34 +00:00
Yongye Zhu and GitHub
11b56b2ff2
[Kernel] Add FlashInferCutedslMxfp8LinearKernel (cute-dsl mm_mxfp8) ( #46393 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-23 12:45:49 -07:00
0d4d164488
[Bugfix] Allow flashinfer_cutlass as a clamped NVFP4 MoE backend ( #46492 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-23 12:43:36 -07:00
Mike G and GitHub
0775b882ba
[NVFP4 MoE/Deepseek V4] Marlin: wire SwiGLU clamp + allow it for clamped models on non-Blackwell ( #45836 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-23 12:21:19 -07:00
7c2e08451a
[Docker] Remove redundant flashinfer download-cubin step ( #46517 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 12:16:51 -07:00
Giancarlo Delfin and GitHub
ef361de916
[Model Runer V2][DFlash] Fix lm head sharing for dflash ( #46435 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-23 19:09:06 +00:00
Yan Ma and GitHub
acce57d8dd
Deprecate old FP8 online MoE quantization class ( #44514 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:53:38 -07:00
68afd78897
[Bugfix][ROCm] Fix cumem sleep and teardown ( #46203 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 02:45:31 +08:00
37a682d392
[Kernel] Extend Marlin thread-tile padding to MoE (WNA16 + FP8/MXFP8) ( #45703 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 11:45:10 -07:00
Rui Yin and GitHub
d8e422ccda
[Bugfix] Parse MiniMax M3 streaming reasoning by text markers ( #45718 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-23 14:43:58 -04:00
fxmarty-amd and GitHub
e368415daa
[AMD][OCP MX][CI] Fix tests to not dispatch on UNFUSED_TRITON backend on MI300, improve w_mxfp4_a_fp8 emulation support ( #46142 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-23 14:25:27 -04:00
Andreas Karatzas and GitHub
ceae5bcbda
[ROCm][CI] Fix nixl tests ( #45219 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 13:11:40 -05:00
6691f087a6
[Minimax-M3] BF16/FP8 Indexer using MSA ( #45892 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-23 10:28:49 -07:00
f4d5f73ffa
[Bugfix]: Fix unquantized gpt-oss weight loading broken by FusedMoE r… ( #45818 )
...
Signed-off-by: priyansh jain <priyansh.jain2@amd.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-23 16:56:17 +00:00
fd50a66015
[CI][ROCm] Skip unsupported test cases on ROCm ( #46160 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 11:35:49 -05:00
84586c9acc
[ROCm][CI] fix fp8 range in vit_fp8_quant ( #46410 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 11:34:21 -05:00
Taneem Ibrahim and GitHub
40e5522121
[Docs] Add Qwen3 forced alignment online example ( #46197 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-23 11:59:45 -04:00
Willow Lopez and GitHub
f3410b3bb1
fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility ( #45404 )
...
Signed-off-by: Oxygen <1391083091@qq.com >
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-23 11:46:23 -04:00
568874fec2
[ROCm][CI] pass merge-base to container for python-only wheel metadata ( #45869 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 15:44:43 +00:00
275b43183c
[MyPy] Fix mypy for vllm/benchmarks ( #39896 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-23 15:22:29 +00:00
Yan Ma and GitHub
547d2c40d7
Add weights padding for fp8 per-block online quantization ( #44763 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:08:17 -04:00
2aaaf3febd
[ROCm][Test] Fix stale test_gfx950_moe MXFP4 oracle tests ( #46260 )
...
Signed-off-by: Spandan Tiwari <23646532+spandantiwari@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 23:07:46 +08:00
Micah Williamson and GitHub
156b12667c
[ROCm][CI] Skip Quark mxfp4 tests unless Quark version is compatible with Torch version ( #46431 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 22:26:39 +08:00
Jee Jee Li and GitHub
9f6f296428
[CI/Build] Remove BaiChuanForCausalLM from the LoRA test ( #46494 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:09:48 +08:00
e51e700470
[LoRA] Gate all_gather on fully_sharded_loras inside _mcp_apply; rewrite regression test ( #45715 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-23 07:08:33 -07:00
f59db63732
[Bugfix] GPT-OSS Autodrop reasoning in Response API and cleanup ( #45048 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 09:36:33 -04:00
Rukhaiya2004 and GitHub
9f5117820f
[HARDWARE][POWER] Enable fp16 support for PowerPC ( #46135 )
...
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
2026-06-23 13:24:49 +00:00
1bf149f334
Filter Pydantic-internal markers from validation error param ( #46457 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Rehmann <114791389+mahadrehmann@users.noreply.github.com >
2026-06-23 13:20:50 +00:00
2a675a7b9f
[Bugfix] Responses API assistant EasyInputMessageParam input ( #44361 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 08:54:46 -04:00
7d47cff933
[Bugfix][KV Offload] Fix swap_blocks_batch on the default stream ( #46379 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-23 05:45:27 -07:00
091bc1026e
[KV Offloading] Add tiering metric plumbing ( #45959 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-23 15:10:36 +03:00
3554ada5d8
[CPU][Bugfix][Speculative Decoding] Accept USE_FP64_GUMBEL in CPU recovered-tokens sampler ( #46069 )
...
Signed-off-by: hillel.darshan <hillel.darshan@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-23 11:54:07 +00:00
wang.yuqi and GitHub
31ca9504b1
[Frontend] Split ServingRender into renderer and entrypoint. ( #44285 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-23 11:19:09 +00:00
d32575a2d2
[ROCm][P/D] Support MoRIIO heterogeneous TP fan-in ( #46332 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-23 10:33:23 +00:00
Juan Pérez de Algaba and GitHub
83fa302ca4
fix(security): prevent infinite loop in split_audio with NaN audio sa… ( #46463 )
2026-06-23 10:24:51 +00:00
frida-andersson and GitHub
20b5af55c1
[ROCm][Perf] DSv3.2: fuse MLA Q concat+fp8-quant in forward_mqa ( #43673 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-06-23 18:12:04 +08:00
Qiming Zhang and GitHub
901a3b091c
fix gpt_oss pp>1 with ep ( #46441 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-23 16:59:11 +08:00
2d721ab5d8
[Rust Frontend] Align Rust allowed_token_ids validation with Python ( #46348 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 08:32:33 +00:00
accaa434f3
[Rust Frontend] Support echo for token-ID completion prompts ( #46219 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-23 08:04:41 +00:00
Sunny Yuan and GitHub
a04654da23
Doc: fix missing GLM-5.x in supported models ( #46452 )
...
Signed-off-by: Sunny Yuan <y.zichen@wustl.edu >
2026-06-23 07:42:27 +00:00
Bugen Zhao and GitHub
25bc3be49c
[Rust Frontend] Correct --reasoning-parser semantics ( #46359 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 15:38:39 +08:00
Ting SUN and GitHub
a46f3eb232
[Bugfix][Model Runner V2] Preserve all allowed_token_ids in the logit bias kernel ( #46245 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-23 07:01:13 +00:00
6c427dd401
[BugFix] Omit empty tool_calls from OpenAI chat responses ( #44105 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-23 13:43:53 +08:00
3ce5823762
[Refactor] Responses API parser state into conversation context ( #46030 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 13:42:58 +08:00
Woosuk Kwon and GitHub
04c2a8deac
[DeepEP V2] Fill invalid recv_topk_idx with -1 ( #46432 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-22 21:45:49 -07:00
7e47fb72b5
[ROCm][P/D] Fix MoRIIO WRITE mode for mixed KV layouts ( #46290 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
2026-06-23 12:12:51 +08:00
a8481be7a9
[Rust Frontend][Perf] Use dedicated runtime for HTTP/request-processing/ZMQ ( #46051 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 04:03:20 +00:00
Kunshang Ji and GitHub
9d3317172c
[XPU][CI]fix xpu kv cache layout test ( #46429 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-23 03:43:29 +00:00
430a95ae3a
[v1][kvcache] Honor prefix-cache retention interval for Mamba/linear attention ( #45845 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 19:51:11 -07:00
Mike G and GitHub
56e5797511
[Quant] Enable modelopt_mixed on Turing (SM75) ( #45375 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-22 19:30:49 -07:00
8db12169a4
fix: stream Qwen3 tool call string arguments ( #46351 )
...
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-06-23 10:26:37 +08:00
33f50773cb
[Doc] Fix typos, grammar, and broken commands across docs ( #46398 )
...
Signed-off-by: MichaelCaoo <a992033227@163.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 02:01:22 +00:00
Micah Williamson and GitHub
fa36f86d77
[CI] Torch 2.11 flaky test_spec_decode_logprobs and gritlm tests ( #45772 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 01:26:54 +00:00
8207ce0850
[Bugfix] Fix humming lm_head crash and FusedMoE weight_shape coercion ( #46420 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 18:19:29 -07:00
e48592066e
[DeepEP V2] Bound num_max_tokens_per_rank in do_expand=False ( #46404 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Roy Wang <jasonailu87@gmail.com >
Co-authored-by: gnovack <novackgm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 18:14:53 -07:00
91ba720b75
[ROCm][CI] Only require q_scale==1.0 for fp8 query in RocmAttention ( #46148 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 18:25:43 -05:00
fxmarty-amd and GitHub
6ead164e52
[CI] Add TP=4 requirement to test_mixed_precision_model_accuracies ( #46161 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-22 18:19:43 -05:00
c97e8f99d6
[ROCm][Quantization][4/N] refactor quark_moe fp8 w/ oracle ( #43721 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:58:03 -07:00
183b5f27ea
[Bugfix][V1][TurboQuant] Reserve workspace before CUDA graph capture ( #44053 )
...
Signed-off-by: Guipeng Zhang <zhangguipeng23z@ict.ac.cn >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 15:47:48 -07:00
ca5b24695b
Fix static actorder handling for compressed-tensors WNA16 MoE ( #41161 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:46:46 -07:00
Charlie Fu and GitHub
6f6bd3b8fe
[ROCm][CI] Increase the max wait time for server startup ( #46417 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-22 17:46:31 -05:00
Andreas Karatzas and GitHub
70ef4d3009
[ROCm][CI] Purging away redundant test group definitions ( #46418 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:42:47 -07:00
e2fe837572
[CI] Fix CPU-Multi-Modal Model Tests timeout by adding a 4th shard ( #46388 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 22:08:00 +00:00
Aarushi Jain and GitHub
fbf9ff7cf4
[CI][ROCm] Restrict MLA cross-layer KV cache test to supported backends on ROCm ( #46401 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-22 17:05:26 -05:00
6cc2c9ba3a
[CI] Add DGX Spark GPQA smoke test ( #39541 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 14:52:38 -07:00
c0b2d8f471
[Bugfix] FusedMoE: coerce shape-(1,) per-tensor scales to 0-D scalar … ( #43362 )
...
Signed-off-by: Varshith <kvarshithgowda@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 13:26:53 -07:00
Mohammad Miadh Angkad and GitHub
d1a38c2762
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend ( #42235 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-22 16:17:18 -04:00
2b4a7491ec
[ROCm][CI] Query total device memory via amdsmi to avoid HIP init ( #46141 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:12:24 -05:00
Saddss and GitHub
82ede09a5a
[Bugfix][KVConnector] Fix SimpleCPUOffloadConnector GPU->CPU store race ( #46278 )
2026-06-22 13:08:47 -07:00
Nick Hill and GitHub
fbf520cf3a
[MRV2] Generalize use of WhisperModelState ( #46096 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-22 12:40:02 -07:00
44d95069e9
Enable DeepSeek V4 and GLM-5.1 on SM120 ( #43477 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-22 11:54:14 -07:00
3ce15fd574
[v1][kvconnector] DecodeBenchConnector: fill list/tuple (Mamba/KDA) KV caches ( #45080 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 11:54:00 -07:00
e4b3da3feb
[Quantization][CI] add humming lm-eval test ( #43752 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:23:55 -07:00
3e6529cc0e
[Bugfix][Spec Decode] Fix EAGLE drafter multimodal encoder cache misses ( #46315 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 18:14:02 +00:00
ac614587f5
[EPLB] Enable nixl eplb communicator for elastic ep ( #45013 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-22 10:54:08 -07:00
f2069b005b
[Pooling] Validate non-negative rerank top_n ( #46119 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:40:47 -04:00
Martin Hickey and GitHub
ccd49f6821
[MyPy] Fix mypy for vllm/lora ( #41722 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-22 10:57:09 -04:00
Li, Jiang and GitHub
1c7bc18318
[Bugfix][CPU] Fix CPU model runner v2 ( #46365 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-22 22:52:05 +08:00
AlexHuang and GitHub
9a938df64e
[Test][KV Offloading] Add unit tests for OffloadingSpecFactory and SecondaryTierFactory ( #46355 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
2026-06-22 17:45:04 +03:00
Liangliang Ma and GitHub
3da4a1b124
[XPU] add awq format for INCXPULinear ( #43404 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
2026-06-22 22:29:13 +08:00
6871738777
[Doc] Document pull request limit ( #46376 )
...
Signed-off-by: simon-mo <simon.mo@hey.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 14:04:56 +00:00
Yifan Qiao and GitHub
aa4990a9a2
[Attention] Re-enable cross-layer KV cache layout for MLA via stride-aware kernels ( #45111 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-22 06:57:02 -07:00
a4610da0c6
[docs] link security docs from AGENTS ( #46373 )
...
Add a security-review routing sentence to AGENTS.md that points agents to SECURITY.md, docs/usage/security.md, and docs/contributing/vulnerability_management.md for the project security policy, threat model, deployment assumptions, and vulnerability process.
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 06:28:25 -07:00
liuzhenwei and GitHub
09cdcf34aa
[XPU] update nixl to v1.2.0 ( #46327 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 20:55:06 +08:00
d2c671c29b
[CPU][RISC-V] Add RVV micro GEMM for WNA16 ( #44324 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-22 12:53:54 +00:00
xiangdong and GitHub
b5a2adec4b
[XPU][CI]Skip v1/spec_decode/test_speculators_correctness.py in intel GPU nightly ( #46356 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 19:30:41 +08:00
78739e3bda
[Bugfix] Reject matryoshka embedding dimensions above hidden size ( #46313 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-22 10:16:35 +00:00
Tuukka Sarvi and GitHub
89accad2cc
[ROCm][DSV4] Disable TileLang MHC dispatch on gfx942 ( #45931 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-06-22 09:26:54 +00:00
3c8e49596c
[Model] ColQwen3.5: fix retrieval correctness (bias + bidirectional) ( #46108 )
...
Signed-off-by: Athrael Soju <athrael.soju@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 17:25:54 +08:00
cec2ec1176
[Bugfix] Avoid racy accepted counts in async spec decode ( #45100 )
...
Signed-off-by: Weiwei Sun <68775773+sunnweiwei@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-22 08:53:16 +00:00
liuzhenwei and GitHub
435f82d61a
[Bugfix] Fix Llama4ForCausalLM initialization test failure ( #46341 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 08:40:43 +00:00
Roger Wang and GitHub
1c4b51b990
Temporarily skip M3 on CI ( #46352 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:35:31 -07:00
2e2c47928b
[Doc] Update MiniMax-M3 ( #45940 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:23:27 -07:00
80abe0de7d
[Rust Frontend] Support thinking_token_budget for chat and completions ( #46137 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-22 16:00:02 +08:00
a9f7b2d41c
[feature][kv_offload] Self-describing KV events for OffloadingConnector ( #43468 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 07:27:46 +00:00
d14e551a53
[Model] Remove MiniMaxText01, MiniMaxVL01, MiniMaxForCausalLM ( #45993 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:20:46 +08:00
68567ef2df
[CPUOffloadingManager] Maintain evictable list in LRUCachePolicy ( #46216 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-22 06:54:44 +00:00
6bc6f2d86d
[1/N][Core] add partial prefix cache primitives ( #45939 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 23:43:10 -07:00
wang.yuqi and GitHub
1eb2cc961e
[Frontend] Refactor ServingTokenization entrypoint. ( #46022 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-22 06:27:58 +00:00
31124749d1
[Bugfix] [Rust Frontend] Fix stop string truncation with repeated matches ( #46113 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-22 14:11:29 +08:00
Ma Jian and GitHub
9037498c22
[DSV4][XPU] Pass gemm1_clamp_limit to XpuFusedMoe ( #44517 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-22 12:57:10 +08:00
db32b53e30
[SpecDecode] Support DFlash with FlashInfer ( #43081 )
...
Signed-off-by: gss <2783977641@qq.com >
Co-authored-by: gss <2783977641@qq.com >
2026-06-22 04:55:30 +00:00
xiangdong and GitHub
b529bfd6c5
[XPU][CI] Add agent_tags for Intel GPU CI ( #45768 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 10:33:17 +08:00
Micah Williamson and GitHub
f3df7a7231
[ROCm][CI] Enable kv_connector unit tests on ROCm ( #45955 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-22 05:08:44 +03:00
485bbe1c6f
[CI] Fix missing tp_size attribute on RoutedExperts ( #46163 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 18:46:49 -06:00
Matt and GitHub
a19ff2218a
[Hardware][AMD][CI] Fix Spec Decode Eagle test group ( #46018 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:40:02 -05:00
Matt and GitHub
4f0d0049a0
[Hardware][AMD][CI] Fix Kernels Attention test groups ( #46080 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:10:51 -05:00
13b83d77ad
[ROCm][CI] skip test_double_aiter_rms_quant_fusion ( #45967 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 16:53:11 -05:00
Matt and GitHub
50241602fd
[Hardware][AMD][CI] Fix gfx942 Kernels MoE test group ( #46298 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 16:45:37 -05:00
Ting SUN and GitHub
12fe2a9aac
[Bugfix][Qwen3-VL] Fix multi-video crash with list-valued fps/num_frames ( #46305 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 14:31:23 -07:00
Benjamin Chislett and GitHub
89bd2c14d3
[Spec Decode] Add Qwen3 architecture support for EAGLE3 ( #43132 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-21 13:55:26 -07:00
ZedongLiu and GitHub
9c450b1027
[Kernel][Bugfix] Fix INT8 per-token-head KV cache rounding in Triton reshape-and-cache ( #45361 )
...
Signed-off-by: ZedongLiu <113341356+Zedong-Liu@users.noreply.github.com >
2026-06-21 15:59:40 -04:00
635c38338a
[Multimodal] Add Qwen2-VL/Qwen2.5-VL processor-mapped video loader ( #45555 )
...
Signed-off-by: Ranran <hzz5361@psu.edu >
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-21 18:56:50 +00:00
c441ad1c07
[KV Offloading] Add labeled metrics support ( #45957 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-21 18:04:01 +00:00
Jee Jee Li and GitHub
745bba5ea8
[Model]Fix MiniMaxM2ForCausalLM perf regression ( #45935 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-22 00:28:52 +08:00
2cac89f9da
[Spec Decode] Support mixed KV page sizes for DFlash ( #45181 )
...
Signed-off-by: Alex Steiner <asteiner@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 22:45:14 +08:00
3e6e33526d
[Disagg] return routed_experts on streaming generate responses ( #44638 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-21 07:37:10 -07:00
b91b7726e0
[ROCm][P/D] Support MiniMax-M3 mixed KV layouts in MoRIIO READ mode ( #46039 )
...
Signed-off-by: Jun Kang Chow <junkangchow@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-21 12:55:19 +00:00
Palaiologos1453 and GitHub
d3ad8e8bcd
[Bugfix] Defer offload reads while transfers are pending ( #46231 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-21 14:30:13 +03:00
b80ce9dd2f
[CI][test] Replace InternVL2-1B with InternVL3-1B in test_pipeline_parallel.py ( #46241 )
...
Signed-off-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-21 15:11:19 +08:00
b5495cc5f9
Fix memory pointer overflow in Mamba state buffers ( #44665 )
...
Signed-off-by: Shifani Rajabose <shifani.rajabose@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-21 14:00:50 +08:00
Ting SUN and GitHub
183a430c13
[Bugfix][Model Runner V2] Fix min_tokens off-by-one in the V2 GPU sampler ( #46243 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 05:06:49 +00:00
Matt and GitHub
a346d589f5
[Bugfix] Fix NVFP4/OCP MX MoE emulation ( #46254 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-20 23:13:10 -05:00
Nick Hill and GitHub
7df3d7dada
[Core] Ensure memory is pinned prior to async h2d copy ( #45424 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 20:02:24 -07:00
8dd1b702f2
[Misc] Fix stale doc URL and docstring module path ( #35530 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-20 23:57:01 +00:00
f57ac274b2
[Render] Add reasoning/tool parsing to /derender + fix byte-fallback FFFD ( #45919 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-20 19:43:32 -04:00
6e919960af
[Perf] Skip/shrink all_token_ids copy in scheduler for non-async and V2 runner ( #45840 )
...
Signed-off-by: amanchugh89 <amanchugh.89@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 22:36:57 +00:00
Jonathan Chen and GitHub
c88d3d4775
[SimpleCPUOffloadConnector] PCP + DCP support ( #39831 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
2026-06-20 15:01:06 -07:00
Yifan Qiao and GitHub
ab7fcbdd5d
[Perf][KVConnector][Mooncake] Compact chunk-hash keys and zero-copy lookup wire format ( #45969 )
2026-06-20 15:00:11 -07:00
3b4a76b63f
[KV-Offloading] : Expose CPU cache usage metric ( #45737 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-20 21:21:55 +00:00
cc22621b51
[KV Offload] Support packed HMA KV cache layout ( #46205 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:19:40 +00:00
77148992cf
[Bugfix] Move extract_layer_index back inside is_v32 guard ( #46199 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 21:19:10 +00:00
891cc4b9c5
[Frontend] Report cache usage in Anthropic /v1/messages API ( #40912 )
...
Signed-off-by: mistral0105 <zhangshuoming17@mails.ucas.ac.cn >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:12:48 +00:00
TJian and GitHub
1bdf9810aa
[ROCm] [Bugfix] Bugfix ROCm Sparse Indexer ( #46222 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-20 13:38:42 -07:00
ebfbcfe46a
Stop setting CUDA_VISIBLE_DEVICES internally in vLLM, add device_ids arg ( #45026 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: kourosh hakhamaneshi <kouroshHakha@users.noreply.github.com >
2026-06-20 13:38:10 -07:00
e9de72fe6c
[Bugfix] Guard model_config access in _log_compilation_config ( #46198 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 19:26:38 +00:00
d272418f45
[Perf] Optimize Qwen3-VL multi-video prompt processing ( #46026 )
...
Signed-off-by: Sirius29 <422058530@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 07:09:18 -07:00
Sumanth R Hegde and GitHub
7ff7f5c8eb
Revert "Fix Stale Encoder Cache After Weight Update" ( #46125 )
2026-06-20 07:09:09 -07:00
dced290769
[Hardware][AMD][CI] Fix e2e core test group ( #46024 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 02:04:35 -05:00
JasonLi314 and GitHub
93bad11912
[Bugfix] Fix gridDim.y overflow for large row counts ( #45255 )
...
Signed-off-by: Jason Li <li.jason.cs@gmail.com >
2026-06-19 23:27:45 -04:00
djramic and GitHub
0fbf42af84
[ROCm] Fix VRAM not freed in test_phi3v ( #46046 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 17:20:59 -05:00
Charlie Fu and GitHub
e6cd8913dd
[ROCm][CI] Skip Qwen3.5-35B-A3B-MXFP4-AITER-TP2 for non gfx950 ( #46109 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-19 17:20:10 -05:00
Ben Browning and GitHub
859e4d436b
[Bugfix][Parser] Fix U+FFFD leak at reasoning-to-content transition in engine parsers ( #46159 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 22:09:28 +00:00
Micah Williamson and GitHub
4a083cc858
[ROCm][CI] Pin test_rocm_compressed_tensors_w8a8 to TRITON_ATTN ( #46180 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-19 15:20:06 -05:00
Vadim Gimpelson and GitHub
ca7e1f2c43
Move CI failure diagnosis docs into ci-fails-buildkite skill ( #45975 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-06-19 20:12:40 +00:00
djramic and GitHub
dec860fb19
[ROCm] Use vLLM's fp8 quant max in AITER hipBLASLt accuracy test ( #46176 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 13:24:02 -05:00
Harry Mellor and GitHub
0a49fb2b13
Fix dead link in docs ( #46181 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-19 18:16:09 +00:00
Ben Browning and GitHub
4a8abf37c7
[Test] Migrate test_openai_schema.py to schemathesis 4.x ( #46173 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 18:05:18 +00:00
01192139bf
[DSv4] Pack KV caches into contiguous per-block allocations for DeepSeek V4 ( #44577 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-19 12:55:42 -04:00
Chris Leonard and GitHub
b9a7cd464c
[12/n] final _C library kernel migration ( #45415 )
2026-06-19 06:57:26 -07:00
69bdd34542
[Bugfix] Fall back to Pydantic loc for param in validation errors ( #46038 )
...
Signed-off-by: professorsab <135441198+professorsab@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-19 19:11:11 +08:00
Kunshang Ji and GitHub
ec67d7ae61
[xpu] bump up vllm-xpu-kernels v0.1.10 and upgrade 2618 umd ( #40367 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-19 15:37:20 +08:00
ecf9d83520
[AMD][CI] Fix Language Models Test (Extended Generation) failures ( #45509 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-19 12:06:56 +08:00
Samuel Shen and GitHub
c9135db27c
[Docs] Update stale LMCache examples ( #45762 )
...
Signed-off-by: Samuel Shen <slshen@tensormesh.ai >
2026-06-19 03:21:36 +00:00
2a6c6b9429
[DeepSeek-V4] Support TEP=16 for the block-FP8 shared expert ( #46001 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:10:12 -07:00
Jared Wen and GitHub
ab66606993
[bugfix]Indexer init skip and MTP TopK share for iteration ( #45895 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-19 09:57:51 +08:00
9ea3a4015b
[Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded ( #42120 )
...
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 18:26:09 -07:00
Flora Feng and GitHub
560fb8b867
[Cohere] Remove dead prepare_structured_tag override in Cohere parser ( #46099 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-19 01:02:11 +00:00
Wentao Ye and GitHub
675cd5d228
[Model Runner V2] Fix MRv2 memory leak test ( #46095 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-19 00:36:40 +00:00
7f616c327d
[Bugfix] [Parser] Fix empty tool block silently dropping subsequent content ( #46091 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-18 23:17:18 +00:00
Ivy Xu and GitHub
c3c6d723fd
[Perf] Remove unused loggers in reasoning/ ( #45988 )
...
Signed-off-by: Ivy <fakeshadow1337@gmail.com >
2026-06-18 22:24:29 +00:00
41dcf49ca5
[Bugfix][KV Connector] Disable Mooncake TP put-striding when DCP > 1 ( #45371 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:13:44 -07:00
35e4dd4a69
[KV Connector][Mooncake] Async lookup to reduce scheduler overhead ( #45659 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-18 21:44:02 +00:00
4ce2d01453
fix(anthropic): auto-detect template support for mid-conversation system messages ( #46025 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:19:11 -04:00
Woosuk Kwon and GitHub
16908e132e
[MRV2] Make FP32 Gumbel sampling more accurate ( #45996 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-18 19:42:09 +00:00
Wentao Ye and GitHub
225936a1dd
[CI Bug] Revert #42379 to fix CI Multi-Modal Models (Extended Generation 1) ( #46070 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:37:39 -07:00
f6ba720963
(security) Upgrade Starlette to >= 1.0.1 to fix CVE-2026-48710 ( #45675 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-18 12:35:13 -07:00
Wentao Ye and GitHub
b53b1c7ffe
[Model Runner V2] Migration to support quantized model by default [5/N] ( #44446 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:20:44 -07:00
79ca54d221
[Bugfix][Quantization] Don't reject fp8_e5m2 KV cache for non-fp8 quantized checkpoints ( #45040 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:18:25 -04:00
Ben Browning and GitHub
09f3cd5c10
[Bugfix] [Parser] Fix Qwen3 latent bug in partial params dropping values containing < ( #46047 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 18:04:06 +00:00
ea6078fe6a
[KV Connector][Offloading] Disable parallel-agnostic fs-tier cache on V2 model runner ( #46044 )
...
Signed-off-by: Itay Etelis <etelis2019@gmail.com >
Co-authored-by: Itay Etelis <etelis2019@gmail.com >
2026-06-18 20:43:35 +03:00
Palaiologos1453 and GitHub
a0df04e477
[Tests] Add Qwen3 streaming parser delta boundary cases ( #45708 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-18 17:37:39 +00:00
stefankoncarevic and GitHub
e2352c2974
[ROCm][Spec Decode] Fix probabilistic draft probs test attention backend ( #45706 )
...
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com >
2026-06-18 11:59:37 -05:00
qli88 and GitHub
25faa1f4cc
[CI]Enable mxfp4 lora test for ROCm platform ( #43802 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-18 16:59:09 +00:00
Humphrey and GitHub
4583630b56
[Bugfix][Kernel] Check output alignment in vectorize_with_alignment (fixes misaligned-address crash for non-multiple-of-8 head sizes) ( #45466 )
...
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com >
2026-06-18 16:58:22 +00:00
Divakar Verma and GitHub
21da47dabe
[ROCm][CI] move lora%N test to mi300 and gate ( #45970 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-19 00:50:32 +08:00
6c379b9e54
[Frontend] Add Streaming Parser Engine and new GLM4.7/GLM5.1/GLM5.2 Parser ( #45915 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-19 00:42:10 +08:00
Rohan Potdar and GitHub
5099474633
[Bugfix][ROCm] Fix rocm_aiter_per_tensor_quant custom op aliasing ( #45747 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-18 11:30:21 -05:00
Yuwen Zhou and GitHub
058cc0a8b6
[Bugfix] Restore is_sym guard for zp in GPTQ/CT MoE to fix symmetric quant regression ( #45656 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-18 16:20:29 +00:00
837db7605e
[Bugfix][Tool Parser] Handle non-finite numbers in coerce_to_schema_type ( #43984 )
...
Signed-off-by: ashishpatel26 <shriganesh.patel@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:00:20 +00:00
Mark McLoughlin and GitHub
bf2a393034
Temporarily remove @markmc from CODEOWNERS ( #46053 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-18 14:15:43 +00:00
d682968aa9
[Model] Remove BambaForCausalLM ( #45990 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-18 06:51:00 -07:00
021cdf72bc
Fix _riscv_supports_rvv_vlen128() to detect RVV on hardware without zvl flags ( #43179 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn >
2026-06-18 21:22:35 +08:00
Ashar and GitHub
4cb5e746b6
[Rust Frontend]: Add /get_world_size route with static parallel size ( #44801 )
2026-06-18 13:10:20 +00:00
Jee Jee Li and GitHub
22cc891108
[Kernel] Add PDL support for DeepGEMM kernel ( #46006 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-18 20:49:01 +08:00
afdcbd5d39
[ROCm][DSv4] Functional fixes for DeepSeek V4 on MI300X/MI325X ( #45681 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Jin Tao <jintao12@amd.com >
2026-06-18 12:21:14 +00:00
8d4f54966c
fix(quantization): Fix AWQ dequantize on Intel XPU and refactor AutoAWQ config ( #42727 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@tencent.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:12:28 +08:00
351c72d6e5
[CPU] Skip Triton kernel monkey-patches when Triton-CPU is available ( #44991 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-18 18:59:30 +08:00
Tahsin Tunan and GitHub
7299e6509e
[Rust Frontend] Return model metadata fields in /v1/models ( #45950 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-18 10:29:21 +00:00
littlecircle0730 and GitHub
08985351f3
Fix Stale Encoder Cache After Weight Update ( #45093 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
2026-06-18 09:32:10 +00:00
Wei Zhao and GitHub
5fd3b276f8
[Mooncake] Skip KV lookup for non-reachable SWA blocks ( #45444 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-18 02:23:20 -07:00
1e9f04da14
fix(anthropic): preserve inline system message position for prefix caching ( #44602 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
2026-06-18 15:58:11 +08:00
702214146c
[Bugfix][Frontend] Fix Anthropic count_tokens decorator order driving server load negative ( #44725 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 23:56:46 -07:00
a331589394
[XPU] Update nixl to v0.10.1 in Dockerfile ( #40287 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:01:26 +08:00
Micah Williamson and GitHub
e945169207
Revert "[Kernel] Add PDL support for DeepGEMM kernel" ( #45999 )
2026-06-17 22:59:48 -07:00
554352a311
[Test][KV Connector] Add request_finished fence population tests for offloading scheduler ( #45679 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@future.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:52 +03:00
421c1ec448
[KV Offloading] Remove dummy worker-side stats from OffloadingConnector ( #45905 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@alexai.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:28 +03:00
b4c80ec0fd
[Refactor] Remove dead cutlass mxfp8 code ( #44681 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-17 21:18:25 -07:00
Ronen Schaffer and GitHub
f428718ffe
[Fix][KV offload] Defer on_request_finished until in-flight transfers drain ( #45823 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-18 07:05:46 +03:00
Jee Jee Li and GitHub
4403af8fb5
[Kernel] Add PDL support for DeepGEMM kernel ( #42996 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-17 20:37:17 -07:00
d57888efa4
[SimpleCPUOffloadConnector]: Add support for reset_cache() ( #39726 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
Signed-off-by: Jonathan <chenleejonathan@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 19:47:12 -07:00
ed938ad7db
[CPUOffloading] Guard CPU eviction check ( #45757 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-18 05:34:59 +03:00
Reid and GitHub
731fb3323d
[Rust Frontend] Validate tokenized bad_words vocabulary range ( #45876 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-18 02:28:45 +00:00
8dd8b6ed78
[XPU] Fix FP8 block-scaled scheme selection on non-CUDA platforms ( #43958 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-18 10:16:20 +08:00
e1a5fc406b
[Rust Frontend][Perf] O(n) argument scan in tool parser ( #45826 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-18 01:42:35 +00:00
Ace Eldeib and GitHub
b4092176b9
[Bugfix] Complete one-shot fused all-reduce PDL at end to avoid NaN ( #45448 )
2026-06-18 00:54:39 +00:00
Jee Jee Li and GitHub
ebbb2d55ac
[CI/Build][Bugfix] Fix SD LoRA ( #45941 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 00:34:15 +00:00
liuzhenwei and GitHub
2959a9273a
[XPU][CI] add model runner v2 into CI ( #44650 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-18 00:28:34 +00:00
1797576237
Revert "[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture" ( #45309 ) ( #45972 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-17 17:20:59 -07:00
0d339cf135
[Bugfix] Fix NixlConnector handshake block_len validation for GQA-replicated KV heads ( #45879 )
...
Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com >
Co-authored-by: waynehacking8 <waynehacking8@gmail.com >
2026-06-17 15:11:29 -07:00
5fd21eb0b2
[BUG] fix hidden states nan for hybrid attention models ( #45849 )
...
Signed-off-by: shanjiaz <hezhao@redhat.com >
Co-authored-by: shanjiaz <hezhao@redhat.com >
2026-06-17 18:02:24 -04:00
Ting SUN and GitHub
9d4b87f4f0
[Bugfix][Model] Validate DefaultModelLoader / LoadConfig and fail with clear errors ( #45196 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-17 21:46:33 +00:00
58b2e89642
[Bugfix][Gemma4] Render reasoning on assistant turns without tool_calls ( #45867 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-17 20:44:15 +00:00
Wentao Ye and GitHub
2659f60a1a
[Refactor] Remove dead quantization code and tests ( #45454 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 16:12:01 -04:00
091386a99b
[Bugfix] MiniMax-M3 (AMD): add packed_modules_mapping and pass swiglu… ( #45794 )
...
Signed-off-by: wangjiaxin99 <jiaxwang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-06-17 19:15:46 +00:00
qli88 and GitHub
d112eb1ac7
[feature] MiniMax-M3-MXFP4 support added ( #45896 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-17 18:50:48 +00:00
Wentao Ye and GitHub
2a47a9ff0f
[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture, 26.8% ~ 27.9% E2E TTFT improvement ( #45309 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 09:34:53 -07:00
Wentao Ye and GitHub
9c7c74bf10
[Log] Update deepgemm log ( #45857 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 15:34:22 +00:00
danisereb and GitHub
5e27b2baf4
[Bugfix] Pass TP group to FlashInfer all-reduce fusion ( #45917 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-06-17 15:24:26 +00:00
zhanqiuhu and GitHub
eb0fdeb1e8
[Bugfix][PD] Fix DSV4 disaggregated serving ( #45831 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-06-17 15:17:14 +00:00
46f74e144b
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_dynamic_per_token_quant ( #34432 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-17 23:03:54 +08:00
Wentao Ye and GitHub
0a7bacdcac
[DSv4 Perf] DSv4 flashinfer sparse index cache for metadata, 2%~4% TTFT improvement ( #45863 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 10:55:48 -04:00
amirkl94 and GitHub
8b2b566ea7
Feature: Enable Flashinfer non-gated MoE bf16 ( #43853 )
...
Signed-off-by: Amir Klein <203507526+amirkl94@users.noreply.github.com >
2026-06-17 14:32:49 +00:00
xaguilar-amd and GitHub
0b131b16c9
[ROCm][AITER][Quark] Tag per-channel FP8 weights as PER_CHANNEL so AITER pre-shuffled GEMM is selected ( #44626 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
2026-06-17 14:05:34 +00:00
bcb518ad7a
[quant][autoround]Refactor INC quantization into package with INCScheme orchestrator ( #40601 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
Co-authored-by: n1ck-guo <heng.guo@intel.com >
Co-authored-by: Zhenzhong1 <zhenzhong.xu@intel.com >
2026-06-17 21:51:32 +08:00
Chaojun Zhang and GitHub
06e1e0885c
[XPU] Fix test_logprobs_e2e import error: pin lm-eval[api]>=0.4.12 ( #44469 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-17 12:26:47 +00:00
Isotr0py and GitHub
1a59078c87
[CI/Build] Avoid duplicate ViT CG test introduced by accident ( #45654 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-17 12:23:44 +00:00
Oğuzhan KIR and GitHub
fa85ead2f3
[MM][Perf][CG] Support ViT full CUDA graph for Kimi-VL ( #41992 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
2026-06-17 12:14:01 +00:00
e28e8c8782
[ROCm][Quant] Minimax-M3: Enable fp8_per_channel for bf16 weights on mi300x ( #45854 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-17 12:02:40 +00:00
Angelo Ruocco and GitHub
ee0fd6984a
docs, kv_offloading: add docs for selective offload ( #45279 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
2026-06-17 14:58:00 +03:00
vllmellm and GitHub
d537122398
[ROCm][Bugfix]: Fallback GFX942 sparse MLA ops to Triton ( #45782 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-17 11:41:29 +00:00
Juan Pérez de Algaba and GitHub
3d20275bb4
fix(security): enforce audio decode duration limit in chat completions path ( #45908 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-17 11:07:13 +00:00
f694d43b33
[Bugfix][test] Use Salesforce/wikitext for ppl tests ( #45913 )
...
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-17 10:37:19 +00:00
Nikhilesh Chhetri and GitHub
3c6084bb0d
[Bugfix][Gemma4] Pre-initialise streaming reasoning state when prompt ends inside an open <|channel> ( fixes #45834 ) ( #45852 )
...
Signed-off-by: nikhilesh-csa <nchhetri@csa1.com >
2026-06-17 06:16:02 -04:00
Joel Smith and GitHub
68ff30d40e
[Bugfix] Fixes MiniCPM-O resampler device placement to avoid tensor device mismatch ( #42332 )
...
Signed-off-by: j9smith <j.smith9103@outlook.com >
2026-06-17 08:35:27 +00:00
6d8fff5698
[KV Connector][Offloading] Avoid blocking the engine to flush offloads on idle ( #45595 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-17 11:35:07 +03:00
e2c58570ea
[Rust Frontend] Support hybrid/external DP LB in Python supervised bootstrap ( #45805 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-17 07:32:40 +00:00
Taneem Ibrahim and GitHub
43fa24e832
[Misc] Validate Cohere Embed Mixed Content Payloads ( #45873 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-17 06:57:12 +00:00
arghyadeep sarkar and GitHub
93bbe94d3a
[Kernel] Add weightless RMSNorm CUDA kernels for has_weight=False ( #41430 ) ( #44109 )
...
Signed-off-by: hello-args <args.sarkar@gmail.com >
2026-06-16 23:45:55 -07:00
Will Eaton and GitHub
17bc144556
[Rust Frontend] Add serde defaults for omit_defaults fields in EngineCoreSamplingParams ( #45848 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-17 06:40:49 +00:00
Sahil Singh and GitHub
295232a26a
[Rust Frontend] Add /abort_requests endpoint ( #44382 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-17 06:40:47 +00:00
Reid and GitHub
56e4345226
[Rust Frontend] Support prompt-only completions ( #44938 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-17 06:38:06 +00:00
Nick Hill and GitHub
e9993a52aa
[BugFix][CI] Fix scheduler plugin test ( #45897 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 06:30:49 +00:00
a46abb7ae6
[Bugfix][Quantization] Reject unsupported compressed tensors KV cache schemes ( #45312 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 05:08:41 +00:00
4c62663315
[M3] Enable FP8 sparse GQA ( #45744 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-16 21:38:03 -07:00
d78650cf97
[CI][NIXL] Pin NIXL to 1.2.0 ( #45843 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
Signed-off-by: Itay Alroy <75032521+itayalroy@users.noreply.github.com >
Co-authored-by: ovidiusm <ovidium@nvidia.com >
2026-06-16 21:29:34 -07:00
5bdc01bcc3
[M3] Tune Triton indexer score decode for spec-decode ( #45743 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 21:07:34 -07:00
liangel-02 and GitHub
20a5f8b43b
[FlexAttention] make custom mask mods fully cudagraphable ( #45232 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-06-17 11:53:12 +08:00
7b5d60cc37
[Bugfix][V1] Clean up compiled-model bytecode hooks on VllmRunner exit ( #45195 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 20:31:17 -07:00
Nick Hill and GitHub
14b438a98b
[ModelRunnerV2] Various model/config compatibility fixes ( #45868 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 03:23:01 +00:00
2785a5e0e6
[Bugfix][ROCm] Fix FP8 per-tensor scale rank mismatch causing Inductor assertion failure ( #44912 )
...
Signed-off-by: nehmathe2 <nehmathe2@gmail.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: nehmathe <nehmathe@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:17:42 -07:00
efd15e192a
[Bugfix][ROCm] Fix MiniMax-M3 FP8 KV cache dtype ( #45720 )
...
Signed-off-by: Cam Quilici <cjquilici@gmail.com >
Signed-off-by: Cameron Quilici <cjquilici@gmail.com >
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-17 03:14:45 +00:00
556b063e45
[XPU] Fix test_spec_decode_logprobs: use FLASH_ATTN for XPU in GPU_DETERMINISM_KWARGS ( #44468 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 11:07:04 +08:00
aa0ac8a661
[CI] Run pre-commit on self-hosted vllm-runners ( #45865 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 19:49:22 -07:00
Federico and GitHub
b831374cf1
[Bugfix][Gemma4] Fix parsing when thinking is disabled ( #45832 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-17 02:41:36 +00:00
71bc19dbdd
[Bugfix] Fix MoE model load OOM in FlashInfer_TRTLLM backend with sleep mode ( #45589 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 19:36:51 -07:00
Kunshang Ji and GitHub
ef2c40dc00
[XPU][CI] fix server test file path ( #45870 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 09:06:25 +08:00
4bf699d310
[Kernel] Support DS Mamba tail copy for MTP align mode ( #45473 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Co-authored-by: Thomas Parnell <tom.parnell@gmail.com >
2026-06-16 22:50:30 +00:00
Stan Wozniak and GitHub
520828789c
Apply LRU policy only to proper cache entries ( #42656 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-16 21:49:15 +00:00
9d4dc4ca2f
[Kernel] Support GLM-5 dimensions for TRT-LLM ragged MLA prefill ( #43525 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 20:49:47 +00:00
Federico and GitHub
b9684d99e9
[Bugfix] Gemma4: skip forced JSON for required/named tool choice ( #45795 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-16 20:38:29 +00:00
Divakar Verma and GitHub
4fadf9c92c
[ROCm][CI] fix multimodel run cmds ( #45858 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-16 15:31:52 -05:00
Nick Hill and GitHub
d8d95998dc
[Core] Add prefill step cadence for better non-PD DP balancing ( #44558 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 13:17:18 -07:00
Flora Feng and GitHub
475a6ad18a
[Misc] Update Mergify tool-calling label ( #45853 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-16 19:08:00 +00:00
Hongxia Yang and GitHub
f2beaa80c8
[ROCm][Quant] mxfp8 moe/linear gfx950 tuning for MiniMax-M3 ( #45725 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-06-16 18:50:40 +00:00
8e27a9c215
[PERF] Fuse multi-group block table staged writes ( #44944 )
...
Signed-off-by: jesse <szxfml@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 10:53:27 -07:00
7d567172fc
[Bugfix] Fix Qwen3 prompt tool-call reasoning false positive ( #45763 )
...
Signed-off-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
Co-authored-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
2026-06-16 17:48:01 +00:00
Chauncey and GitHub
f00e163f35
[Frontend] Add Streaming Parser Engine and new MinimaxM2 Parser ( #45701 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-16 13:38:17 -04:00
44b2512767
[KV Connector][Mooncake] Add cache_prefix to namespace store keys ( #45767 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 10:24:20 -07:00
188c68798e
[KVConnector][MoRIIO] Allow overriding the advertised host IP ( #45488 )
...
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 17:18:37 +00:00
c45f681932
[Bugfix][Core] Fall back when numactl --membind is blocked in constrained containers ( #45438 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-16 09:49:56 -07:00
89e8645a9e
[Model] Remove Dots1ForCausalLM ( #45637 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-17 00:32:18 +08:00
Wentao Ye and GitHub
88a9cdd439
[Model Runner V2] Enable GraniteMOE for MRv2 by default ( #45461 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-16 09:31:32 -07:00
Micah Williamson and GitHub
6f612fbedf
[ROCm][CI] Patch conftest to resolve occasional OOMs ( #45722 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-16 10:00:15 -05:00
Sting Lin and GitHub
506ec6d656
Upgrade tpu-inference to v0.22.1 ( #45793 )
2026-06-16 07:54:57 -07:00
a52205bccf
[Model] Add HrmTextForCausalLM (Hierarchical Reasoning Model — Text) ( #43098 )
...
Signed-off-by: Wuyifei <wuyifei@me.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-16 22:41:41 +08:00
3d34f8cbdc
[ROCm][Cleanup] Remove stale AITER FA hybrid KV-cache TODO ( #44178 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-16 07:28:06 -07:00
Carl Y and GitHub
eb04c769d3
feat: MLA prefill enable FA4 fp8 output ( #43050 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-06-16 07:10:59 -07:00
ce3ef17bec
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_per_block_quant ( #36895 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-16 22:09:52 +08:00
bf5149b516
[Bugfix] Fix FlashMLA sparse accuracy with topk_length and zero-init padding ( #36616 )
...
Signed-off-by: AjAnubolu <anuboluajay@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 07:09:00 -07:00
Tahsin Tunan and GitHub
cca3365b73
[Rust Frontend] Add CORS support ( #45753 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-16 13:47:11 +00:00
040df8f2ea
[CI] Fix attention benchmark smoke test ( #45728 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 13:43:37 +00:00
ced32bb474
[Perf] Add VLLM_TRITON_FORCE_FIRST_CONFIG to skip Triton autotuning ( #42425 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 15:16:45 +02:00
c5e5c33fcd
[Bugfix][MoE] Restore routed output unpadding before shared expert add ( #45707 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 16:06:28 +03:00
Mike G and GitHub
a8c86eeb16
[Quant] Support modelopt_mixed on Ampere (SM80/SM86) ( #45306 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-16 08:43:44 -04:00
Andreas Karatzas and GitHub
7e179e4bc0
[ROCm][CI] Gate incompatible HF references on Transformers v5 ( #41532 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:34:11 +08:00
405c7cf283
[ZenCPU] Add zencpu Platform Runtime Logging and Docs ( #42726 )
...
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-16 08:23:12 -04:00
3f53e2138f
[Refactor] Remove Fp8OnlineLinearMethod as scheduled ( #45463 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 04:35:58 -07:00
Hank Han and GitHub
d53f4593ce
[KV Connector][Mooncake] Pipeline-parallel support for PD-disaggregated serving with Mooncake connector ( #44528 )
...
Signed-off-by: hanhan.hank <hanhan.hank@bytedance.com >
Signed-off-by: Hank Han <hanhan7630@outlook.com >
2026-06-16 04:35:38 -07:00
ad32608e24
[MM][Perf][CG] Support dual-path ViT full CUDA graph for DeepSeek-OCR ( #43586 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 04:35:20 -07:00
Thien Tran and GitHub
b2cfae777d
Add Triton recompile detection ( #45631 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-16 18:25:28 +08:00
wangxiyuan and GitHub
3f1ff1ff14
[Misc]Clean up useless test ( #45792 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-06-16 09:53:08 +00:00
c69c73418a
[XPU][CI] add intel xpu cases for nightly CI ( #44372 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-16 16:35:08 +08:00
Thomas Parnell and GitHub
ebf3a6d705
[Bugfix] Fix trtllm fused allreduce+rms_norm for transformers backend ( #45307 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-16 08:34:27 +00:00
wang.yuqi and GitHub
c4fd9794e9
[Frontend] Remove AsyncMicrobatchTokenizer. ( #45759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-16 08:02:11 +00:00
7ad894c86a
[Bugfix] Prevent cuMemcpyBatchAsync segfault with MTP and KV offloading ( #44784 )
...
Signed-off-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-16 07:58:39 +00:00
Li, Jiang and GitHub
a7fdfeef72
[CPU] Support Gemma Diffusion ( #45690 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-16 14:39:56 +08:00
Jimmy Lee and GitHub
8bf374955f
[Bug Fix] Allow pinned memory for WSL2 ( #41496 )
...
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com >
2026-06-16 05:56:26 +00:00
Cyrus Leung and GitHub
9096659edb
[Cleanup] Remove dead env ( #45777 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-15 22:56:23 -07:00
Taneem Ibrahim and GitHub
81d8f4ebac
[Misc] Added validation for Cohere /v2/embed input field exclusivity ( #45640 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-16 05:42:43 +00:00
a9a8a32dcd
Register parsed config classes before tokenizer init ( #40299 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-16 05:33:08 +00:00
9d808e2309
[Core] Use fastsafetensors ParallelLoader for weight loading ( #40183 )
...
Signed-off-by: Git Bisector <gitbisector@gmail.com >
Signed-off-by: gitbisector <gitbisector@gmail.com >
Signed-off-by: git bisector <gitbisector@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-15 22:32:05 -07:00
Ben Browning and GitHub
f3858d5422
[Frontend] [Parser] Migrate Nemotron V3 to streaming parser engine ( #45755 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-16 05:31:21 +00:00
Bugen Zhao and GitHub
259ff891be
[Rust Frontend] Require ModelConfig.vocab_size to be present ( #45696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 05:30:25 +00:00
6607a80dab
[Bugfix][Gemma4] Fix offline parser truncation, adjust_request token leak, and chat template sync ( #45553 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-16 04:31:53 +00:00
liuzhenwei and GitHub
b8bd773fe4
[XPU] Fix Triton attn fp8/bf16 check failing ( #45758 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-16 12:31:20 +08:00
Ruinan Ma and GitHub
2addbb9cc9
[BugFix] Support async scheduling with prompt embeds for multimodal models ( #45673 )
...
Signed-off-by: Ruinan Ma <r7ma3088@gmail.com >
2026-06-16 04:12:54 +00:00
Isotr0py and GitHub
e3cfea2e1b
[Multimodal] Add Qwen3-VL video loader ( #44412 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 03:45:34 +00:00
Bugen Zhao and GitHub
f99260d2aa
[Rust Frontend] Lower out-of-vocab validation to text layer ( #45685 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 03:37:58 +00:00
Bugen Zhao and GitHub
3f65e21e32
[Rust Frontend] Support max_logprobs validation ( #45674 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 10:57:56 +08:00
xx-thomas and GitHub
b00e76ff72
[Misc][Model] add io processor for query/document embeddings from ColBERT (jinaai/jina-colbert-v2) ( #45210 )
...
Signed-off-by: thomas <thomas.varghese@columbia.edu >
2026-06-16 01:32:32 +00:00
Woosuk Kwon and GitHub
f4359a70f9
[DSV4][Minor] Fix supported KV cache dtypes ( #44892 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-16 00:14:51 +00:00
Itay Alroy and GitHub
3afe659b6b
[EP] Enable DBO with NIXL EP ( #45275 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 23:37:22 +00:00
Itay Alroy and GitHub
16e91176cf
[EP] Query NIXL EP top-k index dtype ( #45298 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:50:18 +00:00
Itay Alroy and GitHub
ab8b0fe338
nixl_ep: Skip post-receive quantization for NVFP4 ( #45606 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:42:05 +00:00
d467a2a7f2
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer ( #45357 )
...
Signed-off-by: llx-08 <2596671364@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-15 21:36:09 +00:00
76a373eff4
[Frontend] Replace legacy Gemma4 parsers with engine-based implementation ( #45588 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 21:34:07 +00:00
Zang Peiyu and GitHub
25ee659db0
Fix parallel_tool_calls: null treated as false instead of default true ( #44955 )
...
Signed-off-by: factnn <166481866+factnn@users.noreply.github.com >
2026-06-15 21:14:10 +00:00
eacff17c8d
[Model Runner V2][Bugfix] Fix MRV2 LoRA warmup ( #35536 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-15 13:17:23 -07:00
Flora Feng and GitHub
cd9078fe59
[Frontend] Skip structural tags for auto tool_choice without strict mode ( #45600 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 19:55:31 +00:00
Wentao Ye and GitHub
e18fe932ca
[Perf] Optimize DSv4 prefill chunk planning, 4.0% E2E Throughput Improvement ( #45061 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-15 19:50:21 +00:00
51ec5cf08f
[Bugfix] Chat Completions Harmony Refactor Clean up ( #45464 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-15 14:45:19 -04:00
7e612a0f06
[KV Offloading] Implement reset_cache for TieringOffloadingManager ( #44541 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 18:42:53 +00:00
+1
0a1c5034f5
[Model] Add MiniMax M3 support ( #45381 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-16 01:01:25 +08:00
RoyWang and GitHub
a3195fab7b
[AMD][Bugfix][Quantization] Honor fused-name match in is_layer_skipped ( #43981 )
2026-06-15 09:37:52 -07:00
Flora Feng and GitHub
0d80979644
[Chore] Consolidate reasoning/tool parser attributes into unified Parser in chat serving ( #45548 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 11:16:45 -04:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
fa63bb9db6
Remove redundant Triton KV cache dtype asserts and enforce architectural support (fp8 >= sm89) ( #43914 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-15 06:49:57 -07:00
5ed15f42b9
Fix the E8M0 scale computation in the MXFP4 (W4A4) MOE CUTLASS kernel ( #43557 )
...
Signed-off-by: Xin He <xin3.he@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-15 06:04:54 -07:00
Juan Pérez de Algaba and GitHub
b997071ec4
(security) Enforce audio upload size limit before full file materialization ( #45510 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-15 10:25:24 +00:00
Martin Kukla and GitHub
6c5872efc5
[Bugfix] Unset HF's default max_new_tokens for DiffusionGemma ( #45417 )
...
Signed-off-by: Martin Kukla <martin.kukla@cantab.net >
2026-06-15 17:31:57 +08:00
wang.yuqi and GitHub
1d88c4dadd
[Docs] Update the online serving docs. ( #45676 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-15 17:23:36 +08:00
vllmellm and GitHub
25c53d1293
[ROCm][Doc] Add installation notes about python version requirement ( #45671 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-15 17:22:55 +08:00
Yejing Lai and GitHub
9872921c5f
[XPU] skip UT test_with_ngram_gpu_spec_decoding ( #44423 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-06-15 08:46:30 +00:00
Reid and GitHub
c17e2f7c84
[Bugfix][Rust Frontend] Make metrics respect --served-model-name ( #45465 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-15 08:05:10 +00:00
FAUST and GitHub
40eac9a9d9
[Rust Frontend] Support parallel_tool_calls = false ( #44760 )
...
Signed-off-by: zhoujinyu <2319109590@qq.com >
2026-06-15 07:50:48 +00:00
b5adb027ad
[Models] Fix MiMo v2.x QKV TP sharding + FP4 support ( #45200 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 15:13:34 +08:00
Sahil Singh and GitHub
64833f8158
[Rust Frontend] Add external→internal request-id map for abort() ( #45137 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-15 06:51:24 +00:00
ddad5dbda2
[Bugfix][Rust] Sync EngineCoreReadyResponse with the Python dataclass ( #45557 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-15 06:49:42 +00:00
Peter Pan and GitHub
ebb0a71ad0
[Bugfix] Reject out-of-range temperature values in SamplingParams ( #44965 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
2026-06-14 23:12:44 -07:00
Ting SUN and GitHub
48df95c43e
[Feature][Frontend] Report multimodal token counts in usage.prompt_tokens_details ( #45458 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-15 05:20:58 +00:00
7df4fe1bd7
[Model] Remove XverseForCausalLM ( #45638 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-14 22:09:00 -07:00
b8336c3c7c
[Bugfix][V1] Split V2 model-runner attention groups on num_heads_q ( #45564 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-14 21:49:46 -07:00
e8d3e22c88
Fix included router missing path for FastAPI >=0.137 ( #45629 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-15 04:28:52 +00:00
c4a3f9d137
[Frontend] Add Streaming Parser Engine and new Qwen3 Parser ( #45413 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 11:59:05 +08:00
Flora Feng and GitHub
e3e3cd5458
[Bugfix][CI] Update Dockerfile dependency graph PNG ( #45602 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 10:35:24 +08:00
Li, Jiang and GitHub
8760f972ca
[CPU] Refine CPU attention frontend ( #45391 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-14 19:26:54 -07:00
b675cb7d0f
[Bugfix][CPU] Honor cgroup memory limit when computing KV cache size ( #45086 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-14 19:26:50 -07:00
Chaojun Zhang and GitHub
2725c84aae
[XPU] Enable sequence parallel support for XPU ( #38608 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-06-14 19:26:46 -07:00
Noa Neria and GitHub
1801fad0ba
[Bugfix] Stream Llama4 weight loading to avoid host-OOM with copy-returning loaders ( #44645 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-06-14 19:23:44 -07:00
Ting SUN and GitHub
3d6ce816f0
[Bugfix][Model] Validate runai_streamer model_loader_extra_config ( #45291 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-14 19:23:30 -07:00
Taneem Ibrahim and GitHub
2c764c089a
Added real /v1/embeddings support for messages + chat_template_kw ( #45173 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-15 09:08:10 +08:00
Michael Ma and GitHub
c621af1690
[BugFix] Fix prompt_embeds for multimodal models ( #45383 )
...
Signed-off-by: ruinan ma <r7ma3088@gmail.com >
2026-06-14 01:44:56 -07:00
Roger Wang and GitHub
e2bf2b3d84
[Perf] Use bisect for mm feature lookup in model runner v2 ( #45566 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-14 00:22:53 -07:00
Amanzhol Salykov and GitHub
725c3bc808
[ROCm][Perf] Enable W4A16 FlyDSL MoE ( #44400 )
...
Signed-off-by: amd-asalykov <asalykov@amd.com >
Signed-off-by: Amanzhol Salykov <asalykov@amd.com >
2026-06-14 00:14:39 -07:00
9548a1887f
[XPU] Support int4 group_size=32 W4A16 MoE ( #45136 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-14 00:14:35 -07:00
Jeff (Junze) Ma and GitHub
9fd737badc
[Bugfix][DCP] Fix illegal memory access in DCP a2a decode under full CUDA graphs ( #45487 )
2026-06-14 00:14:31 -07:00
4ef4492e9b
[V1][Spec Decode] Add Dynamic SD ( #32374 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-14 00:14:27 -07:00
78e7293bb1
[Build] Fix CUDA arch build coverage gaps ( #45277 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Xin Li <xinli-sw@users.noreply.github.com >
Co-authored-by: ShawRong <ShawRong@users.noreply.github.com >
Co-authored-by: Change72 <Change72@users.noreply.github.com >
2026-06-13 22:09:20 -07:00
54bbf51668
[Bugfix] nightly Docker images crash with ImportError: AnthropicOutputConfig since May 28 ( #44795 )
...
Signed-off-by: achyuthan.s <113010327+Achyuthan-S@users.noreply.github.com >
Signed-off-by: Achyuthan S <achyuthan.sivasankar@gmail.com >
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-13 21:45:29 -07:00
Nick Hill and GitHub
cf027b86af
[Core] Simplify MRV2 async output handling ( #45442 )
2026-06-13 18:15:36 -07:00
71b961dd35
[Perf] SM90 cutlass fp8 mm supports odd M by swap_ab, 180~290% kernel performance improvement ( #44572 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 12:05:45 -07:00
521b88c29e
[Bugfix] Reject structured outputs for diffusion decoders with a clear error ( #45468 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-13 12:04:01 -07:00
Harry Mellor and GitHub
b3f0a0a0df
Fix docs build on main ( #45536 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 08:53:23 -07:00
Juan Pérez de Algaba and GitHub
470229c37e
[Security] Fix DoS via prompt_embeds on M-RoPE models ( #45252 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-13 10:17:38 +00:00
2b3006076c
[Security] Add timeout guard for regex compilation in structured outp… ( #45118 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 09:52:56 +00:00
Wentao Ye and GitHub
96fa5cdd9e
[CI Bug] Fix ValueError: There is no module or parameter named 'model.vision_tower.vision_model' ( #45478 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 02:38:37 -07:00
Andreas Karatzas and GitHub
9261dbbc55
Treat null completion max_tokens like the default ( #45491 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 09:34:09 +00:00
Wentao Ye and GitHub
2ecf7d0eb4
[Model Runner V2] Fix openai.InternalServerError: Error code: 500 - 'list index out of range' ( #45467 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 01:44:16 -07:00
midas and GitHub
0d29612292
[Doc] Fix uv dependency resolution failure for setuptools during CPU source builds (x86 & ARM) ( #45412 )
...
Signed-off-by: midas <the.anon.github@gmail.com >
2026-06-13 06:18:58 +00:00
WEI CHENG CHIU and GitHub
5b2943f5a6
[Bugfix] Return the tokenizer from maybe_make_thread_pool so it survives pickling ( #45460 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 06:01:35 +00:00
43f0e024bc
[Render] Add /derender endpoints for disaggregated postprocessing ( #43606 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-13 13:55:33 +08:00
Andreas Karatzas and GitHub
1033ffac2e
[CI] Wait for SSL cert refresher events in the test ( #45489 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 04:57:18 +00:00
ff5a30cfac
[Bugfix] Replace deprecated Qwen2VLImageProcessorFast with Qwen2VLImageProcessor ( #42700 )
...
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-12 21:04:31 -07:00
WEI CHENG CHIU and GitHub
17ee5b1ac5
[Bugfix] Set type/role explicitly in streaming message_start event ( #45376 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 01:40:50 +00:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
Kevin H. Luu and GitHub
e3e31e54b0
[Bugfix][CPU] Don't build triton-cpu on arm64 release image ( #45401 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-06-12 14:51:45 -07:00
badddd254f
[ROCm][DSV4][Perf] Fuse inverse-RoPE and cache bf16 wo_a in o-projection ( #45103 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 15:57:09 -05:00
c90650088d
Add the QuantizedActivation linear-kernel contract ( #44260 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-12 13:48:15 -07:00
Michael Goin and GitHub
9eaacb23ec
[Kernel] Consolidate Marlin thread-tile padding across all dense Marlin paths ( #45295 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-12 13:46:21 -07:00
78739c1946
[Model Runner v2] Migration from v1 to v2, with Qwen and DSv2 MOE models [3/N] ( #42667 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 20:44:52 +00:00
Matthew Bonanni and GitHub
cf567cbc71
[Attention] Improve attention benchmarks: configs and profiling ( #39336 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-12 16:24:25 -04:00
Micah Williamson and GitHub
39cb9bf292
[ROCm] Bump Torch to 2.11 ( #45362 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-12 15:22:26 -05:00
Flora Feng and GitHub
6e4a547176
[Refactor] Deprecate ResponsesParser wrapper, inline parsing into ParsableContext ( #45431 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-12 16:15:41 -04:00
aab639c705
[Core][AMD] Propagate shutdown timeout to MultiprocExecutor ( #43154 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-12 15:13:31 -05:00
efe7adb5e1
[Perf] Use native DSA indexer decode path for next_n > 2 on SM100 ( #45322 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-12 12:54:00 -07:00
Isotr0py and GitHub
6635279d8a
[Migration] Migrate GGUF quantization support to plugin ( #39612 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 12:02:21 -07:00
Jonas I. Liechti and GitHub
d6fd7ce8da
[Model][Dflash] Enable Dflash support for Qwen3NextForCausalLM targets ( #45319 )
...
Signed-off-by: Jonas I. Liechti <j-i-l@t4d.ch >
2026-06-12 10:30:09 -07:00
272c16953e
[Kernel][Helion][1/N] Add Helion kernel for dynamic_per_token_scaled_fp8_quant ( #33790 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-12 12:50:06 -04:00
Yi Zhong and GitHub
053e7daa79
[Model] Add encoder CUDA graph support to Lfm2VL ( #44930 )
...
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com >
2026-06-12 09:17:26 -07:00
5af4aec141
[Rust Frontend] Add standalone granite4 tool parser ( #45216 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 00:16:36 +08:00
Sai Sridhar Tarra and GitHub
a30addc754
[Docs][KV Connector][NIXL] document KV Transfer stat logging and Prometheus metrics ( #44055 )
...
Signed-off-by: Sai Sridhar <tarrasridhar1154@gmail.com >
2026-06-12 15:39:11 +00:00
Chauncey and GitHub
3b8fc3fe6d
[Frontend] Support strict mode for tool calling with ResponsesAPI ( #45396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-12 10:59:59 -04:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
Guan-Ming (Wesley) Chiu and GitHub
c7aa3d2630
[Core] Support structured outputs for beam search ( #35022 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <guanmingchiu@gmail.com >
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-06-12 06:56:25 -07:00
fbc3a1907a
[Bug] Migrate Reset cache for both v2 and v1 model runner ( #42759 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 09:38:12 -04:00
4171ae406c
[V1][Metrics] Add MLA attention metrics for DeepSeek MFU estimation ( #39457 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-06-12 14:28:40 +01:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
8af550b399
[BUGFIX][XPU] Update fa interface for compatibility ( #45394 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-12 11:45:01 +00:00
f1e13f7df9
[Model] Remove Mono-InternVL (InternLM2VEForCausalLM) ( #45129 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 10:41:09 +00:00
88ed636218
[KV Connector]: Support KV push from Prefill to Decode node using Nixl KV Connector ( #35264 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-06-12 10:38:41 +00:00
a014dddbaa
[11b/n] Migrate Machete kernels to torch stable ABI ( #45304 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-12 10:36:49 +00:00
Thomas Parnell and GitHub
a37b4a940e
[Doc] AGENTS.md: add section about coding style ( #45301 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-12 06:23:04 -04:00
Juan Pérez de Algaba and GitHub
f715f25f29
Fix misleading error for audio duration limit rejection ( #45113 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-12 09:58:08 +00:00
Fynn Schmitt-Ulms and GitHub
462ef83d58
Update hidden states extraction integration test triggers ( #45294 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-06-12 01:05:19 -07:00
1ae1051b4b
[Bugfix][Rust Frontend] Return 400 for prompt-validation submit errors ( #45286 )
...
Signed-off-by: xiaguan <751080330@qq.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 07:53:11 +00:00
2043258dec
[Frontend] Support strict mode for tool calling ( #45003 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: cjackal <44624812+cjackal@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 07:51:48 +00:00
bd59c913bc
[CI] ci-fetch-log.sh: fetch all failed jobs from a build URL or PR number ( #45274 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 00:42:18 -07:00
04cec9e4d8
[XPU][DeepSeek-V4] Fix MTP: sync with upstream fixes #44821 and #43746 ( #45240 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 15:41:36 +08:00
Will Eaton and GitHub
87b98d6d6c
[Rust Frontend][Bugfix] Forward --shutdown-timeout and --disable-log-stats to the managed Python engine ( #45300 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-12 07:39:27 +00:00
Yuwen Zhou and GitHub
0cd9b7af25
[CPU] Support CPU W4A16 INT4 MoE ( #43409 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-12 07:12:37 +00:00
Isotr0py and GitHub
a2c72d4388
[Bugfix] Fix Dockerfile dependency graph pre-commit error ( #45374 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 07:10:18 +00:00
Rohan Potdar and GitHub
fe04238292
[ROCm][gpt-oss] Pass GateMode.INTERLEAVE for MXFP4 W4A16 fused MoE ( #44893 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-12 01:02:04 -05:00
39dee1114a
[MM][Perf][CG] Support ViT full cudagraphs for mllama4 ( #40660 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-11 22:17:55 -07:00
+1
eb28452b10
[Model] Add DiffusionGemma Support ( #45163 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Martin Kukla <martin.kukla@cantab.net >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Dipika Sikka <dsikka@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
Co-authored-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Alec Kohlhoff <134344302+aleckohlhoff@users.noreply.github.com >
Co-authored-by: Porras Huang <20535584+porrashuang@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: scoootscooob <167050519+scoootscooob@users.noreply.github.com >
2026-06-11 22:17:35 -07:00
Divakar Verma and GitHub
1ce3cdc5c1
[ROCm][CI] fix fp8 support for test_deepep_moe ( #45302 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-12 00:16:14 -05:00
Dao007forever and GitHub
6fbfdd1831
[NIXL] Per-region KV transfer classification for mixed full-attn + MLA groups ( #44583 )
2026-06-11 21:42:41 -07:00
Chris Leonard and GitHub
7021be66e8
[11a/n] Migrate Marlin kernels to torch stable ABI ( #45176 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-06-11 21:22:37 -07:00
Ekagra Ranjan and GitHub
226ba9fc9e
[ASR] Add Long Audio benchmark and correctness test ( #44587 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-06-12 04:11:16 +00:00
b927004c44
[Bugfix] Mamba CPU Offloading ( #44599 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-11 21:07:35 -07:00
e0b9fb1290
[ASR] Optimize CPU preproc to get 2.5x RTFx via multi-threading ( #44612 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 21:05:11 -07:00
42ae5e7ac6
[Bugfix] Fix --enable-prompt-tokens-details omitting zero cached tokens ( #44383 )
...
Signed-off-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-11 20:37:42 -07:00
Nick Hill and GitHub
2263f8a3de
[CI][BugFix] Fix broken test_mamba_prefix_cache.py due to stale mock ( #45345 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 03:26:17 +00:00
Ting SUN and GitHub
c1076839c9
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads ( #45308 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-11 20:21:46 -07:00
fcf5115c45
[ROCm][DSv4][Perf] Flash-decode split-K decode attention kernel ( #44899 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-12 03:17:52 +00:00
4bc83323f2
[Bugfix] OffloadingConnector: respect skip_reading_prefix_cache flag ( #44592 )
...
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: littlecircle0730 <43994952+littlecircle0730@users.noreply.github.com >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-06-12 02:20:39 +00:00
yzong-rh and GitHub
e0871ad225
[Refactor] Chat Completions Streaming Harmony Refactor and Bugfixes ( #45104 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-12 01:09:47 +00:00
6f573f486b
[Bugfix] Initialize missing attributes in mistral eagle ( #45217 )
...
Signed-off-by: jpwang <jpwang@smail.nju.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 08:21:01 +08:00
Neil Schemenauer and GitHub
9bbf42be26
Make mistral_common optional by deferring MistralToolCall import ( #45305 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
2026-06-11 22:59:11 +00:00
8a91228dbe
[Bugfix][KVConnector][Mooncake] Close MooncakeDistributedStore on connector teardown ( #45206 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 14:33:48 -07:00
yzong-rh and GitHub
f712fd0d7d
[Refactor] Chat Completions Harmony Refactor, non-streaming path. ( #45171 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-11 21:18:30 +00:00
Wentao Ye and GitHub
5a6c7b7ab5
[Bug] Fix test flashmla for DSv4 ( #45052 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 16:22:26 -04:00
c9340e6f35
[Model] Remove InternLMForCausalLM registry alias ( #45128 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 20:02:51 +00:00
Ben Browning and GitHub
235b63c004
[Bugfix] Fix Anthropic tool_use content handling dropping args ( #45287 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-11 20:01:29 +00:00
3b03a2cf47
[Rust Frontend] Support continuous_usage_stats stream option ( #43965 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 17:50:59 +00:00
wentian-byte and GitHub
b8142294b7
[Bugfix] Restrict FlashInfer cuDNN FP8 ViT attention gate to Blackwell (SM 100) ( #45251 )
...
Signed-off-by: Wentian Byte <3400259131@qq.com >
2026-06-11 16:39:24 +00:00
2ec6594db9
[Kernel][Helion][1/N] Add Helion kernel for per_token_group_fp8_quant ( #36902 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 08:59:08 -07:00
vraiti and GitHub
79f8c5bd8c
[Metrics] Scope unregister_vllm_metrics() to strictly "vllm:" metrics ( #42331 )
...
`unregister_vllm_metrics()` currently uses "vllm" in `collector._name` to decide
which collectors to remove from the Prometheus registry, removing every even
metrics registered by other subsystems or downstream extensions like "vllm_omni:"
Signed-off-by: vraiti <vraiti@redhat.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-11 15:43:14 +00:00
Jiangyun Zhu and GitHub
f81daf8880
[Attention] add triton diff-kv backend for mimo ( #41797 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-11 11:36:31 -04:00
4085ff7cb4
[Core] Add kvcache watermark to reduce preemptions ( #44594 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-11 08:27:31 -07:00
23eb7c8fbb
[Bugfix] Fix NixlEPAll2AllManager's dependency on --enable-elastic-ep to function ( #44422 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-11 08:14:49 -07:00
wineandchord and GitHub
c2b4cd39ac
[Doc][Attention] Fix MLA top-of-file comments ( #37047 )
...
Signed-off-by: wineandchord <guoqizhou19@gmail.com >
2026-06-11 08:14:45 -07:00
Kai K. and GitHub
f1d8d99717
[Bugfix] CohereModel.load_weights: skip modelopt _quantizer.* keys ( #43495 )
...
Signed-off-by: Kai Köhler <kai.koehler@web.de >
2026-06-11 08:14:21 -07:00
Nicolò Lucchesi and GitHub
750aab5b8e
[Bugfix] Fix CPU memory leak related to not cleaning up old remotes data ( #44424 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-11 07:54:52 -07:00
5edf7ff489
[Core] Release cached device memory under pressure on UMA GPUs during weight loading ( #45179 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 17:49:50 +03:00
b78fc47f05
[Docs] Add redirect for moved lmcache examples page ( #45218 )
...
Signed-off-by: nataliepjlin <nataliepjlin@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:41:08 -04:00
Harry Mellor and GitHub
03878d1c22
Deprecations for v0.23 and v0.24 ( #44992 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 14:35:38 +00:00
55911db580
[PD][Core] Fix Mamba prefix cache hit rate in PD disaggregation ( #44243 )
...
Co-authored-by: lHrHenry233 <2381623149@qq.com >
Co-authored-by: underfituu <hzhucong@163.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-06-11 14:10:25 +00:00
cc640ee8bc
[Rust Frontend][Metrics] Export vllm:lora_requests_info from frontend ( #45030 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 06:45:03 -07:00
ebc6ef971a
Hidden states extraction improvements ( #43805 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 09:44:45 -04:00
tc-mb and GitHub
ab3a1fd2e6
minicpmv4_6: fix ImageSize (W,H) order for placeholder token calculation ( #45244 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-11 13:43:56 +00:00
c3662b36ea
[KV offload] Parallel-agnostic fs-tier cache for single full-attention group ( #44733 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-06-11 15:48:37 +03:00
Juan Pérez de Algaba and GitHub
e62d00ab73
docs: add fix disclosure policy to SECURITY.md ( #45253 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 12:48:00 +00:00
1f60771c74
fix: guard flash-attn rotary import ( #42679 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-11 08:43:31 -04:00
05d9848267
[Build] Upgrade CUDA Dockerfiles from GCC 10 to GCC 12 for C++20 compatibility ( #44923 )
...
Signed-off-by: Richard Barnes <rbarnes@meta.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:26:52 +00:00
jasen and GitHub
ef67071b21
[Build] Skip spinloop extension on Python < 3.11 ( #44783 )
...
Signed-off-by: Jasen2201 <yajizhan@amd.com >
2026-06-11 11:23:21 +00:00
x41lakazam and GitHub
3508cb78d4
[Bugfix] Fix broken profile_modular_kernel.py ( #43300 )
2026-06-11 12:17:23 +01:00
Harry Mellor and GitHub
432905d5d6
Only enable PR docs builds manually ( #45262 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 03:14:29 -07:00
1f9dd7900d
[Bugfix][Rust Frontend] Validate out-of-vocab token ids in request params ( #44680 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 03:14:11 -07:00
9492362972
[Security] Apply sanitize_message to Anthropic and STT error paths ( #45119 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 10:05:34 +00:00
7852e50e4d
[docs] Document --scheduler-cls base class requirement (extend AsyncScheduler, not Scheduler) ( #43724 )
...
Signed-off-by: Georgii Kliukovkin <kliukovkin@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:49:51 +01:00
Reid and GitHub
0d657e44dc
[Rust Frontend] Fix DeepSeek V3.2 continue_final_message rendering ( #45155 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-11 09:34:19 +00:00
aa1df36c53
Fix/minicpmv46 missing version ( #44980 )
...
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-11 09:20:45 +00:00
f06aefb4e3
[CPU] Add missing scalar fallback for CPU W4A8 INT4 GEMM ( #44523 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
2026-06-11 08:52:01 +00:00
Julien Denize and GitHub
1c3a72b8b2
[Bugfix] Add fetch_images to MistralCommonImageProcessor ( #45180 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-11 16:13:01 +08:00
Juan Pérez de Algaba and GitHub
d598d23973
[Security] Reject non-finite temperature and repetition_penalty values ( #45116 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 01:12:14 -07:00
Juan Pérez de Algaba and GitHub
f219788f91
[Security] Fix info disclosure via int32 truncation in GGUF dequantize kernels ( #44971 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 08:05:14 +00:00
6e64c1bab1
[10c/n] Migrate MoE kernels to torch stable ABI ( #44565 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-10 23:02:26 -07:00
Kevin H. Luu and GitHub
2f2c5cf4f1
[release] Always block release images to dockerhub ( #45236 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-10 22:53:04 -07:00
Mohammad Miadh Angkad and GitHub
40e065e86a
[Docker] Fix CUTLASS DSL cu13 install order in Dockerfile ( #45204 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-11 05:19:36 +00:00
0b995f8609
Use std::bit_cast for type punning in CPU kernels ( #45089 )
...
Signed-off-by: Yuanyuan Chen <cyyever@outlook.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-10 22:07:44 -07:00
Bugen Zhao and GitHub
43914dd743
[Rust Frontend] Add Python bridge for Rust tool parsers ( #44624 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 04:51:06 +00:00
3501324957
[Build] fix self-contradictory precompiled-flag orthogonality test ( #44942 )
...
Signed-off-by: pjdurden <prajjwalchittori1@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:49:08 +08:00
Flora Feng and GitHub
3a04061701
[Refactor][Parser] Unify Response API to use parser.parse() like Chat Completion API ( #45190 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-11 04:37:51 +00:00
Yifan Qiao and GitHub
f272dfdce1
[KV Connector] Mooncake store: prefix-cache retention interval for sparse attention ( #44774 )
2026-06-10 21:36:34 -07:00
velonica0 and GitHub
f31bc2ea60
[CPU][RISC-V] Enable oneDNN W8A8 INT8 to run on RISC-V ( #44478 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-06-11 04:09:05 +00:00
248e33c40d
[Bugfix][Responses API] Set id on function_call item in streaming done event ( #44608 )
...
Signed-off-by: Aniruddh Krovvidi <aniruddh.krovvidi@oracle.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-11 03:52:42 +00:00
Bugen Zhao and GitHub
5d5591d99b
[Rust Frontend] Populate cached_token_count in responses ( #44887 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 20:50:05 -07:00
Wentao Ye and GitHub
85a0ffae42
[CI Bug] Remove qwen test ValueError: No example model defined for Qwen/Qwen-7B-Chat ( #45194 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-10 20:11:00 -07:00
Harry Mellor and GitHub
18d87a87dc
Deprecate Transformers v4 support ( #45161 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 11:04:01 +08:00
Flora Feng and GitHub
b038a2f73b
[CI][Bugfix] Update Dockerfile dependency graph PNG ( #45209 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 19:40:25 -07:00
Ting SUN and GitHub
2d481f8a94
[Bugfix][Rust Frontend] Stop unescaping XML-style tool-call parameter values ( #45025 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-10 19:05:23 -07:00
7920ccb97c
[Bugfix]: Fix Quark gpt-oss weight loading broken by FusedMoe refactor ( #45067 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 18:17:46 -07:00
Wentao Ye and GitHub
86111c00c7
[Chore] Add Github notification for MRv2 for @yewentao256 ( #45191 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 09:01:49 +08:00
qizixi and GitHub
e2db0222e9
[Perf][Attention] Pin MLA chunked-context metadata tensors so H2D copies are truly non-blocking ( #45074 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-10 15:56:49 -07:00
82d6b59f04
[CI/Build] Skip test_use_trtllm_attention on non-CUDA platforms ( #44687 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 18:18:42 -04:00
Andreas Karatzas and GitHub
16282a9c4e
[ROCm][CI] Moving MI300 tests to MI325 until cluster is stabilized ( #45170 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 20:26:17 +00:00
5b6b536fdc
[ROCm][Bugfix] Make intermediate_pad TP-aware in rocm_aiter_fused_experts ( #44679 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 15:10:50 -05:00
12f3f19c19
feat(qwen3-asr): support prompt parameter in v1/audio/transcriptions ( #35415 )
...
Signed-off-by: Nathan Price <nathan@abridge.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 19:54:59 +00:00
Ilya Markov and GitHub
6471ec75bd
[EPLB] Reject NCCL-based EPLB communicators with async EPLB ( #44978 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
2026-06-10 19:51:27 +00:00
3d300aecb1
[Doc] Switch K8S examples to default MP mode ( #39400 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
Signed-off-by: Peter Pan <peter.pan@daocloud.io >
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-10 18:17:11 +00:00
ffce72c041
[Model Runner V2] Fix v2 AttributeError: 'CohereASRDecoder' object has no attribute 'embed_input_ids' ( #44568 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 11:06:01 -07:00
TJian and GitHub
bfe1001ab6
[Bugfix] [DSV4] [ROCm] Pin apache-tvm-ffi version to 0.1.10 ( #45169 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-10 17:41:15 +00:00
fa8c868a3c
[Bugfix] Fix Llama4 weight loading ( #45047 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-10 13:40:45 -04:00
Ben Browning and GitHub
d1bcb4b44c
[Bugfix] Fix tool parsing crash with non-function tool types (e.g. WebSearchTool) ( #45147 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-10 17:17:16 +00:00
bnellnm and GitHub
29026682cb
[Bugfix] Fix nemotron accuracy drop introduced by #41184 ( #45037 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 13:16:25 -04:00
Stan Wozniak and GitHub
dc66e01a70
[Hybrid] Marconi-style admission policy for hybrid cache ( #37898 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-10 10:03:13 -07:00
Yongye Zhu and GitHub
2ba68d9bf7
[Test] Fix one-sided MNNVL alltoall test workspace under-reservation ( #44946 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-11 00:43:12 +08:00
Julien Denize and GitHub
2131b597b1
[CI] Ping Mistral team for ministral/voxtral/mixtral/pixtral changes ( #45153 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-10 08:48:00 -07:00
0bae1d3848
[MRV2][Spec Decode] DFlash ( #44586 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-10 08:47:46 -07:00
Yufeng He and GitHub
4673ca1d78
fix: prefix DeepSeek V4 MTP projections ( #44821 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
2026-06-10 08:47:04 -07:00
Angela Yi and GitHub
de900fa7e5
fix: AOT compile cache collision for dataclass-based HF configs ( #45059 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-06-10 08:05:29 -07:00
Divakar Verma and GitHub
166d14e9bf
[bugfix] skip conch kernel for g_idx reordering ( #45072 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-10 23:04:19 +08:00
af65e08fc5
KV-Cache multi-tier offloading async batched lookup ( #44193 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 14:59:30 +00:00
Harry Mellor and GitHub
3cc9fecd58
Deprecated 1st generation Qwen and QwenVL models ( #45131 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 14:55:33 +00:00
ccc05de038
[Bugfix] Fix missing sequence_lengths in EXAONE-4.5 vision encoder ( #45073 )
...
Signed-off-by: Jongsu Liam Kim <jongsukim8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 15:44:34 +01:00
6ec7dcd641
[Frontend][Metrics] Add vllm:tool_call_parser_invocations_total Prometheus metric ( #44448 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 10:29:11 -04:00
c9e5bf8135
[Bugfix] Fix layerwise reload dropping params after a composed weight loader ( #44814 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-06-10 06:42:05 -07:00
Roberto L. Castro and GitHub
6850839c6f
[Perf] Fix dsv3_router_gemm heuristic ( #44217 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-10 06:08:41 -07:00
87c15d46e3
[Bugfix] Lazily import the humming quantization backend ( #44921 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 06:06:17 -07:00
4882fd7632
[Bugfix][Reasoning] Nemotron V3: surface reasoning as content when thinking is unterminated ( #39091 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-06-10 05:58:19 -07:00
77f42d9725
[Model] Remove obsolete ERNIE models ( #45127 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 20:54:30 +08:00
9dfc313bdc
Feature/offloading manager stats ( #35669 )
...
Signed-off-by: Sriusa4414@gmail.com
Signed-off-by: srinivas_oo7 <Sriusa4414@gmail.com >
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 12:44:55 +00:00
9ad08c4d15
[Bugfix][Rust Frontend] Fix missing added tokens in hf/fastokens tokenizer ( #44683 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 03:52:41 -07:00
Shantipriya Parida and GitHub
a1ec011a83
[Bugfix] Add deepseek_v32 to Quark dynamic MXFP4 model type check ( #39498 )
...
Signed-off-by: Shantipriya Parida <shantipriya.parida@amd.com >
2026-06-10 02:52:33 -07:00
Bugen Zhao and GitHub
fdfb2566c0
[Rust Frontend] [CI] Unify Rust artifact builds with setuptools-rust ( #44981 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 17:48:34 +08:00
Juan Pérez de Algaba and GitHub
8a5cf1ccd6
[Security] Fix remote DoS via invalid recovered token reinjection ( #44744 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 02:31:43 -07:00
Kunshang Ji and GitHub
fe1d923afc
[BUGFIX][XPU] fix xpu flash_attn_varlen_func interface ( #45110 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-10 17:07:40 +08:00
32daf56b42
[Refactor] Rename rocm_moe.py to rocm_moe_rdna.py ( #45011 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 17:02:09 +08:00
Andreas Karatzas and GitHub
82a42234be
[ROCm][CI] Defer AITER sampler import and isolate server test PYTHONPATH ( #44823 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 08:56:11 +00:00
Harry Mellor and GitHub
af9f583344
Revert "[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures" ( #45029 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 01:37:03 -07:00
yiheng and GitHub
bd2d83ff31
[SpecDecode] Reduce TP communication for large-vocab draft models speculative decoding ( #39419 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
2026-06-10 07:59:24 +00:00
xiaohuguo2023 and GitHub
bb78168b21
[ROCm][gpt-oss] Hybrid CDNA4 swizzle gate for A8W4 MoE ( #44804 )
...
Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com >
2026-06-09 23:59:44 -07:00
89c6a41001
[Bench] Add BFCL dataset for vllm bench serve tool-calling workloads ( #42457 )
...
Signed-off-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-09 23:59:18 -07:00
7fdfa6441d
Model/colbert autoweightsloader ( #44999 )
...
Signed-off-by: Furkan Fidan <dev@yufufi.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 23:58:50 -07:00
Harry Mellor and GitHub
e9b728de8a
Change from owning configs to owning config utils ( #45058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 06:40:25 +00:00
5828a205ef
Fix Harmony tool descriptions for optional fields ( #44686 )
...
Signed-off-by: Varun Shenoy <varun.vinayak.shenoy@oracle.com >
Co-authored-by: Codex <codex@openai.com >
2026-06-09 23:29:22 -07:00
Yaoming Zhan and GitHub
7a74f31d2e
[Rust Frontend] Add seed_oss and step3p5 reasoning parsers ( #44552 )
...
Signed-off-by: yzhan1 <zhanyaoming2014@gmail.com >
2026-06-09 23:01:33 -07:00
47930b59ca
[Bugfix] Handle HWC images in ImageProcessorItems.get_image_size ( #45057 )
...
Signed-off-by: YellowFoxH4XOR <yellowfoxh4xor@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 05:35:50 +00:00
Flora Feng and GitHub
6aec99f030
[Refactor] Remove dead states from chat completion serving ( #45081 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-09 22:20:15 -07:00
bnellnm and GitHub
f4966f8b3d
[Bugfix] Fix weight loading issues caused by #41184 ( #45054 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 01:20:13 -04:00
Mohammad Miadh Angkad and GitHub
2c9c07c85e
[Bugfix][CI/Build] Fix Rust frontend build after chat conversion refactor ( #45085 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 20:04:41 -07:00
Change72 and GitHub
320c52b134
[Bench] benchmark_serving_multi_turn: make non-standard conversation_id payload opt-in ( #43756 )
...
Signed-off-by: Change72 <cguo51@asu.edu >
2026-06-09 19:41:56 -07:00
6deb05e0e4
[Core][Model] Gemma4: Unified FA4 for all layers + FlashAttention mm_prefix support ( #42175 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-09 17:45:39 -07:00
Flora Feng and GitHub
d82ac00923
[Refactor][Mistral] Extract parsing logic into MistralParser ( #44596 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 00:12:23 +00:00
Bugen Zhao and GitHub
dac9e9a640
[Rust Frontend] Extract shared options in route helper params ( #44884 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 17:02:35 -07:00
Wentao Ye and GitHub
d7607ad273
[Bug] Fix deepseek v4 OOM issue ( #44914 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-09 15:47:06 -07:00
d955745d58
[ROCm][CI] fix test_rope_kvcache_fusion.py ( #44678 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-09 21:53:46 +00:00
Micah Williamson and GitHub
e1ed89dbee
Revert "[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuf… ( #45066 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 14:12:06 -07:00
1c2ffc6f88
feat(multi-turn-bench): add api_key and custom headers for multi turn benchmark ( #44516 )
...
Signed-off-by: Jimmy <jinmingyi1998@sina.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: simon-mo <simon.mo@hey.com >
2026-06-09 14:00:07 -07:00
Jiangyun Zhu and GitHub
ca4cfd8731
[Bugfix] fix qwen3.5 ep weight loading ( #45002 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 13:55:30 -07:00
Micah Williamson and GitHub
c9c1540e61
[ROCm][V2] Fix failed assertion in Llama models when using EAGLE with ROCM_AITER_FA ( #44936 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 13:30:52 -05:00
c1d754d681
[Mooncake] Use all HCAs on multi-NIC hosts instead of GPU-indexed RNIC selection ( #43799 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-06-09 11:05:36 -07:00
01d8cd92dd
[ROCm][Perf] Use fused softplus-sqrt-topk router under AITER fused-MoE ( #44945 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-09 17:53:05 +00:00
a4b14b98c6
[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuffle reduction + vectorized I/O ( #44173 )
...
Signed-off-by: SII-yangdian <yangdian@sii.edu.cn >
Co-authored-by: SII-yangdian <yangdian@sii.edu.cn >
2026-06-09 10:41:26 -07:00
Juan Pérez de Algaba and GitHub
cf1c906724
[Security] Fix image EXIF orientation and tRNS transparency handling ( #44974 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-09 09:34:44 -07:00
766ce2bb6b
Fix MiDashengLM TP>1 crash in audio encoder attention ( #44408 )
...
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-09 09:29:14 -07:00
3d119f78f7
[Docs] Add KV offloading usage guide (single- and multi-tier) ( #44415 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 19:20:23 +03:00
Juan Pérez de Algaba and GitHub
1b1359c332
[Security] Fix DoS via audio decompression bomb in speech-to-text endpoint ( #44970 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 00:18:53 +08:00
Tyko Niemi and GitHub
cad4ca12b8
[Bugfix] Add X-Session-ID from conversation_id in multi-turn benchmark ( #44663 )
...
Signed-off-by: Tyko Niemi <tyko.niemi@amd.com >
2026-06-09 08:57:00 -07:00
Andreas Karatzas and GitHub
b697119800
[ROCm][CI] Stabilize ModernBERT token-classification parity against Hugging Face ( #44040 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:52:36 +01:00
Kunshang Ji and GitHub
b4c6dc6454
[WIP][XPU] upgrade torch-xpu to 2.12 ( #42262 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-09 15:51:39 +00:00
Raushan Turganbay and GitHub
2ee5106372
Remove raw_inputs from transformers backend ( #39425 )
...
Signed-off-by: raushan <raushan@huggingface.co >
2026-06-09 15:01:04 +00:00
Jiangyun Zhu and GitHub
7a89b72564
[Perf] fuse qk rmsnorm rope gate for qwen3.5 ( #44176 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 22:12:17 +08:00
Jee Jee Li and GitHub
dc10e467a9
[Bugfix] Fix minimax_qk_norm_fusion ( #44983 )
2026-06-09 06:43:46 -07:00
Terrence Zhao and GitHub
ee4d7df2b5
[Cohere] Cohere2 moe parser fix ( #44907 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:32:18 -07:00
Terrence Zhao and GitHub
3e8afdf785
[Cohere] Fix Cohere2MoE weight loading when using Transformers ≥5.10 ( #44747 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:27:40 -07:00
Nicolò Lucchesi and GitHub
6690a0c4de
[PD][Bugfix] Fix KV Cache sharing with HMA ( #44629 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 06:10:06 -07:00
Maria Guevara and GitHub
1c23c42030
[Rust Frontend] Support Kimi K2 tool call IDs ( #44901 )
2026-06-09 05:31:26 -07:00
xiangdong and GitHub
b12e42d132
[XPU][CI] Refine docker image build and pull/create lock mechanism in Intel GPU CI ( #44481 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-09 20:20:32 +08:00
69fdaffbcd
[Rust Frontend] Add /tokenize and /detokenize endpoints ( #44222 )
...
Signed-off-by: Tan Ngoc Do <darkknightkhtn2008@gmail.com >
Signed-off-by: TanNgocDo <darkknightkhtn2008@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 05:11:37 -07:00
80e2c4462d
[ROCm][Compile] Fuse AR + RMSNorm + per-group FP8 quant (+ DSv3.2 indexer fan-out) ( #42864 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
2026-06-09 12:06:56 +00:00
Sage and GitHub
5b3807e862
[KV Events] Switch event structs from array to map encoding ( #42892 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-06-09 11:39:52 +00:00
Qiuyang Yue and GitHub
59401ac9f1
[Kernel][Perf] Tune fused_moe FP8 config for Qwen3-Next-80B tp=4 on H100 (+25% at batch 96-512) ( #44830 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-09 04:15:51 -07:00
d841386d27
[Rust Frontend] Support API key authentication ( #44321 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 10:15:20 +00:00
Mohammad Miadh Angkad and GitHub
fff9210b2a
[CI/Docs] Remove stale disagg prefill links ( #44918 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 03:05:53 -07:00
Ma Jian and GitHub
70db1488c5
[DSV4][XPU] Add MHC fused_post_pre support ( #44144 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-09 17:23:17 +08:00
Andreas Karatzas and GitHub
2385e140d6
[ROCm][CI] Stabilize sleep-mode memory release ( #43022 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:51:12 +08:00
Nicolò Lucchesi and GitHub
dab60fc658
[Bugfix][CI] Fix test_offloading_connector.py::test_fs_tiering_offloading ( #44903 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 00:57:34 -07:00
wang.yuqi and GitHub
996222f4bf
[CI] Reorganize entrypoints CI ( #44947 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 00:46:11 -07:00
e6fc848d4f
[Bugfix][MiniCPM-o] Fix cuda/cpu device mismatch in Resampler2_5 pos_embed ( #43844 )
...
Signed-off-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 23:28:26 -07:00
Andreas Karatzas and GitHub
f843ac1a1c
[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures ( #44952 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 05:49:30 +00:00
7c2aa3108a
fix: prevent MM cache hang from stale LRU order keys ( #43595 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:48:31 -07:00
ebf53ba373
[Bugfix][Rust Frontend] Set a structured-output backend so requests do not 500 ( #44729 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 22:30:54 -07:00
baacbfcebf
[ROCm][MLA][Bugfix] Reserve FP8 prefill workspace before lock for Kimi-K2.5 ( #42978 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 22:25:52 -07:00
d8218b1ee7
[Bugfix] Propagate ImportError from load_audio_pyav when vllm[audio] … ( #44750 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
2026-06-09 04:24:52 +00:00
9f153aa781
[MM][Perf][CG] Support ViT full CUDA graph for glm4_1v image and video inference ( #40576 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-09 11:13:56 +08:00
Kunshang Ji and GitHub
d3de61502f
[XPU][CI] fix test case path ( #44940 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-08 20:02:31 -07:00
Lanze Liu and GitHub
540aaf2140
[Bugfix][Model] Qwen3-Omni: move cu_seqlens to GPU before VIT attention ( #44264 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-08 20:02:27 -07:00
Lanze Liu and GitHub
4128605ad4
[Docs] Remove broken link to deleted disaggregated_prefill.sh ( #44929 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-09 01:40:06 +00:00
e2f993dc41
[WideEP] Integrate DeepEP v2 ( #41183 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-08 18:07:29 -07:00
Andreas Karatzas and GitHub
05cb606cad
[ROCm][CI] Re-route NixlConnector jobs ( #44809 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 18:57:11 -05:00
3f627ebef7
[Misc] usage_stats: report more engine, spec-decode, and EP config ( #44595 )
...
Signed-off-by: Zach Xi <zachary.xi@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-08 15:20:00 -07:00
Bugen Zhao and GitHub
bc941f375d
[Rust Frontend] [Refactor] Refine utility call interfaces ( #44856 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 15:13:08 -07:00
Michael Goin and GitHub
6afa25000c
[Bugfix] Canonicalize FP8 weight layout to (K, N) at the source ( #44735 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-08 14:37:36 -06:00
Mohammad Miadh Angkad and GitHub
823a0ab754
[Bugfix][MoE] Fix fused MoE expert mapping helper call sites ( #44897 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-08 13:35:04 -07:00
Wentao Ye and GitHub
2c27c294c0
[Model Runner V2] Fix mrv2 mm lora issue ( #44450 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-08 14:30:09 -04:00
ba94a3b998
[Attention] Extract KV-cache update from CPU attention backend ( #40470 )
...
Signed-off-by: Diego Maniloff <diego.maniloff@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-08 15:43:05 +00:00
bnellnm and GitHub
dc68bd8c41
[MoE Refactor] FusedMoE/MoERunner inversion refactor ( #41184 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-08 10:42:58 -04:00
753e9d55e6
[Quantization] add online fp8 ptpc ( #44132 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:42:11 +08:00
akii96 and GitHub
ac3409d162
[Benchmark] Auto-detect and correct client/server tokenizer mismatch for random dataset ( #44708 )
2026-06-08 06:10:20 -07:00
wang.yuqi and GitHub
93ee4cd47f
[CI] Consolidate multimodal entrypoint tests. ( #44819 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-08 04:48:08 -07:00
Li, Jiang and GitHub
980796cd07
[CI/Build][CPU] Fix flaky CI image build failure and unexpected warnings ( #44852 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-08 11:10:06 +00:00
Nicolò Lucchesi and GitHub
5add018beb
[Connector] Remove P2pNcclConnector ( #44854 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-08 18:58:29 +08:00
d5fe994e79
[CPU][Spec Decode] Warn about throughput loss when libiomp5 is not preloaded ( #44419 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <bigpyj64@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-08 03:45:08 -07:00
Chaojun Zhang and GitHub
fa662b1a8b
[XPU] Cap topk/topp Triton BLOCK_SIZE to 4096 to fix Top-p mask difference failures ( #44470 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-08 09:36:51 +00:00
3c0b4432be
[Rust Frontend] Add /pause, /resume, /is_paused endpoints ( #44499 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 17:28:37 +08:00
Sungjae Lee and GitHub
469f3dcf1d
[BugFix] Use served model name in gemma4 audio-tower error message ( #44828 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
2026-06-08 06:58:31 +00:00
xiangdong and GitHub
94fcdd007f
[XPU][CI] Add more test cases in Intel GPU CI ( #43663 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-08 06:21:24 +00:00
Andreas Karatzas and GitHub
d9ff7e4e9a
[ROCm][CI] Stabilizing teardown and timeout of flaky tests to prevent rare OOMs ( #44761 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 14:11:17 +08:00
Andreas Karatzas and GitHub
967c5c3bc3
[ROCm][CI] Stage C mirrors ( #42793 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 23:00:59 -07:00
Yan Ma and GitHub
54c660c3a6
[XPU][Minor] format moe kernel name and add in kernel list ( #44771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-08 13:58:16 +08:00
Shanshan Shen and GitHub
8fb0274415
[MM][CG] Simplify ViT CUDA graph interfaces ( #44484 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-08 05:57:06 +00:00
Ma Jian and GitHub
eebce65756
[XPU]feat: add DeepSeek-V4 XPU attention decode path ( #42953 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-08 13:27:12 +08:00
303916e93d
[Bugfix]: Fix assertion in MambaManager.allocate_slots() ( #39562 )
...
Signed-off-by: Holworth <kangqihan17@mails.ucas.ac.cn >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-08 00:34:37 -04:00
Taneem Ibrahim and GitHub
5633405964
Added extra_repr() to pooler classes to improve debuggability ( #44805 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-08 03:19:31 +00:00
6124a98a9b
[Bugfix] Fix FunASR-Nano crash during initialization ( #44215 )
...
Signed-off-by: SunskyXH <sunskyxh@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-07 20:00:02 -07:00
Agata Dobrzyniewicz and GitHub
2ed0a9627b
[Kernel][Test] Make kernel tests for mamba dual-HW (CUDA + XPU) ( #42736 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-08 08:22:47 +08:00
4dcd10eb0d
[1/N][KV-Cache Layout Refactor] Refactor DSV4 KV cache config construction ( #44454 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-07 14:53:37 +00:00
Charlie Fu and GitHub
228bcc436b
[ROCm][Kernel] Enable permute_cols for ROCm ( #44674 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-07 09:50:03 +00:00
3d3ba460a2
Modify torch dependency in xpu.txt ( #43087 )
...
Signed-off-by: Bram Vanroy <2779410+BramVanroy@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-07 16:33:50 +08:00
Mohammad Miadh Angkad and GitHub
66ecfd0568
[Dependency] Remove stale cuDNN frontend upper bound ( #42599 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-07 16:09:25 +08:00
Andreas Karatzas and GitHub
f0f6805d8a
[CI] Stabilize the multi-audio OpenAI server path ( #44051 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 15:54:32 +08:00
15652a6b70
[Doc] Fix multimodal torch.compile troubleshooting to not use removed VLLM_TORCH_COMPILE_LEVEL ( #44378 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-07 00:34:07 -07:00
Yifan Qiao and GitHub
51ef688831
[Bugfix][Mooncake] Fix per-group block_size/block_hash and group_idx in MooncakeStoreConnector KV events ( #44103 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-07 07:12:43 +00:00
Jared Wen and GitHub
6ac69203e8
[videoloader] implement glm46v video loader ( #44417 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-07 06:27:20 +00:00
1505b3d8a1
[Cohere] Enable Cohere Mini Code model and update Command A-plus test registry ( #44707 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:44:16 -07:00
32f34d3935
[feature] add index share feature for DSA MTP ( #44420 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:04:14 -07:00
Qiuyang Yue and GitHub
9c7f7741d4
[Bugfix] Fix benchmark_moe.py after inplace mechanism removal ( #44041 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-07 00:32:00 -04:00
6181e80fe0
[XPU] add xpu branch in compressed_tensors_moe_w4a4_mxfp4 ( #44540 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Co-authored-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-07 12:27:34 +08:00
Yan Ma and GitHub
3bb46975bd
[XPU][Feature] transparent sleep mode support for XPU platform ( #37149 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-07 10:45:31 +08:00
Chaojun Zhang and GitHub
810966453a
[XPU] Support cpu kv offloading and tiering offloading on XPU platform ( #36423 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-07 09:59:28 +08:00
Woosuk Kwon and GitHub
2a983c79ac
[DSV4] Decouple DS V4 Sparse MLA Metadata from DS V3.2 ( #44699 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-06 20:37:56 -04:00
bc5745a00f
[ROCm][MLA] Replace torch.cat in sparse-MLA forward_mqa with fused concat_mla_q ( #42838 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 18:20:50 -05:00
Nick Hill and GitHub
3b3d5287fa
[BugFix] Resolve multiple async kv load deadlock ( #44560 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-06 23:05:47 +00:00
062b05ff3a
[ROCm][Perf] Fused MoE W4A16 HIP kernel for AMD RDNA3 (gfx1100) ( #44075 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 15:30:39 -05:00
fa27d4e9cf
[PERF] [Qwen3.5] Split mixed prefill+decode batches: route decodes to the recurrent kernel ( #44700 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:13:50 +08:00
Vadim Gimpelson and GitHub
67d3792d99
[Bugfix] Fix Qwen3.5-FP8 nightly fail. Guard fused_add_rms_norm input/weight dtype mismatch in RMSNorm + quant fusion ( #44694 )
2026-06-06 08:46:14 -04:00
00d1fb7747
[Bugfix][ROCm] ApplyRotaryEmb: fall back to native when flash_attn rotary grid would exceed the HIP per-dim limit ( #43684 )
...
Signed-off-by: vLLM ROCm fix <noreply@example.com >
Signed-off-by: amd-fuweiy <fuweiy@amd.com >
Co-authored-by: vLLM ROCm fix <noreply@example.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 02:29:16 -07:00
c9b4b184b4
[Bugfix][Voxtral] Add fetch_audio to MistralCommonFeatureExtractor (transformers>=5.10 compat) ( #44559 )
...
Signed-off-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 07:58:09 +00:00
f87df1df9e
[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig ( #44613 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-05 23:14:41 -07:00
Taneem Ibrahim and GitHub
eafbb06331
[Misc] Replaced asserts with proper exceptions to improve UX for pooling ( #44593 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-06 05:57:26 +00:00
ec0a31d4aa
[Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 ( #44692 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 10:44:21 +08:00
Devin Lai and GitHub
c8beda4cc3
[Rust Frontend] Add Phi-4 mini JSON tool parser ( #44213 )
2026-06-06 10:40:00 +08:00
2f27c9a150
Preserve layout-changing clones ( #44574 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-05 20:45:24 -04:00
4765f0f189
[Bugfix] Fix sequence_parallel_chunk_impl custom op aliasing its input ( #44130 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-05 23:56:36 +00:00
Terrence Zhao and GitHub
a50e675b0d
[Cohere] fix RoutingMethodType ( #44021 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-05 16:25:53 -07:00
Daoyuan Li and GitHub
f6a708ab2b
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models ( #44435 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-05 16:04:32 -07:00
4200f62147
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path ( #42832 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 16:22:19 -05:00
Walter Beller-Morales and GitHub
c73b0d0db9
[Core][Engine] allow DP ray placement groups to be set on specific nodes ( #44669 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-06-05 20:07:47 +00:00
Harry Mellor and GitHub
e28e369f78
Male Mergify comment less spammy ( #44666 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 10:56:52 -07:00
yzong-rh and GitHub
703fb17b13
[Bugfix] GPT-OSS instruction rendering ( #44330 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-05 13:52:32 -04:00
Sting Lin and GitHub
b593396c7a
Upgrade tpu-inference to v0.21.0 ( #44621 )
...
Signed-off-by: StingLin <sting.lin@cienet.com >
2026-06-05 16:12:49 +00:00
Flame and GitHub
91e17d4315
Fix sarvam forward compatibility with transformers v5 ( #38804 )
...
Signed-off-by: vikrantpalle <vikrantpalle@gmail.com >
2026-06-05 11:51:44 -04:00
TJian and GitHub
aa6fb8a329
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm ( #44648 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-05 15:51:17 +00:00
Effi Ofer and GitHub
6a894574bf
Add objectstore as a secondary tier to multi-tier kv cache offloading ( #41968 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
2026-06-05 18:05:41 +03:00
Yan Ma and GitHub
7f003a1285
Support MiniCPMV batched preprocessing ( #44609 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-05 15:05:31 +00:00
Harry Mellor and GitHub
ef0df7dbd6
[CI] Bump mypy version 1.19.1 -> 1.20.2 ( #44647 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:56:27 +00:00
Harry Mellor and GitHub
a80af24356
Speed up docs build ( #44635 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:51:44 +00:00
Harry Mellor and GitHub
c66b19800b
[CI] Bump mistral-common ( #44649 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:18:50 +00:00
6a11d72df7
[Reasoning][Structured Outputs] Add Command A plus tags for structural tags ( #44588 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-05 06:51:20 -07:00
Woosuk Kwon and GitHub
02d2da0748
[DSV4] Move more ops out of eager breakpoint ( #44561 )
2026-06-05 06:42:41 -07:00
adhithyamulticoreware and GitHub
bbb6c274c8
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call ( #44615 )
...
Signed-off-by: ADHITHYA BALAKRISHNAN <adhithya.balakrishnan@multicorewareinc.com >
2026-06-05 12:47:56 +00:00
62215e72c6
Remove KV cache scale boilerplate from model weight loading methods ( #43167 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 05:19:04 -07:00
7fe7800fa4
[BUG] Fix FP64 Gumbel precision coverage ( #43150 )
...
Signed-off-by: tianyu-z <zhangtianyupro@gmail.com >
Signed-off-by: Tianyu Zhang <53099276+tianyu-z@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-05 19:04:14 +08:00
8a83e6f2d7
[Rust Frontend] Batch auto-abort requests by engine ( #44591 )
...
Signed-off-by: Hugh Ryan <197298026+HueCodes@users.noreply.github.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-05 02:59:09 -07:00
Chunyang Wen and GitHub
efc347f1b2
docs: fix tokenizer optimization typo ( #44066 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-06-05 02:12:49 -07:00
Nicolò Lucchesi and GitHub
d98b8f371c
[NixlConnector] Initiate deprecation cycle for kv_both role ( #43874 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-05 11:08:17 +02:00
Chao-Ju Chen and GitHub
e64237ae82
[Rust Frontend] Support include_reasoning=false ( #44391 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-06-05 16:47:50 +08:00
d61d8566ec
[Bugfix] Update mistral tokenizer test for continue_final_message fix ( #44622 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 16:13:26 +08:00
Uranus and GitHub
d2f70da116
fix: pad dummy run query_start_loc ( #44603 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-05 00:43:04 -07:00
6542d48964
[Bugfix] Fix test_invocations flaky failure with newer openai SDK ( #44618 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 07:36:20 +00:00
Ting SUN and GitHub
ca73293fa6
[Bugfix][Rust Frontend] Fix UTF-8 char-boundary panic in incremental detokenizer ( #44620 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-05 07:36:17 +00:00
Vic Wen and GitHub
ef3af56a97
Fix LLM.wait_for_completion output type docstring ( #44617 )
...
Signed-off-by: viiccwen <viiccwen@gmail.com >
2026-06-05 00:16:38 -07:00
b4a6f26c90
[ROCm][perf] Use workspace manager for sparse indexer allocations ( #41002 )
...
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:46:29 -07:00
165b7864d0
[ROCM] [FEAT] Integrate Aiter hipBLASLt GEMM online tuning ( #40426 )
...
Signed-off-by: hanlin12 <hanlin12@amd.com >
Signed-off-by: Han Lin <hanlin12@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:45:36 -07:00
Li, Jiang and GitHub
c505cd93ef
[CI/Build] Disable CPU-Compatibility Tests ( #44605 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-05 13:14:26 +08:00
qizixi and GitHub
96229fa99e
[KVConnector][1/N] PP-aware handshake aggregation and intermediate-PP output plumbing ( #43720 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-04 22:04:19 -07:00
da1daf40bf
[Bugfix] Exclude vision embedder from quantization in Gemma4 Unified ( #44571 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-04 20:47:38 -07:00
Woosuk Kwon and GitHub
4efd6ffde0
[DSV4] Refactor DeepseekV4Attention ( #44569 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-04 20:23:07 -07:00
Chris Leonard and GitHub
56aff0dd15
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to torch stale ABI. ( #44334 )
2026-06-04 20:14:43 -07:00
zofia and GitHub
063ce98fb7
[XPU][MoE] support block_fp8_moe on xpu ( #42139 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
2026-06-05 08:36:58 +08:00
Bugen Zhao and GitHub
62d6f06e3d
[Rust Frontend] Skip loading multimodal processor if --language-model-only is specified ( #44500 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-04 17:02:54 -07:00
Schwinn Saereesitthipitak and GitHub
b7c5baf63d
fix: keep DeepSeek V4 RoPE cache on inv_freq device ( #43926 )
...
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com >
Signed-off-by: Schwinn Saereesitthipitak <17022745+galletas1712@users.noreply.github.com >
2026-06-05 02:30:29 +04:00
Jiangyun Zhu and GitHub
a55fccfc7c
[mamba] unify KDA conv states into one cache to match 2-state SSM layout ( #44539 )
2026-06-04 20:38:05 +02:00
Wentao Ye and GitHub
41a4829f22
[Logs Refactor] Optimize shutdown logs, easier to follow and consistent ( #43707 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-04 14:36:32 -04:00
38fd2405f3
use split_group for pytorch process group creation ( #41980 )
...
Signed-off-by: Tushar Jain <tushar00jain@users.noreply.github.com >
Co-authored-by: Tushar Jain <tushar00jain@users.noreply.github.com >
2026-06-04 14:36:07 -04:00
Agata Dobrzyniewicz and GitHub
a947f7a420
[Kernel][Test] Extend lightning_attn and awq_triton kernel tests to XPU ( #43307 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-04 14:25:59 -04:00
bnellnm and GitHub
439203d32c
[Bugfix] Fix test_cutlass_moe.py ( #44380 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-04 14:18:52 -04:00
Taneem Ibrahim and GitHub
8d9536a775
[Misc] Add unit tests for pooler head classes ( #44471 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-04 17:59:25 +00:00
Fadi Arafeh and GitHub
3da29aa4a5
[DOC] Add INT8 W4A8 docs and Arm's supported quantization schemes ( #34894 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-04 16:27:17 +00:00
06f94633e7
[ROCm][CI] Add test for Aiter unified attn kernel ( #44436 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 16:15:05 +00:00
99ef652907
[Bugfix] Reject non-positive values for ParallelConfig int knobs ( #44057 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-04 11:46:50 -04:00
Tyler Michael Smith and GitHub
4cc78c9d5d
[Core] Freeze garbage collector in workers after model initialization ( #44363 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 08:39:04 -07:00
tc-mb and GitHub
3dbb4e0ace
[Bugfix] MiniCPM-V-4.6 video inference crash: placeholder count mismatches visual embedding count ( #44509 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-04 08:22:30 -07:00
b21443e23c
Add model support for granite speech plus ( #43519 )
...
Signed-off-by: Zvi Kons[WSL] <zvi@il.ibm.com >
Signed-off-by: Zvi Kons (BlueVela) <zvi@il.ibm.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
2026-06-04 14:47:48 +00:00
Michael Goin and GitHub
06ee2d8433
[Quant] Support compressed-tensors WNA8O8Int linears and WNInt embeddings ( #44340 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-04 07:40:33 -07:00
Yongye Zhu and GitHub
b5235fca2e
[DSv4] Adding TRTLLM gen attention kernel ( #43827 )
2026-06-04 07:35:09 -07:00
Andreas Karatzas and GitHub
3e77036768
[ROCm][CI] Specifying time outs for the lm eval models ( #44255 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:35:00 +08:00
Andreas Karatzas and GitHub
6f68ca3e91
[ROCm][CI] Stabilize memory-release in the Hybrid model generation tests ( #44046 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:34:24 +08:00
Turner Jabbour and GitHub
0c96dd64fb
[ROCm] Bump fastsafetensors to v0.3.2 from PyPI, remove git source build ( #43625 )
...
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com >
2026-06-04 07:30:57 -07:00
Nicolò Lucchesi and GitHub
68f5e565c9
[PD][Nixl] Mamba prefix caching mode support ( #42554 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 06:41:46 -07:00
QiliangCui2023 and GitHub
9354fb1ba5
[Bugfix][Compile] Guard per_token_group_fp8_quant lookup on non-CUDA platforms ( #44476 )
2026-06-04 09:31:50 -04:00
Harry Mellor and GitHub
f35b557239
Add GH token to docs build pre run check ( #44534 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-04 05:43:49 -07:00
Dipika Sikka and GitHub
e68988a248
Refactor CT NVFP4 linear to use a single class ( #42443 )
2026-06-04 08:25:08 -04:00
4b87b3e845
[Bugfix] fix EVS for qwen3-vl ( #44205 )
...
Signed-off-by: Rui "Garry" Gao <garrygaogg@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-04 11:06:51 +00:00
90619351e3
[Attention] Mamba attention module refactor - LINEAR ( #43556 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 18:45:29 +08:00
d0975a4b50
[perf] Add gemma RMS AR fusion ( #42646 )
...
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-04 01:33:59 -07:00
Kevin_Xiong and GitHub
1bdc60ed53
Fix Kimi-K2.5 FlashInfer ViT metadata ( #44493 )
...
Signed-off-by: Kevin-XiongC <kevin_xiong1997@outlook.com >
2026-06-04 08:14:35 +00:00
a6183563b6
[Prefix Caching] DeepSeekv4 - Support selective prefix-cache retention for sliding-window KV cache ( #43447 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-04 00:48:31 -07:00
Andreas Karatzas and GitHub
22c2e87555
[CI] Reverted gitignore changes ( #44497 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 00:37:44 -07:00
wang.yuqi and GitHub
d01d0b4646
[Frontend] Consolidate online serving utils. ( #44479 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-04 06:49:31 +00:00
b4b4aaa70e
[Inductor] Fast-path Inductor fallback for vllm::*/vllm_aiter::* custom ops ( #42129 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-04 00:03:52 -05:00
Andreas Karatzas and GitHub
5e2af28838
[CI] Resolve release V2 docker build after ROCm CI wheels change ( #44463 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-03 21:35:40 -07:00
4f423bd5bc
[EPLB] Nixl communicator optimization. Zero-copy transfers ( #41633 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 03:40:34 +00:00
f0cd590d62
optimize the compressor 128 split cutedsl kernel ( #44230 )
...
Signed-off-by: Jie Fang <jief@nvidia.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-03 20:22:57 -07:00
e6018c644a
[Refactor] Remove dead code in tests and parallel_state ( #41471 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 19:32:39 -07:00
f25952e59b
[MM][Perf][CG] Support ViT full CUDA graph for InternVL ( #41759 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-04 10:24:25 +08:00
maobaolong and GitHub
b58e082d95
[KV Connector] Update lmcache kv_offloading_backend to use LMCacheMPConnector ( #42865 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-06-03 19:23:55 -07:00
Ted Mostly and GitHub
0c1e6f63f5
[Bugfix] Fix VLLMNotFoundError when using LoRA adapter name in poolin… ( #44410 )
...
Signed-off-by: Ted Mostly <wanghenshui@qq.com >
2026-06-04 02:22:03 +00:00
Giancarlo Delfin and GitHub
ceb0111a90
[Model Runner V2][Spec Decode] Add Gemma4 MTP support ( #43241 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-04 00:51:06 +00:00
0414d75410
[XPU] skip unapplied UT in test_gpu_model_runner.py ( #44289 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 08:48:17 +08:00
128adabfe0
[Bugfix] Fix Gemma4 MTP block_table batch_size mismatch under concurrent load ( #43982 )
...
Signed-off-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
Co-authored-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
2026-06-03 17:11:10 -07:00
bdbf08fc02
Bump actions/stale from 10.1.1 to 10.2.0 ( #35078 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 14:14:41 -07:00
Woosuk Kwon and GitHub
6bad553f4e
[Minor] Remove FlashInfer version check in topk_topp_sampler ( #44442 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 21:06:00 +00:00
91945b6e4a
[Bug Fix][Model Runner V2][Spec Decode] Warmup & capture with different attention states for speculator prefill ( #44253 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 13:32:40 -07:00
2b237c7a41
[Bugfix] Honor tool_choice="none" in Chat Completions streaming ( #42752 )
...
Signed-off-by: hoobnn <111053672+hoobnn@users.noreply.github.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 13:27:45 -07:00
Wentao Ye and GitHub
dad95e34d8
[Feature] Support batch invariant rms norm with residual ( #42453 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 15:22:01 -04:00
a248b45d05
[Model] Add Gemma4 Unified (encoder-free) support ( #44429 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-03 12:01:39 -07:00
linitra24 and GitHub
271328e256
[LoRA] Fix dedup for post-replacement module aliases ( #44413 )
...
Signed-off-by: bk-201 <joy25810@foxmail.com >
2026-06-03 18:23:23 +00:00
Wentao Ye and GitHub
2b91012650
[Refactor] Remove dead code fp quant ( #44122 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 14:22:23 -04:00
JartX and GitHub
5b2a2beade
[ROCm][CI] Move Model Executor test step from MI250 to MI300 (gfx942) ( #44370 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-03 12:23:51 -05:00
59d0236193
[10b/n] Migrate custom all-reduce, DeepSeek V4 fused MLA, MiniMax reduce-RMS, and MXFP8 MoE to libtorch stable ABI ( #44365 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-04 00:29:46 +08:00
0a5cbf633e
Handle spinloop ext load failure gracefully ( #43659 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-03 16:09:52 +00:00
Willow Lopez and GitHub
51e0c579b0
fix(config): validate max_num_scheduled_tokens >= 0 on all paths ( #44207 )
...
Signed-off-by: Oxygen56 <1391083091@qq.com >
2026-06-03 16:06:45 +00:00
0c6631f02a
[KVCache] Support Pluggable KVCacheSpec ( #37505 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 09:05:16 -07:00
Nicolò Lucchesi and GitHub
df7252c343
[CI] Align PD tests to HMA on by default ( #44174 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 00:04:30 +08:00
Jee Jee Li and GitHub
4d1fd13613
[CI/Build] Fix LoRA testing ( #44425 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-03 08:58:06 -07:00
Nick Hill and GitHub
ec8d60bea8
[Model Runner V2] Use FlashInfer sampler ( #42472 )
2026-06-03 07:59:31 -07:00
27f1d34a23
[Frontend][Responses API] Move developer-to-system conversion into HF renderer ( #43590 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: kdcyberdude <kdsingh.cyberdude@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-03 14:52:24 +00:00
Flora Feng and GitHub
e3e132d2dd
[Refactor] Suppress SyntaxWarning from ast.literal_eval in tool parsers ( #44346 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 10:42:19 -04:00
e5232679a3
[XPU] Add XPU block-scaled W8A8 fp8 path ( #39968 )
...
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com >
Co-authored-by: Yuxiang <yuxiang.liang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 20:16:19 +08:00
309385a359
[Rust Frontend] Add /server_info to Rust frontend ( #43942 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 04:30:47 -07:00
3d76f395e3
[SharedOffloadRegion] Align blocks to page-size ( #43689 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 14:25:57 +03:00
Li, Jiang and GitHub
823d271c0d
[Attention][CPU] Standardize kv layout to blocks first ( #44393 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-03 19:03:09 +08:00
Andy Lo and GitHub
95b1615ec9
[Perf] Improve multimodal item handling from O(n) to O(log n) per step ( #44212 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-06-03 11:00:26 +00:00
1fa9ea09f6
[Perf] Triton fast path for small CPU→GPU swap_blocks_batch in the offloading connector ( #42212 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 13:38:17 +03:00
02564b4de0
[XPU]fallback to TRITON_ATTN for vit attn on xpu when use float32 dtype ( #43759 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 03:20:21 -07:00
Flora Feng and GitHub
209709a8c1
[Bugfix] Fix unstreamed tool call args dropped in Responses API streaming ( #44348 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 03:19:08 -07:00
ace95c9cf8
[Bugfix] Update TrtLLM MoE routing methods ( #44347 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 02:56:43 -07:00
Shanshan Shen and GitHub
0e2b13103b
[Doc] Update ViT CUDA graph interfaces ( #44388 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-03 01:20:59 -07:00
Bugen Zhao and GitHub
449be4f934
[Rust Frontend] Fix several hf chat template rendering issues ( #44311 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 01:04:43 -07:00
6550ff12f2
[Rust Frontend] Add dynamic LoRA endpoints ( #43778 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:55:29 +00:00
4aaed4ca22
[Rust Frontend] Add server router extension hook ( #43774 )
...
Signed-off-by: NolanHo <kujyo.eia.serias@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:45:31 +00:00
7268457999
[KV Offloading] Enable HMA models for Tiering Offloading ( #44287 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 10:03:00 +03:00
9af53a3c13
[Perf] Add tuned selective_state_update configs for H200 and RTX PRO … ( #44251 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-06-02 23:59:01 -07:00
Andreas Karatzas and GitHub
87954eb50e
[ROCm][CI] Optimize ROCm Docker build: registry cache, DeepEP, and ci-bake script ( #36949 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 23:43:07 -07:00
Charlie Fu and GitHub
71df063c49
Enable perf_token_group_quant/_C_stable_libtorch for ROCm ( #42758 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-02 23:23:28 -07:00
Albert Cheng and GitHub
e0081ef8cf
[Benchmark] Enable reasoning-model (thinking) benchmarking via --chat-template-kwargs for client-rendered datasets ( #44244 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
2026-06-02 22:49:51 -07:00
f0204358d9
[Bugfix] fix crash in postprocess for null tool args ( #43862 )
...
Signed-off-by: William-Rom <william.rom@intility.no >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 22:17:26 -07:00
Willow Lopez and GitHub
597bc15936
fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init ( #44236 )
...
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-03 01:07:10 -04:00
Rotem Shavitt and GitHub
3f0a91bb96
Nit Changes in Tiered KV Offload ( #44293 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-06-02 21:53:21 -07:00
Flora Feng and GitHub
e67063826b
[CI] Add missing vllm/parser/ CI trigger and fix test_parse.py ( #44352 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 21:05:19 -07:00
Andreas Karatzas and GitHub
53b88d1dfc
[CI] Reject out-of-vocabulary before they reach the GPU logprob path ( #44042 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 22:27:52 -05:00
JartX and GitHub
7b476c8f14
[ROCm][CI] Skip fp8 reload tests on gfx90a (MI250) ( #44369 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:27:14 -05:00
JartX and GitHub
4454a18695
[ROCm][CI] Fix stale wvSplitK GEMM fallback test for N=5 ( #44368 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:00:25 -05:00
02a01496fc
[Platform] Add is_cumem_allocator_available ( #43838 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 10:54:50 +08:00
Kevin H. Luu and GitHub
27a93cd426
[docker] Stop using extra-index-url for flashinfer-jit-cache ( #44366 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-02 18:58:22 -07:00
Wei Zhao and GitHub
969aec4bc8
[Bugfix] Fix Deepseek v4 non-mega-moe model init error ( #44356 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-02 18:26:30 -07:00
ca17b6b17d
[Perf] Apply single-pass min_larger finding and binary search in Triton Top-p path. ( #42191 )
...
Signed-off-by: js_park <cakeng@naver.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 17:57:26 -07:00
Woosuk Kwon and GitHub
b254e0456c
[DSV4] Minor cleanup for DeepseekV4MegaMoEExperts ( #44367 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 17:54:27 -07:00
Daoyuan Li and GitHub
bd98e97557
[Misc] Remove dead VLLM_RPC_TIMEOUT env var and fix profiling doc that references it ( #44128 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-03 00:22:10 +00:00
a4ac746405
[MoE/b12x] Accept W4A16 (kNvfp4Static, None) in FlashInferB12xExperts supports check ( #43332 )
...
Signed-off-by: Junhao Shen <junshen@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-02 15:20:37 -07:00
8b3b71ee9d
[CI/Build] Bump flashinfer to v0.6.12 ( #44036 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-02 15:19:05 -07:00
Siddharth Bedekar and GitHub
0917a009d3
Fix sparse NCCL weight transfer test construction ( #44345 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
2026-06-02 21:51:21 +00:00
3099de3617
[Kernel][MoE] Add GELU_TANH to CPU, CUTLASS, and WNA16 MoE backends ( #42027 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
2026-06-02 17:12:08 -04:00
Nick Hill and GitHub
e15f20258b
[ModelRunnerV2] Avoid pipeline parallel bubbles ( #42187 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 14:02:01 -07:00
557781131a
[Misc] Remove stray empty file ( #44350 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-02 12:53:03 -07:00
Yifan Qiao and GitHub
e9e08c49b9
[Bugfix] Cache the EAGLE/MTP lookahead block in the SWA prefix-cache mask ( #44082 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 12:21:07 -07:00
Woosuk Kwon and GitHub
e4a2e584e5
[MRV2] Remove assignment of graph_pool in cudagraph_utils ( #44338 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 11:50:27 -07:00
b8b49e2395
Bump actions/github-script from 8.0.0 to 9.0.0 ( #39667 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-02 11:26:57 -07:00
da107a59e5
[MRV2] Also enable MRV2 for Llama and Mistral dense models ( #43458 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: yewentao256 <zhyanwentao@126.com >
2026-06-02 11:18:46 -07:00
ed9a7526b6
[Anthropic] Support system role messages inside messages array ( #44283 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
Co-authored-by: Ang Kah Min, Kelvin <syraxius@hotmail.com >
2026-06-02 18:13:54 +00:00
2427094152
[Feature] Support EPLB for DeepSeek v4 Mega Moe ( #43339 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-lyris01.lyris.clusters.nvidia.com >
2026-06-02 10:56:44 -07:00
Kartavya sonar and GitHub
fe32e7830b
[Bugfix] flashinfer: fail fast when --kv-cache-dtype nvfp4 used on unsupported arch ( #43669 )
...
Signed-off-by: Kartavya Sonar <sonarkartavya@gmail.com >
2026-06-02 10:50:00 -07:00
afcb580715
[BugFix] Fix Humming MoE deploy error ( #43100 )
...
Signed-off-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Signed-off-by: Alireza Dadgarnia <49554709+adotdad@users.noreply.github.com >
Co-authored-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Co-authored-by: Jinzhen Lin <linjinzhen@hotmail.com >
2026-06-02 09:32:50 -07:00
3f3e2702c2
[XPU] Enable rms_norm/act quant fusions ( #43963 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 16:14:41 +00:00
Flora Feng and GitHub
478b49ddec
[Refactor] Remove dead code from parser infrastructure ( #44279 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 12:08:27 -04:00
Nick Hill and GitHub
cab5c9a2a9
[Core] Move max_concurrent_batches to VllmConfig ( #44274 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 08:57:25 -07:00
Brian Dellabetta and GitHub
774e552397
[compressed-tensors] Asymmetric support for MoE WNA16 marlin ( #44025 )
...
Signed-off-by: Brian Dellabetta <bdellabe@redhat.com >
2026-06-02 08:51:45 -07:00
XiaoZ and GitHub
53fa09d085
[Misc] Support local image encoding in benchmarks ( #43843 )
...
Signed-off-by: xiaoz <Sukra1@outlook.com >
2026-06-02 15:15:06 +00:00
Chris Leonard and GitHub
4d93bc35c9
Migrate header files to torch stable abi ( #44013 )
2026-06-02 08:09:52 -07:00
Bugen Zhao and GitHub
586201ebdc
[Rust Frontend] Cover different thinking modes in roundtrip tests ( #44320 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:51:25 -07:00
pschlan-amd and GitHub
88f172188b
[ROCm] Fix AITER RMSNormQuantFusion for Kimi-Linear ( #44308 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-06-02 14:50:21 +00:00
Bugen Zhao and GitHub
880fc032f4
[Rust Frontend] Support recursive tool parameter conversion ( #44299 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:45:35 -07:00
6314de8bad
[XPU] [Bug] remove xpuw4a16 output size check ( #44168 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 22:26:20 +08:00
IdoAtadTD and GitHub
c91a87f01a
[BugFix] [GDN] Read linear_key_head_dim from hf_text_config for multimodal models ( #43978 )
...
Signed-off-by: IdoAtadTD <ido.atad@twodelta.com >
2026-06-02 17:17:55 +03:00
Matthew Bonanni and GitHub
ea0d045a05
[FlashAttention] Sync FA with upstream ( #44065 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-02 07:15:37 -07:00
0bdfd5eb84
[Bugfix] Vendor MiniCPMV/MiniCPMO processors to unblock Transformers v5 ( #44282 )
...
Signed-off-by: guanwei-wu <b08901019@ntu.edu.tw >
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: guanwei-wu <b08901019@ntu.edu.tw >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-02 07:14:38 -07:00
0cbc48c4f9
Support ModelOpt MXFP8 non-gated MoE ( #42958 )
...
Signed-off-by: tbarnatan <tbarnatan@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-02 13:56:03 +00:00
2fd0e52252
[Bugfix] Fix Gemma4 startup crash with recent transformers multimodal processor ( #44232 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-02 13:42:40 +00:00
654bd2bca4
[Bugfix] Sync block_size from EngineCore to frontend for hybrid Mamba… ( #42967 )
...
Signed-off-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-02 13:41:00 +00:00
wang.yuqi and GitHub
b623f7ea95
[Frontend] Consolidate dev entrypoints. ( #44170 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-02 06:30:21 -07:00
Shreyas Kulkarni and GitHub
0eeba5eec1
Fix DFlash prefix cache corruption due to missing lookahead block ( #42971 )
...
Signed-off-by: Shreyas Kulkarni <shreyas.gp269@gmail.com >
2026-06-02 12:06:33 +00:00
f69ede495b
[XPU][Mamba] Triton-based selective scan forward op for XPU ( #43421 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:50:26 -07:00
Ronen Schaffer and GitHub
2a2b5ca791
[KV Offload] Add on_schedule_end() hook to separate step lifecycle from event draining ( #44206 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 13:42:52 +03:00
689b0eeb9e
[HARDWARE][POWER] Enable SHM communicator support for PowerPC ( #43754 )
...
Signed-off-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
Co-authored-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Co-authored-by: Akash kaothalkar <61960177+Akashcodes732@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 18:06:32 +08:00
Isotr0py and GitHub
f8e9c56d15
[Multimodal] Automatically select registered video loader for VLM ( #44126 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-02 09:09:47 +00:00
alberto and GitHub
e30313220c
[Parser] Migrate ResponsesParser to unified Parser interface ( #42977 )
...
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com >
2026-06-02 08:50:05 +00:00
d247a9dc13
[EC Connector] Non blocking EC Connector lookup ( #41627 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 08:48:25 +00:00
Yifan Qiao and GitHub
7c37096620
[Core][Refactor]: thread scheduler_block_size into KVCacheManager and KVCacheCoordinator ( #44165 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 01:14:44 -07:00
Maria Guevara and GitHub
b817b23f7b
[Rust Frontend] add --enable-request-id-headers flag support. ( #43883 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-06-02 16:08:37 +08:00
Ronen Schaffer and GitHub
93da882e73
[kv_offload] Add @override decorators to subclass method implementations ( #44177 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 08:07:47 +00:00
0b25cf4419
[CPU][Perf] Enable fused kernels for GDN's gated delta rules ( #43534 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 08:00:48 +00:00
Jiangyun Zhu and GitHub
dcdfe66bfa
[Perf] use triton moe backend on hopper by default ( #44220 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-02 15:52:30 +08:00
Flora Feng and GitHub
68dafcca75
[Refactor] Unify reasoning + tool-call parsing behind Parser.parse() ( #44267 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 15:11:42 +08:00
zhrrr and GitHub
1edfd09ffd
[Model Runner V2] Use actual batch max_seq_len for attn metadata ( #43991 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 06:07:56 +00:00
zhrrr and GitHub
8a9eb40808
[Model Runner V2] Support zeroing freshly allocated KV blocks for hybrid + fp8 KVCache ( #43990 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 05:56:53 +00:00
f91fb2fcf3
[Bugfix] Convert Gemma4-MM ViT linear layers to vllm native impl ( #43798 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: ZiTian Zhao <zitian.zhao@tencentmusic.com >
Co-authored-by: B-201 <Joy25810@foxmail.com >
2026-06-01 21:41:16 -07:00
JooHo Lee and GitHub
a045c7425f
[MM][CG] Profile encoder CUDA graph pool memory ( #41714 )
...
Signed-off-by: JooHo Lee <jooho414@gmail.com >
2026-06-02 12:27:34 +08:00
a3a5a5ece5
[XPU][Bugfix] Fix per_token_group_fp8_quant missing dummy args on XPU ( #43930 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:09:21 +00:00
Or Ozeri and GitHub
480fadab1b
[BugFix][kv_offload]: Prevent offloading stale sliding window blocks ( #42959 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-06-02 05:59:48 +03:00
279d25f5cb
[BugFix] Fix TypeError in MiniCPM-O audio feature unpadding ( #38053 )
...
Signed-off-by: Krishna Chaitanya Balusu <krishnabkc15@gmail.com >
Signed-off-by: wjinxu <1299461899@qq.com >
Signed-off-by: Kc Balusu <kcbalusu@users.noreply.github.com >
Co-authored-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Kc Balusu <kcbalusu@users.noreply.github.com >
2026-06-01 19:57:28 -07:00
Andreas Karatzas and GitHub
54d0c36fff
[CI] Stabilize OpenAI schema fuzzing for malformed structural tags ( #44131 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-01 19:56:15 -07:00
Flora Feng and GitHub
9affc17a05
[Refactor] Move unstreamed tool-arg flush from serving layer to parser ( #44017 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 10:37:43 +08:00
Alec and GitHub
816cc73a9b
[Bugfix][CI] Normalize NIXL connector CUDA wheel installs ( #44266 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
2026-06-01 19:34:05 -07:00
Micah Williamson and GitHub
2588ec4f0a
[ROCm] Upgrade AITER to v0.1.13.post1 ( #44265 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-02 01:48:59 +00:00
d68f0b220e
[Bugfix][Mooncake] Release GPU pin on failed store in MooncakeStoreConnector ( #43742 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-01 18:29:18 -07:00
Woosuk Kwon and GitHub
517e74a964
[DSV4] Refactor RoPE initialization ( #44262 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 01:26:58 +00:00
JartX and GitHub
48c0d13e65
[ROCm][CI] Skip unbacked dynamic shapes tests on PyTorch < 2.11 ( #44256 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-01 19:09:01 -05:00
Woosuk Kwon and GitHub
8c3cc98cff
[DSV4] Remove unncessary classes & functions ( #44246 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-01 14:43:00 -07:00
Nick Hill and GitHub
e4cbc4385d
[Test][BugFix] Fix double-BOS in PD+specdec acceptance test ( #44234 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:31:12 -07:00
Nick Hill and GitHub
6f8b40a23f
[BugFix][CI] Fix added _has_module tests ( #44248 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:23:12 -07:00
266b9d9c64
[Frontend][Core] Add sparse NCCL weight transfer support for in-place updates ( #40096 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-01 15:37:30 -04:00
182c67daf1
[Rust Frontend] Support streaming generate endpoint ( #43779 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 19:30:55 +00:00
fd9e91d7e4
[ROCm][CI] Fix and stabilize EAGLE3 acceptance tests ( #41294 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
2026-06-01 12:40:01 -05:00
Yongye Zhu and GitHub
035733515f
[Kernel][DSv4] Optimize sparse FP8 compressor kernels ( #44161 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-02 00:18:32 +08:00
023808c23d
[Feature] Add support for JetBrains' Mellum v2 code generation model ( #43992 )
...
Signed-off-by: Madeesh Kannan <madeeswaran.kannan@jetbrains.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-01 10:11:35 -04:00
985c97a6a8
[Perf] Optimize cutlass fp8 scaled mm bypassing padding, 20% kernel performance improvement ( #43706 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-01 09:05:21 -04:00
Chaojun Zhang and GitHub
bd0aecdc08
[XPU][CI] Fix test_audio_in_video flake by using module-scoped server fixture ( #44146 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-01 11:21:36 +00:00
8796838910
[Bugfix] fix wrong partial_rotary_factor calculation for bailing_moe model. ( #43770 )
...
Signed-off-by: zzt <zengzetang.zzt@antgroup.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-01 02:42:49 -07:00
de21863419
[Rust Frontend] Add InternLM2 tool parser ( #43481 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 08:58:46 +00:00
wang.yuqi and GitHub
0910f7e0e1
[Frontend] Resettle generative scoring entrypoint. ( #44153 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-01 07:54:59 +00:00
Uranus and GitHub
1f6048abe5
fix: glm5.1 pp model loading ( #42944 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-01 15:14:47 +08:00
98f1279815
[CPU][RISC-V] Add missing RVV cpu_types helpers for WNA16 ( #42730 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-01 14:56:41 +08:00
Isotr0py and GitHub
1fd8bd02a4
[Docs] Replace broken video url in examples ( #44159 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-01 06:01:10 +00:00
29d69332aa
[BugFix] Fix _has_module to verify native deps via trial import ( #44035 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: esmeetu <jasonailu87@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-31 22:06:33 -07:00
Lucas Wilkinson and GitHub
4721bb3aa4
[MRV2] Remove Eagle's dedicated CUDA graph pool ( #44078 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-05-31 22:00:33 -07:00
Umut Polat and GitHub
f46e6be169
[Misc] Use VLLMValidationError consistently in chat completion and completion protocol validators ( #36254 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
2026-06-01 04:04:11 +00:00
8b8546da1c
docs: fix MLA attention docstring examples ( #44118 )
...
Co-authored-by: nightcityblade <nightcityblade@gmail.com >
2026-05-31 12:28:38 -07:00
Jee Jee Li and GitHub
6bdabbad5b
[CI/Build] Enable Step3p7ForConditionalGeneration testing ( #43956 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-31 05:16:12 +00:00
3fd9d2d357
[CPU][Zen] Route W8A8 and W4A16 linear inference through zentorch on AMD Zen CPUs ( #41813 )
...
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Harshal Adhav <harshal.adhav@amd.com >
Signed-off-by: Aakar Dwivedi <aadwived@amd.com >
Co-authored-by: R <Ganesh.R@amd.com >
Co-authored-by: Harshal Adhav <harshal.adhav@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-30 14:17:21 -05:00
Woosuk Kwon and GitHub
27fa5aa3b9
[MRV2] Support breakable CUDA graph ( #44050 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-30 09:40:52 -07:00
e1105064b2
[Bug] Fix gemma4 MTP IMA issue when TP>1, CUDA error: an illegal memory access was encountered ( #43909 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-30 10:34:33 -04:00
Bugen Zhao and GitHub
50c80d7923
[Governance] Add @BugenZhao as Rust frontend code owner ( #44047 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-30 22:23:54 +08:00
3becc5db40
[ROCm] Add attention sink support to AITer flash attention backend ( #43817 )
...
Signed-off-by: Xiaoran Chen <xiaoran@fb.com >
Co-authored-by: Xiaoran Chen <xiaoran@fb.com >
2026-05-30 18:13:18 +08:00
124fac10cb
[Bugfix] Fix RMSNorm kernels to multiply in weight's native dtype ( #42379 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 23:16:53 -07:00
e9499996df
[BugFix][Platform] Fix import vllm.platforms.rocm error on non-CUDA test_gpt_oss.py ( #43571 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 23:16:49 -07:00
c0056b19bf
[ROCm] cmake: support PYTORCH_FOUND_HIP for torch 2.13 native HIP language support ( #43881 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 22:16:57 -07:00
Andreas Karatzas and GitHub
ef8840adc7
[ROCm][CI] Fix failure in the Phi3V pooling test ( #44028 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-30 12:14:37 +08:00
Flora Feng and GitHub
1a096d8208
[Refactor] Remove dead current_tool_name_sent assignments from tool parsers ( #43997 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 21:45:15 -04:00
Gagan Dhakrey and GitHub
1e2ce5d11a
offload prompt_embeds decode in render_prompts_async to avoid blocking ( #43792 )
...
Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com >
2026-05-30 01:36:34 +00:00
559d6710bf
[PERF]MiniMax-M2 gate kernel ( #38445 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Co-authored-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
2026-05-29 18:28:34 -07:00
bnellnm and GitHub
187457a952
Revert "[MoE Refactor] Migrate MoeWNA16Method quantization to MK orac… ( #44033 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 16:45:29 -07:00
8fad266507
[CI] Fix smoke test step key to bypass block gate ( #43974 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 16:28:32 -07:00
Flora Feng and GitHub
8c6daf6e2f
[CI] Remove duplicate Harmony test coverage ( #44023 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 22:52:46 +00:00
bnellnm and GitHub
7b98f498cd
[MoE Refactor] Remove supports_expert_map ( #43108 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 17:26:56 -04:00
106aa92f04
[MoE Refactor] Migrate MoeWNA16Method quantization to MK oracle ( #42647 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 17:19:31 -04:00
yzong-rh and GitHub
46409fd2a1
[Fronten] Clean up stop_token_ids override for Harmony ( #44009 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-29 13:28:06 -07:00
38b864d81d
[Metrics] Exclude KV transfer tokens from iteration_tokens_total ( #43346 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 19:56:44 +00:00
Wentao Ye and GitHub
5dbf1605a0
[Feature] SSL support for dp supervisor ( #43688 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 19:28:12 +00:00
Kevin H. Luu and GitHub
acbc203340
Add @khluu to CODEOWNERS ( #44019 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-29 12:24:29 -07:00
Flora Feng and GitHub
6de08e8b46
[CI] Remove redundant test_chat_with_tool_reasoning.py ( #44011 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 19:23:56 +00:00
6aabe221a5
[CI] Make Model Executor test hangs fail fast with a traceback ( #43971 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 11:58:25 -07:00
Wentao Ye and GitHub
739096a028
[Bug] Fix torch device issue for MOE permute ( #44005 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 18:55:00 +00:00
czhu-cohere and GitHub
8b9deeec4b
[Bugfix] Fix Ray placement group allocation with grouped nodes ( #43998 )
...
Signed-off-by: <conway.zhu@cohere.com >
Signed-off-by: root <conway.zhu@cohere.com >
2026-05-29 12:51:05 -06:00
d07ad0693b
[Bugfix] Use storage_block_size in KV cache reshape for compressed specs (DeepSeek V4) ( #43988 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-05-29 11:14:25 -07:00
4aaba00f92
[EPLB] Make async EPLB default ( #43219 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-29 18:07:16 +00:00
84b2a8a7e7
[MoE Refactor] WNA16 MoE backend selection into oracle module ( #42553 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 13:11:17 -04:00
4ff865c38e
[Bugfix] Disable allreduce_rms_fusion when pipeline_parallel_size > 1 ( #43616 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 22:57:43 +08:00
5502c3b52d
[Misc] added unit tests for the core pooling methods ( #43818 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-29 14:40:31 +00:00
Chunyang Wen and GitHub
f191d5630e
docs: clarify ITL acronym in optimization docs ( #43922 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-29 07:40:05 -07:00
11dfa3169d
Add vLLM library info to Hugging Face Hub requests ( #43857 )
...
Signed-off-by: Wauplin <lucainp@gmail.com >
Signed-off-by: Lucain Pouget <lucain@huggingface.co >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-29 14:04:58 +00:00
Li, Jiang and GitHub
3f6f508e14
[Bugfix][CPU] Remove invalid extra deps ( #43977 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-29 22:02:09 +08:00
Harry Mellor and GitHub
0585b5ba2e
Skip docs build if PR doesn't affect docs ( #43972 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-29 12:09:52 +00:00
Thien Tran and GitHub
d2889722ff
[Bugfix] Corrupted MLA + linear attention ( #43961 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-05-29 05:00:51 -07:00
0b56815a24
[ROCm][Perf] DSv3.2 MI355X TP4 decode-step orchestration cleanup (3 micro-opts) ( #42982 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-29 04:26:57 -07:00
ab12aab127
[Bugfix] [ROCm] [DSV4] Fix AITER MXFP4 MoE weight loading and shuffle… ( #42595 )
...
Co-authored-by: MHYangAMD <MHYangAMD@users.noreply.github.com >
2026-05-29 04:08:33 -07:00
JartX and GitHub
0cff0741ff
[Kernel][ROCm] Native W4A16 kernel for AMD RDNA3 (gfx1100) — fp16 + bf16 ( #41394 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-05-29 11:04:40 +00:00
60a7a2214f
[Bugfix] Fix Step3 pipeline parallel KeyError for residual tensor ( #37622 )
...
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-29 03:04:02 -07:00
Nicolò Lucchesi and GitHub
7ebc0ec104
[CI] Nixl+SimpleCPUOffloadingConnector unit tests ( #43871 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-29 02:40:42 -07:00
e8b5199973
[XPU] support MTP of gdn attention ( #43565 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 17:10:24 +08:00
Simon Danielsson and GitHub
b7fb747d8d
[CI][ROCm] Don't skip MoRI-IO Connector tests ( #43703 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-29 17:06:23 +08:00
Kunshang Ji and GitHub
30c6289b8e
[XPU] fix xpu install document triton-xpu version ( #43947 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 02:05:12 -07:00
Andreas Karatzas and GitHub
ff990d0d32
[ROCm][CI] Fix AITER unified attention for encoder-decoder cross-attention ( #43945 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-29 16:43:39 +08:00
Chauncey and GitHub
87f12e5c7c
[Frontend]Responses API supports chat_template_kwargs ( #43761 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-29 07:58:19 +00:00
kliuae and GitHub
ab7521d77c
[ROCm][DSv4] Remove device pipeline stall in sparse attention ( #43898 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
2026-05-29 15:42:40 +08:00
94d3f4d205
[CPU Backend] CPU top-k and top-p sampling kernels using Triton ( #43633 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 15:02:39 +08:00
04516eabc8
[XPU] add gelu_tanh to xpu moe backend supported activations ( #42822 )
...
Signed-off-by: yintong-lu <yintong.lu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 14:37:20 +08:00
648c3ebee6
[CI] Separate non-root smoke tests from image build step ( #43712 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 23:34:16 -07:00
22a58640b4
[9/n] Migrate attention and cache kernels to torch stable ABI (continued) ( #43717 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 04:44:45 +00:00
710f077617
[Refactor] Remove dead code ( #43234 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 00:29:56 -04:00
d63108fb18
[kv_offload] Skip decode-phase blocks in CPU offload ( #43797 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-05-29 06:39:43 +03:00
9636709372
[XPU] add scale transpose to prepare_fp8_moe_layer_for_xpu and bump up kernels ( #43277 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 03:22:51 +00:00
Weida Hong and GitHub
dfe8ba7c80
Adjust design around encoder_cudagraph_forward ( #42288 )
...
Signed-off-by: Weida Hong <wdhongtw@google.com >
2026-05-29 03:02:52 +00:00
212deff2ec
[feat] add GlmgaProcessor specific logits in glm4_1v.py ( #43575 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-05-29 02:56:02 +00:00
Woosuk Kwon and GitHub
7bd45da585
[DSv4] Move mHC tilelang kernels & Don't use CustomOP in dsv4/nvidia ( #43905 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-29 10:25:02 +08:00
bf18d7e0b4
[Misc][NUMA] Auto-bind to PCT priority cores on DGX B300 + widen EngineCore across shard NUMA nodes ( #43270 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <noreply@cursor.com >
2026-05-29 10:07:44 +08:00
Bugen Zhao and GitHub
1521173c17
[Rust Frontend] Add /version endpoint using engine-reported value ( #43854 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-29 00:32:27 +00:00
b690b2bb67
[Model]Support Step-3.7-Flash ( #43859 )
...
Signed-off-by: luotingdan <luotingdan@stepfun.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: luotingdan <luotingdan@stepfun.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Yu Huang <yuhuang@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-28 17:01:48 -07:00
yzong-rh and GitHub
325a1ec4fb
[CI] Enable prefix caching in BFCL benchmark ( #43925 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-28 23:36:31 +00:00
69c9f19957
fix(frontend): Add multimodal placeholders to Gemma4 tool message template ( #41459 )
...
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-05-28 14:48:12 -07:00
rasmith and GitHub
9769e2df2a
[AMD][CI][BugFix] Fix Distributed Compile Unit Tests (2xH100-2xMI300) group ( #43120 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-28 14:39:01 -07:00
Michael Goin and GitHub
03f03f9630
Refactor output filename handling in ci-fetch-log.sh ( #43901 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-05-28 14:20:12 -07:00
Benjamin Chislett and GitHub
9202ea6fda
[Spec Decode] Allow causal DFlash ( #43445 )
2026-05-28 21:18:44 +00:00
Woosuk Kwon and GitHub
69b8956dcd
[Model Refactoring] Remove unncessary torch op registration for DSv4 ( #43891 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:04:55 -07:00
a3ed5ab10c
[KV Offload] Add per-request offloading policy via on_new_request lifecycle hook ( #43205 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 20:45:18 +00:00
7e53283b1c
[Core] Cleanup KVConnector handling with PP + fix MRV2 ( #43732 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 13:12:03 -07:00
9090368b65
[Feat] Add support for per GPU worker RDMA NIC selection ( #42083 )
...
Signed-off-by: Raj Joshi <rajjoshi@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-28 12:45:23 -07:00
Harry Mellor and GitHub
085ac221a3
Deprecate JAISLMHeadModel ( #43784 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 18:29:12 +00:00
Hua Huang and GitHub
9006204e90
[MM][CG] Avoid over-padding Qwen2.5-VL encoder cudagraph window metadata ( #42796 )
...
Signed-off-by: Hua Huang <huah@nvidia.com >
2026-05-28 11:22:56 -07:00
ed7fe831da
[ROCm] Enable the aiter top-k/top-p sampler by default ( #43331 )
...
Signed-off-by: John Qin <yanyuan.qin@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-28 13:19:59 -05:00
Nicolò Lucchesi and GitHub
5b115bb8a3
[Attention][AMD] Standardize kv layout to blocks first for AMD ( #43660 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-28 12:28:50 -05:00
53a2088675
Allow native KV cache dtype in Triton cache update ( #43330 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-05-28 16:51:40 +00:00
Chao-Ju Chen and GitHub
099024762c
[Rust Frontend] Optimize multimodal prompt expansion ( #43670 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-05-28 09:46:18 -07:00
9aa131f944
Add Cosmos3 Reasoner model ( #43356 )
...
Signed-off-by: Maciej Bala <mbala@nvidia.com >
Signed-off-by: MaciejBalaNV <mbala@nvidia.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-28 09:43:55 -07:00
Micah Williamson and GitHub
1b5437cec8
[ROCm] Bump ROCm to 7.2.3 ( #43136 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-28 09:42:43 -07:00
3207e7680e
[XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) ( #41426 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-28 16:30:48 +00:00
Matthias Gehre and GitHub
a9ec46d4b7
[ROCm][Perf] Support N=5 in wvSplitK skinny GEMM kernels for speculative decoding ( #40687 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-05-28 16:28:21 +00:00
Ronen Schaffer and GitHub
4bfa0f2b14
[KV Offload] Rename SecondaryTierManager.get_finished() to get_finished_jobs() ( #43870 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-28 16:00:18 +00:00
Vadim Gimpelson and GitHub
5d126dd155
[Bugfix] Exclude Ray DP from #42585 's deferred port allocation ( #43864 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-28 15:55:14 +00:00
c08ebebf30
[Perf] Add do_not_specialize to Mamba SSD chunk kernels ( #43803 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
2026-05-28 15:40:02 +00:00
Wentao Ye and GitHub
be4062fd6c
[Bug] Fix tests/distributed/test_elastic_ep.py - assert False ( #43813 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 11:00:56 -04:00
577d693838
[rust] fix: aggregate is_sleeping and reset_prefix_cache across DP engines ( #43429 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-28 07:56:56 -07:00
Bugen Zhao and GitHub
61a1e30473
[Rust Frontend] Reduce Gemma4 tool parser args scan complexity ( #43850 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:52:29 +00:00
Bugen Zhao and GitHub
3a282230ee
[Rust Frontend] Add hy_v3 tool parser ( #43872 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:42:47 +00:00
Li, Jiang and GitHub
20d69d100a
[CPU] Migrate cpu_awq into awq_marlin ( #43841 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-28 22:36:31 +08:00
Simon Danielsson and GitHub
552eb81918
[Bugfix][ROCm] Resolve MoRI connector hangs at high concurrency ( #40344 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-28 14:30:21 +00:00
Woosuk Kwon and GitHub
9957e4d240
[Model Refactoring] Remove torch compile dependency in DSv4 ( #43746 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:26:25 +00:00
864990e8d9
Add token-offset based selective offload in OffloadConnector ( #39983 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-05-28 14:11:02 +00:00
f3b2a819f7
[Perf][KDA] Fuse gate softplus, chunk-local cumsum, and RCP_LN2 scaling ( #43667 )
...
Signed-off-by: haojiangzheng <justineric096@gmail.com >
Co-authored-by: haojiangzheng <justineric096@gmail.com >
2026-05-28 13:47:08 +00:00
Wentao Ye and GitHub
64e1218673
[Perf] Optimize moe permute by pre-allocate buffer, 9~14% kernel performance improvement ( #43014 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 06:18:26 -07:00
Julien Denize and GitHub
02606b0b09
[BUGFIX] Multimodal benchmark with MistralTokenizer ( #42965 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-05-28 05:36:24 -07:00
19af4e6dd4
Fix OlmoHybridForCausalLM not initialising ( #43846 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 05:33:31 -07:00
omerpaz95 and GitHub
811d805195
[EC Connector] Add shutdown API to EC Connector. ( #42423 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
2026-05-28 12:28:01 +00:00
Vadim Gimpelson and GitHub
c1c4db8b4b
Log dummy DP step in iteration details ( #41406 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-28 12:18:39 +00:00
Chauncey and GitHub
d692b89c2c
[Feature] Add structured output and effort support to Anthropic Messages API ( #42396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-28 12:06:48 +00:00
Bugen Zhao and GitHub
8e0580f4ee
[CI] Auto-apply rust label to relevant PRs ( #43866 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 11:57:22 +00:00
61288b5458
[Bugfix] Fix HyperCLOVAX CI failure after upstream removed remote code ( #43860 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 03:37:36 -07:00
a583c84e2b
[Bugfix][ROCm] Fix Accuracy Drop in Sparse Indexer on gfx950 ( #43781 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-28 03:37:15 -07:00
4ec2817313
[Model][Bugfix] Rename weight_mapper to hf_to_vllm_mapper in LlamaNemotronVL pooling models ( #43581 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
Co-authored-by: opencode <noreply@opencode.ai >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-05-28 03:32:22 -07:00
Wei Zhao and GitHub
f2caefe226
[UX] Increase DP Coordinator startup timeout from 30s to 120s ( #42343 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-28 03:31:45 -07:00
Animesh Trivedi and GitHub
bfb9ebc211
[Feature] Add support for timed trace replay in vllm bench serve to replay Moonshot and Alibaba workload traces ( #39795 )
...
Signed-off-by: Animesh Trivedi <Animesh.Trivedi@ibm.com >
2026-05-28 03:31:34 -07:00
Andreas Karatzas and GitHub
a9bc0ad8e4
[ROCm][CI] Move workload from MI300 to MI325 ( #43824 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 03:31:29 -07:00
b372ad3e90
[Bugfix] Stream DeepSeek DSML tool-call argument deltas incrementally ( #42879 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-28 17:50:23 +08:00
Harry Mellor and GitHub
2a781756a1
Restore Literal for WeightTransferConfig.backend ( #43183 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 09:39:41 +00:00
Woosuk Kwon and GitHub
a04afd76aa
[DSV4] Remove AMD/XPU path in deepseek_v4/nvidia ( #43829 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 08:00:52 +00:00
6cc8577421
[Kernel] Marlin MoE: include SM 12.x in default arch list ( #40923 )
...
Signed-off-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-28 15:30:26 +08:00
d6b48f928f
[BugFix] Fix hard-coded timeout for multi-API-server startup ( #43768 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:09:13 -07:00
Rotem Shavitt and GitHub
1b16f2ddc9
change name of fs_python secondary tier to fs. ( #43600 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-05-28 07:05:48 +00:00
TJian and GitHub
0ba46d4b11
[ROCm][DSV4] Enable Tilelang MHC replacing torch/triton mhc ( #43679 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-28 07:05:28 +00:00
JINO ROHIT and GitHub
e1814f822d
minor docs: fix incorrect example path ( #43830 )
...
Signed-off-by: JINO-ROHIT <find.jinorohit@gmail.com >
2026-05-27 22:58:43 -07:00
7909f82a45
[Bugfix][Frontend] streaming tool-call serializer drops first args chunk when name and args share a DeltaMessage ( #42683 )
...
Signed-off-by: ignaciosica <mignacio.sica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-28 05:20:55 +00:00
Nick Hill and GitHub
626fa9bba5
[BugFix] Fix blocked reasoning parsing with MRV2 ( #43808 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 04:59:34 +00:00
Thien Tran and GitHub
e54eff769d
[Bugfix] Pass routed_scaling_factor to FlashInfer TRTLLM BF16 MoE ( #43769 )
2026-05-27 21:29:14 -07:00
05ac829629
fix: parse Qwen3 XML JSON arguments first ( #43243 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-28 03:35:59 +00:00
Andreas Karatzas and GitHub
33e94fc3ad
[ROCm][CI] Stabilize Cargo cache and pre-test image checks ( #43815 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 11:24:44 +08:00
413ac5c070
[Misc][Rocm] Remove redundant AiterUnifiedAttentionBackend block size log ( #43664 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 22:19:11 -05:00
Yongye Zhu and GitHub
2d2c660104
[MoE] Remove inplace fused experts mechanism ( #43727 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 20:00:19 -07:00
Benjamin Bartels and GitHub
05eec7120e
Fix RunAI streamer tensor buffer reuse during weight loading ( #43464 )
...
Signed-off-by: bbartels <benjamin@bartels.dev >
2026-05-27 19:16:52 -07:00
Bugen Zhao and GitHub
c87f62ccf8
[Rust Frontend] Introduce mock engine for benchmark baseline ( #43469 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 01:40:35 +00:00
1223732dda
[ModelRunnerV2][Hybrid model] Support kernel block size in hybrid model ( #38831 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:55:55 +00:00
amitz-nv and GitHub
381edde1b9
[Bugfix][Kernel] TRTLLM NVFP4 MoE chunking ( #43599 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-28 00:36:21 +00:00
Andreas Karatzas and GitHub
094124af15
Add @AndreasKaratzas to CODEOWNERS ( #43740 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-27 16:14:50 -07:00
Dakai An and GitHub
5963c19478
Fix Qwen3-VL and Qwen3-omni-thinker accuracy degradation from deepstack inputs under torch.compile ( #43617 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-05-27 15:34:08 -07:00
7fb9c0197a
[Bugfix][DFlash]allocate the proper number of lookahead slots ( #43733 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@gmail.com >
2026-05-27 21:45:34 +00:00
Harry Mellor and GitHub
2c2c966669
Validate against some config fields being set to 0 ( #43794 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 21:14:49 +00:00
Harry Mellor and GitHub
2616f67faa
Remove Transformers forward/backward compatibility tests ( #43785 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 12:46:36 -07:00
206b72c982
[Quantization] Fix Humming RoutedExperts import ( #43540 )
...
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 10:51:56 -07:00
284e6f543d
[8/n] Migrate merge_attn_states, mamba, sampler to torch stable ABI (continued) ( #43361 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-27 09:35:24 -07:00
jatseng-ai and GitHub
05c50c721e
[ROCm] mori: add InterNodeV1LL inter-node kernel selection via VLLM_MORI_INTERNODE_KERNEL ( #41751 )
...
Signed-off-by: jatseng-ai <jatseng@amd.com >
2026-05-28 00:33:32 +08:00
Harry Mellor and GitHub
41688e2dc7
Fix early CUDA init ( #43791 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 09:30:11 -07:00
Chunyang Wen and GitHub
49a3510266
[Docs] Fix the duplicate doc icon issue ( #43546 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-27 16:09:58 +00:00
Injae Ryou and GitHub
165460941f
[BugFix] HFValidationError with cloud storage URIs when HF_HUB_OFFLINE=1 ( #39155 )
...
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
2026-05-27 10:53:32 -05:00
Yongye Zhu and GitHub
03d9cc2fe2
[misc] Bump cutedsl version to 4.5.2 ( #43745 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 08:25:36 -07:00
52a31ccecc
[Bugfix] Map reasoning_effort to enable_thinking in chat template kwargs ( #43401 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 05:39:49 -07:00
2272062471
[Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes ( #43731 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 18:36:27 +08:00
Mohammad Miadh Angkad and GitHub
158289e0fc
[Docs] Fix MLA prefill backend default docs ( #43697 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-27 10:13:22 +00:00
Bugen Zhao and GitHub
396c8fee50
[Rust Frontend] Align tool parser fallback behavior between streaming & non-streaming paths ( #43662 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 10:13:12 +00:00
ad464e16c0
[Doc] Add Ascend NPU tab to the quickstart installation guide ( #43550 )
...
Signed-off-by: Aditya Singh <adisin650@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 08:41:29 +00:00
akii96 and GitHub
de12f5ca0b
[ROCm][GPT-OSS] Avoid repeated compile-time cos_sin_cache.to(bf16) casts in rotary path ( #42833 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-27 16:22:27 +08:00
683033d4ba
[Frontend] Add MiniCPM5 XML tool call parser ( #43175 )
...
Signed-off-by: zhangtao <zhangtao2@modelbest.cn >
Signed-off-by: zhangtao2 <zhangtao2@modelbest.cn >
Co-authored-by: zhangtao <zhangtao2@modelbest.cn >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 00:39:35 -07:00
8c94938cfb
[MRV2][BugFix] Fix KV connector handling in spec decode case ( #43719 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 06:37:56 +00:00
Nico Holmberg and GitHub
7b54690244
[ROCm][Perf] Expose AITER MoE sorting dispatch policy via env var ( #39177 )
...
Signed-off-by: nholmber <nholmber@users.noreply.github.com >
2026-05-27 13:11:02 +08:00
1fc2cee50a
[KVConnector][Mooncake] Wire reset_cache cascade end-to-end ( #42694 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Signed-off-by: Ao Shen <aoshen@inferact.ai >
Co-authored-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 20:52:35 -07:00
Angela Yi and GitHub
0fa3114ae1
Fix test_aot_compile for torch 2.12 ( #43695 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-05-26 23:12:49 -04:00
Woosuk Kwon and GitHub
adaa5e455a
[DSv4] Refactor compressor & Fix ROCm compatibility ( #43710 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 19:56:46 -07:00
c02c758ea4
[Deprecation] Deprecate functions as scheduled for v0.21.0 ( #43358 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 19:56:21 -07:00
Matthew Bonanni and GitHub
aa6138169f
[MLA][Attention] Add OOT MLA prefill backend registration mechanism ( #43325 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-26 19:56:09 -07:00
7e33081cee
[Attention] Make FlexAttention and FlashAttention use num-blocks first layouts ( #42095 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-26 19:55:56 -07:00
Xin Yang and GitHub
d8eebe6d97
[Perf] Optimize Fp8BlockScaledMMLinearKernel input_scale tensor using new_empty() ( #43677 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-05-26 19:55:52 -07:00
Andreas Karatzas and GitHub
5bdb181df5
[ROCm][CI] Fix ROCm multimodal Qwen2.5-VL activation compile and Phi4MM ragged image mask handling ( #43647 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 19:53:34 -07:00
Bugen Zhao and GitHub
0b68f21e7c
[Rust Frontend] Add reasoning/tool parser & renderer roundtrip tests ( #43582 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 00:49:30 +00:00
dede691c95
[Bugfix] Split attention groups by num_heads_q for spec-decode drafts ( #43543 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 00:11:01 +00:00
e19b9b1045
[ci] Add arm64 ci image ( #41303 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 14:38:09 -07:00
812e7e7364
[Bugfix][V1] Fix TOCTOU race causing intermittent EADDRINUSE on multi-API-server DP startup ( #42585 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 14:06:00 -07:00
d98cbf472b
[KV Connector] MooncakeStore: drop dead discard_partial_chunks parameter ( #43627 )
...
Signed-off-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 13:40:21 -07:00
Jee Jee Li and GitHub
6e503868ca
[Kernel] Porting fuse_minimax_qk_norm to manual fusion ( #43410 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-26 13:16:03 -07:00
49b4882779
[CI] Soft-fail AMD entrypoints mirror tests ( #43709 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 13:08:48 -07:00
Woosuk Kwon and GitHub
193ce8812e
[DSv4] Drop _get_compressed_kv_buffer in DeepseekCompressor ( #43690 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 10:11:25 -07:00
3aea37d28e
[Doc] Add line limit to AGENTS.md ( #43635 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-05-26 09:31:23 -07:00
Wei-Ming Chen and GitHub
6f5b533241
Add LM head quantization support for ModelOpt ( #42124 )
...
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com >
2026-05-26 09:21:05 -07:00
Woosuk Kwon and GitHub
c8414a8271
[ROCm] Remove MegaMoE integration in deepseek v4 ( #43629 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 08:56:04 -07:00
f51bbc694d
[MoE Refactor] W4a8 int8 oracle ( #42789 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:15:42 -04:00
b226ddacfd
[MoE Refactor] Migrate ModelOptMxFp8FusedMoE to oracle ( #42768 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:14:14 -04:00
Yongye Zhu and GitHub
6ab6ffb428
[Feat][DSV4] Fuse q pad into deepseek v4 fused kernel ( #43162 )
2026-05-26 05:12:54 -10:00
Andreas Karatzas and GitHub
445ded18c1
[ROCm][CI] Extend ROCm quick reduce coverage ( #40990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 21:57:13 +08:00
d565357a90
[Docs][ROCm] MoRI-IO Connector Usage Guide ( #43603 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Signed-off-by: Simon Danielsson <70206058+simondanielsson@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 21:52:30 +08:00
Mohammad Miadh Angkad and GitHub
a970fb5a1a
Fix CuPy runtime deps and restore humming ( #43530 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-26 05:59:40 -07:00
Chaojun Zhang and GitHub
861b97765d
[XPU] Fix fused MoE LoRA kernel crash on XPU by using platform-agnos num_compute_units ( #43646 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-05-26 03:40:32 -07:00
ebd0692f80
[Model] Use AutoWeightsLoader for InternLM2 ( #38278 )
...
Signed-off-by: Jesus De Jesus <dejesus.9297@gmail.com >
Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-26 03:39:26 -07:00
739af5c7e1
[Reasoning] [Bugfix] Reject invalid thinking_token_budget values ( #43402 )
...
Signed-off-by: linzm1007 <linzm1007@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 03:37:30 -07:00
Thibault Castells and GitHub
5d09f471f4
[Misc] Support interleaved custom image benchmark datasets ( #43636 )
...
Signed-off-by: ThibaultCastells <thib.castells@icloud.com >
2026-05-26 03:37:25 -07:00
681d7dd38b
[Misc][Refactor][ROCm] Convert MoRI-related envvars to extra config args ( #43303 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-26 03:33:35 -07:00
Ethan Feng and GitHub
755043cf3c
[KV Transfer] Enable HMA by default for connectors that support it ( #41847 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-26 12:28:51 +02:00
97e4022c6c
[Bugfix] Apply fc_norm in Eagle3DeepseekV2 combine_hidden_states ( #43482 )
...
Signed-off-by: Yubo Wang <yubowang2019@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-26 00:46:10 -07:00
Hank_ and GitHub
b3269454b1
[chores][log] change registry log from warning to debug ( #43045 )
...
Signed-off-by: Hank <hcc.mayday@gmail.com >
2026-05-26 00:13:46 -07:00
a37e47100c
Add CuTe DSL sparse compressor support ( #43584 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-26 00:11:12 -07:00
Sting Lin and GitHub
e6adbd7834
Upgrade tpu-inference to v0.20.0 ( #43394 )
2026-05-25 20:26:25 -10:00
zhao, zhenhui and GitHub
771e1e48b1
[CPU] Enable non-divisible GQA for decode workitems in mixed batches ( #43032 )
...
Signed-off-by: zhejiangxiaomai <zhenhui.zhao@intel.com >
2026-05-26 14:15:47 +08:00
Thien Tran and GitHub
d56612c621
[GDN] GDN Prefill kernel for SM100 ( #43273 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-05-26 14:02:11 +08:00
6f955986e1
[Bugfix][Model] Fix GPT2ForSequenceClassification sub-module prefix ( #43579 )
...
Signed-off-by: QingZhou-YangHY <3868850350@qq.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-25 22:43:19 -07:00
d5cf7b4a2c
[Frontend] Split the offline inference APIs and utils. ( #43553 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 05:20:24 +00:00
Yan Ma and GitHub
f815c99954
[Bugfix] fix device mismatch in MiniCPM-o-4_5 resampler ( #43194 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-05-26 13:12:50 +08:00
Dao007forever and GitHub
c2a4005c70
[KV Connector] Propagate MooncakeStore load failures ( #42788 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
2026-05-25 22:12:15 -07:00
7966fc7233
[KV Connector][Bugfix] MooncakeStore: don't double-apply Eagle prune in load_mask ( #43516 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-25 22:11:57 -07:00
Woosuk Kwon and GitHub
aa2b56ffb0
[DeepSeek V4] Move MegaMoE input prep kernel to nvidia/ops ( #43632 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-25 21:08:29 -07:00
Jee Jee Li and GitHub
ec5de7fa7d
[LoRA] Add one shot triton kernel For MoE LoRA ( #42290 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-25 19:47:04 -07:00
71d810bbf4
[XPU] Ensure RNG offset alignment with PyTorch requirements in XPU sampler ( #43028 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 02:01:30 +00:00
Jee Jee Li and GitHub
d4004455d2
[Kernel] Remove NormGateLinear ( #43554 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-25 09:49:19 +00:00
Nicolò Lucchesi and GitHub
716d5294e6
[Misc] Print accuracy value for PD tests even on success ( #43583 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-25 02:10:01 -07:00
873758c13a
[KV Connector] Handle Mooncake finish after preemption ( #43281 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-25 01:58:38 -07:00
5c1aec3dc0
Reduce memory usage for granite_speech. ( #42933 )
...
Signed-off-by: Yihuki <wangbovbvb@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-25 14:12:57 +08:00
Roy Wang and GitHub
0c942c69d6
[Doc] Add section on escalating stalled contributions ( #43568 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-05-25 14:11:01 +08:00
Yifan Qiao and GitHub
81252d4e24
[Feat][KVConnector] Support DSV4 in SimpleCPUOffloadBackend ( #42296 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-25 14:04:30 +08:00
3df1c7c43e
[Docker] Non-root support for vllm-openai; add opt-in vllm-openai-nonroot target ( #40275 )
...
Signed-off-by: TheDuyIT <nduy250299@gmail.com >
Signed-off-by: dtnguyen <dtnguyen@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-25 13:45:31 +08:00
1b26fa361e
[Docs] Reorganize offline inference docs. ( #43552 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-25 13:44:39 +08:00
weizhoublue and GitHub
6cbe448eed
fix: MoE model using shared routed experts crashes on AMD GPUs ( #42373 )
...
Signed-off-by: weizhou.lan@daocloud.io <weizhou.lan@daocloud.io >
2026-05-25 12:03:05 +08:00
Jee Jee Li and GitHub
b06813e872
[Kernel] Add mhc_pre_big_fuse_with_norm_tilelang ( #43474 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-25 01:19:45 +00:00
d0a100c87a
File system secondary tier implemented in python ( #41735 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-05-24 18:14:44 +00:00
d56285c747
Tuning script and configs for Triton Mamba SSU kernel ( #43083 )
...
Signed-off-by: Banani Ghosh <bg2502@nyu.edu >
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
Co-authored-by: Banani Ghosh <bg2502@nyu.edu >
2026-05-24 20:12:44 +03:00
TJian and GitHub
1806d1adfc
[ROCm] [DSv4] [Perf] Support DeepSeek v4 MTP ( #43385 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-24 18:43:08 +08:00
Andreas Karatzas and GitHub
5940590855
[ROCm][CI] Stabilize 400 error return code for invalid schema inputs ( #43016 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-24 10:06:49 +00:00
Or Ozeri and GitHub
357fddf614
[kv_offload]: Add DSv4 support ( #43142 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-24 11:10:12 +03:00
0902d8e62f
[KV Connector] Keep MooncakeStore full hits block-aligned ( #43494 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-23 23:15:03 -07:00
Wentao Ye and GitHub
33d7cbe02c
[Model Runner v2] Force v1 runner for tests ( #43233 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-23 16:37:24 -07:00
Flora Feng and GitHub
b32fe416ea
[Bugfix] Fix reasoning dropped on streaming boundary deltas ( #42691 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-23 16:18:30 -07:00
Michael Goin and GitHub
10d264a2b9
Revert "[Misc] add humming to dependencies" ( #43492 )
2026-05-23 14:21:13 -07:00
TJian and GitHub
46f95b2ec2
[ROCm][Critical] Fix the GDN import bug ( #43486 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-23 21:12:58 +00:00
Dao007forever and GitHub
819c610f9b
[Mooncake] Add metrics for MooncakeStoreConnector operations ( #43392 )
2026-05-23 13:34:40 -07:00
4438b6e7dc
[MoE] Migrate W4A8 CT to oracle kernel setup ( #42680 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-05-23 13:56:01 -04:00
Holegots and GitHub
8737e4a857
[Docs] Fix stale version number in token_classify.md ( #43489 )
...
Signed-off-by: holegots <ikun3.1415927@gmail.com >
2026-05-23 10:42:20 -07:00
Holegots and GitHub
7c2ff1f819
[Docs] Fix stale version number in token_embed.md ( #43488 )
...
Signed-off-by: holegots <ikun3.1415927@gmail.com >
2026-05-23 10:06:56 -07:00
a0be71ee47
[MM] Enable FlashInfer metadata support for Qwen2.5-VL vision attention ( #42787 )
...
Signed-off-by: Hua Huang <huah@nvidia.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-23 16:08:40 +00:00
d8b385b7ea
[Bugfix][Frontend] Fix input_audio parsing when uuid is present ( #43414 )
...
Signed-off-by: ffggs <314137448@qq.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-23 09:03:19 -07:00
Andreas Karatzas and GitHub
2a7d5b7324
[ROCm][CI] Remove benchmarks test group and shard long test groups ( #41669 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 23:31:46 +08:00
5bb8d2767a
[Kernel] Batch invariant NVFP4 linear using cutlass ( #39912 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-23 09:41:12 -04:00
GuangYaoZheng and GitHub
3f3e862681
fix(eagle3): read norm_before_fc from eagle_config for NVIDIA checkpoint ( #42143 )
...
Signed-off-by: FERRARIZHENG <popkart06@gmail.com >
2026-05-23 08:21:34 +00:00
Gabriel Wu and GitHub
82536acc54
Keep scheduler alive for delayed KV connector frees ( #43433 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-05-23 06:23:32 +00:00
Wei-Ming Chen and GitHub
09a219c075
[ModelOpt] Support Qwen3.5/3.6 VLM quantized prefix mapping ( #42546 )
...
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com >
2026-05-23 06:23:31 +00:00
d19db10974
[Bugfix] Fix native Triton top-k/top-p kernel assumes contiguous logi… ( #42739 )
...
Signed-off-by: xiaogang.zhou <xiaogang.zhou@bytedance.com >
Co-authored-by: xiaogang.zhou <xiaogang.zhou@bytedance.com >
2026-05-22 22:56:16 -07:00
Taneem Ibrahim and GitHub
3a1c062151
[Misc] Added missing return type annotations to improve mypy and IDE tooling ( #43383 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-23 13:28:22 +08:00
a7be0f342d
[7/n] Migrate pos_encoding and norm kernels to libtorch stable ABI (continued) ( #43209 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-23 13:20:00 +08:00
54d153637b
[XPU] reudce host overhead of XPU MOE ( #42915 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 13:09:34 +08:00
a5bbd81e2e
[XPU]feat: enable FP8 block-scaled quantization on XPU ( #42952 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 12:33:18 +08:00
Andreas Karatzas and GitHub
d28bdf9344
[ROCm][CI] Fix ROCm LoRA Transformers fallback with full CUDA graphs ( #41577 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 04:31:32 +00:00
84e351555a
[Bugfix] Auto-raise max_num_batched_tokens for prefix-LM multimodal models ( #43051 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-05-22 21:23:50 -07:00
Andreas Karatzas and GitHub
76ea1d5d2f
[ROCm][CI] Stabilize Granite tool-use and test URL construction ( #43017 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 12:21:11 +08:00
Andreas Karatzas and GitHub
6a4723a2e0
[ROCm][CI] Stabilize runner teardown between sampler tests ( #43023 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-23 12:19:54 +08:00
Yongye Zhu and GitHub
367cb81966
[DSV4] More multi-stream enablement for c4a ( #42925 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-23 09:22:27 +08:00
3cb83c9592
Add model to WeightTransferEngine.__init__ ( #42922 )
...
Signed-off-by: SumanthRH <sumanthrh99@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-22 17:52:15 -07:00
Duncan Moss and GitHub
552bbe6f4e
[Attention] Add head_dim=512 support for FlashInfer trtllm attention backend ( #38822 )
2026-05-22 20:27:35 -04:00
Itay Alroy and GitHub
6d30655b13
elastic_ep: stage/commit MoE quant method on reconfigure ( #40881 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-05-22 18:57:26 -04:00
8de5cabeb7
[XPU]fix: add XPU platform guards to DeepSeek-V4 ops ( #42950 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-23 06:29:45 +08:00
4e2eba28be
[Perf] Optimize hidden state extraction logic ( #37374 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-22 18:23:08 -04:00
gnovack and GitHub
f743254143
DSv4 fused Q-norm kernel grid refactor ( #42353 )
2026-05-22 15:21:33 -07:00
Nick Hill and GitHub
47d4407d7c
[Model Runner V2] Support sharing kv cache layers ( #35045 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 22:18:23 +00:00
Juhi Mittal and GitHub
e203006a8b
[Quantization][ModelOpt] W4A16 NVFP4 fused MoE + mixed-precision dispatch ( #42566 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
2026-05-22 20:51:49 +00:00
08cb46789d
mhc_post - remove sts & add vectorized copies ( #43437 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-22 13:44:29 -07:00
4e597b7491
[Bugfix] Clear error message for FP8 torchao quantization on unsupported GPUs ( #36854 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-22 20:09:17 +00:00
Artem Perevedentsev and GitHub
23f7b11bf4
[Bugfix] Detect wrong libcute_dsl_runtime.so variant in FlashInfer GDN ( #43427 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-22 19:33:33 +00:00
977703aa94
[RFC][EPLB][ #32028 ] Remove dead torch.accelerator.synchronize() from sync path ( #40733 )
...
Signed-off-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com >
Co-authored-by: SandishKumarHN <3078999+SandishKumarHN@users.noreply.github.com >
2026-05-22 15:19:24 -04:00
2b94d1c0ca
[Frontend] Simplify AuthenticationMiddleware path extraction ( #43426 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-22 11:59:14 -07:00
Yongye Zhu and GitHub
843715739b
[Refactor] Extract DeepSeek V4 sparse MLA impl into model folder ( #43149 )
2026-05-22 10:06:31 -07:00
b21f3d56d4
[KV Connector] MooncakeStore: don't co-queue save with load to avoid double delayed-free ( #43371 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-22 16:14:11 +00:00
c7624bea5e
[Bugfix] Source num_qo_heads from Attention layers in Flashinfer/Triton metadata builders ( #42650 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
2026-05-22 16:10:03 +00:00
Bugen Zhao and GitHub
91f5b92438
[Rust Frontend] [Refactor] Extract a newtype for utility call ID ( #43405 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-22 08:22:11 -07:00
Isotr0py and GitHub
f0feb15e7f
[Multimodal] Simplify ViT CUDA graph interfaces ( #41234 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-22 22:31:00 +08:00
sychen52 and GitHub
fb21d8b4f9
Add NVFP4 MOE support for Deepseek V4. ( #42209 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-22 07:21:51 -07:00
haosdent and GitHub
a377631d21
[CI] Fix AMD docker build tests ( #43329 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 14:06:24 +00:00
d3a563501b
[EPLB] Change default EPLB communicator ( #43110 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-05-22 09:43:27 -04:00
Jee Jee Li and GitHub
15f7cd33dc
[LoRA] Reduce memory of 2D weights when EP is set ( #42737 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-22 06:41:56 -07:00
79ff0ffa98
[BugFix] wire make_empty_intermediate_tensors on AyaVision and Voxtral ( #43118 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
2026-05-22 05:26:41 -07:00
Tobias Wasner and GitHub
4658bf882b
[Bugfix] Clear P0 mm sender cache on sleep/pause to fix mm_hash desync ( #43001 )
...
Signed-off-by: Tobias Wasner <wasnertobias@gmail.com >
2026-05-22 03:54:29 -07:00
b3c7ffcab8
[Misc] Replace assert with proper exceptions for security and validation in pooling ( #43286 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 18:43:33 +08:00
d3d1cf6972
[XPU]feat: add XPU fallback for MoE topk routing and MXFP4 backend ( #42951 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-22 10:22:45 +00:00
wangxiyuan and GitHub
7e1b45a092
[Attention] Mamba attention module refactor ( #41126 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-05-22 17:13:12 +08:00
Li, Jiang and GitHub
65b7a812a2
[CPU] Experimentally enable Triton and MRV2 ( #43225 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-22 01:48:17 -07:00
2380bfc210
[Docs] Note image preprocessing difference between qwen_vl_utils and vllm. ( #43393 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-22 01:43:14 -07:00
mrjunwan-lang and GitHub
a761697717
Fix the docker build failure in tpu-inference ( #43360 )
...
Signed-off-by: mrjunwan-lang <mrjunwan@google.com >
2026-05-22 01:36:17 -07:00
Nick Hill and GitHub
694d9a81bb
[BugFix] Fix setuptools-rust dep in requirements files ( #43377 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-22 15:25:10 +08:00
Weida Hong and GitHub
6bb8753db1
Correcting the mock classes for MM GC tests ( #43321 )
...
Signed-off-by: Weida Hong <wdhongtw@google.com >
2026-05-22 15:21:35 +08:00
haosdent and GitHub
025d4f5cd2
[CI] Fix "test_awq_load[gemma4-moe-*]" failure ( #43296 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 07:13:59 +00:00
5ea76fa89a
[CI] Fix test_lora_with_spec_decode on V2 model runner ( #43314 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-22 14:24:18 +08:00
tc-mb and GitHub
fa1ff88b31
[Model] Fix MiniCPM-V 4.6 vit_merger qkv weight loading ( #43213 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-05-21 22:44:06 -07:00
Furkan F and GitHub
e746a2eebf
[Model] Use AutoWeightsLoader for Voyage ( #42972 )
...
Signed-off-by: Furkan Fidan <dev@yufufi.com >
2026-05-22 05:28:23 +00:00
haosdent and GitHub
1fe3303983
[CI] De-flake renderers/test_hf.py::test_resolve_content_format_fallbacks[Qwen/Qwen-VL-string] ( #43064 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-22 12:15:22 +08:00
8c8b1825eb
[XPU] Enable multiple key kernels for sparse attention ( #37888 )
...
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com >
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-22 12:02:51 +08:00
18a27cc9a3
[Bugfix] Make CuMemAllocator free callback stream-aware ( #43020 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-22 03:36:22 +00:00
0ddd7dd656
[Frontend] DP Supervisor ( #40841 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: robertgshaw2-redhat <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 20:33:16 -07:00
60af5c16ee
[Frontend] Add truncation side to OpenAI endpoints ( #43260 )
...
Signed-off-by: Rui Zhang <rza21.bc@gmail.com >
Signed-off-by: Rui Zhang <rui.zhang@globalrelay.net >
Co-authored-by: Rui Zhang <rui.zhang@globalrelay.net >
2026-05-21 20:32:31 -07:00
Divakar Verma and GitHub
35d0141a0b
[ROCm][CI] add warmup to mem_util test before measurement ( #43236 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-22 03:17:54 +00:00
Simon Danielsson and GitHub
86ccef7d44
[ROCm] Add XGMI backend for MoRI Connector ( #41753 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-22 03:06:40 +00:00
2998a047aa
[Bugfix] Fix DSV4 Base model swiglu limit issue in FP8 path ( #42855 )
...
Signed-off-by: Chengze Fan <chengze@meta.com >
Signed-off-by: Chengze Fan <fancz2002@gmail.com >
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com >
2026-05-21 19:43:01 -07:00
Isotr0py and GitHub
ba369b7eb5
[CI] Fix dockerfile dependency graph failure for pre-commit ( #43378 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-22 10:26:05 +08:00
39910f2b25
[Rust Frontend] Move code from vllm-frontend-rs ( #43283 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Eric Curtin <eric.curtin@docker.com >
Signed-off-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com >
Signed-off-by: Will.hou <1205157517@qq.com >
Signed-off-by: Will.hou <willamhou@ceresman.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Eric Curtin <eric.curtin@docker.com >
Co-authored-by: Dev-X25874 <283057883+Dev-X25874@users.noreply.github.com >
Co-authored-by: Will.hou <1205157517@qq.com >
Co-authored-by: Will.hou <willamhou@ceresman.com >
Please see https://github.com/Inferact/vllm-frontend-rs for full original commit history.
2026-05-21 17:21:48 -07:00
Lanze Liu and GitHub
39d5fa96a7
[Bugfix] Zero stale is_prefilling in padded CUDA graph rows for Mamba ( #41873 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-05-21 15:42:42 -07:00
Nick Hill and GitHub
565b745ec5
[BugFix] Use correct logprobs for logprob_token_ids ( #43125 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 15:42:20 -07:00
e26e1f0928
[Feature] Add --cpu-distributed-timeout-seconds CLI Option for CPU Process Group Timeout ( #42968 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Signed-off-by: zWaNg3 <389750525@qq.com >
Co-authored-by: zWaNg3 <389750525@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-21 15:42:07 -07:00
Nick Hill and GitHub
0f66623b0d
[Frontend] Rework fastokens integration ( #43168 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 15:36:58 -07:00
0b59fc45dd
Disable build isolation to bypass CUDA related deps for vllm-tpu ( #43038 )
...
Signed-off-by: Ylang Tsou <ylangt@google.com >
Co-authored-by: Ylang Tsou <ylangt@google.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-21 18:00:52 -04:00
17b69828a0
[Core] Add native ModelExpress load format ( #43105 )
...
Signed-off-by: Zheng Luo <zheluo@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-21 16:05:01 -04:00
Wentao Ye and GitHub
b29cbf0652
[Perf] zeros -> empty to remove additional fill ( #42988 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-21 16:00:29 -04:00
Michael Goin and GitHub
9b54e50e2c
[Deprecation] Mark env vars covered by --moe-backend / --linear-backend ( #43148 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-05-21 12:51:12 -07:00
1c78f76c29
[Bugfix] Add early validation to reject incompatible runner types for embedding models ( #43079 )
...
Signed-off-by: anish <anishesg@users.noreply.github.com >
Signed-off-by: Your Name <ak8686@princeton.edu >
Signed-off-by: anish <145943060+anishesg@users.noreply.github.com >
Co-authored-by: anish <anishesg@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-21 11:07:46 -04:00
haosdent and GitHub
9b9d5dbaab
[CI] Fix CPU tests failing on tl.exp2 import ( #43311 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 14:28:34 +00:00
b730c46352
[Perf] [Hybrid] Fused Triton kernel for GPU-side Mamba state postprocessing ( #40172 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-21 04:50:54 -07:00
c68c55d43e
[CPU][RISC-V] Add VLEN=256 support to RVV attention kernels ( #42943 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
Signed-off-by: velonica0 <47554626+velonica0@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-21 04:50:49 -07:00
5ecd8e9c70
[XPU][CI]Fix Docker image pull-to-run race in Intel GPU CI ( #43266 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-21 10:41:38 +00:00
haosdent and GitHub
caf69823d6
[CI] Pin protoc binary in rust-build stages ( #43292 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 03:38:07 -07:00
68e07d5916
[Bug] Fix ci issue assert output_size is not None AssertionError ( #43261 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-05-21 16:58:09 +08:00
ebbfb34e3e
[Test] Replace zephyr-7b-beta (7B) with SmolLM2-135M in tokenization test ( #43085 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-21 01:57:47 -07:00
zhangxin81 and GitHub
edafea3555
Fix FlashInfer TRTLLM NvFP4 monolithic MoE routing ( #43223 )
...
Signed-off-by: zhangxin81 <115389973+zhangxin81@users.noreply.github.com >
2026-05-21 01:17:12 -07:00
b719b1635b
Update KDA chunk prefill decay to use exp2 semantics ( #43195 )
...
Signed-off-by: zexplorerhj <19794632+zexplorerhj@users.noreply.github.com >
Co-authored-by: zexplorerhj <19794632+zexplorerhj@users.noreply.github.com >
2026-05-21 01:16:27 -07:00
Kunshang Ji and GitHub
0a54df2847
[XPU] add setuptools-rust for xpu dependency ( #43287 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-21 00:14:13 -07:00
haosdent and GitHub
a950e9447e
[CI] De-flake test_models for bigscience/bloom-560m ( #43197 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-21 06:30:14 +00:00
050611a3dd
[Bugfix] Fix glm4_moe_tool_parser._is_string_type for /v1/responses FunctionTool format ( #39601 )
...
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 22:58:59 -07:00
yzong-rh and GitHub
905b97adfa
[Benchmark] Add num-warmup to vllm bench throughput ( #43245 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-21 05:13:15 +00:00
Daoyuan Li and GitHub
a6682d1d25
[Bugfix] Warn when renderer_num_workers has no effect on offline LLM ( #42905 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-05-20 21:35:08 -07:00
f2ace1d57d
[Frontend][RFC] Rust front-end integration ( #40848 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-05-21 12:24:48 +08:00
d97ba29fdc
[ToolParser][Bugfix] Re-land: Fix anyOf/oneOf/$ref type resolution in Qwen3CoderToolParser ( #37831 ) ( #38973 )
...
Signed-off-by: AAISSJ <maze0717@g.skku.edu >
Signed-off-by: <>
Signed-off-by: sejung-son <sejung.son@nhn.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: 세덩 <saison@sedeong-ui-MacBookAir.local >
Co-authored-by: sejung-son <sejung.son@nhn.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-21 12:24:08 +08:00
Flora Feng and GitHub
6441cf4a44
[Refactor] Use shared coerce_to_schema_type in Seed-OSS tool parser ( #43140 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 21:24:06 -07:00
346cf163a1
[Frontend] Normalize reasoning_content to reasoning for client compatibility ( #42664 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-20 21:23:47 -07:00
haosdent and GitHub
7e5070934e
[CI] Fix "test_vit_cudagraph_[image|video][step3_vl]" failure ( #43082 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-20 21:22:10 -07:00
2b75a73b8e
[Perf][Gemma4] Batch vision encoder calls for image and video processing ( #43169 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-20 21:22:06 -07:00
e45df8c3f7
[Bugfix] Fix Qwen3.5 GatedDeltaNet in_proj_ba Marlin failure at TP>=2 ( #36329 )
...
Signed-off-by: Adi McM Sonus Flow <biuro@sonusflow.pl >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-20 21:22:01 -07:00
Jee Jee Li and GitHub
ee05e8137e
[Minor] Bigger overlap for FI AR ( #43103 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-20 21:20:57 -07:00
Louie Tsai and GitHub
5d041cc1fe
update GPU json file based on h200 recipes ( #43262 )
...
Signed-off-by: louie-tsai <louie.tsai@intel.com >
2026-05-21 03:57:48 +00:00
9640970de2
[Model Runner V2] Fix lora Triton Error [CUDA]: device-side assert triggered ( #43139 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-21 01:00:30 +00:00
63ea11709b
[CI] Add composed-schema regression tests for DeepSeek V3.2/V4 parsers ( #43255 )
...
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-21 00:36:16 +00:00
akii96 and GitHub
bde560ed6e
[ROCm] Add QuickReduce min-size override and codec threshold ( #41675 )
...
Signed-off-by: <>
2026-05-20 17:46:51 -05:00
Jiangyun Zhu and GitHub
6dc0a71843
[Misc] downgrade nvidia-cutlass-dsl to 4.5.0 ( #43230 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-20 14:19:50 -07:00
Michael Goin and GitHub
5774aad9c5
[Perf][gpt-oss] Downgrade triton_kernels to v3.5.1 ( #43135 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-20 14:13:12 -07:00
Douglas Lehr and GitHub
452baa860b
Add dllehr-amd to CODEOWNERS and committers list ( #42772 )
...
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com >
2026-05-20 16:10:44 -05:00
Flora Feng and GitHub
2a43b407c5
[Bugfix][CI] Add missing import of pad_nvfp4_activation_for_cutlass in flashinfer ( #43237 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 11:59:12 -07:00
53ff50fcd3
[Perf] Optimize CutlassFP8ScaledMMLinearKernel when padding needed by pre-weight processing, 13.5% TTFT improvement ( #42651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-20 11:57:42 -07:00
363fc84407
Integrate flashinfer b12x MoE and FP4 GEMM kernels for SM120/121 ( #40082 )
...
Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-05-20 17:21:11 +00:00
f2d5e3d3ae
[CI] Lower granite-4.0-h-tiny gsm8k threshold for Hybrid SSM NixlConnector PD accuracy tests (4 GPUs) ( #43186 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-20 17:00:24 +00:00
2d6b3489b9
[R3] Add routed experts to openai entrypoint ( #38939 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-20 09:07:59 -07:00
Vadim Gimpelson and GitHub
9c78c99995
[MISC] Fix symm_mem cap-equal gate; log AR backend selection ( #42993 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-20 08:50:24 -07:00
Flora Feng and GitHub
a10d69116c
[Bugfix] Use shared coerce_to_schema_type in DeepSeekV32 tool parser ( #43019 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-20 10:21:00 -04:00
644b2a28e7
[Bugfix] Use enable_sm120_family for per-tensor FP8 CUTLASS kernels on SM12.1 ( #41215 )
...
Signed-off-by: j9smith <j.smith9103@outlook.com >
Signed-off-by: Joel Smith <j.smith9103@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-20 14:10:01 +00:00
ded871201a
[Bug][Structured Outputs] Fix bug that leads to unconstrained generations with structural tags ( #42452 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-20 07:08:58 -07:00
Dipika Sikka and GitHub
df84fb07a6
Remove additional dead code as a follow-up to #42889 ( #43144 )
...
Signed-off-by: Dipika Sikka <dipikasikka1@gmail.com >
2026-05-20 10:01:45 -04:00
Benjamin Chislett and GitHub
0a508743d4
[Spec Decode] Support non-MTP speculation for NemotronH ( #43130 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-20 09:15:52 -04:00
Kebe and GitHub
19cf334207
[Feature] Support manually enabling the cumem allocator ( #33648 )
...
Signed-off-by: Kebe <mail@kebe7jun.com >
2026-05-20 08:58:30 -04:00
87e31455b0
[Doc] Sync CLI guide with actual help modes and launch subcommand ( #40326 )
...
Signed-off-by: Rui Wang <raygorous@gmail.com >
Co-authored-by: Rui Wang <raygorous@gmail.com >
2026-05-20 02:32:03 -07:00
cb600d1cdb
[Frontend] Forward X-data-parallel-rank header on /inference/v1/generate ( #42330 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-20 08:58:46 +00:00
xiangdong and GitHub
6f21558da1
[XPU][CI] Add 2 server model test files in Intel GPU CI ( #42499 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-05-20 16:54:58 +08:00
Artem Perevedentsev and GitHub
1cb224430b
[GDN] Enable FI Blackwell GDN prefill kernel ( #40717 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-20 01:46:55 -07:00
Harry Mellor and GitHub
9b343dd4f5
Enable mermaid diagrams in the docs ( #43192 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-20 08:10:00 +00:00
07aeaf9d4d
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) ( #42663 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-20 00:18:12 -07:00
Nicolò Lucchesi and GitHub
40651c0207
[Docs][PD][NIXL] Bidirectional kv-cache transfer ( #43097 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-20 09:02:36 +02:00
Nicolò Lucchesi and GitHub
7e4bc2cecb
[Docs][PD][NIXL] Lease extension mechanism for blocks on P ( #43099 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-20 08:58:25 +02:00
Kevin H. Luu and GitHub
85959567c3
[ci] Revert model executor test back to L4 ( #43188 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-19 23:01:41 -07:00
Ronen Schaffer and GitHub
4f940896a3
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers ( #43076 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-20 03:32:08 +00:00
Michael Goin and GitHub
cd0ff26e7a
[CI] Add DSV4-Flash to gsm8k moe-refactor/config-b200.txt ( #42111 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-19 20:21:01 -07:00
Izik Golan and GitHub
2ae910ed88
[Perf] Avoid forward scan for async output placeholders ( #42938 )
2026-05-19 20:16:07 -07:00
fadf5d332c
add enqueue all option to throughput benchmark ( #42975 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Signed-off-by: pmaybank <113125070+pmaybank@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 20:16:02 -07:00
Benjamin Chislett and GitHub
c628a93a64
[Perf][Bugfix] Update dflash aux layer indexing ( #40727 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-19 20:15:57 -07:00
Terrence Zhao and GitHub
5774aaed0c
[Cohere] Enable Cohere MoE ( #43143 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-05-19 19:32:06 -07:00
Nick Hill and GitHub
39bba710be
[MRV2][BugFix] Fix default-stream CG capture in P/W LoRA case ( #43160 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-19 19:19:05 -07:00
Aaron Hao and GitHub
73dd2f33b7
[bug] fix WeightTransferConfig.backend to allow for all strings ( #43121 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
2026-05-19 21:01:29 -04:00
Fadi Arafeh and GitHub
be16785998
[CPU][DOC] Fix installation commands for Arm CPUs ( #43115 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-05-19 23:31:15 +00:00
117afeea46
Fix error in Dynamic NTK scaling ( #41277 )
...
Signed-off-by: Max de Bayser <mbayser@br.ibm.com >
Signed-off-by: Max de Bayser <maxdebayser@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-19 17:27:54 -04:00
Doğaç Eldenk and GitHub
1242196295
[Model] Support post-norm architecture for EAGLE-3 supeculators ( #42764 )
...
Signed-off-by: Doğaç Eldenk <dogacel@gmail.com >
2026-05-19 13:39:00 -07:00
Kevin H. Luu and GitHub
a65093c1a3
[ci] Move language models tests (hybrid) back to L4 ( #43129 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-19 11:51:34 -07:00
9aaf83ef50
[CI failure] Temporarily disable using persistent cache for flashinfer autotune ( #43119 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-19 11:44:32 -07:00
tomeras91 and GitHub
f54721bcc3
[Bugfix][MoE] FlashInfer one-sided: workspace union across heterogeneous layers ( #42976 )
...
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com >
2026-05-19 14:43:04 -04:00
aed2eb355a
[Docs] Fix MooncakeStoreConnector role in disaggregated example ( #42994 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-19 11:14:43 -07:00
Dom Brown and GitHub
d247a931cc
[feat] Add FP8 per-tensor Q scale support to Triton attention backend ( #42080 )
...
Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com >
2026-05-19 09:02:05 -07:00
Jinzhen Lin and GitHub
8200fbe1ac
[Misc] add humming to dependencies ( #42540 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
2026-05-19 08:36:47 -07:00
Flora Feng and GitHub
42b4f1fdf7
[Refactor] Extract extract_types_from_schema utility from Minimax M2 tool parser ( #43025 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-19 11:21:12 -04:00
Wang Yiwen and GitHub
1c6158083a
[Model] Openvla support ( #42654 )
...
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com >
2026-05-19 08:17:42 -07:00
Xinyu Chen and GitHub
d740e2c029
[XPU] update xpu graph usage ( #43043 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
2026-05-19 23:09:07 +08:00
Nick Hill and GitHub
b82e908b4c
[Perf][4/n] Eliminate various GPU<->CPU syncs ( #42347 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-19 10:35:54 -04:00
Sage and GitHub
a78b842d0e
[Bugfix] Fix top logprobs token placeholders in /inference/v1/generate ( #42887 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-05-19 10:21:49 +00:00
129019f334
[CI] Add MTP + PD disagg test for Qwen3.5 ( #42677 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-19 11:44:33 +02:00
Shanshan Shen and GitHub
ef54a4d604
[Misc][MM] Remove redundant code in CLIPAttention ( #43046 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-05-19 08:43:16 +00:00
Woosuk Kwon and GitHub
07beaed842
[Model Refactoring] Rename deepseek_v4.py to model.py [4/N] ( #43077 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 01:12:46 -07:00
Yifan Qiao and GitHub
056bc2e166
[KVConnector][DSV4] HMA support for Mooncake store connector ( #42828 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-19 01:07:46 -07:00
f34623bf3c
[bug] AsyncScheduler drops first post-resume token after pause_generation + clear_cache ( #42117 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 01:06:21 -07:00
Woosuk Kwon and GitHub
b14be81c1f
[Model Refactoring] Move deepseek_v4_ops to models/deepseek_v4 [3/N] ( #43073 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-19 00:52:54 -07:00
wang.yuqi and GitHub
301d986473
[Frontend] Consolidate beam search by BeamSearchMixin. ( #42946 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-19 07:37:40 +00:00
257af77bc2
[Docs] Reorganize online serving docs. ( #41907 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-19 14:43:18 +08:00
Taneem Ibrahim and GitHub
4a4fdabe28
[Misc] Aligning tokwise pooler heads for consistency ( #43041 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-19 06:16:42 +00:00
f1e3f0e6d6
[XPU] Use custom op collective behavior ( #41354 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-19 14:14:59 +08:00
9fd8487d2f
[Docs] Add SVG images for pooling models. ( #42626 )
...
Signed-off-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Gracie Guo <gracieguo@Gracies-MacBook-Pro.local >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-18 22:50:38 -07:00
27f4ba9481
fix: use keyword arguments for shard_id and expert_id in weight_loade… ( #42671 )
...
Signed-off-by: junyanxu <junyanxu5513@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-19 05:29:04 +00:00
6e889b582b
[ci] Route 28 gpu_1_queue tests to h200_35gb queue ( #43030 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-18 21:58:36 -07:00
fab07e4d0f
[Bugfix][KV Connector] Fix SimpleCPUOffloadScheduler TOCTOU between Phase A and Phase B ( #42289 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
Co-authored-by: gemini-code-assist <noreply@google.com >
2026-05-18 21:22:33 -07:00
3ca8db2ef8
add cutedsl dsv4 indexer fp8 kernel ( #42899 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-18 21:17:56 -07:00
Woosuk Kwon and GitHub
87b08c5f64
[Model Refactoring] Move DeepSeek V4 layers to models/deepseek_v4/ [2/N] ( #43039 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-18 21:00:58 -07:00
fba010dd74
[Bugfix][MRV2] Fix KVCache tensor explicit kernel_block_size dim ( #42766 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-18 20:25:41 -07:00
Mohammad Miadh Angkad and GitHub
da03e549b3
[UX] Add a persistent cache for FlashInfer autotuning ( #42537 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-18 20:25:37 -07:00
Kunshang Ji and GitHub
36dcaf25d8
[XPU] add gptq(int4) support ( #37844 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-19 11:17:09 +08:00
Ofir Zafrir and GitHub
8f16c4a5c0
[BugFix][CPU][Spec Decode] Fix Eagle implementation on CPU backend ( #42468 )
...
Signed-off-by: Ofir Zafrir <ofir.zafrir@intel.com >
2026-05-19 03:16:07 +00:00
afd7b1dce9
[Bugfix] Use platform-agnostic device in example_connector load ( #42926 )
...
Signed-off-by: Revital Sur <eres@il.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-19 03:12:04 +00:00
Woosuk Kwon and GitHub
287471b994
[Model Refactoring] Migrate DeepSeek V4 to vllm/models/ [1/N] ( #43004 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-18 19:50:02 -07:00
239b5ff30c
[Frontend] Add --spec-method/--spec-model/--spec-tokens CLI aliases ( #42476 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-18 17:22:27 -07:00
Artem Perevedentsev and GitHub
f85c76d701
[CI/Build] Bump nvidia-cutlass-dsl to 4.5.1 ( #42991 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-05-18 16:58:15 -07:00
shanjiaz and GitHub
a171e6b52d
Add parallel drafting to v2 model runner unsupported features ( #43010 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-05-18 16:39:09 -07:00
Wentao Ye and GitHub
37ece593c1
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement ( #42774 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 16:38:12 -07:00
Flora Feng and GitHub
57fef4e0bf
[Refactor] Extract shared coerce_to_schema_type utility from Minimax M2 tool parser ( #43006 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-18 17:55:39 -04:00
haosdent and GitHub
0191354827
[Perf][MLA] Enable FULL cudagraph capture for TRITON_MLA decode ( #42885 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-18 14:29:10 -07:00
Wentao Ye and GitHub
cd49a05d5a
[Refactor] Remove dead code ( #42889 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 16:41:22 -04:00
Ronen Schaffer and GitHub
84747489de
Tier offload followup ( #42529 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-18 19:41:58 +00:00
Tuukka Sarvi and GitHub
8fc1c284b9
[ROCm] Guard AITER GDN decode fast path by layout ( #42880 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-05-18 11:56:22 -07:00
Amit Portnoy and GitHub
ce88f01c9a
[Docs] update attribution to reflect EDEN foundation ( #41666 )
...
Signed-off-by: amitport <1131991+amitport@users.noreply.github.com >
2026-05-18 11:22:56 -07:00
Wentao Ye and GitHub
00e20e76f7
[Refactor] Remove dead cuda kernels ( #42767 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 11:14:21 -07:00
czhu-cohere and GitHub
9758a6e5c5
[BugFix] support PP for Cohere vision model ( #42819 )
...
Signed-off-by: <conway.zhu@cohere.com >
Signed-off-by: root <conway.zhu@cohere.com >
2026-05-18 11:12:06 -07:00
Bowen Bao and GitHub
a2c8fc6657
[ROCm][Quantization][3/N] Refactor quark_moe w4a4 w/ oracle ( #41436 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-18 13:46:13 -04:00
6859ca7615
[Bugfix] fix swiglu limit issue for humming backend + deepseek v4 ( #42541 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-18 17:32:26 +00:00
Mohammad Miadh Angkad and GitHub
67f58ce23f
[Bugfix] Fix DSV4 MTP after ROCm mHC integration ( #42930 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-18 17:02:01 +00:00
Wei Zhao and GitHub
8c296de63b
[Perf] Re-enable flashinfer autotune by default and cleanup ( #42857 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-18 09:12:27 -07:00
Harry Mellor and GitHub
b12745e4f3
Fix --convert passed without --runner on causal models ( #42935 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 15:56:09 +00:00
Wentao Ye and GitHub
e26736973a
[Model Runner V2] Fix prompt logprobs calculation Sizes of tensors must match error ( #42778 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-18 08:27:21 -07:00
Netanel Haber and GitHub
47829b1159
[Bugfix] mamba: run single-token extends as decodes ( #42430 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-18 15:26:00 +00:00
Blanc Swan and GitHub
4a39b4f553
[Model] Add Apertus Tool Parser ( #41154 )
...
Signed-off-by: Blanc <swan.blanc@infomaniak.com >
2026-05-18 11:20:04 -04:00
78e7a7b9b0
Refactor AWQ Marlin MoE onto modular WNA16 oracle ( #42483 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Signed-off-by: Siddharth Bedekar <104613085+bedeks@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 08:02:43 -07:00
f5d3dc7115
[Model Runner v2] Support update_config ( #42783 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-18 10:26:07 -04:00
1ac10f159a
Revert "[torch.compile] Add patch for fullgraph compilation" ( #42686 ) ( #42913 )
...
Co-authored-by: Luka Govedič <luka.govedic@gmail.com >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-18 09:02:51 -04:00
e5417657e5
[KV Connector][Offloading] Flush all pending jobs on last step ( #42611 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 12:59:42 +00:00
xiangdong and GitHub
2e40faf08b
[XPU][CI] Temporarily skip test_moe_lora_align_block_size_mixed_base_and_lora[1] in Intel GPU CI ( #42954 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-05-18 20:34:48 +08:00
Nicolò Lucchesi and GitHub
69c91d010a
[MRv2] Default to MRv1 when a connector is present ( #42955 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-18 20:34:16 +08:00
roikoren755 and GitHub
737bfa3a43
[Bugfix][Hybrid][NemotronH] Fix mamba_cache_mode=all + speculative decoding crash ( #41233 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-05-18 14:54:00 +03:00
Kfir Toledo and GitHub
e414e1f1c0
[Bugfix][KV Offload] count appended GPU blocks in store group_sizes ( #42945 )
...
Signed-off-by: Kfir Toledo <kfir.toledo@ibm.com >
2026-05-18 11:36:02 +00:00
df852ed503
fix: remove unused norm for dpskv4 ( #41710 )
...
Signed-off-by: inisis <desmond.yao@buaa.edu.cn >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-18 18:33:29 +08:00
Yuwen Zhou and GitHub
88a860d754
[CPU] Add MXFP4 W4A16 MoE support ( #41922 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-05-18 03:04:45 -07:00
cac81b6eda
[CPU Backend] Improve cpu thread utilization ( #42666 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 03:04:41 -07:00
Li, Jiang and GitHub
b4601ad43f
[CPU] Add fused GDN support for AMX CPU platform ( #42707 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-18 03:04:36 -07:00
Jee Jee Li and GitHub
2267f70070
[Kernel] Pack topk id/weights triton kernel ( #42527 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-18 03:04:31 -07:00
965d076148
[CPU] Specify required KV cache layout for CPU attention backend ( #42740 )
...
Signed-off-by: Tony Lin <tony.lin@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-18 17:38:54 +08:00
c38bed4248
delete xpu ci ( #42582 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-18 16:36:45 +08:00
Xin Yang and GitHub
998714b21b
[Perf] Add do_not_specialize in fused FP8 RoPE kernel ( #42849 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-05-18 01:32:46 -07:00
Harry Mellor and GitHub
9537542537
Revert checkpoint specific workaround in Transformers modelling backend ( #42923 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 17:31:06 +09:00
Rishapveer Singh and GitHub
5ab6d1b3fd
[Model] [Perf] Use flatten for Qwen3.5's GDN output projection ( #42311 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-05-18 16:14:36 +08:00
7d5b033782
[LoRA] Support 2D and 3D MoE LoRA adapter at the same time ( #42242 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-18 15:22:26 +08:00
e3aeee5ff8
[Bugfix] moe lora align kernel grid ( #40131 )
...
Signed-off-by: TheDuyIT <nduy250299@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: dtnguyen <dtnguyen@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-18 00:17:53 -07:00
Harry Mellor and GitHub
c1f7854342
Improve logging when docs build is skipped ( #42929 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-18 06:33:32 +00:00
gaozihao-shy and GitHub
23c15acd77
[BugFix] Kimi-K2.5: skip vision tower dtype conversion when using quantization ( #42869 )
...
Signed-off-by: gaozihao-shy <gaozihao-shy@users.noreply.github.com >
Signed-off-by: gaozihao <gaozihao3@huawei.com >
2026-05-18 05:07:16 +00:00
Andreas Karatzas and GitHub
b50646e5ef
[ROCm][CI] Stabilize ROCm pooling and multimodal CI ( #42909 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-18 03:57:59 +00:00
Soyaazz and GitHub
990f49bdcb
[MM][CG] Enable encoder Cudagraph for Step3VL ( #42224 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
2026-05-17 20:19:13 -07:00
107210442d
[CI] Add NIXL EP import canary ( #42567 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-05-17 19:11:46 -07:00
03ddc1c9bc
[Perf] Wire silu_and_mul_per_block_quant into TritonFP8MoE (MiniMax-M2) ( #42497 )
...
Signed-off-by: qianlihuang <yiliu.dong@qq.com >
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Co-authored-by: qianlihuang <yiliu.dong@qq.com >
2026-05-17 21:57:04 -04:00
Luka Govedič and GitHub
966903eb93
[torch.compile] Add patch for fullgraph compilation ( #42686 )
...
Signed-off-by: Luka Govedič <luka.govedic@gmail.com >
2026-05-17 19:49:16 +00:00
TJian and GitHub
599e75f432
[ROCm] [Bugfix] Fix DeepSeek V4 Functionality and Accuracy ( #42810 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-17 12:18:50 -04:00
Taneem Ibrahim and GitHub
1c8e9c0399
Refactor: Pass num_labels explicitly to PoolerClassify instead of reading from global config ( #42851 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-17 14:40:21 +00:00
0fa888465e
[XPU] fix weight scale shape ( #42725 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-17 16:55:10 +08:00
liuzhenwei and GitHub
ff712f6447
[MRV2][XPU] add Model Runner V2 log ( #42710 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-17 04:15:50 +00:00
Qi Zhou and GitHub
504a26ce2b
Support bf16 for mamba ssm cache ( #41680 )
...
Signed-off-by: Qi Zhou <qizzzh@google.com >
2026-05-16 17:54:58 -07:00
weizhoublue and GitHub
a94189295b
Fix Weight loading for Qwen3.5-MTP and Qwen3-VL using runai_streamer ( #42716 )
...
Signed-off-by: weizhoublue <weizhou.lan@daocloud.io >
2026-05-16 17:54:27 -07:00
0867497368
[CI/Build] Bump flashinfer to v0.6.11.post2 ( #41711 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-16 14:55:12 -07:00
36e74c9ea4
[KV Connector] Support disk offloading in MooncakeStoreConnector ( #42689 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-16 13:34:15 -07:00
Taneem Ibrahim and GitHub
787bc0d031
Add unit tests for pooler activation functions ( #42824 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-16 14:58:16 -04:00
weizhoublue and GitHub
d1586e1a12
Fix: Propagate pinned model revisions into Ultravox secondary weight loading ( #42830 )
2026-05-16 17:02:54 +00:00
Jiangyun Zhu and GitHub
8a56da3845
[Experimental] Breakable CUDA graph ( #42304 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-16 22:04:12 +08:00
Andreas Karatzas and GitHub
4db300e95f
[ROCm][CI] Removed problematic command override mechanism ( #42807 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-16 17:35:05 +08:00
657b42b592
[Docker][KVConnector] Build mooncake-transfer-engine from source ( #42114 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: khluu <khluu000@gmail.com >
2026-05-16 00:26:25 -07:00
Jee Jee Li and GitHub
32b7177909
[LoRA][Bugfix] Dedup LoRA wrapping for modules referenced from multiple attribute paths (MoE gate) ( #42757 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-16 11:22:35 +08:00
39c67d714e
fix: add API key authorization to /v2 endpoints ( #42594 )
...
Signed-off-by: DustHunter <dusthunter@126.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-16 01:29:27 +00:00
87a2adcb43
[Misc] Add common random prefix option to structured-output serving benchmark ( #41632 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-16 00:44:48 +00:00
Michael Goin and GitHub
852f567444
[Bugfix] Respect explicit --kv-cache-dtype over checkpoint kv_cache_scheme ( #42782 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-15 17:15:52 -07:00
Michael Goin and GitHub
b2a27b82d9
[Kernel][UX] Add --linear-backend arg for linear kernel selection ( #39538 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-15 17:07:39 -07:00
Keyi Li and GitHub
d0921bafef
[Bugfix] Unwrap VLM wrappers for EPLB on Model Runner V2 ( #42706 )
2026-05-16 07:20:33 +08:00
1ccdf87507
[Bugfix] Fix layerwise reload alias-buffer corruption ( #42481 )
...
Signed-off-by: rasdani <73563550+rasdani@users.noreply.github.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-15 15:20:53 -07:00
Rita Brugarolas and GitHub
bd9dbe6060
[ROCm][Bugfix] Fix fused_mla_dual_rms_norm for AITER API rename _fused_qk_rmsnorm ( #42606 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-05-15 14:50:03 -06:00
de2d76f352
[Build] Switch CUDA 12.9 wheel builds to PyTorch manylinux_2_28 base ( #41668 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-15 13:46:16 -07:00
9a7a273dfe
Add HumanEval and GSM8K benchmarks to datasets ( #42648 )
...
Signed-off-by: southfreebird <yvorott@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-15 13:01:21 -07:00
b2c58ee942
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants ( #42685 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-15 15:34:59 -04:00
frida-andersson and GitHub
4d67d3bde2
[ROCm] Restore fast top_k_per_row kernels for sparse MLA when topk_tokens=2048 ( #42072 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-05-15 19:02:57 +00:00
06d020bb6e
[Bugfix] Fix SM121 (DGX Spark) exclusion from Marlin/CUTLASS FP8 paths ( #35568 )
...
Signed-off-by: Blake Ledden <blake@secondnaturecomputing.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Pavani Majety <pmajety@nvidia.com >
2026-05-15 10:59:00 -07:00
chunxiaozheng and GitHub
f45c210885
[LMCacheMPConnector] Prioritize importing the lmcache_mp_connector from lmcache ( #42596 )
...
Signed-off-by: idellzheng <idellzheng@tencent.com >
2026-05-15 17:46:31 +00:00
akii96 and GitHub
be7a03ea65
[ROCm] Widen AITER fused AR RMSNorm 1-stage gate ( #42409 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-15 17:44:38 +00:00
6147c70224
[Model Runner v2] Support reload weights (sleep mode) ( #42673 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-15 16:41:23 +00:00
0162596603
[Model Runner V2] FP32 gumbel sampling. ( #41775 )
...
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
2026-05-15 09:20:08 -07:00
46a95815d3
[ROCm][MLA] FP8 ASM prefill for AITER dense MLA backend on gfx950 ( #42509 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: clintg6 <clint.greene@amd.com >
Co-authored-by: frida-andersson <frida.andersson@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-15 23:56:58 +08:00
BadrBasowid and GitHub
fb5bd03f51
[Perf] Set IR Op Priority Once at Worker Init ( #42631 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-05-15 15:56:13 +00:00
Mohammad Miadh Angkad and GitHub
ee58665aac
[Bugfix] Fix DeepGEMM context lens contiguity in MLA indexer ( #42135 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-15 23:29:58 +08:00
Wentao Ye and GitHub
491e8d8539
[Perf] Optimize MLA attention _v_up_proj bmm by removing additional copy ( #42561 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-15 08:14:26 -07:00
Wentao Ye and GitHub
af9616d845
[Model Runner V2] Fix kv_connector pre_forward order ( #42676 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-15 08:13:59 -07:00
d792d993c1
[ROCm] Widen OAI Triton MoE capability range to include gfx12 (RDNA4) ( #37826 )
...
Signed-off-by: L.B.R. <lbr@mmonad.com >
Co-authored-by: L.B.R. <lbr@mmonad.com >
2026-05-15 07:59:57 -07:00
Aaron Hao and GitHub
e0a45f1455
[Feat][RL] IPC weight sync optimizations: multigpu support and chunked packed tensors ( #37476 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Signed-off-by: hao-aaron <ahao@anyscale.com >
2026-05-15 22:53:06 +08:00
Benjamin Chislett and GitHub
0fe7550254
[Bugfix] DFlash FP8 KV-Cache ( #42692 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-15 08:29:45 -06:00
Li, Jiang and GitHub
95cfe102a5
[Bugfix] Ensure embeding model compilation on CPU ( #42709 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-15 18:58:19 +08:00
1dc3fe08ea
gemma3 multi-gpu bug-fix ( #42630 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-15 02:32:05 -07:00
d26a28ab03
fix: propagate revision/code_revision pins to all artifact boundaries ( #42616 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-15 02:31:54 -07:00
Andreas Karatzas and GitHub
d735968f6d
[ROCm][CI] Stage B gating ( #42025 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-15 01:49:27 -07:00
ccde9540be
DeepSeekV4-Pro enable cuda graph full and piecewise mode ( #42604 )
...
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-15 01:45:30 -07:00
wang.yuqi and GitHub
75fd68c7a5
[Entrypoints] Split the pooling offline API into PoolingOfflineMixin. ( #42267 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-15 08:05:57 +00:00
Yifan Qiao and GitHub
4b364f810e
[Core][DSV4] Skip caching SWA blocks that can never serve a prefix-cache hit ( #42258 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-05-15 15:59:18 +08:00
31fa757cf9
[Misc] Make it simpler to replace out-of-tree layer classes with related LoRA layers. ( #42306 )
...
Signed-off-by: paulyu12 <507435917@qq.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-15 15:20:42 +08:00
Cyrus Leung and GitHub
2676ab1e0b
[Deprecation] Remove old locations of get_tokenizer and resolve_hf_chat_template ( #35024 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-15 00:13:32 -07:00
27b85d2084
[Bugfix] Clarify CPU backend memory error messages reference shared flag ( #42479 )
...
Signed-off-by: daniel-devlab <282598346+daniel-devlab@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-15 06:35:05 +00:00
Louie Tsai and GitHub
e30f39c4f1
Update Intel Xeon model list and vLLM Benchmark Suite BKMs ( #42607 )
...
Signed-off-by: louie-tsai <louie.tsai@intel.com >
2026-05-15 05:14:03 +00:00
bf610c2f56
[Bugfix] Fix inverted condition causing thinking_token_budget to be silently ignored ( #41674 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-15 12:48:49 +08:00
faa4b76afa
[Model] Support InternS2 Preview ( #42705 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: zxy <46674730+CUHKSZzxy@users.noreply.github.com >
2026-05-14 21:30:26 -07:00
f351455f0f
[CPU][RISC-V] Add RVV-optimized attention kernels for RISC-V Vector Extension ( #40119 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-15 12:08:23 +08:00
Cyrus Leung and GitHub
56434e8651
[Bugfix] Fix incorrect chat template format for Qwen3.5 ( #42660 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-14 20:52:52 -07:00
0d4d334eaa
Bump llguidance to 1.7 ( #42150 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 20:35:27 -04:00
fa2a33b893
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py ( #38288 )
...
Signed-off-by: Chengyi Nie <cnie@roblox.com >
Co-authored-by: Chengyi Nie <cnie@roblox.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-15 08:25:52 +08:00
Giancarlo Delfin and GitHub
3b6a204789
[Model Runner V2][Bug Fix][DSV4] Ensure lazy attention state initializations happen during cudagraph capture ( #42444 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-05-14 16:16:17 -07:00
f8848b2f2d
[Bugfix] Add swiglu limits to deepgemm fp8 methods ( #41986 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-14 15:43:13 -07:00
Charlie Fu and GitHub
4cfcc0866f
[CI][ROCm] Remove unsupported cases in test_fusion.py ( #38680 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-05-14 17:37:18 -04:00
f887aa1a53
[Aiter][ROCm] RMSNormGated+GroupedQuantFP8 fusion ( #40710 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Tres Popp <trespopp@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-14 15:37:09 -04:00
Matthew Bonanni and GitHub
9898f94abe
[Attention] Remove deprecated MLA prefill arguments ( #42555 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-14 10:34:06 -07:00
ae4f59f0ec
[Model Runner v2] Oracle for model runner v2 - qwen3 dense model by default [1/N] ( #39337 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-14 10:02:33 -07:00
Ranran and GitHub
f3d5360591
[Bugfix][Multimodal] PyAV video backend returns keyframes labeled as targets ( #42586 )
...
Signed-off-by: Ranran <hzz5361@psu.edu >
2026-05-14 08:56:59 -07:00
Baorun (Lauren) Mu and GitHub
a7737cb4f3
[Fix] Misc Fixes in ViT CUDA Graph ( #38040 )
...
Signed-off-by: Baorun Mu <bmu@nvidia.com >
2026-05-14 23:49:06 +08:00
Cyrus Leung and GitHub
b8a25d0e12
[Bugfix] Fix LM detection for Nemotron Parse ( #42641 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-14 23:42:10 +08:00
frida-andersson and GitHub
f07b1da797
[ROCm] Enable gluon paged MQA logits on gfx950 (MI355X) ( #42062 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-05-14 15:39:26 +00:00
f60c6b33a5
[V1][DP][LB] Publish request counts at the start of each engine step ( #41626 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-14 15:39:24 +00:00
24337fb860
PD disagg with NIXL Connector: GDN support (Qwen3.5) ( #41869 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-14 16:33:01 +02:00
c7560af424
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support ( #39568 )
...
Signed-off-by: xhx1022 <1737006628@qq.com >
Signed-off-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: arlenxu <arlenxu@tencent.com >
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com >
2026-05-14 14:12:30 +00:00
Mohammad Miadh Angkad and GitHub
2317682f95
[Bugfix] Fix TRTLLM ragged MLA prefill workspace warmup ( #42112 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-14 09:48:56 -04:00
5bd8c71e79
[kv_offload] Implement reset_cache() for the offloading connector ( #41956 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com >
2026-05-14 16:00:10 +03:00
Wentao Ye and GitHub
6548560496
[Compile] Fix compile warning with topk softplus sqrt ( #41261 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-14 05:12:50 -07:00
0a65d46628
[DSV4] Fuse norm and router for low latency scenario ( #41263 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io >
Co-authored-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 05:11:02 -07:00
Zhenzhong Xu and GitHub
1ea9401364
[Quantization][Autoround][Toolkit] Add W4A16 Support ( #39778 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
2026-05-14 19:18:49 +08:00
9946c38b7f
[XPU] Fix double-transpose in XPUFP8ScaledMMLinearKernel for W8A8 quant method ( #41689 )
...
Signed-off-by: Libin Tang <libin.tang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-14 17:17:39 +08:00
23c85343fb
[Bug] Fix DeepSeek V4 AttributeError: module 'cutlass.cute.nvgpu' has no attribute 'LoadCacheMode' ( #42342 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-14 02:00:20 -07:00
rasmith and GitHub
768f4a6f26
[CI][AMD][BugFix] Prevent triton compiler error when running test_moe_layer with use_ep = True on ROCm ( #40857 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-14 08:44:22 +00:00
rasmith and GitHub
addef3299c
[CI][AMD] Skip tests where models have problems or fails on both HW types ( #42126 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-14 08:21:06 +00:00
ce29c26b31
Update Dockerfile.rocm for AINIC & Thor NIC ( #40453 )
...
Signed-off-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu >
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu >
Co-authored-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
Co-authored-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-14 15:24:27 +08:00
aoshen02 and GitHub
8c79ad6580
Revert "[Core] Replace routing replay with device cache and async D2H pipeline" ( #39917 ) ( #42434 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
2026-05-13 23:49:01 -07:00
0d2732dd91
[MLA Attention Backend] Add TOKENSPEED_MLA backend for DSR1/Kimi K25 prefill + decode on Blackwell ( #41778 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-13 23:48:02 -07:00
Rebecca Lee and GitHub
fd7d858c8a
Use hidden_pad and intermediate_pad from vLLM #34301 ( #42098 )
...
Signed-off-by: Rebecca Lee <Rebecca.Lee@amd.com >
2026-05-14 14:21:04 +08:00
liuzhenwei and GitHub
b26558d4a3
[CI][XPU] skip ut of offload connector ( #42598 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-14 13:13:53 +08:00
Sarah Salah and GitHub
bf0d2dc6d7
[Misc] Fix mypy error in parser_manager type narrowing ( #42441 )
...
Signed-off-by: Sarah-Salah <11881117+Sarah-Salah@users.noreply.github.com >
2026-05-14 02:48:59 +00:00
ca60a4e84f
[Fix] Weight loading for qwen3_5 using runai_streamer ( #42521 )
...
Signed-off-by: Harsh Shah <iharsh@google.com >
Co-authored-by: Harsh Shah <iharsh@google.com >
2026-05-14 10:36:20 +08:00
Roy Wang and GitHub
77e1421a68
[Bugfix] Fix EPLB initialization for VLM wrapper models ( #39805 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-05-14 02:26:15 +00:00
Kunshang Ji and GitHub
751b9f14bd
[XPU][CT] Support mxfp8 moe model ( #41918 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-14 09:47:10 +08:00
Krish Gupta and GitHub
70c00163ff
[Feature] Add instruction support for score/rerank chat templates ( #42412 )
...
Signed-off-by: KrxGu <krishom70@gmail.com >
2026-05-14 09:41:22 +08:00
Siddharth Bedekar and GitHub
f51f6844f9
[Bugfix][Spec Decode] Wire draft_probs into probabilistic draft_model rejection ( #40269 )
2026-05-13 21:04:03 -04:00
longguo and GitHub
665f9c4253
[Bugfix] Fix Gemma4ToolParser streaming float corruption ( #42128 )
...
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-05-13 18:03:30 -07:00
Flora Feng and GitHub
1087676a90
[Refactor] Use shared utils in hermes tool parser ( #42570 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-13 20:35:45 -04:00
63cc8a55a9
fix(tool-parser): preserve "none"/"nil" strings as valid enum values in minimax_m2 ( #39599 )
...
Signed-off-by: Yiyang Liu <yiyangliu@microsoft.com >
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-13 20:35:34 -04:00
Divakar Verma and GitHub
ca7e4546da
[CI] set max transformers version for skywork model ( #42104 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-13 16:53:49 -07:00
b2198670b1
[Bugfix] V1: support tuple model outputs in ubatch wrapper (dbo + spec decode) ( #40789 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-13 15:47:51 -07:00
Mohammad Miadh Angkad and GitHub
f1cc7aad3c
[Bugfix] Fix DeepSeek V4 MTP HC state handling ( #42320 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-13 15:44:52 -07:00
Lukas Geiger and GitHub
597ed13803
[Core][MM] Do not use urllib3 to parse data URLs ( #42535 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-05-13 22:21:01 +00:00
liangel-02 and GitHub
6b5c389ee3
expose flex block size for batch invariant mode ( #41252 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-05-13 14:11:57 -07:00
Michael Goin and GitHub
8efd508204
[Quantization] Rework quantization_config to use QuantKey and allow for activation override ( #41566 )
2026-05-13 16:58:32 -04:00
ovidiusm and GitHub
cca32d55a2
[PD] Fix broken NIXL EP installation ( #42542 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
2026-05-13 13:55:51 -07:00
Walter Beller-Morales and GitHub
873910d608
[Frontend] add support for thinking_token_budget in completions ( #42116 )
2026-05-13 16:01:52 -04:00
Wentao Ye and GitHub
3f611f6106
[CI] Fix pre-commit issue ( #42563 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-13 12:37:26 -07:00
Nick Hill and GitHub
a505cf807e
[ModelRunner V2] Share identical MTP weights ( #42538 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-13 18:57:04 +00:00
40330967ab
[Quark] Support loading Quark NVFP4 checkpoints in vLLM ( #35859 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Signed-off-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-05-13 11:17:36 -07:00
Fynn Schmitt-Ulms and GitHub
ab1ad0d7a9
Remove verifier model type check in speculative config ( #42536 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-05-13 18:14:39 +00:00
Ben Browning and GitHub
0f69128a37
[Bugfix] Handle real-world gpt-oss tool call output in Harmony parsing ( #42454 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-05-13 17:54:46 +00:00
b3c69595a6
[MM][CG] Support ViT CG for Qwen2-VL ( #41736 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-14 01:52:35 +08:00
2f821faeae
[Spec Decode] Support hybrid attention models in extract_hidden_states ( #39949 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-13 10:45:53 -07:00
5794c65f8c
[Bugfix][Model] Gemma4 MoE routing closure captures per_expert_scale, breaking functional_call substitution ( #42250 )
...
Signed-off-by: Noelia <noeliabentancor1@gmail.com >
Signed-off-by: Noelia Bentancor <71080743+NoeliaBentancor@users.noreply.github.com >
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 17:43:12 +00:00
CynicDora and GitHub
256dbcaabf
[Feature] Support custom callable proposer backend for speculative decoding ( #39487 )
...
Signed-off-by: 524031910363 <hyzhyzsh@sjtu.edu.cn >
Signed-off-by: CynicDora <hyzhyzsh@sjtu.edu.cn >
2026-05-13 16:53:01 +00:00
Wentao Ye and GitHub
e35c0d4c63
[Feature] Support compile mode for batch invariance on SM80 ( #42456 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-13 11:02:39 -04:00
Ronen Schaffer and GitHub
11f6b545d4
[kv_offload] Add multi-tier KV cache offloading framework ( #40020 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-13 17:21:43 +03:00
a8887c208f
[Bugfix] [ROCm] [DSV4] [Perf] Add aiter mhc support ( #41946 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-13 21:43:15 +08:00
0ddaf6dffa
[XPU] [CT] Enable CT W4A4MxFp4 path and add xpu kernel ( #38896 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-13 06:43:00 -07:00
Marek Wawrzos and GitHub
67671692ac
[CI] Re-enable Nemotron Parse parity test and switch testing to nemotron-parse v1.2 ( #42498 )
...
Signed-off-by: <mwawrzos@nvidia.com >
2026-05-13 21:05:27 +08:00
hissu-hyvarinen and GitHub
0a62f5eec9
[AMD] skip machete tests for rocm ( #42326 )
...
Signed-off-by: Hissu Hyvarinen <hissu.hyvarinen@amd.com >
2026-05-13 12:11:03 +00:00
PikaPikachu and GitHub
3b1ef03be4
[Bugfix][Quark] Fix W8A8 INT8 garbage outputs on Step-3.5-Flash (and other 3-key fused-MoE Quark exports) ( #41892 )
...
Signed-off-by: kangletian <kangletian@hotmail.com >
2026-05-13 11:59:49 +00:00
3c413a5481
Triton attention: add USE_TD constexpr for tensor descriptor Q/K/V load/store ( #40327 )
...
Signed-off-by: Artur Fierka <artur.fierka@intel.com >
Co-authored-by: quinnlp <quinnlp@users.noreply.github.com >
2026-05-13 13:57:41 +02:00
Ronen Schaffer and GitHub
79fd1bc7ed
[kv_offload] Add req_id to ReqContext for per-request tracking ( #42507 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-13 11:11:10 +00:00
SILONG ZENG and GitHub
cee6751e54
[Bugfix][Qwen3-VL] Fix pipeline-parallel deepstack initialization ( #42394 )
...
Signed-off-by: MrZ20 <2609716663@qq.com >
2026-05-13 10:58:42 +00:00
16863072ca
[Bugfix] Fix scipy audio resampling ratio ( #42233 )
...
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com >
2026-05-13 18:52:41 +08:00
Andreas Karatzas and GitHub
d628a3c5cb
[ROCm][CI] Skip ROCm batch invalid-input test pending torch fix ( #41572 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-13 18:50:47 +08:00
akii96 and GitHub
74dffae666
[ROCm] Run AITER RMSNorm pad fusion before AR RMS fusion ( #42411 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-13 18:35:12 +08:00
97c4317bf5
[Bugfix][Frontend] Default max_tokens server-side on /inference/v1/generate ( #42329 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-13 11:16:46 +02:00
f6e868fbdf
[CI] Use uv with Python 3.12 for PyPI wheel upload ( #42470 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-13 02:12:06 -07:00
13bf242100
[Feat][KVConnector] Add bind_gpu_block_pool() to KVConnectorBase_V1 ( #39654 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 02:10:29 -07:00
Jiangyun Zhu and GitHub
140dc2ec30
[Bugfix] Install nvidia-cutlass-dsl[cu13] extra on CUDA 13 platforms ( #42438 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-13 01:57:21 -07:00
9ce74042d3
[Bugfix][SimpleCPUOffloadBackend] Dedup in-flight CPU offload stores across scheduler steps ( #41289 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-13 01:53:32 -07:00
sychen52 and GitHub
a8c13d2837
Patch SlidingWindowSpec.real_page_size_bytes for nvfp4 kv ( #42464 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-13 01:46:30 -07:00
Shanshan Shen and GitHub
92def124bc
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3.5 ( #42151 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-05-13 16:00:32 +08:00
85b2fecab7
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) ( #42339 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
2026-05-13 07:24:39 +00:00
503697c9ce
[chore] Refactor pooling metadata token ID accessors ( #42368 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-13 06:08:01 +00:00
Nicolò Lucchesi and GitHub
71bcd02ef3
[Bugfix][PD] Fix multi-node TP (TP>8) ( #39907 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-12 22:20:57 -07:00
Matthew Bonanni and GitHub
dcacdf9a88
[Attention] Sync FA with upstream ( #41052 )
2026-05-12 23:34:18 -04:00
bnellnm and GitHub
18f6bf5a21
[MoE Refactor] Add sequence parallel tests to test_moe_layer.py ( #41299 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-12 21:52:19 -04:00
Alec and GitHub
07534b8782
[PD] Bump NIXL connector dependency to 1.x ( #42364 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
2026-05-12 18:05:01 -07:00
Wentao Ye and GitHub
3d635c58c0
[Perf] Optimize MLA compute_prefill_context memory allocation ( #42460 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-12 16:23:46 -07:00
+4
ebeb09d822
[KV Transfer] Add MooncakeStoreConnector for KV cache offloading via Mooncake distributed store ( #40900 )
...
Signed-off-by: leichao.lc <leichao.lc@antgroup.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: leichao.lc <leichao.lc@antgroup.com >
Co-authored-by: ivanium <yifanqiao@inferact.ai >
Co-authored-by: aoshen524 <aoshen@inferact.ai >
Co-authored-by: Dao007forever <daole@inferact.ai >
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com >
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com >
Co-authored-by: foraxe <1055696449@qq.com >
Co-authored-by: Skywalker-EP <173423846@qq.com >
Co-authored-by: fems14 <1804143737@qq.com >
Co-authored-by: jianzs <zheng.shoujian@outlook.com >
Co-authored-by: baxingpiaochong <771405853@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 16:09:10 -07:00
Michael Goin and GitHub
184577ae46
[Build] DeepGEMM: trim comments, add integration notes + TODOs ( #42429 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-12 15:57:58 -07:00
Kevin H. Luu and GitHub
8c4fc4202a
[CI] Inline build artifact annotations in release pipeline ( #42357 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 15:57:43 -07:00
Nick Hill and GitHub
fe8b42e80c
[CI] Fix test_async_scheduling.py flakiness ( #42455 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-12 21:38:32 +00:00
Giancarlo Delfin and GitHub
fe5b4e0fe7
[Model Runner V2] Apply synthetic mode to probabilistic rejection sampler ( #41035 )
2026-05-12 13:37:03 -07:00
0ce6613b9c
platforms: add uses_cpu_device() hook to Platform for DeviceConfig ( #42313 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-12 12:39:17 -07:00
379f0ec369
[CI] Migrate 6 verified jobs from gpu_1_queue to h200_18gb MIG ( #42446 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-12 11:52:01 -07:00
KaivalyaMDabhadkar and GitHub
67c89fe40a
[Model][Bugfix] Fix Step3-VL image_embeds input path ( #42333 )
...
Signed-off-by: Kaivalya Dabhadkar <kdabhadkar@nvidia.com >
2026-05-12 18:47:55 +00:00
d9b4990783
[MoE Refactor] EPLB refactoring for FusedMoE ( #41055 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-12 14:16:31 -04:00
4d591db470
[MoE Refactor] Introduce RoutedExperts alias for FusedMoE and don't store SharedExperts in MK ( #40735 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-05-12 13:37:44 -04:00
yzong-rh and GitHub
6ff7405b81
[Bugfix] [Frontend] Responses API, fix merging of messages ( #42189 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Yifan <yzong@redhat.com >
2026-05-12 16:09:59 +00:00
Yan Ru Pei and GitHub
bcb9c133ba
feat(kv-events): emit KV cache metadata ( #40984 )
...
Signed-off-by: PeaBrane <yanrpei@gmail.com >
2026-05-12 15:58:48 +00:00
c8a6e272e0
[CPU] Fix rotary embedding for CPU without flash-attn ops ( #42225 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-12 15:05:35 +00:00
Wentao Ye and GitHub
a1b2d87498
[Refactor] Clean up pooling models build_tok_params logic ( #42341 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-12 15:05:05 +00:00
Martin Hickey and GitHub
418ba8ef14
[kv_offload][BugFix] Fix store deferral ( #41945 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-12 18:04:44 +03:00
5a6a9fc6f6
[docs] Added one new contact to the Vulnerability Management team ( #42145 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 10:59:59 -04:00
289cee0473
[vLLM IR] Minor improvements ( #39362 ) ( #39558 )
...
Signed-off-by: Avishek Goswami <avishek.goswami@ibm.com >
Co-authored-by: Avishek Goswami <avishek.goswami@ibm.com >
2026-05-12 10:58:36 -04:00
shanjiaz and GitHub
6ccb10d794
Added peagle speculators support ( #41826 )
...
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com >
2026-05-12 07:55:57 -07:00
7a9cc5e7f0
[Model] Support MiniCPM-V 4.6 ( #41254 )
...
Signed-off-by: caitianchi <caitianchi@tc-mb.com >
Signed-off-by: tc-mb <157115220+tc-mb@users.noreply.github.com >
Co-authored-by: caitianchi <caitianchi@tc-mb.com >
2026-05-12 14:28:10 +00:00
d077622d60
[Build] Build bundled DeepGEMM _C per-Python so the wheel imports on every CPython ( #41516 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-12 10:27:29 -04:00
dd6b3a5ef5
[Perf] Use 2D-grid to eliminate divmod in W8W8 group quant ( #42153 )
...
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-12 10:01:30 -04:00
593d5a4033
[Bugfix] Fix mismatched kernel-per-logical blocks in NIXL HMA transfer ( #42097 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-12 15:53:30 +02:00
bnellnm and GitHub
6427603ae8
[MoE Refactor] Move remaining experts classes to experts directory ( #42334 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-12 09:19:46 -04:00
206eaed08d
[MoE Refactor] Move expert map related code into ExpertMapManager class ( #41046 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-05-12 09:18:27 -04:00
8f89381fc6
[Hybrid] Warmup Mamba2 SSD kernel ( #39822 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-12 12:46:22 +00:00
Dipika Sikka and GitHub
a7b801e26d
[MXFP4] Support for linear layers + compressed-tensors integration ( #41664 )
2026-05-12 07:49:33 -04:00
Kunshang Ji and GitHub
4df1be9547
[XPU] bump up vllm-xpu-kernels to v0.1.8 ( #42410 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-05-12 11:47:37 +00:00
bc03f280c8
[XPU] keep generator state of sycl kernel align with pytorch ( #41771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Qiming Zhang <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 11:44:47 +00:00
997132911e
[Doc] Fix typo in llm-d documentation link ( #42397 )
...
Signed-off-by: Florian Woerner <florian.woerner@onmyown.io >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-05-12 04:26:46 -07:00
haosdent and GitHub
fc8bf6eedb
[CI] De-flake Language Models Test (Extended Generation) test_models(False-False-5-32-bigcode/starcoder2-3b) ( #42392 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-12 10:46:48 +00:00
liuzhenwei and GitHub
07a40ede19
[UT][XPU] fix test_parallel_sampling due to global random state ( #42388 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-12 18:03:23 +08:00
Kevin H. Luu and GitHub
e1c8776e90
[CI] Move DockerHub and PyPI publish steps to end of release pipeline ( #42355 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 09:17:42 +00:00
Kevin H. Luu and GitHub
1ff9d33535
[CI] Migrate remaining B200 jobs to b200-k8s with test fixes ( #42387 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 02:00:37 -07:00
7f65f84428
[Bugfix] Fix empty channel/recipient in harmony for /v1/responses ( #35540 )
...
Signed-off-by: kg6-sleipnir <christopherhazen42@gmail.com >
Signed-off-by: chazen <45186108+kg6-sleipnir@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-12 08:45:51 +00:00
amitz-nv and GitHub
ef34592a1a
[Bugfix] Fix double reduce in flashinfer_nvlink_two_sided and flashinfer_nvlink_one_sided backends ( #41382 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-12 07:47:47 +00:00
Kevin H. Luu and GitHub
f69644caf8
[CI] Migrate more B200 jobs to b200-k8s queue ( #42356 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-12 00:38:31 -07:00
d37e25ffbe
[Frontend] Consolidate Speech to Text entrypoints. ( #42370 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 07:06:57 +00:00
8517cdaf90
[XPU] update dp rank w/o env-var isolation ( #39856 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-12 14:49:54 +08:00
Lucas Kabela and GitHub
4e498b5e5c
[Bugfix][Performance Improvement] Improve penalties triton kernel performance ( #40657 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-12 05:47:20 +00:00
28ee78af54
Implement custom dataset class for ASR benchmarking ( #41576 )
...
Signed-off-by: Yasmin Moslem <48152713+ymoslem@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 12:17:58 +08:00
ZiTian Zhao and GitHub
630492da30
[Fix] Gemma4 Mixed-Resolution Image Co-Batching Crash ( #42217 )
...
Signed-off-by: zitian.zhao <zitian.zhao@tencentmusic.com >
2026-05-12 03:13:03 +00:00
Chauncey and GitHub
920bf3ec84
[Bugifx] [Qwen3CoderTool] Restore supports_required_and_named for required tool_choice ( #42292 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-12 02:09:56 +00:00
pschlan-amd and GitHub
39dff5ff39
Add VLLM_USE_SPINLOOP_EXT to use more efficient busy polling ( #36517 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 16:11:49 -07:00
d7af6b34d8
[Model Runner V2] Bug fix: logprob dtype int64/int32 issue ( #41761 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 21:55:43 +00:00
bbee532988
[Perf][1/n] Eliminate various GPU<->CPU syncs ( #41429 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 20:36:03 +00:00
53181384e0
[Bugfix] Fix DSV4 swiglu_limit on marlin backend ( #42287 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 13:03:56 -07:00
wang.yuqi and GitHub
a0dc7a0f36
[CI] Consolidate Speech to Text tests ( #42274 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-11 19:50:17 +00:00
56e5810ff1
[BugFix] Prevent orphaned process on NCCL destroy ( #39846 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-11 15:25:26 -04:00
Flora Feng and GitHub
639cbfd274
[CI] Add tests/parser to CI coverage ( #41877 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-11 19:08:54 +00:00
a721315488
[ROCm][Perf] Fix RMSNorm+Quant fusion for gfx950 (non-fnuz) ( #41825 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Chuan Li <chuali@amd.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: Frida Andersson <frida-andersson@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-11 15:00:51 -04:00
6fdb49392e
[Bugfix] Fix int32 overflow in DeepGEMM SiLU/mul FP8 Triton kernel ( #42201 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Signed-off-by: Vensen <vensenmu@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 14:52:31 -04:00
cf0d279142
[Docs] Add Apple Silicon documentation for vLLM-Metal GPU support ( #41987 )
...
Signed-off-by: alexagriffith <agriffith96@gmail.com >
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com >
2026-05-11 11:34:25 -07:00
5497ffbf7c
Add documentation about vLLM FIPS compliance ( #42190 )
...
Signed-off-by: Vinay Damodaran <vrdn@hey.com >
Signed-off-by: Vinay R Damodaran <vrdn@hey.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
2026-05-11 18:17:02 +00:00
Nick Hill and GitHub
9af6a5ed75
[Model Runner V2] Fix seq_lens_cpu_upper_bound ( #42202 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-11 10:37:50 -07:00
Hexiang Wang and GitHub
7863fff6e5
[ROCm][DSv4] implement flash sparse mla with triton kernels ( #41812 )
...
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
2026-05-11 09:27:11 -07:00
Wentao Ye and GitHub
0d453e2336
[Perf] Batch invariance with Cutlass fp8 support, 28.9% E2E latency improvement ( #40408 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-11 12:20:58 -04:00
Wentao Ye and GitHub
3f9c0c25b3
[Bug] Fix kimi dtype issue with mm_projector_forward ( #42081 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 11:45:24 -04:00
Vadim Gimpelson and GitHub
a2e776d716
[Bugfix] Accept canonicalized modelopt_* quant_method in _extract_modelopt_quant_algo ( #42181 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-11 11:10:57 -04:00
4955990f1b
[kv_offload] Move FilterReusedOffloadingManager logic to CPUOffloadingManager ( #41727 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-11 18:09:29 +03:00
Wentao Ye and GitHub
4b64fc2cbf
[Refactor] Cleanup batch invariant dead code ( #41993 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-11 10:48:39 -04:00
pschlan-amd and GitHub
5f1b313900
[ROCm] Clean up a bit the AITER FA backend ( #41942 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-05-11 22:45:18 +08:00
724ed2fc35
[DSv4] Improved dequant gather K cache kernel ( #42236 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-11 10:41:12 -04:00
a51376b3f0
[Performance][DSR1]: Fused RoPE+KVCache+q_concat for MLA ( #40392 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
Co-authored-by: ElizaWszola <ewszola@redhat.com >
2026-05-11 14:10:50 +00:00
Martin Hickey and GitHub
8415bf2cdb
[kv_offload] Set offloading connector to prefer HND layout ( #41928 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-11 15:05:41 +03:00
Noa Neria and GitHub
ac062147fa
Avoid silent weights corruption when loading Nemotron Nano VL with reusable-buffer loaders like runai distributed streaming ( #42244 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-05-11 12:03:14 +00:00
Chauncey and GitHub
617239b70c
[Frontend]Responses API supports chat_template_kwargs ( #42272 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-11 11:59:39 +00:00
Kyungmin Lee and GitHub
27ae676364
Fix EXAONE-4.5 to align with Transformers update ( #42246 )
...
Signed-off-by: lkm2835 <lkm2835@gmail.com >
2026-05-11 10:25:31 +00:00
haosdent and GitHub
17ed5e61f5
[CI] Make Python-only Installation optional ( #42293 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-11 09:47:16 +00:00
Nicolò Lucchesi and GitHub
5672d100ed
[KV Connector][NIXL][Bugfix] Fix NIXL handshake failures not honoring kv_load_failure_policy ( #40364 )
...
When NIXL handshake fails (e.g., due to compatibility hash mismatch
between prefill and decode instances), requests fail with "engine dead"
error instead of gracefully falling back to local recomputation as configured
by kv_load_failure_policy='recompute'.
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:37:21 +00:00
Nicolò Lucchesi and GitHub
770e9bd6b3
[Nixl][PD] Lease renewal TTL KV blocks on P ( #41383 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-11 09:27:30 +00:00
Cyrus Leung and GitHub
9efdddca28
[Model] Fix missing maybe_prefix ( #42280 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-05-11 09:04:06 +00:00
Qiu and GitHub
b1b59720b2
bugfix(flashinfer,dcp): remove kv_cache_layout for BatchDCPPrefillWrapper._new_tokens. ( #38895 )
...
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com >
2026-05-11 08:11:49 +00:00
f9f770ca0b
fix nixl side-channel host selection ( #41806 )
...
Signed-off-by: Shahar Mor <smor@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 07:40:37 +00:00
Haoqing Wang and GitHub
5cba6839e6
Document MolmoWeb hf_overrides ( #42163 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-10 23:58:22 -07:00
Jee Jee Li and GitHub
05d610e5cd
[CI/Build] Reduce LoRA model tests. ( #42266 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-11 14:49:08 +08:00
581b5e9afc
[Frontend] Return rendered prompt text in chat completion response ( #42052 )
...
Signed-off-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Wang, Zhipeng | RASIA <zhipeng.wang@rakuten.com >
Co-authored-by: Cursor <cursor@cursor.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-11 13:53:39 +08:00
wangxiyuan and GitHub
5536fc0c01
[Misc] Replace mamba_type string literals with MambaAttentionBackendEnum ( #41188 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-05-11 03:59:36 +00:00
vllmellm and GitHub
7f95e66a11
[ROCm][Bugfix]: dynamically align BLOCK_DMODEL with Lv in MLA decode kernel ( #41119 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-11 11:14:19 +08:00
yzong-rh and GitHub
b1687527b8
[Bugfix] Gemma 4 chat template crash with missing tool name and tool id ( #42188 )
...
Signed-off-by: Yifan <yzong@redhat.com >
2026-05-11 03:07:45 +00:00
171019ab19
add fused mhc_post_pre kernel ( #41536 )
...
Signed-off-by: george <george@inferact.ai >
Co-authored-by: george <george@inferact.ai >
2026-05-10 19:56:52 -07:00
Haoqing Wang and GitHub
879a8c3180
Fix Molmo2 image token metadata ( #42162 )
...
Signed-off-by: Haoqi Wang <78337154+hqhq1025@users.noreply.github.com >
2026-05-11 01:19:21 +00:00
1b57eb41f2
[MoE] Move various experts classes to fused_moe/experts/ ( #41979 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Jackmin801 <56836461+Jackmin801@users.noreply.github.com >
2026-05-11 07:54:33 +08:00
Mohammad Miadh Angkad and GitHub
21943d4c25
[Performance] Make safetensors checkpoint prefetch settings configurable ( #41499 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 15:55:15 +00:00
f396bee56f
[DSV4] Add PP support for deepseek-v4 ( #41694 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: qizixi <22851944+zixi-qi@users.noreply.github.com >
2026-05-10 15:47:26 +00:00
215e2f7990
[Bugfix][Mamba] IMA in causal_conv1d kernel for long sequences ( #41617 )
...
Signed-off-by: vensen <vensenmu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 12:38:28 +00:00
Ronen Schaffer and GitHub
e175192d33
[KV Offload] Pass ReqContext to touch(), complete_load(), and complete_store() ( #41366 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-10 15:09:25 +03:00
a54f0d1049
[CPU] Fix spec decode kernel signatures for synthetic mode compatibility ( #41932 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-05-10 12:07:15 +00:00
Isotr0py and GitHub
48698b1b9b
[Bugfix] Fuse Qwen3.5 in_qkvz_proj forwarding with LoRA enabled ( #37912 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
2026-05-10 10:59:02 +00:00
Andreas Karatzas and GitHub
0a309b5ee9
[ROCm] Cap Triton paged attention block size to fix ROCm shared memory OOM ( #38502 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 10:03:00 +00:00
Jee Jee Li and GitHub
84f7a55340
[CI] Trigger LoRA test when changing MoE code. ( #42196 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-10 01:26:09 -07:00
Ethan Feng and GitHub
a2c9d548d7
[Docs] Fix broken local links ( #42160 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-10 01:15:38 -07:00
Yongye Zhu and GitHub
301305c093
Add @zyongye to CODEOWNERS ( #42200 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-10 16:07:32 +08:00
Mohammad Miadh Angkad and GitHub
efd0e7789d
Fix mypy failure on main ( #42197 )
...
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu >
2026-05-10 07:55:57 +00:00
a5d0a5afba
[Frontend][Bugfix] Abort ASR engine requests on cancellation ( #41266 )
...
Signed-off-by: abdulrahman-cohere <abdulrahman.abdulrazzag@cohere.com >
Signed-off-by: <>
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-09 23:51:11 -07:00
Andreas Karatzas and GitHub
f2840120f6
[ROCm][CI] Fix NIXL spec-decode acceptance startup and diagnostics ( #41313 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 14:50:16 +08:00
Dao007forever and GitHub
3f5bd482f5
[Bugfix][KV Transfer][NIXL] Notify P node on pre-admission rejection to free stranded KV blocks ( #41269 )
2026-05-09 22:52:09 -07:00
Andreas Karatzas and GitHub
fb1ac806c5
[ROCm][CI] Stabilize ROCm shutdown and distributed compile CI ( #41573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-10 03:47:40 +00:00
Wei Zhao and GitHub
986edc858a
[Bugfix] Fix DeepSeek v4 topk numerical issue for unaligned max-model-len ( #42169 )
2026-05-09 20:30:08 -07:00
27d3bac272
docs: clarify Gemma 4 assistant speculative decoding ( #42180 )
...
Signed-off-by: AbhiOnGithub <abhiOnGithub@users.noreply.github.com >
Co-authored-by: AbhiOnGithub <abhiOnGithub@users.noreply.github.com >
2026-05-09 20:08:44 -07:00
00b0618a03
Use CU_MEMCPY_SRC_ACCESS_ORDER_ANY for batch KV cache swaps ( #39306 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <etelis2019@gmail.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Itay Etelis <etelis2019@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 05:57:09 +03:00
0d382ecde8
Handle optional bool-or-string CLI args in get_kwargs ( #40951 )
...
Signed-off-by: Christian Van <cvan20191@gmail.com >
Co-authored-by: Christian Van <cvan20191@gmail.com >
2026-05-09 19:47:21 -07:00
Isotr0py and GitHub
1029e5ef28
[CI/Build] Use modelscope's international site for regression test ( #42176 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-09 19:47:09 -07:00
0b272a6e01
[Bugfix] Fix SP pass for multimodal models and PP+SP residual handling ( #33322 )
...
Signed-off-by: Xingran Wang <wangxingran123456@outlook.com >
Signed-off-by: Hongjian Zhang <hirokenovo@gmail.com >
Co-authored-by: Hongjian Zhang <hirokenovo@gmail.com >
2026-05-09 19:44:16 -07:00
dcb3135af7
Fix: Nemotron 3 rescue whitespace-only final_content, not just None ( #41846 )
...
Signed-off-by: Nave Assaf <nassaf@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 02:07:58 +00:00
bc5fdc1e6a
Add NVFP4 all-gather GEMM fusion for AsyncTP ( #41882 )
...
Signed-off-by: roG0d <baonudesifeizhai@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 01:13:22 +00:00
aoshen02 and GitHub
006af4b956
[Bugfix] Skip routed-experts hot path when disabled ( #42148 )
2026-05-09 18:01:04 -07:00
Wentao Ye and GitHub
ea0e501bb1
[KV Connector] Remove compat support for pre-v0.12.0 constructor signatures without KVCacheConfig ( #39832 )
...
The v0.12.0 release contained initial support for HMA in KV Connectors. As part
of these changes, a KVCacheConfig argument was added to KV connector
constructors. Backwards compatibility support for out-of-tree connectors was
included in this change, with a very prominent warning. See #25712 and #27887 .
Since the warning has been around for over 5 months, we can safely remove
the support of it.
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-09 23:39:46 +00:00
Wentao Ye and GitHub
f80aa53c9d
[Refactor] Nixl util using lazy init ( #41392 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-09 17:46:52 -04:00
Juhi Mittal and GitHub
7a2b596982
[Quantization] Add ModelOpt NVFP4 W4A16 (4-bit weights, fp16/bf16 activations) support ( #41769 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
2026-05-09 21:15:50 +00:00
Jiangyun Zhu and GitHub
2ee8c2a56e
[SpecDecoding] extend mtp support for mimo 2.5 ( #41905 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-05-09 18:22:59 +00:00
SoluMilken and GitHub
cd74911d92
[Model] use AutoWeightsLoader for DeepSeekV2 ( #41706 )
...
Signed-off-by: SoluMilken <ypiheyn.imm02g@g2.nctu.edu.tw >
2026-05-10 01:55:25 +08:00
SoluMilken and GitHub
25abddc1a5
[BugFix] Fix Gemma4 'layers.0.moe.experts.0.down_proj_packed' KeyError issue ( #40708 )
...
Signed-off-by: SoluMilken <ypiheyn.imm02g@g2.nctu.edu.tw >
2026-05-09 17:20:44 +00:00
171d59ae8d
[Bugfix][PD] Fix DSv4 Disaggregated ( #41957 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: ZhanqiuHu <zhu@redhat.com >
2026-05-09 16:48:24 +00:00
3dda9aeb54
[Bugfix] Remove nested torch.compile in GDN rearrange_mixed_qkv causing CUDA graph capture failure ( #42070 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-05-09 08:30:55 -07:00
Kermit and GitHub
adb6d96516
[Bugfix] Fix GDN KKT precision loss on Hopper GPUs by aligning tl.dot operand layout with WGMMA ( #42076 )
...
Signed-off-by: kermit <ckeming@outlook.com >
2026-05-09 13:08:46 +00:00
Thien Tran and GitHub
530d371302
[DSv4] Improved fused Indexer Q quant kernel ( #41428 )
2026-05-09 01:20:32 -07:00
Micah Williamson and GitHub
34ab4f2565
[ROCm] Upgrade aiter to v0.1.13-rc5 ( #42113 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-09 08:13:45 +00:00
Jee Jee Li and GitHub
ecd0b60aad
[LoRA] Initial EP support for LoRA ( #40867 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-09 00:31:23 -07:00
d6563d693c
Require C++20 for compatibility with PyTorch ( #40380 )
...
Signed-off-by: Richard Barnes <rbarnes@meta.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-08 22:04:43 -07:00
Rishapveer Singh and GitHub
f6490a2841
[Bugfix] Preserve leading/trailing whitespace in GLM non-streaming tool parser ( #42026 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-05-08 21:49:15 -07:00
a2812becd6
[Models] Cohere Eagle + fix to Cohere MoE ( #42078 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-08 21:46:26 -07:00
e8f9038ebd
[ROCm][Bugfix] Re-tag AITER MoE weights as preshuffled after replace_parameter ( #42061 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-08 21:42:07 -07:00
df2636a9d8
[Bugfix] Fix LOGITPROC_SOURCE_ENTRYPOINT test to use spawn-compatible dist-info registration for XPU/ROCm ( #42040 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
Signed-off-by: David Zheng <153074367+dzhengAP@users.noreply.github.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 12:32:04 +08:00
Shengqi Chen and GitHub
97cc7685c4
Add @Harry-Chen in CODEOWNERS ( #42130 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-09 04:08:22 +00:00
haosdent and GitHub
e934e459e6
[CI][Bugfix] Make test_gpt2_cache_hit observable across V1 EngineCore ( #42037 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-09 11:53:15 +08:00
David Zheng and GitHub
845ca327ce
[Bugfix] Fix test_whisper distributed test process handling ( #42038 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-09 11:37:21 +08:00
4f6fa6341d
[XPU] update supported models on XPU ( #41911 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 10:44:03 +08:00
Ethan Feng and GitHub
a43bc34baf
[Docs] Update server entrypoint examples ( #42077 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-09 02:03:52 +00:00
Ethan Feng and GitHub
236bf9d152
[Docs] Fix RLHF example links ( #42073 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-09 02:03:42 +00:00
Lucas Wilkinson and GitHub
b1728c1e66
[Attention][Cleanup] Remove tree attention ( #42121 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-05-08 18:36:19 -07:00
be0dcc29dc
[XPU] remove q/k/v force contiguous for flash_attn ( #40356 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-09 01:19:05 +00:00
Sumanth R Hegde and GitHub
e3b65a5ba0
[feat] Add explicit /start_weight_update and /finish_weight_update APIs for weight transfer ( #39212 )
2026-05-08 18:03:33 -07:00
Harry Mellor and GitHub
30f519e947
Use pre-commit / pre-run-check to gate docs build too ( #42053 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-09 00:02:51 +00:00
Roy Wang and GitHub
60851b1d22
[Bugfix][KV Transfer] Reject NixlConnector + expandable_segments:True ( #41237 )
2026-05-08 16:47:33 -07:00
Michael Goin and GitHub
8bcd8a260c
[Bugfix] Fix FlashInfer CUTLASS MXFP4-MXFP8 MoE by restoring swizzled scale ( #42089 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-08 15:59:06 -07:00
John Calderon and GitHub
8a2fc80b84
[CUDA][CUTLASS] Enable cutlass scaled mm for non-compatible sizes ( #41868 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
2026-05-08 15:58:05 -07:00
6881c754e1
use HIP_VERSION variables to guard against duplicate atomicAdd definitions ( #41802 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-08 18:44:37 -04:00
Kevin H. Luu and GitHub
0c2e9d4892
[CI] Narrow misc.yaml source dependencies ( #42059 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 15:10:12 -07:00
Kevin H. Luu and GitHub
d2f22dfc9f
[CI] Narrow engine.yaml source dependencies ( #42055 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:55:33 -07:00
Kevin H. Luu and GitHub
f4dd5c116c
[CI] Narrow Platform Tests (CUDA) source dependencies ( #42054 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:54:06 -07:00
Kevin H. Luu and GitHub
f47ccc8b1c
[CI] Narrow pytorch.yaml compile job source dependencies ( #42057 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-08 14:43:17 -07:00
dbd86a67e3
[Bugfix][Gemma4] Fix infinite loop and array boundary issues in tool parser ( #41991 )
...
Signed-off-by: David Oy <david.oy@baseten.co >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-08 17:24:37 -04:00
2c6b59b807
[ROCm][Perf] Add Fused Shared Expert (FSE) support for Qwen3-Next ( #39280 )
...
Signed-off-by: nholmber <nholmber@users.noreply.github.com >
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Doug Lehr <douglehr@amd.com >
Co-authored-by: nholmber <nholmber@users.noreply.github.com >
Co-authored-by: Tres <tpopp@users.noreply.github.com >
Co-authored-by: Tres Popp <tres.popp@amd.com >
Co-authored-by: Doug Lehr <douglehr@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-05-08 15:38:00 -04:00
44e6b44a21
[CI][Elastic EP] Fix Elastic EP Scaling Test Failure ( #41792 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-08 15:17:44 -04:00
Hiroaki Mikami and GitHub
90f145aaf7
[Models][Gemma3/Gemma4] Support hidden_act variants in gated MLP ( #40588 )
...
Signed-off-by: Hiroaki Mikami <hiroaki8270+github@gmail.com >
2026-05-08 11:29:11 -07:00
Ethan Feng and GitHub
4140faa4a5
[Docs] Fix OpenAI batch model argument examples ( #42066 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-08 14:02:46 +00:00
liuzhenwei and GitHub
f2bbd575e2
[CI][XPU] Skip fork-dependent logits processor test ( #42013 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-05-08 06:10:19 -07:00
haosdent and GitHub
52458b60a8
[CI][Examples][RLHF] Disable async scheduling in rlhf_async_new_apis ( #42042 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 04:58:48 -07:00
Harry Mellor and GitHub
630820a59b
Make docs environment deterministic ( #41926 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-08 10:13:03 +00:00
Chaojun Zhang and GitHub
19df11f5d1
[CI][XPU]Ignore some lora tests from LoRA Intel CI pipeline ( #42010 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
2026-05-08 17:34:27 +08:00
haosdent and GitHub
36b2c79d4b
[CI][Bugfix] Drop duplicated examples/ prefix in tensorize_vllm_model command ( #42039 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 02:23:22 -07:00
haosdent and GitHub
160858cba4
[CI][Bugfix] Surface subprocess output in spawn_new_process_for_each_test ( #41943 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-08 10:39:37 +02:00
Simon Danielsson and GitHub
f9b9bf3bbb
[CI][ROCm] Ship RIXL with vllm/vllm-openai-rocm ( #41634 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-08 07:05:17 +00:00
445d747434
[Bugifx] Missing Renderer for fastokens mode ( #41984 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-07 23:45:14 -07:00
wang.yuqi and GitHub
77b13b9602
[Docs] Reorganize examples docs. ( #41082 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-07 23:23:44 -07:00
ed582b6a4c
[Aiter][ROCm] gdn_linear_attn kernel fusion ( #40711 )
...
Signed-off-by: Tres Popp <tres.popp@amd.com >
Signed-off-by: Chuan Li <chuali@amd.com >
Co-authored-by: hellozhuo <zhuo.su@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-07 23:11:37 -07:00
David Zheng and GitHub
1acd67a795
[Bugfix] Fix XPU/ROCm compatibility in spawn_new_process_for_each_test ( #41895 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-08 00:47:22 -04:00
0b99971352
[Kernel][Helion] Optimize Helion config parsing latency ( #40850 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-05-07 20:27:34 -07:00
baf068d8be
enable persistent mla for sparse mla backend ( #41990 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
2026-05-07 20:10:50 -07:00
SamareshSingh and GitHub
01b0f3adab
fix: default TILELANG_CLEANUP_TEMP_FILES=1 to avoid shared /tmp conflicts ( #41486 )
...
Signed-off-by: Samaresh Kumar Singh <ssam3003@gmail.com >
2026-05-07 19:59:00 -07:00
Nick Hill and GitHub
989c176c0a
[Perf][3/n] Eliminate GPU<->CPU syncs in attention impls ( #41434 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-07 19:44:24 -07:00
cd58e30872
[Perf] Use numpy zero-copy path for embedding float response serialization ( #41681 )
...
Signed-off-by: Shrinav Loka <lokashrinav@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 19:42:21 -07:00
1d694e78c9
[Examples][last/6] Resettle examples. ( #41084 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-07 19:42:12 -07:00
haosdent and GitHub
57c2f724c1
[CI][Bugfix] Fix CI failures for "PyTorch Compilation Unit Tests" ( #41940 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-05-07 19:42:00 -07:00
5f6a02812a
[CI][Bugfix] Fix failure CI step "PyTorch Fullgraph Smoke Test" ( #41953 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-07 19:41:56 -07:00
50f2db2555
add: LFM2/2.5 Tool Parser ( #39243 )
...
Signed-off-by: Jonathan Buchanan <jonathan.buchanan@liquid.ai >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-08 09:58:17 +08:00
09a7cc5ba9
[KV Connector] Opt DecodeBenchConnector into SupportsHMA ( #41770 )
...
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-07 16:10:00 -07:00
Nick Hill and GitHub
10ebb40d62
[Core] Avoid using extra thread in UniProcExecutor ( #40891 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-07 15:33:00 -07:00
54f548e9e5
[Bugfix] Restore moe_forward output shape invariant on TRTLLM MXFP4 path ( #41646 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-07 15:26:06 -07:00
Kyle Sayers and GitHub
c1819ca283
[Compressed Tensors] Allow configs with non-explicit ignores ( #41965 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
2026-05-07 14:03:45 -07:00
969fbfb4a9
Laguna xs dflash support ( #41880 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-07 13:31:16 -07:00
akii96 and GitHub
3af561ec0a
[ROCm] Fix AITER AR+RMSNorm no-residual fusion ( #41972 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-07 13:14:58 -07:00
akii96 and GitHub
c936548ce6
[ROCm][DeepSeek] Enable V3.2 TP4 AITER MLA ( #41835 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-07 15:10:57 -05:00
TomerBN-Nvidia and GitHub
8189a15914
[Core] Replace routing replay with device cache and async D2H pipeline ( #39917 )
...
Signed-off-by: Tomer Barnatan <tbarnatan@nvidia.com >
2026-05-07 11:24:56 -07:00
Flora Feng and GitHub
8eb401134e
[Refactor] Consolidate required/named tool_choice streaming into DelegatingParser ( #41876 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-07 09:50:59 -07:00
Nicolò Lucchesi and GitHub
9d6500b89d
[Misc] Delay EPLB Nixl import until needed ( #41805 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-07 09:43:07 -07:00
zhrrr and GitHub
7a08b34fbf
[Model Runner V2] support qwen35 / mamba hybrid model ( #35520 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-05-07 09:31:05 -07:00
noobHappylife and GitHub
06a60d3dd0
Fix spec decode benchmark metrics ( #41916 )
...
Signed-off-by: noobhappylife <aratar1991@hotmail.com >
2026-05-07 09:23:21 -07:00
2a16ece2d3
tokenizer: Add fastokens support ( #41741 )
...
Signed-off-by: AlonKejzman <alonkeizman@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-07 22:49:42 +08:00
Andreas Karatzas and GitHub
003159d98b
[ROCm][CI] Avoid duplicate ROCm AITER norm-quant patterns ( #41534 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-07 06:33:30 -07:00
2a84da3b17
[XPU] Implement out-of-place all-reduce functionality ( #41808 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 05:58:05 -07:00
Chaojun Zhang and GitHub
805e9f7b77
[XPU] Fix lora bugs & enable UTs under tests/lora ( #38206 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
2026-05-07 05:58:00 -07:00
s-yanev and GitHub
75f0d516c4
[Bugfix] Fix GLM4-MoE weight loading for NVFP4 quantized checkpoints ( #41755 )
...
Signed-off-by: Stoyan Yanev <stoyan.yanev@cleverpine.com >
2026-05-07 05:55:52 -07:00
f650ace6de
[MM][Gemma4] Use video profiling hints in encoder budget ( #41837 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
2026-05-07 05:46:04 -07:00
Li, Jiang and GitHub
b3945cc316
[CPU] Bump up to the latest CPU kernels ( #41924 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-07 05:45:59 -07:00
ffee741626
[Model] Use AutoWeightsLoader for AXK1 ( #41901 )
...
Signed-off-by: liwenyi <lwy.lwy@163.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-07 05:40:29 -07:00
Jee Jee Li and GitHub
9c0812ffd0
[Bugfix] Fix FusedMoEWithLoRA has no attribute runner ( #41889 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-07 04:53:14 -07:00
Fadi Arafeh and GitHub
b20731d0ae
[CI][Arm] skip e2e model tests if HF_TOKEN is not set ( #41919 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-05-07 11:31:50 +00:00
d4b0048404
Eliminate redundant MoE buffer copies in AITER fused experts (without dependency on AITER changes) ( #41713 )
...
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-07 03:46:40 -07:00
6e6d182d18
[Bugfix] Fix OOM in tensorizer LoRA deserialization ( #41845 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-07 02:17:47 -07:00
tej and GitHub
8a4888be21
[ROCm] Profiler api support for ROCm MORI toy proxy server in PD Disaggregation ( #40264 )
...
Signed-off-by: Tej Kiran <kiran.tej@amd.com >
2026-05-07 16:58:38 +08:00
Yuwen Zhou and GitHub
713b28bd0b
[CPU] Add FP8 W8A16 MoE support ( #41314 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-05-06 23:17:07 -07:00
51f22dcfd0
[Feat][CPU] Enable Gated DeltaNet Attention (Qwen 3.5 / 3.6) ( #41025 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-07 12:57:09 +08:00
20cac26b19
[ROCm] Enable SimpleCPUOffloadConnector on ROCm ( #40549 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-06 20:52:02 -07:00
Russell Bryant and GitHub
5a0a8fc1ea
[Docs] add cache directory security guidance ( #38920 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
2026-05-06 16:54:29 -07:00
Micah Williamson and GitHub
7a576e2c72
[ROCm][CI] Remove TORCH_NCCL_BLOCKING_WAIT=1 After Bugfix In ROCm 7.2 ( #41840 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-06 16:37:11 -07:00
Yongye Zhu and GitHub
80d5e7d103
[Bugfix] Fix condition to clear persistent topk so that it can be captured regardless ( #41665 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 16:17:48 -07:00
95582868ef
[Bugfix] DeepSeekV32/v4: respect string='true|false' attribute andunwrap arguments/input wrapper ( #41801 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
2026-05-06 21:48:01 +00:00
50acdc5b5c
Fix Qwen3 streaming content routing ( #40820 )
...
Signed-off-by: xy3 <120182408@qq.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-06 17:22:01 -04:00
JackyLiu and GitHub
deb737e323
[Doc] Add ModernBertForSequenceClassification to scoring.md cross-en… ( #41832 )
...
Signed-off-by: JLiu4Coding <lzwgre@126.com >
2026-05-06 14:17:56 -07:00
Flora Feng and GitHub
f3f8efa73a
[CI] Enable gemma4 parser test on CI ( #41857 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-06 20:25:34 +00:00
Johnny Yang and GitHub
ca3e62d336
Upgrade tpu-inference to v0.19.0 ( #41844 )
...
Signed-off-by: Johnny Yang <johnnyyang@google.com >
2026-05-06 11:41:37 -07:00
Benjamin Chislett and GitHub
38e16678ba
[Bugfix] Align block table for TRTLLM MLA edge-case ( #39324 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-05-06 11:17:02 -07:00
27702f6d08
[Bugfix] Fix token loss in PP mode which causes degraded accuracy ( #41133 )
...
Signed-off-by: Jing Wang <jingwang96@qq.com >
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-06 14:07:32 -04:00
Divakar Verma and GitHub
22a3cbe152
[ROCm] aiter_unified_attn fp8 q scale refactor ( #38296 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-05-06 16:11:36 +00:00
Viktor Pus and GitHub
d5b31c954d
[Bugfix] Account for truncate_prompt_tokens when computing max_tokens ( #41800 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
2026-05-06 16:10:17 +00:00
David Zheng and GitHub
ee38750a75
[Bugfix] Fix spawn_new_process_for_each_test silently swallowing test failures ( #41423 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-06 11:17:15 -04:00
27e0057aed
[Spec Decode] Add Gemma4 MTP speculative decoding support ( #41745 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-06 22:39:29 +08:00
Ronen Schaffer and GitHub
f39bcf1e30
[KV Offload] Return None from lookup() for in-flight blocks ( #41795 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-06 17:31:21 +03:00
6467213a9f
fix(openai): tolerate empty content in forced tool choice ( #40148 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-06 07:16:03 -07:00
df8e63f4ed
nixl refactor: new transfer design ( #40731 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-06 06:16:25 -07:00
242afc6bf4
[MM][Gemma4] Respect max_soft_tokens in encoder budget ( #41799 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: gemini-code-assist <gemini-code-assist@google.com >
2026-05-06 05:54:42 -07:00
lyd1992 and GitHub
5d0fd87038
[CPU][RISC-V] Auto-bind OMP threads and harden nobind path ( #40569 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-05-06 11:38:08 +00:00
Harry Mellor and GitHub
d8deb5b7ad
Fix some legacy checkpoints with deprecated rope_type values ( #41734 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-06 11:13:12 +00:00
2e777d21a8
[Bugfix][Rocm]Aiter MoE re-uses existing tensor addresses after weight update. ( #40390 )
...
Signed-off-by: Yuankai Chen <yuankach@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-06 10:32:26 +00:00
Nicolò Lucchesi and GitHub
e43a791284
[Bugfix][CI] Fix Disaggregated test area path ( #41794 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-06 17:41:24 +08:00
66d1cc0c77
fix(rocm): remove workaround causing invalid argument on Qwen3.5 with TP=2 ( #40686 )
...
Co-authored-by: Test User <test@example.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-06 01:38:32 -07:00
1c58876618
[XPU] Disable CUDA graph memory estimate on XPU platform ( #41344 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-06 16:38:18 +08:00
51c1ee9b7c
[Examples] Resettle Disaggregated examples. ( #40759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-06 01:20:38 -07:00
Lucas Kabela and GitHub
213f10bfdd
[Bugfix] Fix codegen for unqualified names ( #40726 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-06 01:11:37 -07:00
e87e09a50a
[Feat] dnnl build for AVX2 W8A8 Int8 ( #41318 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-06 15:28:02 +08:00
Yuwen Zhou and GitHub
809b98e5b7
[CPU] Add FP8 W8A16 linear support ( #41186 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-05-06 07:05:27 +00:00
wi-adam and GitHub
b53c507bc9
[Bugfix] Skip PP sampled-token receive on last rank during async scheduling ( #40749 )
...
Signed-off-by: Adam Winstanley <adam@winstanley.industries >
2026-05-06 05:31:14 +00:00
2d7d6cf765
[Spec Decode] Allow multimodal models with a warning ( #41752 )
...
Signed-off-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Li Zhang <lzhanga@amazon.com >
2026-05-05 22:16:44 -07:00
Andreas Karatzas and GitHub
91740ca5ea
[ROCm][CI] Refine gating tests ( #37243 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-05 22:05:20 -07:00
e47c98ef7a
[Fix] Add missing stubs from cpu fp8 attention changes ( #41387 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-06 12:16:27 +08:00
aee190ac37
[Build] Fall back to system libgomp when torch has no vendored copy ( #40575 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-06 11:42:03 +08:00
16e336491e
[Mistral Tokenizer] allow more leniency in apply_chat_template ( #41658 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-05 19:56:15 -07:00
Chauncey and GitHub
c7aa186d67
[Frontend] Supports resubmitting output items with missing fields in Responses API ( #41355 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-05 22:21:33 -04:00
f653761252
[CI] Route part of B200 jobs to b200-k8s ( #41453 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-05-05 19:00:30 -07:00
Andreas Karatzas and GitHub
4a8ae26e53
[ROCm][CI] Use vLLM generation defaults for DeepSeek prefetch-offload eval ( #41575 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-06 01:08:12 +00:00
Kevin H. Luu and GitHub
1333864408
[CI] Automate Docker Hub release image publishing ( #40415 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-06 00:15:23 +00:00
Matthew Bonanni and GitHub
01b9b5af67
[Attention] Minor refactor: layer takes ownership of the MLA prefill backend ( #41744 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-05 23:22:41 +00:00
8c57b6e7bc
Bump model-hosting-container-standards to >= 0.1.14 ( #39755 )
...
Signed-off-by: EC2 Default User <ec2-user@ip-172-31-20-13.us-west-2.compute.internal >
Co-authored-by: EC2 Default User <ec2-user@ip-172-31-20-13.us-west-2.compute.internal >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-05 19:09:57 -04:00
Lanze Liu and GitHub
79246b5ea6
[Spec Decode] Fix max_model_len logging in speculative config for draft model ( #41571 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-05-05 21:56:06 +00:00
48954de237
Fix DeepGEMM ep_scatter output address overflow ( #39213 )
...
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-05 18:56:56 +00:00
Julien Denize and GitHub
c6235ed180
[BUGFIX] Support streamed_args_for_tool in MistralToolParser ( #41730 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-05-05 17:48:53 +00:00
628c436301
[New Model][ROCm] Add AMD support for DeepSeek V4 ( #40871 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: tjtanaavllm <tunjian.tan@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaavllm <tunjian.tan@amd.com >
2026-05-05 08:55:37 -07:00
Canlin Guo and GitHub
2228fe6868
[Attention] Move FA3→FA4 upgrade into get_flash_attn_version() ( #40815 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-05-05 15:43:03 +00:00
Harry Mellor and GitHub
84bd8a3c1e
Remove unnecessary runtime asserts from linear layers ( #41729 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-05 14:42:56 +00:00
Lidang Jiang and GitHub
b786ec8e74
[Bugfix] Suggest upgrading Transformers for tokenizer class errors ( #38099 )
...
Signed-off-by: Lidang-Jiang <lidangjiang@gmail.com >
2026-05-05 14:10:45 +00:00
20dcd984f9
[Bugfix] Fix RuntimeError: Already borrowed by adding thread-safe Hugging Face fast-tokenizer wrappers ( #41181 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-05 14:04:01 +00:00
Martin Hickey and GitHub
6fca518157
[BugFix][MyPy]: Module has no attribute "sched_getaffinity" [attr-defined] ( #41465 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-05 13:20:37 +00:00
98661fe012
[Bugfix][KVConnector] Support DCP/PCP in OffloadingConnector ( #41549 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-05-05 14:54:29 +03:00
Harry Mellor and GitHub
b0765bee17
Fix DeepSeek-OCR for Transformers v4 ( #41460 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-05 11:11:21 +00:00
0a201b60cf
[Model] support Qianfan-OCR model ( #40136 )
...
Signed-off-by: bairongz <baiyuu.cs@gmail.com >
Signed-off-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-05 10:51:25 +00:00
8b9ea2f881
[Feature] Add Triton kernel JIT compilation monitor for inference ( #40137 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-05-05 14:08:57 +04:00
Kunshang Ji and GitHub
2ceea42958
[XPU] use xpu topk topp sample kernel ( #39285 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-05 18:05:17 +08:00
bee126165f
[P/D][Mooncake] Add KVConnectorStats for transfer observability ( #40414 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-05 02:17:38 -07:00
BitToby and GitHub
27cc676be3
[Model] Use AutoWeightsLoader for Plamo2 ( #41699 )
...
Signed-off-by: bittoby <218712309+bittoby@users.noreply.github.com >
2026-05-05 08:56:24 +00:00
4845aee6b7
[Benchmark] Add --trust-remote-code flag to multi-turn benchmark ( #41661 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-05 01:00:37 -07:00
BitToby and GitHub
0c620d2e08
[Model] Use AutoWeightsLoader for CohereMoe ( #41690 )
...
Signed-off-by: bittoby <218712309+bittoby@users.noreply.github.com >
2026-05-05 04:44:15 +00:00
6bb924bbf3
[Model] Fix Gemma4 MoE activation mismatch ( #41574 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-05 04:34:11 +00:00
czhu-cohere and GitHub
eaec7be446
[BugFix] Preserve max_seq_len in ubatch metadata during CUDA graph capture ( #40961 )
...
Signed-off-by: root <conway.zhu@cohere.com >
Signed-off-by: <conway.zhu@cohere.com >
2026-05-05 04:29:34 +00:00
Jeffrey Wang and GitHub
f04fd1677b
[Ray] Enable RayExecutorV2 by default ( #41421 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
2026-05-05 04:27:34 +00:00
420b0a5c95
[Hardware][Power]Add Power VSX Attention Backend and fix l2 Cache Crash ( #40451 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-04 20:51:09 -07:00
Bowen Bao and GitHub
1e9500410a
[ROCm][Quantization][2/N] Refactor quark_moe w4a8 w/ oracle ( #39136 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-04 19:50:38 -07:00
Nick Hill and GitHub
416f9cdede
[Perf][2/n] Eliminate GPU<->CPU syncs in pooling code ( #41433 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-05 02:43:25 +00:00
685bf811d6
[XPU] enable is_act_and_mul for xpu ( #37481 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-05 01:07:39 +00:00
Giancarlo Delfin and GitHub
e1e4646b06
[Model Runner V2] Rebuild attn metadata between draft decode steps ( #41162 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-05-05 00:44:55 +00:00
4f2af1a7c0
[Feature] TurboQuant: support hybrid models and uniform quantization ( #39931 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Sandermage <sandermage@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-04 20:14:01 -04:00
Wentao Ye and GitHub
577b9623e6
[Bug] Fix status update address for non-MOE model within external dp mode ( #40839 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-04 16:37:16 -07:00
Andreas Karatzas and GitHub
1cb0838721
[ROCm][CI] Fix MLA prefill scale for DeepSeek GSM8K ( #41569 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 16:32:55 -07:00
Matthew Bonanni and GitHub
be5983b874
[Docs] Add non-causal support to attention backend docs ( #41643 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-04 20:35:15 +00:00
fxmarty-amd and GitHub
9c07342fdc
[NVFP4][fix] Fix layer.weight -> w13 typo in NVFP4 MOE emulation kernel preparation ( #41630 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-05-04 20:13:37 +00:00
844df54269
feat: update xgrammar==0.2.0 to use structural tags for strict tool calling + reasoning for more models ( #40894 )
...
Signed-off-by: Yuchuan <yuchuan.7streams@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Ubospica <ubospica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Ubospica <ubospica@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-04 12:45:24 -07:00
422dd02598
[bugfix] Fix prompt logprobs on request eviction during chunked prefill ( #41411 )
...
Signed-off-by: Joachim Studnia <joachim@mistral.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-04 11:46:00 -07:00
8c780943b4
Fix Nano Nemotron text-only weight loading ( #41205 )
...
Signed-off-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Signed-off-by: Baekpica <35071468+Baekpica@users.noreply.github.com >
Signed-off-by: sunghoon.baek <seanbb93@gmail.com >
Co-authored-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-04 21:43:07 +03:00
e724b0ea8d
[ROCm] ROCm7.2.2 + profiler fix + AITER 0.1.12.post2 ( #41386 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-05-04 13:07:19 -05:00
712ad0286c
[Bugfix] KimiK2ReasoningParser: guard against buffered end-token in streaming ( #41068 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-04 17:42:05 +00:00
Ekagra Ranjan and GitHub
321fa2d6d1
Limit gpu utils and lower max BS on test_transcription_api_correctness.py ( #41649 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-05-04 10:30:02 -07:00
Wentao Ye and GitHub
3e1ad4435f
[Bug] Fix tests/compile/test_config.py AttributeError: 'NoneType' object has no attribute 'dtype' ( #41288 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-05 00:22:07 +08:00
Netanel Haber and GitHub
8decbfa02c
Test nemotron nano-v2 and nemotron nano-v3 separately, disable super-omni redundant tests ( #41616 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-04 16:31:37 +03:00
Stefano Castagnetta and GitHub
62ba7516e8
Revert "[Doc] Fix RTD build: pytorch.org/docs/stable/objects.inv returns 404" ( #41618 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-05-04 04:47:42 -07:00
Stefano Castagnetta and GitHub
6f53753fc9
[Bugfix] Apply ruff-format to hyperclovax.py ( #41620 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-05-04 03:37:16 -07:00
Andreas Karatzas and GitHub
6ec9bbec38
[CI] Stabilize cpu offload compressed tensors test ( #41102 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 05:22:42 +00:00
Andreas Karatzas and GitHub
01d4d1ad37
[ROCm][CI] Align spec decode logprob test prefill settings ( #41335 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 04:33:29 +00:00
c103c02a1a
[Transformers v5] Vendor HCXVisionConfig for compatibility ( #38447 )
...
Signed-off-by: Fang Han <fhan0520@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-04 04:19:52 +00:00
Andreas Karatzas and GitHub
67058ca326
[CI] Clean up remote servers on pytest parent exit ( #41570 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 03:11:22 +00:00
Akim Tsvigun and GitHub
894a02500b
[Bench] Forward --seed to CustomDataset and CustomMMDataset shuffle ( #40788 )
...
Signed-off-by: akimtsvigun <akimtsvigun@gmail.com >
2026-05-04 00:39:10 +00:00
66dfee7121
[Bugfix] Fix degenerate KV cache stride causing TMA cudaErrorIllegalInstruction ( #40737 )
...
Signed-off-by: David Oy <david@baseten.co >
Signed-off-by: David Oy <58150256+the-david-oy@users.noreply.github.com >
Signed-off-by: David Oy <david.oy@baseten.co >
Co-authored-by: David Oy <david@baseten.co >
Co-authored-by: Claude <claude@anthropic.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-03 23:52:18 +00:00
Alex Brooks and GitHub
db9a84e0cd
[Bugfix] Fix FP8 Bias Loading ( #41424 )
...
Signed-off-by: Alex Brooks <albrooks@redhat.com >
2026-05-03 20:30:04 +00:00
tomeras91 and GitHub
cb03fee32b
[Bugfix][Ray] Fix RayExecutorV2 actor name collision with DP > 1 ( #40398 )
...
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com >
2026-05-03 13:00:41 -07:00
Wei Zhao and GitHub
c51df43005
Disable flashinfer autotune temporarily due to correctness issues ( #41524 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-03 16:19:59 +00:00
Taneem Ibrahim and GitHub
54dc64d5d3
[Doc] Add Qwen3-30B-A3B-Thinking-2507-FP8 to batch invariance verified models ( #41513 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-05-03 08:47:55 -04:00
Woosuk Kwon and GitHub
e6ff3e9c83
[MRV2] Add shutdown() method ( #41297 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-02 21:06:30 -07:00
Jinzhen Lin and GitHub
08834cc3ce
[Quantization] add humming mxfp4 moe backend ( #41083 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
2026-05-02 18:36:03 -07:00
856ec4804a
[DSv4] Tune default value of VLLM_MULTI_STREAM_GEMM_TOKEN_THRESHOLD ( #41526 )
...
Co-authored-by: Copilot <copilot@github.com >
2026-05-02 18:32:09 -07:00
Yongye Zhu and GitHub
1c607d7b2c
[DSV4] Guard megamoe flag with Pure TP ( #41522 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-02 16:41:40 -07:00
4f7309fcc0
[CI] Add ci-fetch-log.sh helper for Buildkite job logs ( #41517 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-02 15:23:59 -07:00
Michael Goin and GitHub
0a9362d6ab
Revert "[Build] Make bundled DeepGEMM wheel portable across Python versions" ( #41512 )
2026-05-02 09:42:41 -07:00
cfd2573f23
[Build] Switch CUDA 13.0 wheel builds to PyTorch manylinux_2_28 base ( #41416 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-02 05:51:28 -07:00
c3ad791e1a
[Bugfix][Gemma 4] Clamp soft-token estimate to max_soft_tokens ( #40796 )
...
Signed-off-by: Hoang Nguyen <118159510+hnt2601@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-02 06:34:59 +00:00
Matthew Santiago and GitHub
8586369f61
Refactor Step3Text loading to use AutoWeightsLoader ( #41492 )
...
Signed-off-by: Matthew Santiago <carag.matthew@gmail.com >
2026-05-02 06:22:14 +00:00
Chauncey and GitHub
ae3b4deb8a
[Doc] Add Codex usage example ( #41358 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-01 22:27:43 -07:00
Rita Brugarolas and GitHub
c293ccc58e
[ROCm][Bugfix] Fix init-time bias dtype cast when gate.out_dtype is None ( #41405 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-05-02 00:13:15 -04:00
Luka Govedič and GitHub
d58c42e19c
[vLLM IR] 2/N fused_add_rms_norm and maybe_inplace overload ( #36823 )
...
Signed-off-by: Luka Govedič <lgovedic@redhat.com >
Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-05-01 23:41:15 -04:00
Ekagra Ranjan and GitHub
3e49479c4b
Limit concurrency on test_transcription_api_correctness.py ( #41478 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-05-02 03:19:07 +00:00
John Calderon and GitHub
964a4bc2a5
[MM][CG] Support ViT CG for Qwen2.5-VL ( #40830 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
2026-05-02 11:10:14 +08:00
FredericOdermatt and GitHub
c408fdd663
[Fix] Sync gemma4 chat template from hf ( #39570 )
...
Signed-off-by: Frederic Odermatt <frederic.odermatt@44ai.ch >
2026-05-02 03:06:54 +00:00
Andy Lo and GitHub
5737770c6c
Re-enable allreduce rms fusion for DP / PP ( #41458 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-05-01 19:01:37 -04:00
Michael Goin and GitHub
0c99629ede
[Build] Make bundled DeepGEMM wheel portable across Python versions ( #41476 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-01 14:45:03 -07:00
Yongye Zhu and GitHub
edd60ac93a
[Bugfix] Fix persistent_topk inter-CTA init race on RadixRowState ( #41444 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-01 14:42:52 -07:00
Yongye Zhu and GitHub
bcf5cac9fb
[DSV4] Add knob to enable pre-attn gemm ( #41443 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-01 12:23:17 -07:00
a9484dac7b
[Perf] Intergrate Tile Kernels head_compute_mix_kernel for Deepseek-V4 ( #41255 )
...
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-01 15:01:17 -04:00
f3fef12350
[Attention] Abstract the MLA prefill backends and eliminate cuDNN ( #32623 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 13:36:20 -04:00
51295793a2
[Model Runner V2] Add logprob_token_ids support ( #40559 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-01 10:02:03 -07:00
Michael Goin and GitHub
3ccc1ff495
[Eval][CI] Add basic mrcr eval to tests/evals/ ( #40164 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-01 12:00:38 -04:00
529c671e80
[ROCm][FEAT] AITER Fused Allreduce + RMSNorm ( #37646 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
Signed-off-by: junkang1991 <junkangchow@gmail.com >
Co-authored-by: Rita Brugarolas <Rita.BrugarolasBrufau@amd.com >
Co-authored-by: junkang1991 <junkangchow@gmail.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-01 23:07:18 +08:00
bc635fad23
[ROCm][Deepseek] dsv3.2 further optimization ( #41217 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-05-01 23:06:00 +09:00
c3e64696cd
[Perf] Warmup forward_native sampler kernel ( #41375 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-01 18:04:11 +04:00
sungsoo ha and GitHub
4f7bde572a
[Kernel] Pack output and LSE in DCP A2A ( #41160 )
2026-05-01 09:01:17 -04:00
Or Ozeri and GitHub
2fa1f8ec00
[kv_offload+HMA][13/N]: Enable HMA support ( #41445 )
...
This is the final PR in a series to enables HMA support for the
offloading connector. The connector advertises `SupportsHMA`
and is validated with unit tests and e2e tests.
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-01 12:30:03 +01:00
raviguptaamd and GitHub
7075df79b3
[ROCm] Enable DBO (Dynamic Batch Optimization) on ROCm ( #34726 )
...
Signed-off-by: raviguptaamd <ravi.gupta@amd.com >
2026-05-01 09:18:30 +00:00
Yuyi Ao and GitHub
0dbaf9daad
Refractor longcat loading to use AutoWeightsLoader ( #41448 )
...
Signed-off-by: George-ao <yuyiao772@gmail.com >
2026-05-01 09:07:23 +00:00
a3ec4a35f5
[Bugfix][Metrics] Fix RayPrometheusMetric.labels() returning shared labeled child ( #40840 )
...
When vLLM runs with Ray Prometheus `vllm:request_success{finished_reason=...}`
only ever increments the repetition bucket regardless of the request's actual finish
reason; stop, length, abort, and error stay at zero. Root cause was `labels()` mutated
the wrapped Ray metric's default tags in place and returned self, so every `.labels(...)`
call on a given wrapper returned the same object.
Co-authored-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Signed-off-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
2026-05-01 08:43:39 +01:00
Andreas Karatzas and GitHub
32964e7700
[ROCm][CI] Upgraded UCX and RIXL ( #41210 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-01 16:40:47 +09:00
a07642667d
[Bugfix] Pass reasoning parser kwargs to structured output ( #41199 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-30 23:38:02 -07:00
baonudesifeizhai and GitHub
c3868bbbe4
[compile] Add FlashInfer FP8 async TP fusion and preserve allreduce fusion ordering #27893 ( #39505 )
...
Signed-off-by: baonudesifeizhai <baonudesifeizhai@gmail.com >
Signed-off-by: baonudesifeizhai <85092850+baonudesifeizhai@users.noreply.github.com >
Signed-off-by: roG0d <baonudesifeizhai@gmail.com >
2026-05-01 05:08:34 +00:00
sychen52 and GitHub
947138b6c2
Add nvfp4 kv cache support ( #40177 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-01 04:55:16 +00:00
Or Ozeri and GitHub
941fb50835
[kv_offload+HMA][12/N]: Scheduler-side support for sliding window groups ( #41228 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-01 06:59:17 +03:00
6b6ac6c3c7
[Kernel][MoE] Support GELU on TRT-LLM NvFP4 fused MoE for Gemma4 ( #41050 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 03:37:43 +00:00
Stefano Castagnetta and GitHub
b542bdf7fb
[Bugfix] Disable FlashInfer CUTLASS MoE on SM110 (Jetson Thor AGX) ( #40808 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-30 20:08:49 -07:00
Ronen Schaffer and GitHub
415a879899
[KV Offload] Use Collection instead of Sequence/Iterable for OffloadingManager key parameters ( #41361 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-01 05:18:38 +03:00
Dong W and GitHub
7198940b39
[Model] Add Moondream3 model support(only query and caption skills) ( #32325 )
...
Signed-off-by: Dong Wang <dongw2019@gmail.com >
2026-05-01 10:06:48 +08:00
14043dfecd
feat: Enable prompt_embeds Content Part Support in vLLM Chat Completions API ( #40720 )
...
Signed-off-by: Luis Robaina <luis@protopia.ai >
Signed-off-by: Luis Robaina 🚀 <luisfabian1545@gmail.com >
Signed-off-by: LuisRobaina <luis@protopia.ai >
Co-authored-by: Andrew Sansom <qthequartermasterman@gmail.com >
2026-05-01 10:05:55 +08:00
Andreas Karatzas and GitHub
1adaa5056b
[ROCm][CI] Add ROCm score absolute tolerance floor ( #41341 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-30 18:59:35 -07:00
4d5c89295b
(bugfix): block_size check for flex attn ( #41363 )
...
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-30 18:59:26 -07:00
Nick Hill and GitHub
dd5506a157
[Core] Simplify handling of scheduler_reserve_full_isl option ( #41064 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-30 18:10:00 -07:00
a3c83ff2fd
Faster per-token fp8 group quant packed kernel for blackwell ( #41326 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-30 18:09:55 -07:00
Woosuk Kwon and GitHub
9c61864bf8
[DeepSeek] Use torch.mm for bf16xbf16->fp32 gemm ( #41300 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-04-30 16:28:57 -07:00
Tran Le and GitHub
71725f6730
[Bugfix] Fix RoutedExpertsCapturer for Gemma 4 MoE (top_k_experts) ( #41401 )
...
Signed-off-by: Tran Le <tranle@fireworks.ai >
2026-04-30 16:19:59 -07:00
b4806c8ee1
[DSV4] Add BF16 and MXFP8 A2A support for flashinfer a2a one sided ( #40960 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com >
Co-authored-by: Zijing Liu <liuzijing2014@users.noreply.github.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-04-30 15:33:12 -07:00
Wentao Ye and GitHub
526927be94
[Model Runner v2] Fix v2 compile counter num_gpu_runner_capture_triggers and num_cudagraph_captured ( #41285 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-30 15:20:11 -07:00
Michael Goin and GitHub
75a4c166f2
Fix typo in log message for indexer cache ( #41419 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-30 15:02:14 -07:00
2917d6363a
[NVFP4][Hopper/AMD Instinct] Add Triton kernels for NVFP4 dequantization and QDQ emulation ( #40033 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-30 17:35:48 -04:00
Stefano Castagnetta and GitHub
efb4cdf2b8
[CI/Build] Skip Prithvi/Terratorch model-registry tests when terratorch is missing ( #41389 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-30 12:47:55 -07:00
92a7c121b6
[CI] Add MTP coverage: Qwen3.5 correctness + no-sync spec decode ( #40472 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-30 12:24:09 -07:00
Jee Jee Li and GitHub
307b17ce33
[DSV4] Avoid redundant dtype conversion. ( #41374 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-30 09:57:27 -07:00
3ca6ca210f
xpu docker: pin oneAPI to 2025.3 and avoid unintended 2026 upgrade ( #41380 )
...
Signed-off-by: wendyliu235 <wenjun.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-30 16:02:23 +00:00
Stefano Castagnetta and GitHub
10558f5f46
[CI/Build] Skip terratorch + torchgeo while PyPI has lightning quarantined ( #41377 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-30 07:59:07 -07:00
121dbe7a22
[ROCm] ROCm DeepEP API updated to latest ( #39721 )
...
Signed-off-by: Tej Kiran <vpolamre@amd.com >
Signed-off-by: tej <37236721+itej89@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: HAIAI <39548240+HAIAI@users.noreply.github.com >
2026-04-30 07:46:59 -07:00
Matthew Bonanni and GitHub
f03d82efdd
[UX][Bugfix] Fix OOM by setting PyTorch max_split_size_mb during model loading ( #41268 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-30 07:46:54 -07:00
a7fb008510
[EPLB] Optimize memory overhead in Nixl communicator ( #40013 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-30 07:46:49 -07:00
Harry Mellor and GitHub
ff449b6426
Stop mergify labelling from skipping pre-commit ( #41362 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-30 05:48:38 -07:00
3527229517
[Doc] Fix RTD build: pytorch.org/docs/stable/objects.inv returns 404 ( #41353 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-30 05:06:44 -07:00
b55b26520c
[MoE] Make MoERunnerInterface a PluggableLayer for OOT support ( #35178 )
...
Signed-off-by: wxsIcey <1790571317@qq.com >
Signed-off-by: Icey <1790571317@qq.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-30 03:31:08 -07:00
snadampal and GitHub
3179e53135
[P/D] Prefill compute optimizations with bi-directional KV cache transfers between P and D nodes ( #32553 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
2026-04-30 10:14:20 +00:00
Nicolò Lucchesi and GitHub
efdc95674d
[KVConnector] MultiConnector SupportsHMA ( #39571 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-30 02:10:50 -07:00
54146a9bf9
[Bugfix] correct h matrix layout in chunk_kda output kernel ( #40956 )
...
Signed-off-by: ChenxiQian <chenxi.qian.cq@outlook.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-30 16:22:41 +08:00
ca97f7b9bb
Fix Gemma4 MoE expert weight remapping ( #41206 )
...
Signed-off-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Co-authored-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-30 00:12:42 -07:00
Ekagra Ranjan and GitHub
a04e0cf3b8
Fix Cohere ASR after HF upgrade ( #40582 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-04-29 23:39:04 -07:00
cb1b02d0e8
[Frontend] Add VLLM_SKIP_MODEL_NAME_VALIDATION environment variable ( #34676 )
...
Signed-off-by: Dhruv Singal <dhruvsingalabc@gmail.com >
Signed-off-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local >
Signed-off-by: Your Name <you@example.com >
Signed-off-by: vLLM Assistant <assistant@vllm.ai >
Signed-off-by: Simon Mo <simon.mo@hey.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local >
Co-authored-by: Your Name <you@example.com >
Co-authored-by: OpenCode <noreply@openai.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-04-29 23:19:09 -07:00
a749a33d8d
[Bugfix] Fix persistent_topk cooperative deadlock at TopK=1024 ( #41189 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-04-29 21:03:45 -07:00
c42981d034
[Refactor][kv_offload] KV Offloading maintainability improvements ( #40538 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-30 05:55:31 +03:00
Wei Zhao and GitHub
0ff1bf9bb1
[Bugfix] Fix failure to allocate KV blocks error ( #41282 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-04-29 18:44:07 -07:00
0ab67c0222
[CI] Add key field to all test_areas pipeline steps ( #41201 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-29 16:59:16 -07:00
Rohan Potdar and GitHub
3795d7acf4
[ROCm][Bugfix][GPTOSS]: fix input_ids and expert_map args for quark w4a8 gptoss ( #41165 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-29 16:39:01 -07:00
Nick Hill and GitHub
18599bfdf2
[Ci][BugFix] Fix slow DP tests due to bad teardown logic ( #41166 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-29 19:31:00 -04:00
Thien Tran and GitHub
296741d025
[DSv4] Use cvt PTX for FP32->FP4 conversion ( #41015 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-04-29 16:16:40 -07:00
a966aaed30
[Bugfix][MLA] Size arange_buffer to max_num_batched_tokens to prevent CUDA IMA ( #39277 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-04-29 16:14:50 -07:00
Hemanth Acharya and GitHub
6841f5dc77
[ROCm] Add env flags to disable dynamic MXFP4 quant and enable AITER tuned GEMMs for Attention Projection Layers ( #39987 )
...
Signed-off-by: Hemanth Acharya <heachary@amd.com >
2026-04-29 16:07:46 -07:00
roikoren755 and GitHub
c2fb013312
[Bugfix][Compile] Fix gc.collect/empty_cache patch arity in CUDAGraphWrapper ( #41235 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-29 21:59:18 +00:00
ccfb620c62
Create tests/distributed/test_mnnvl_alltoall.py ( #35241 )
...
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Claude <claude@anthropic.com >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Claude <claude@anthropic.com >
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-29 21:56:56 +00:00
0335316a9b
[BUG] Two phase pause to prevent deadlock ( #39366 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Signed-off-by: Aaron Hao <ahao@anyscale.com >
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-29 17:51:03 -04:00
Rohan Potdar and GitHub
944e138bcf
[ROCm][Bugfix]: W4A4 MOE using emulation instead of AITER on MXFP4-supported hardware ( #41175 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-04-29 16:39:03 -05:00
b58669cb42
[Perf][Spec Decode] Avoid per-step numpy allocation in prepare_next_t… ( #41043 )
...
Signed-off-by: wangluochao902 <wangluochao902@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-29 14:20:13 -07:00
Isotr0py and GitHub
1628239eb2
[Multimodal][Render] Skip mm processor initialization and warmup for text-only mode ( #41246 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-29 14:16:19 -07:00
yzong-rh and GitHub
93da1fe97a
[CI] Add temperature to bfcl eval, default greedy ( #41059 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-04-29 14:01:57 -07:00
Andrew Barnes and GitHub
169988a3c0
[ROCm] Use quant_dtype in per_token_quant instead of hardcoded FP8 ( #39121 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-29 20:46:01 +00:00
faab189554
[Feature]: IndexCache support for DSA models ( #37735 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-29 15:15:35 -04:00
Laith Sakka and GitHub
6f20f81cbf
Replace shape_invariants with simpler apprach in dynamic_arg_dims utilizing shape_id property. ( #36194 )
...
Signed-off-by: Laith Sakka <lsakka@meta.com >
2026-04-29 18:32:15 +00:00
danisereb and GitHub
d1a75e303d
Fix timeout when using LoRA adapters with Nemotron Super ( #40916 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-04-30 01:39:49 +08:00
Cyrus Leung and GitHub
4a42aba380
[CI/Build] Enable FP8 on NVIDIA Thor ( #39712 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-29 09:48:52 -07:00
Avshalom Manevich and GitHub
a80d6f150c
better logging for large uncachable items ( #41145 )
...
Signed-off-by: h-avsha <avshalom.manevich@hcompany.ai >
2026-04-29 09:48:47 -07:00
Terrence Zhao and GitHub
91a2d39014
[Models] Cohere MoE ( #40817 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-04-29 15:54:54 +00:00
Frederik Gossen and GitHub
a05848e255
[Bugfix] Report compile time for in-memory cache hit path ( #41023 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-29 15:32:03 +00:00
51fda1ba44
[Model Runner v2] Fix block table IMA issue ( #40648 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-29 08:30:33 -07:00
Wentao Ye and GitHub
39a7f4f4e2
[Perf] Optimize AllPool.forward by slicing first, 51% faster in the method level benchmark ( #41163 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-29 08:11:04 -07:00
Artem Perevedentsev and GitHub
b92ef9ec5a
[Perf] Enable FlashInfer top-k/top-p sampler by default ( #40376 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-04-29 19:10:34 +04:00
5560cac7e2
[Bugfix][CPU] Backport PT cpp codegen indirect_assert scalar-mask fix ( #40973 )
...
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-29 10:21:55 -04:00
5b39b268f5
hf_name argument for vllm bench throughput CLI ( #41012 )
...
Signed-off-by: Philip Maybank <pmaybank@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-29 12:57:58 +00:00
22524f7a92
[Feat] CPU fp8 attn for AMX/AVX-512 ( #39445 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-04-29 20:43:21 +08:00
9d8ad5b408
[Bugfix] Fix repeated DSv4 RoPE cache initialization ( #41148 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-29 20:29:55 +08:00
11b69129e2
[Frontend] Add defer_loading and tool_reference support for Anthropic and OpenAI APIs ( #40190 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-29 04:35:50 -07:00
Bugen Zhao and GitHub
33f36d4260
[DSV4] Support max reasoning effort ( #40982 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-04-29 11:03:47 +00:00
Ronen Schaffer and GitHub
37e288214b
[KV Offload] Tighten keys type from Iterable to Sequence in OffloadingManager ( #41200 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-04-29 13:50:42 +03:00
5371d6fb40
Fix PP in Gemma4 ( #40786 )
...
Signed-off-by: Rohit kumar Singh <rksingh@habana.ai >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-29 03:17:51 -07:00
Jiangyun Zhu and GitHub
6d7d4da99e
[Bugfix] BailingMoeV2.5: rotate full qk_rope_head_dim in MLA RoPE ( #41185 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-04-29 18:08:55 +08:00
3f1a4bb639
build: embed image provenance metadata in vLLM containers ( #40653 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-29 03:07:41 -07:00
Chauncey and GitHub
762022cafb
[Bugfix] DSV32/V4 add missing type conversion for non-streaming tool calls ( #41198 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 09:55:07 +00:00
Chauncey and GitHub
3885d340a4
[Frontend]Responses API supports Tool/Function calling with streaming with named tool/function ( #41110 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 09:11:27 +00:00
haosdent and GitHub
ef70057ca7
[CI][CPU] Split CPU-Distributed Tests into per-scenario labels ( #41203 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-29 01:28:45 -07:00
e48cb85185
[CI/Build] Auto-detect manylinux ABI tag for nightly wheels ( #41149 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-29 00:37:14 -07:00
Chauncey and GitHub
92879e12ba
[CI] fix test_rotary_embedding_opcheck format error ( #41202 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 00:32:37 -07:00
68dd7db810
[Reasoning] Support for speculative decoding with thinking budget ( #34668 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-29 06:14:52 +00:00
8a8c9b564e
[KV Offload] Per-job store completion for CPU offloading connector ( #39186 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-29 08:52:55 +03:00
Jee Jee Li and GitHub
a269744e9f
[Bugfix] Fix rope ( #41113 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-28 22:42:35 -07:00
8b49cf3a37
[Bugfix] Fix max_num_batched_token not captured in cuda graph ( #40734 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-bia02.bia.clusters.nvidia.com >
2026-04-28 21:33:06 -07:00
Jiangyun Zhu and GitHub
2ae73c758c
[Bugfix] fix inductor error for dpsk v4 ( #41135 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-04-28 21:18:46 -07:00
Fadi Arafeh and GitHub
d95d03c719
[BugFix][CPU] fix error on CPU runner shutdown ( #41034 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-04-28 21:08:35 -07:00
Wei Zhao and GitHub
803b9d7881
[Bugfix] Fix Deepseek V4 import error due to AOT compile cache loading ( #41090 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
2026-04-28 21:08:16 -07:00
Walter Beller-Morales and GitHub
1312f07531
[Feature] add cohere reasoning and tool parsers ( #40422 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-04-28 21:07:53 -07:00
fa1b9840f6
[BE][Torch 2.12] Remove workaround code for fixed cublas issue ( #40845 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
Signed-off-by: Lucas Kabela <lucasakabela@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-04-28 21:07:24 -07:00
916e56c05c
[QeRL] Add warnings for extra memory buffering ( #40309 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-28 21:06:54 -07:00
a085b5257d
[Docs] [QeRL] Layerwise Reloading Documentation ( #40317 )
...
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-28 21:06:38 -07:00
liangel-02 and GitHub
7fd05e05ae
uncomment flex backend for batch invariant mode ( #40842 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-04-28 21:05:14 -07:00
99255f3cb5
[UX] Allow enable/disable model weights loading tracking by config ( #41086 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Copilot <copilot@github.com >
2026-04-28 21:04:49 -07:00
haosdent and GitHub
75a7cf2c10
[CI] De-flake test_chat_completion_n_parameter_non_streaming ( #41147 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-29 03:23:59 +00:00
haosdent and GitHub
4b95e9cec4
[CI] Return HTTP 400 for unsupported chat content part type ( #41121 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-29 10:23:26 +08:00
rasmith and GitHub
856b15c62c
[CI][AMD][BugFix] Patch has_flashinfer decorator for test_select_rocm_aiter_backend ( #41072 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-29 02:12:17 +00:00
qizixi and GitHub
6fb3f7b46b
[DSV4] Align aux stream API with DeepseekV4DecoderLayer ( #41171 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-04-28 17:22:03 -07:00
chelnnexy and GitHub
d109eacd05
[Bugfix][ROCm] Fix gemm_a4w4 call to use updated AITER API signature ( #40754 )
...
Signed-off-by: cheiluno <cheiluno@amd.com >
2026-04-29 09:04:53 +09:00
Nick Hill and GitHub
e68fa1b90a
[Core] Account for num_gpu_blocks_override in max_model_len checks ( #41069 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-28 15:44:09 -07:00
Russell Bryant and GitHub
f05f3664c3
[Doc] Add missing API endpoints to security documentation ( #40532 )
...
Signed-off-by: Russell Bryant <rbryant@redhat.com >
2026-04-28 21:53:19 +00:00
Julien Denize and GitHub
e9f8f31e9a
[FEATURE] Add EagleMistralForCausalLM ( #41024 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-28 12:22:20 -07:00
de3fe8dc62
[Bugfix] release KV blocks for skipped P-ranks to prevent invalid KV errors and timeouts when P_tp > D_tp and MLA ( #40449 )
...
Signed-off-by: yangruize <yangruize7@163.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-28 11:38:43 -07:00
0899f436aa
[New Model] Laguna XS.2 implementation ( #41129 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-28 14:23:00 -04:00
rasmith and GitHub
358a755e43
[CI][AMD][BugFix] Update request URL in test_moriio_connector to match vllm-router compatibility changes ( #41076 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-28 13:14:59 -05:00
Benoit Tigeot and GitHub
a60883644b
[Build] Defer flashinfer cubin download to avoid ~2.5 GB (decompressed) layer duplication ( #41134 )
...
Signed-off-by: Benoit Tigeot <benoit.tigeot@lifen.fr >
2026-04-28 10:27:18 -07:00
Yongye Zhu and GitHub
5aa371dc8e
[DSV4] Enable Multi-stream for Pre-Attn GEMM ( #41061 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-28 09:08:55 -07:00
zhangxin81 and GitHub
de3da0b97c
Add tuned triton fused_moe configs on H100 for gpt-oss ( #39904 )
...
Signed-off-by: zhangxin81 <115389973+zhangxin81@users.noreply.github.com >
2026-04-28 03:38:48 -07:00
Roy Wang and GitHub
9e92de51c6
[Bugfix] Exclude numa_bind fields from ParallelConfig DP hash ( #41098 )
...
Signed-off-by: yasong <yasong.wang@inferact.ai >
2026-04-28 15:52:54 +08:00
bde0efdbb7
[Bugfix][Granite4Vision] Fix deepstack buffer causing decode slowdown in compiled mode ( #40917 )
...
Signed-off-by: artemspector <artems@il.ibm.com >
Co-authored-by: artemspector <artems@il.ibm.com >
2026-04-28 07:43:30 +00:00
zhrrr and GitHub
ea74f701db
Bugfix: fix SpecBench sample argument error ( #40927 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-04-28 00:33:49 -07:00
wang.yuqi and GitHub
a8208e6a81
[Examples] Resettle features examples. ( #40995 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-28 00:33:41 -07:00
anthonsu and GitHub
76c9cccc36
[Core] Fix redundant None append in StepPool.forward for chunked prefill ( #41049 )
...
Signed-off-by: Anthony Su <xsuanthony@gmail.com >
2026-04-27 23:42:47 -07:00
JiangWeixiang and GitHub
ed57f77192
[Bugfix ] fix bailing_moe_linear ( #40859 )
...
Signed-off-by: ghphotoframe <854746559@qq.com >
2026-04-27 22:39:23 -07:00
7a1eb8ac2e
[Model] update for mimo v25 ( #41029 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Copilot <copilot@github.com >
2026-04-27 21:52:54 -07:00
Isotr0py and GitHub
c2e88a281c
[Bugfix] Fix broken example opeanai client ( #41088 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-28 04:43:04 +00:00
Matthew Bonanni and GitHub
fd74c90d9c
[Attention][Spec Decode] Allow independent drafter attention backend selection ( #39930 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-27 19:38:09 -07:00
Chauncey and GitHub
146f44b77d
[Frontend]Responses API supports Tool/Function calling with streaming with required ( #40700 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-27 19:36:58 -07:00
0d4f714208
[Bugfix] Remove tokenizer encode/decode calls from Olmo3 reasoning parser ( #40855 )
...
Signed-off-by: Yifan <yzong@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-27 19:36:54 -07:00
Angela Yi and GitHub
03aeed802f
[Test] Fix test_dynamic_shapes_compilation for torch 2.12 ( #40743 )
...
Signed-off-by: Angela Yi <angelayi@meta.com >
2026-04-27 17:51:15 -07:00
Jee Jee Li and GitHub
2c8b76c5cb
[Model][DSV4] Support base model ( #41006 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-28 08:16:55 +08:00
Kunshang Ji and GitHub
407b34be26
[xpu] bump up vllm-xpu-kernel v0.1.7 ( #41019 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-28 08:04:54 +08:00
Giancarlo Delfin and GitHub
4c7c69b4e0
[Model Runner V2] Skip attention metadata rebuild before draft prefill ( #40410 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-27 15:38:05 -07:00
Andreas Karatzas and GitHub
5e2c37facd
[ROCm][CI] Add missing quantization methods and fix online quant test failures ( #39801 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-27 15:08:57 -05:00
c8bbe05189
[Perf] Update TRTLLM supported MoE routing methods ( #39141 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: root <root@bia0030.bia.clusters.nvidia.com >
Co-authored-by: root <root@bia0036.bia.clusters.nvidia.com >
2026-04-27 14:16:22 -04:00
6232fb4b66
[Docker] Install numactl CLI in CUDA runtime image ( #41032 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-04-27 10:58:06 -07:00
Moritz Sanft and GitHub
2c06cf3486
[Bugfix] use served_model_name for multimodal error message ( #41003 )
...
Signed-off-by: Moritz Sanft <58110325+msanft@users.noreply.github.com >
2026-04-27 08:22:35 -07:00
e6f710a87f
Deprecate support for Transformers v4 ( #40389 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-27 08:19:57 -07:00
c245d35ff4
[Model] Add MiMo-V2.5 support ( #40967 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: zjy0516 <zhujiangyun@inferact.ai >
Co-authored-by: yasong <yasong.wang@inferact.ai >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Copilot <copilot@github.com >
2026-04-27 13:26:51 +00:00
f8ac0c7cf0
[Bugfix] Fix k_norm weight sharding in MiniMaxM2Attention when total_num_kv_heads < tp_size ( #38191 )
...
Signed-off-by: wxsIcey <1790571317@qq.com >
Signed-off-by: Icey <1790571317@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-27 05:57:13 -07:00
ebf862c351
Add system_fingerprint field to OpenAI-compatible API responses ( #40537 )
...
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-27 16:17:52 +08:00
wang.yuqi and GitHub
8d8062d0a7
[Examples] Resettle generate examples. ( #36464 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-27 07:48:37 +00:00
985961345a
[Bugfix] Install libcublas-dev in Dockerfile for FlashInfer CuTe DSL JIT ( #39855 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-27 15:47:39 +08:00
Yongye Zhu and GitHub
706a04d34b
[DSV4] Add silu clamp limit to shared expert ( #40950 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-04-27 00:37:43 -07:00
Isotr0py and GitHub
22631f80a0
[Bugfix] Remove invalid deepstack boundary check for Qwen3-VL ( #40932 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-27 07:27:06 +00:00
Bhoomit and GitHub
2cc008e7b4
[Attention][TurboQuant] Share dequant buffers, eliminate float16_copy ( #40941 )
...
Signed-off-by: Bhoomit Vasani <bhoomit.2010@gmail.com >
Signed-off-by: Vasani Bhoomit <bhoomit.2010@gmail.com >
2026-04-27 13:48:36 +08:00
5d5c776444
[Perf] FP8 FlashInfer Attn for ViT ( #38065 )
...
Signed-off-by: Zhanda Zhu <zhandazhu@gmail.com >
Co-authored-by: Yubo Gao <ybgao-nvidia@users.noreply.github.com >
2026-04-27 13:44:15 +08:00
ojhaanshika and GitHub
592ae6805c
Cutlass W4A16 (Machete) Tests ( #35450 )
...
Signed-off-by: Anshika Ojha <anshikao@nvidia.com >
2026-04-27 05:15:29 +00:00
7b1bc0a3eb
[Bugfix] Cap SWA/chunked-local runtime admission to startup pool-sizing bound ( #40946 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-27 04:33:13 +00:00
Silu Panda and GitHub
c0879d9483
[Tests] Gate Isaac under Transformers v5 ( #40907 )
...
Signed-off-by: Silu Panda <31051721+SiluPanda@users.noreply.github.com >
2026-04-26 19:26:51 -07:00
Giancarlo Delfin and GitHub
f5f9878514
[Model Runner V2] Fix rejection sampling acceptance rate gap vs MRV1 ( #40651 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-26 19:12:08 -07:00
2ce95a761b
Auto-disable expandable_segments around cumem memory pool ( #40812 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-27 09:37:22 +08:00
+8
4d51588e23
[Feat] DeepSeek V4 Rebased ( #40860 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: qizixi <zixi@inferact.ai >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Yongye Zhu <yongye@inferact.ai >
Co-authored-by: Simon Mo <simon@inferact.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Roy Wang <yasong.wang@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: youkaichao <youkaichao@gmail.com >
Co-authored-by: Zhewen Li <jerven.vllm@gmail.com >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
Co-authored-by: khluu <khluu000@gmail.com >
Co-authored-by: qizixi <zixi@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-04-26 18:31:08 -07:00
32e45636e3
[torch.compile]: Disable Sequence Parallelism (SP) for piecewise compilation ( #38373 )
...
Signed-off-by: SouthWest7 <am1ao@qq.com >
Signed-off-by: Xinan Miao <1403572259@qq.com >
Co-authored-by: SouthWest7 <am1ao@qq.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Wang Xingran <72983099+wangxingran222@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-26 17:44:42 +00:00
b39c266dae
[KV Offload] Offload all KV blocks when doing prefill in P/D ( #40346 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Signed-off-by: omerpaz95 <73347585+omerpaz95@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-04-26 15:06:01 +03:00
Dao007forever and GitHub
9558f43903
[Bugfix] Size FlashInfer NVLink MNNVL workspace to EP group ( #40893 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
2026-04-26 01:26:34 -07:00
Jee Jee Li and GitHub
8cd174fa35
[LoRA] MoE LoRA Refactor ( #40338 )
2026-04-26 01:55:19 +00:00
c798593f0d
[Bugfix] Fix the DSML token leakage in DSV4/3.2 ( #40806 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Windswithyou 1694599440@qq.com
2026-04-26 08:58:50 +08:00
12a3f6454b
[Bugfix][MoE] Only unpad routed output before shared expert add or routed output transform ( #40865 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-25 20:50:12 +00:00
Or Ozeri and GitHub
60cd878a3b
[kv_offload+HMA][11/N]: Support store with multiple KV groups ( #39403 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-25 20:00:46 +03:00
rasmith and GitHub
1e9f19ca3f
[CI][AMD]BugFix] Fix deadlock occuring in test_moe_layer ( #40767 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-25 09:34:14 -04:00
labAxiaoming and GitHub
6646c0c7e0
[Opt] Optimize deepstack buffer handling for multimodal Qwen3 models ( #40145 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
2026-04-25 21:04:26 +08:00
Andreas Karatzas and GitHub
95995bbef8
[ROCm][Engine] Fix GPU memory leaks in engine shutdown and test workaround for async KV prefix cache reset ( #38503 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 05:25:20 +00:00
07351e0883
[Feature] Warm up readonly multimodal processor during renderer startup ( #40797 )
...
Signed-off-by: Chenguang ZHENG <645327136@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-25 03:57:41 +00:00
Andreas Karatzas and GitHub
428b988c98
[ROCm][CI] Fix trust_remote_code AttributeError in EAGLE3 acceptance length test ( #40306 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 02:59:31 +00:00
Andreas Karatzas and GitHub
e54894fc85
[ROCm][CI] Fix TestSiluMulGroupFp8QuantModel after W8A8 block linear refactor ( #39799 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-25 11:20:59 +09:00
Angela Yi and GitHub
bc2ae5a3d6
[Test] Increase qwen2_vl num_logprobs to fix torch 2.12 update ( #40818 )
...
Signed-off-by: Angela Yi <angelayi@meta.com >
2026-04-25 00:59:20 +00:00
Wentao Ye and GitHub
a474da2813
[Refactor] Remove unused dead code ( #40640 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-25 07:28:18 +08:00
Lucas Kabela and GitHub
ce6a199ecc
[BE][Bugfix] Respect TORCH_COMPILE_DISABLE env var at the vLLM config level for torch 2.12 ( #40715 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-24 16:25:03 -07:00
Ignacio Sica and GitHub
f88763efc3
[Bugfix] add seq_lens_cpu_upper_bound to CommonAttentionMetadata in mla_runner.py ( #40844 )
...
Signed-off-by: ignaciosica <mignacio.sica@gmail.com >
2026-04-24 23:13:52 +00:00
Artem Perevedentsev and GitHub
333529deae
[EPLB] Fix replica selection bias in fused_moe router ( #40810 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-04-24 22:06:41 +00:00
Zhang Jian and GitHub
8825608205
[Bugfix][CI] Fix wrong residual shape in TestFusedAddRMSNorm.example_inputs that causes flaky test ( #40629 )
...
Signed-off-by: Zhang Jian <jianmusings@gmail.com >
2026-04-24 16:40:07 -04:00
qli88 and GitHub
095d2f87e8
[Bug] Fix GLM-5.1 running error on ROCm platform ( #40763 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-04-24 19:54:40 +00:00
21792520e7
[Build] Add Python 3.14 to supported version list. ( #34770 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-04-24 10:24:05 -07:00
Alex Brooks and GitHub
5e11b40365
[Frontend] Delegate to vLLM Omni When --omni Passed ( #40744 )
...
Signed-off-by: Alex Brooks <albrooks@redhat.com >
2026-04-24 12:30:00 -04:00
f768b4473e
[Docs] Add docs for context extension using the yarn method ( #37430 )
...
Signed-off-by: xiaoming <1259730330@qq.com >
Signed-off-by: labAxiaoming <34019940+labAxiaoming@users.noreply.github.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-24 08:26:09 -07:00
JartX and GitHub
914d0464c1
[Refactor] Unify 2D/3D kernels in triton_unified_attention ( #40631 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-04-24 17:18:06 +02:00
Jinzhen Lin and GitHub
9f771b3ab9
[Quantization] add humming quantization kernel ( #34556 )
2026-04-24 09:29:44 -04:00
Itay Alroy and GitHub
c9d3c6e6af
fused_moe: treat NIXL EP as batched experts ( #40412 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-04-24 08:05:31 -05:00
Or Ozeri and GitHub
51adca74e6
[kv_offload+HMA][9/N]: Support lookup with multiple KV groups ( #39401 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-24 15:32:29 +03:00
Netanel Haber and GitHub
e8eb0490ce
[Bugfix][MoE] Unpad routed output before shared expert add [ Fixes #35949 ] ( #40794 )
...
Signed-off-by: Netanel Haber <nhaber@nvidia.com >
2026-04-24 11:53:23 +00:00
Jiangyun Zhu and GitHub
e8ee2a78db
[Attention] use diff kv backend for mimo v2 flash ( #40045 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-04-24 11:25:55 +00:00
2ec18f5df4
[Bugfix][Parser] Fix Mistral tool parser for HF tokenizers ( #39294 )
...
Signed-off-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: thomasmaindron <thomasmaindron@users.noreply.github.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-24 19:01:56 +08:00
Dmitry Tokarev and GitHub
6dec49f27e
[Build] Bump CUDA to 13.0.2 to match PyTorch 2.11.0 ( #40669 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com >
2026-04-24 10:27:11 +00:00
Shanshan Shen and GitHub
b5587e1013
[CI/Build] Add e2e test for ViT CUDA graph ( #40780 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-24 18:12:14 +08:00
milesial and GitHub
9ad5abe772
Fix Nano Nemotron VL static image inputs ( #40724 )
...
Signed-off-by: Alexandre Milesi <milesial@users.noreply.github.com >
2026-04-24 09:18:55 +00:00
Woosuk Kwon and GitHub
7d3195ea9f
[Bugfix] Fix IMA in DSA + MTP ( #40772 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-04-24 01:40:20 -07:00
512f522192
[Model] Gemma4: add bidirectional vision attention for sliding layers with window guard ( #40534 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Luciano Martins <lucianomartins@google.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-24 08:27:46 +00:00
4c34b2f6fc
[XPU] Enable torch.compile for XPU GDN attention ( #39466 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-24 16:26:16 +08:00
Xin Yang and GitHub
cf8a613a87
Support only half types for concat_mla_q kernel ( #37892 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-04-23 23:51:05 -07:00
xiangdong and GitHub
01acf96c6f
[XPU][CI] Fix Docker cleanup races on Intel CI runners ( #40761 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-24 14:08:45 +08:00
079a4cf399
[MoE] Move cutlass moe to fused_moe/experts/ ( #40574 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-24 06:05:49 +00:00
9744b699ba
[Deprecate] Deprecate LLM.reward offline api, use LLM.encode instead. ( #40688 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-24 05:37:50 +00:00
c662b4359e
[Bugfix] Avoid mutating chat_template_kwargs in HYV3ReasoningParser initialization ( #40713 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 13:08:58 +08:00
lyd1992 and GitHub
100c7b65e7
[Platform] Fix RISC-V platform detection (lscpu parsing + non-NUMA meminfo) ( #40427 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-04-24 04:33:05 +00:00
Neil Schemenauer and GitHub
56bdf85e10
[Feature] Avoid eager import of the "mistral_common" package. ( #40043 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
2026-04-24 02:49:16 +00:00
Vinayak Kumar and GitHub
eba73068ea
[Doc] fix capitalization consistency in README (vLLM, Hugging Face) ( #40729 )
...
Signed-off-by: Vinayak Mishra <vinayakmishra448@gmail.com >
2026-04-24 02:23:54 +00:00
Nick Hill and GitHub
e9f331d72e
[MRV2] Ensure warmup covers prefill path ( #40746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-24 01:33:26 +00:00
c9bf77df92
[BUG]: fix HF tokenizer concurrent borrow in tool parsers ( #40059 )
...
Signed-off-by: Yifan <yzong@redhat.com >
Co-authored-by: timon0305 <timon0305@outlook.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-04-23 18:20:30 -07:00
3041344287
[Misc] Added curl retries in install_python_libraries.sh ( #36700 )
...
Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 01:19:30 +00:00
Doug Campos and GitHub
92762edc53
[Bugfix] Treat <tool_call> as implicit reasoning end in Qwen3 parser ( #35687 )
...
Signed-off-by: Doug Campos <qmx@qmx.me >
2026-04-24 09:10:04 +08:00
626daa2076
[Feat] Unified Synthetic Acceptance Rate for V1 and V2 ( #40662 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-24 00:48:08 +00:00
Nick Hill and GitHub
fe85a92e86
[Core] Avoid seq_lens_cpu GPU->CPU sync ( #40654 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-24 00:35:55 +00:00
Sage Moore and GitHub
62b1bbe470
[EPLB] Remove asyncio infrastructure from Async EPLB ( #40730 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-24 00:21:15 +00:00
Hemanth Acharya and GitHub
fa4b70555b
[ROCm] Cast score correction bias tensor during model construction for DeepSeek/Kimi-K2 ( #39999 )
...
Signed-off-by: Hemanth Acharya <heachary@amd.com >
2026-04-24 09:02:12 +09:00
447c372ac5
[MoE] Move remaining PrepareAndFinalize to prepare finalize folder ( #39009 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-23 20:00:53 -04:00
ff2c2bd80a
[Docs]Add documentation for bench serve visualization arguments ( #40539 )
...
Signed-off-by: Sophie du Couédic <sop@zurich.ibm.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-23 15:48:29 -07:00
Matthew Bonanni and GitHub
cde8d24710
[Spec Decode] Move SpecDecodeBaseProposer out of eagle.py ( #40732 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-23 22:28:27 +00:00
bnellnm and GitHub
4a6dd1c3cc
[Bugfix] Fix DeepSeek V2-Lite Accuracy drop ( #40673 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-23 18:11:37 -04:00
7ff65b1900
[Bugfix] Fix workspace resize leaking reserved GPU memory ( #39226 )
...
Signed-off-by: root <conway.zhu@cohere.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 20:50:05 +00:00
Johnny and GitHub
7f95a66cbf
[NVIDIA] Add sm_110 (Jetson Thor) to CUDA 13.0 build targets ( #39233 )
2026-04-23 15:42:14 -04:00
1b1c01de39
[MoE] Move xpu moe to fused_moe/experts/ ( #40568 )
...
Signed-off-by: Jackmin801 <ongjackm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 13:38:10 -04:00
e9ba519f45
[DP][Ray] Pin DP control bundle to same node as first GPU bundle ( #39167 )
...
Signed-off-by: Shahar Mor <smor@nvidia.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 17:21:13 +00:00
Or Ozeri and GitHub
5ef33ab250
[kv_offload+HMA][10/N]: Support load with multiple KV groups ( #39402 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-23 20:00:45 +03:00
bnellnm and GitHub
1c2c1eb8b9
[MoE Refactor] Rename FusedMoE.make_expert_params_mapping to fused_moe_make_expert_params_mapping ( #40671 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-23 11:22:34 -04:00
Nicolò Lucchesi and GitHub
8824f50f1f
[CI] Split disaggregated tests into own test-area ( #40623 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-23 23:20:12 +08:00
0098db9ec1
[ROCm] Implement GPU-to-NUMA-node detection ( #40015 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-23 10:08:48 -05:00
Kunshang Ji and GitHub
53ecc807c0
[XPU] Upgrade torch 2.11 for xpu ( #37947 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 10:07:35 -05:00
b7a2605020
[Bugfix] Make Attention Backend Auto-Selection Batch-Invariance-Aware ( #40193 )
...
Signed-off-by: Srreyansh Sethi <srreyansh.sethi@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-23 14:57:03 +00:00
d0009ddb0b
[Model] Support Hy3 preview ( #40681 )
...
Signed-off-by: stevenkuang <stevenkuang@tencent.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-23 22:08:26 +08:00
Richard Zou and GitHub
424033f4fc
[Bugfix] Include inductor and functorch configs in compilation cache key ( #40627 )
...
Signed-off-by: Richard Zou <zou3519@gmail.com >
2026-04-23 09:52:59 -04:00
Isotr0py and GitHub
da1e7311ca
[Misc] use model arch converter for bidi models identification ( #40701 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-23 13:42:52 +00:00
xiangdong and GitHub
01cb41dcf5
[XPU][CI]Temporary disable 3 cases on Intel GPU in CI ( #40683 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-04-23 21:42:22 +08:00
2f314bc5e6
[CPU] Added faster exp routine for lower precision data types. ( #38112 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Co-authored-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-04-23 13:14:44 +00:00
BadrBasowid and GitHub
2196bac135
[Compilation] Refactor SiluMul activation+quant Fusion Pass ( #39684 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-04-23 09:10:36 -04:00
Matthias Gehre and GitHub
4b7869d6bc
[ROCm] Add gfx1102/gfx1103 support ( #40037 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-04-23 01:32:04 -07:00
liuzhenwei and GitHub
4a79262e0f
[UT][Hardware] let torchrun example tests use the default backend ( #39879 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-04-23 16:22:28 +08:00
3ed5231c6a
[Build] Switch default CUDA to 13.0, update CUDA architecture lists, clean up stale build-args ( #39878 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-23 15:51:28 +08:00
Nicolò Lucchesi and GitHub
9c2492e501
[Misc] Support Human-readable (k/K/m/M..) json cli arg ( #40473 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-23 09:42:23 +02:00
Shanshan Shen and GitHub
fe57be7809
[MM][CG] Support --enable-vit-cuda-graph option for VLM examples ( #40580 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-22 22:46:14 -07:00
8317cedc77
[Responses] Add tool_choice/tools validation to match OpenAI behavior ( #40399 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 22:46:10 -07:00
Zhengxu Chen and GitHub
98a242ff61
[compile] Skip FX graph deserialiaztion on loading, further reducing warm compile time. ( #40151 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-04-23 13:43:18 +08:00
e4ee48da2d
[MoE refactor] refactor GPTQMarlinMoEMethod with MK ( #37990 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-23 05:21:47 +00:00
Kunshang Ji and GitHub
342c58bc54
[BugFix]fix Qwen3 MoE call gate twice ( #40664 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 05:04:41 +00:00
fe9c3d6c5f
[TurboQuant] enable FA3/FA4 for prefill paths ( #40092 )
...
Signed-off-by: 墨楼 <huangzhilin.hzl@antgroup.com >
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Codex <codex@openai.com >
2026-04-23 07:35:24 +03:00
ccaf5ffaa3
[XPU] disable fusion pattern support on XPU platform ( #39789 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-23 10:07:45 +08:00
Lucas Kabela and GitHub
0283f303d8
[BE] Fix compile time message to be consistent (use monitoring) ( #40641 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-23 00:12:08 +00:00
ac58e2a170
[Fix][MoRI] Align MoRI-IO message format with P2pNcclConnector and vllm-router ( #39565 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: Matvei Pashkovskii <mpashkov@amd.com >
2026-04-23 08:06:31 +09:00
Lucas Kabela and GitHub
b8401a9bf4
[Bugfix] Fix RMS norm + quant fusion on DeepGEMM UE8M0 path for B200 ( #40552 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-04-22 22:04:42 +00:00
Honglin Cao and GitHub
9c271f9403
[gRPC] Add standard gRPC health checking (grpc.health.v1) for Kubernetes native probes ( #38016 )
...
Signed-off-by: Honglin Cao <Caohonglin317@hotmail.com >
2026-04-22 21:31:00 +00:00
22fa63cfe8
[Bugfix][Torch 2.12] Fix batch_invariant test with allow_override for torch 2.12 upgrade ( #40562 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 13:48:55 -07:00
8f87eb4622
[Refactor] Clean up log once scope="local" ( #40540 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-22 16:42:43 -04:00
cfa49213d7
[Bugfix][Parser] Fix Mistral pre-v11 tool parser failing on trailing model output ( #40531 )
...
Signed-off-by: dougbtv <dosmith@redhat.com >
Signed-off-by: Doug Smith <dougbtv@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-22 16:35:00 -04:00
29f64c5f5e
FlexAttention non-causal support ( #40394 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 13:22:57 -07:00
Angela Yi and GitHub
eb6661d522
Fix test_startup.py for torch 2.12 ( #40636 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-04-22 19:31:41 +00:00
d622e27d2b
[NVFP4] NVFP4 MOE emulation fallback for H100/MI300/MI350, standardize TritonExperts usage for OCP MX emulation ( #35737 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Signed-off-by: fxmarty-amd <felmarty@amd.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-04-22 08:58:54 -07:00
5f76b3fb30
[MoE] Convert CT W8A8 To Oracle Structure ( #39187 )
...
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-22 14:53:30 +00:00
bnellnm and GitHub
809d83c2dc
[MoE Refactor] Combine MoERunnerBase + DefaultMoERunner ( #40560 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-22 14:43:17 +00:00
Nicolò Lucchesi and GitHub
33ef1941e2
[Bugfix][CI] Fix v1/kv_connector/unit/test_nixl_connector_hma.py::test_fewer_blocks_with_hma ( #40597 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-22 14:21:02 +01:00
Hank_ and GitHub
a4905133f3
[xpu][rocm] Update current_platform.supports_fp8() for TritonExperts ( #40132 )
...
Signed-off-by: Hank <hcc.mayday@gmail.com >
2026-04-22 13:39:40 +02:00
ecbe42e991
[Doc] Clarify supported keys for --speculative-config ( #40455 )
...
Signed-off-by: Wangxiaoxiaoa <Wangxiaoxiaoa@users.noreply.github.com >
Co-authored-by: Wangxiaoxiaoa <Wangxiaoxiaoa@users.noreply.github.com >
2026-04-22 04:36:17 -07:00
a250f1bd5f
[Bugfix] LoRA for DeepSeek V3.2 ( #35077 )
...
Signed-off-by: Hollow Man <hollowman@opensuse.org >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-22 19:33:50 +08:00
lyd1992 and GitHub
04eac6ba24
[Bugfix][CPU][RISC-V] Clamp exp() input to prevent NaN ( #40428 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
2026-04-22 09:38:18 +00:00
9047288b68
support hotwords for FunASR model ( #39674 )
...
Signed-off-by: zixiao <shunli.dsl@alibaba-inc.com >
Co-authored-by: zixiao <shunli.dsl@alibaba-inc.com >
2026-04-22 02:25:06 -07:00
Johnny Yang and GitHub
ed6d30377d
upgrade tpu-inference to v0.18.0 ( #40395 )
2026-04-22 01:33:45 -07:00
6aa057c9d7
[Multimodal] Support custom video metadata for pre-extracted frame sequences ( #40133 )
...
Signed-off-by: storyicon <storyicon@foxmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 15:50:04 +08:00
Chauncey and GitHub
a2bd09c960
[Bugfix] [Reasoning] Add reasoning_start_str/reasoning_end_str properties to reasoning parsers ( #40566 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-22 07:27:44 +00:00
philip-essential and GitHub
123674879e
[Model] Add block-local attention and YaRN for local layers to Gemma3 ( #39823 )
...
Signed-off-by: Philip Monk <169196560+philip-essential@users.noreply.github.com >
2026-04-21 23:34:50 -07:00
Carl Y and GitHub
4254aeb56f
[fix] flaky test_mla_attn_quant_fusion.py ( #40530 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-04-22 06:29:58 +00:00
aad88f8486
[kv_offload+HMA][8/N]: Support multi-group worker transfer ( #38453 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-22 08:44:00 +03:00
Bugen Zhao and GitHub
0210024ae7
[Bugfix] Pass effective chat template kwargs to reasoning parsers ( #40460 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-04-21 22:17:51 -07:00
4eafc72928
[Audio] Bundle get_generation_prompt() params into SpeechToTextParams ( #36268 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
2026-04-22 12:24:18 +08:00
Micah Williamson and GitHub
6d09769700
[ROCm] Support non-causal attention in ROCM_ATTN ( #40176 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-04-22 12:57:12 +09:00
4506319a28
[compile] mla + group fp8 fusion ( #38877 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 23:16:58 -04:00
9b60e2ffaa
[Bugfix] Fix quantized model initialization failure with prefetch offloading ( #40432 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 20:15:58 -07:00
Martin Hickey and GitHub
3951d3eacd
[MyPy] Enable mypy for vllm/model_executor/layers/ ( #40159 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-04-21 20:15:02 -07:00
6f2c71be8f
[Multimodal] Add PyAV video backend for concurrent video decoding ( #39986 )
...
Signed-off-by: Jaseel Muhammad <jaseel.muhammad@mbzuai.ac.ae >
Signed-off-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-21 20:14:57 -07:00
rasmith and GitHub
2463f00fb6
[AMD][CI][BugFix] Override normalize_e4m3fn_to_e4m3fnuz for fnuz machines in test_moe_layer_no_parallel ( #40550 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-22 02:21:02 +00:00
f946659fff
[Bugfix] Fix W4A8_FP8 MoE tp>1 correctness and view() TypeError ( #40310 )
...
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 21:58:33 -04:00
Soila Kavulya and GitHub
f90aa44662
[NIXL][XPU]Fix nixl import on XPU ( #40430 )
...
Signed-off-by: Soila Kavulya <soila.p.kavulya.intel.com>
2026-04-22 09:26:33 +08:00
rasmith and GitHub
cefa5281a7
[ROCm][P/D][MORI][BugFix] Ensure correct api is used when making requests to prefill / decode nodes ( #39835 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-22 09:48:25 +09:00
Jhao-Ting Chen and GitHub
46794958f0
test: add nan/inf clamp regression test for fused_topk_bias ( #40553 )
...
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
2026-04-22 00:46:53 +00:00
Khushali Desai and GitHub
6ff8dea075
[Bugfix] avoid warmup if text only expectation in multi_modal run ( #40409 )
...
Signed-off-by: khushali9 <khushali.desai9@gmail.com >
2026-04-22 00:19:50 +00:00
TJian and GitHub
583e6f2226
[ROCm] [Wheel] [Bugfix] [Critical] Remove any packages installed from github from rocm.txt e.g fastsafetensors as it is incompatible with uv pip ( #40461 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-04-22 00:18:07 +00:00
96a85c5750
[Startup][UX] Enable CUDAGraph memory profiling by default ( #38284 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-04-21 18:16:59 -04:00
9db4650e5e
[MoE Refactor] Add more MoE layer tests ( #39349 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-21 18:12:36 -04:00
bnellnm and GitHub
5e584ce9ec
[MoE Refactor] Remove SharedFusedMoE class ( #35782 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-21 18:12:12 -04:00
Wentao Ye and GitHub
1842447c09
[Refactor] Remove unused param ( #39750 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 14:59:20 -07:00
Wentao Ye and GitHub
16688b26a6
[Perf] Optimize batch invariant with fused rms norm, 2.1% E2E latency improvement ( #40413 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 19:51:03 +00:00
Jakub Zakrzewski and GitHub
6fbec8ed47
[Bugfix][Kernel] nvfp4 cutlass MoE: fix nvfp4 experts quant out-of-bounds read for expert counts not divisible by 4 or 16 ( #40351 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
2026-04-21 19:06:09 +00:00
5544f8c18b
[Performance] Add is_reasoning_end_streaming() override to GptOssReasoningParser ( #35745 )
...
Signed-off-by: Fergus <fergus.barratt00@gmail.com >
Signed-off-by: fergus barratt <fergus.barratt00@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-21 18:31:27 +00:00
9f39b380d0
[Bugfix] Fix spec decode test failures on Blackwell (SM100+) ( #39546 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-21 18:21:19 +00:00
Zijing Liu and GitHub
9a6a66f3b8
[MRv2]fix: model accuracy regression caused by reusing the stale last_sampled_tokens and draft_tokens ( #39833 )
...
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com >
2026-04-21 16:30:32 +00:00
67eb6083e3
Revert "[Misc] Move pyav and soundfile to common requirements" ( #40276 )
...
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-21 09:08:06 -07:00
Harry Mellor and GitHub
6ee081d1d0
Add new tp plan styles to the Transformers modelling backend ( #40467 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-04-21 08:51:30 -07:00
66cc3fa559
[Model Runner V2] Multiple prompt logprobs support ( #39937 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-21 15:49:05 +00:00
Vadim Gimpelson and GitHub
6d85b36a9f
Revert #38730 and #38791 ( #40032 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-04-21 11:44:11 -04:00
Matthew Bonanni and GitHub
ab5666eb7c
[UX] Bump version in CG memory profiling log message ( #40465 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-21 15:26:06 +00:00
roikoren755 and GitHub
f819265a4a
Default to 'align' mamba cache mode for Mamba-based models when speculative decoding is enabled ( #40454 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-21 14:51:43 +00:00
Shanshan Shen and GitHub
936e0b79aa
[MM][CG] Optimize default max_frames_per_batch auto-infer for ViT CUDA graph video inference ( #40445 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 14:47:53 +00:00
b2a5518679
[XPU][CI] Add misc, engine and lora cases on Intel GPU in CI ( #39887 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-21 22:30:46 +08:00
ℍ𝕠𝕝𝕝𝕠𝕨 𝕄𝕒𝕟 and GitHub
908a713488
[Bugfix] LoRA: extend expert base_layer loading to Qwen3.5 and Step3.x ( #37114 )
...
Signed-off-by: Hollow Man <hollowman@opensuse.org >
2026-04-21 14:17:03 +00:00
ec5ef0ac73
[Doc] Add Qwen3 AWQ models to documentation ( #40034 )
...
Signed-off-by: Yusuf <yusufmohammad@live.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 09:37:41 -04:00
7b1e0b07d0
[Bugfix] Fix dataset name and path argument validation bug in vllm bench serve ( #40288 )
...
Signed-off-by: talora <talora@nvidia.com >
Signed-off-by: Talor Abramovich <talor19@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-21 06:14:28 -07:00
d249a9e90e
Add Granite 4.1 Vision as built-in multimodal model ( #40282 )
...
Signed-off-by: Artem Spector <artems@il.ibm.com >
Signed-off-by: artemspector <artems@il.ibm.com >
Co-authored-by: artemspector <artems@il.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-21 05:43:39 -07:00
d2e2e856ad
[Frontend] Remove frontend pooling multi task support. ( #37861 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-21 12:27:44 +00:00
766cb65d00
feat(multimodal): support externally processed mm_kwargs with cache injection ( #39502 )
...
Signed-off-by: Krish Hung <krishung5@gmail.com >
Signed-off-by: krishung5 <krish@nvidia.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-21 11:31:09 +00:00
Jhao-Ting Chen and GitHub
28c222157b
fix: clamp NaN/Inf in topk_softmax to prevent duplicate expert IDs ( #39391 )
...
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com >
2026-04-21 15:04:41 +04:00
wang.yuqi and GitHub
3975eb6de6
Revert "[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm" ( #40438 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-21 08:47:18 +00:00
Zeyu Zhang and GitHub
5a94a19824
[Bugfix] Normalize malformed dict prompts that carry token IDs in prompt ( #40339 )
...
Signed-off-by: Alchuang22-dev <2584829494@qq.com >
2026-04-21 07:44:36 +00:00
hangy-amd and GitHub
f95c11a848
[Feat] dflash support for ROCm ( #39703 )
...
Signed-off-by: Hang Yang <hangy@amd.com >
2026-04-21 14:58:20 +08:00
milesial and GitHub
257015d5e5
[MoE] Triton MoE Perf regression - restore low latency path ( #39016 )
2026-04-21 02:37:11 -04:00
Shanshan Shen and GitHub
b47840019e
[MM][Misc] Support image+video mixed inputs (per prompt) for VLM examples ( #40335 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 03:43:25 +00:00
SeongJun Lee and GitHub
989cc12d88
[Fix] Add missing space in IP fallback warning ( #40359 )
...
Signed-off-by: lesj0610 <lesj0610@gmail.com >
2026-04-20 20:26:06 -07:00
Wentao Ye and GitHub
301024aa9c
[Deprecation] Deprecate cprofile and cprofile_context ( #39100 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-21 11:25:22 +08:00
Simon Mo and GitHub
8256833fe6
[Startup] Parallelize torch/transformers import + weight prefetch + forkserver prewarm ( #40331 )
...
Signed-off-by: simon-mo <simon@inferact.ai >
2026-04-21 10:49:32 +08:00
Shanshan Shen and GitHub
8097591286
[Doc] Update ViT CUDA graph doc for mixed (image+video) inputs ( #40355 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-04-21 02:31:09 +00:00
20d3743491
[Bugfix] Gemma4: fix multimodal embedder norm order to match HF reference ( #40411 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-04-21 02:28:26 +00:00
Chauncey and GitHub
18563f2072
[Misc] Reduce attention logging levels ( #40086 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-21 02:09:25 +00:00
0e884fe638
[Bugfix] Fix _CONFIG_REGISTRY types getting wrong config class when on-disk model_type differs ( #39554 )
...
Signed-off-by: Misa <misaAle@users.noreply.github.com >
Signed-off-by: Misael Casarez <misacasa@amazon.com >
Co-authored-by: Misael Casarez <misacasa@amazon.com >
2026-04-20 19:04:48 -07:00
fe5c115ee4
[vLLM IR] Add IR op testing and benchmarking infrastructure ( #40167 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Theresa Shan <Theresa.Shan@amd.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-21 00:23:03 +00:00
6867bcd076
[Bugfix] Replace code that disabled shared expert overlap ( #39222 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-20 19:36:16 -04:00
c075702eae
[Misc][UX] Suppress confusing num_gpu_blocks log lines ( #40402 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-20 22:32:46 +00:00
Rita Brugarolas and GitHub
21b086d0aa
[ROCm] Hotfix: guard MLA dual RMS norm fusion against older AITer versions ( #40386 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-04-20 16:20:05 -05:00
Sage Moore and GitHub
3173441b0f
[EPLB] Consolidate is_unchanged/is_received_locally into TransferMetadata ( #37341 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-20 21:12:42 +00:00
Cao Qian and GitHub
8b1f3bebca
[LMCache MP Connector] Add num_lmcache_extra_cached_token in KVTransferParams ( #39843 )
...
Signed-off-by: aeon-x <talexcao@gmail.com >
2026-04-20 20:42:49 +00:00
2390caf157
Enable building MoRI with AMD AINIC stack ( #38371 )
...
Signed-off-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com >
Signed-off-by: Theresa Shan <theresa.shan@amd.com >
Co-authored-by: Theresa Shan <thshan@smci355-ccs-aus-n08-21.prov.aus.ccs.cpe.ice.amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-20 11:17:59 -07:00
Frederik Gossen and GitHub
87805fa11e
[Core] Cache InductorPass.hash_source with functools.cache ( #39328 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 14:06:15 -04:00
Nicolò Lucchesi and GitHub
304d5ba1a0
[Bugfix][CI] Fix tests/distributed/test_torchrun_example_moe.py ( #40349 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-20 11:05:44 -07:00
Tyler Michael Smith and GitHub
81d954f454
[WideEP] Remove naive all2all. Use allgather_reducescatter instead ( #33728 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-04-20 17:53:55 +00:00
Frederik Gossen and GitHub
47fcb8ca68
[Core] Pass donate_graph_module=True to standalone_compile ( #39733 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 17:40:52 +00:00
bai and GitHub
191e3fdaa1
Update flashinfer to 0.6.8 ( #39959 )
...
Signed-off-by: bai <v@gor.io >
2026-04-20 10:37:23 -07:00
Frederik Gossen and GitHub
b9cf629bd0
[Core] Label torch trace logging overhead with dynamo_timed ( #39329 )
...
Signed-off-by: Frederik Gossen <frgossen@meta.com >
2026-04-20 17:31:03 +00:00
3461c8b027
[EPLB] Refactor Async EPLB synchronization logic ( #37601 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-04-20 17:05:41 +00:00
726efe177b
[MoE Refactor] Move the shared/fused expert output sum into MoERunnerBase ( #35949 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-04-20 12:28:46 -04:00
Yan Ma and GitHub
595562651a
[XPU] fix MoE triton backend in online fp8 quantization ( #40109 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-20 11:31:39 -04:00
Hashem Hashemi and GitHub
3a30eaa1d7
Properly enable wvSplitK fp8 path for RDNA ( #37712 )
...
Signed-off-by: Hashem Hashemi <hashem.hashemi@amd.com >
2026-04-20 10:09:24 -05:00
Rita Brugarolas and GitHub
fb5635d3f9
[ROCm] Add MLA dual RMS norm fusion (Q, KV) pass for DeepSeek/Kimi-K2 ( #39242 )
...
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
2026-04-20 14:56:27 +00:00
Wentao Ye and GitHub
b42e878ec0
[Bug] Fix dcp error message ( #40053 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-20 10:52:32 -04:00
7243e02aa1
[ROCm][Feature] Enable AITER MLA attention backend to work with Eagle3 speculative decoding on ROCm ( #39616 )
...
Signed-off-by: larryli2-amd <larryli2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-04-20 09:44:43 -05:00
Sage Moore and GitHub
def8f52200
[CI][EPLB] Add Async EPLB end-to-end integration test to CI ( #40168 )
...
Signed-off-by: Sage Moore <sage@neuralmagic.com >
2026-04-20 10:22:54 -04:00
Vasiliy Kuznetsov and GitHub
38fa87caca
mxfp8 online quant move to new frontend ( #40152 )
...
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com >
2026-04-20 06:26:12 -07:00
a023edfa5b
[bugfix] Use only onlines CPUs in lscpu ( #40161 )
...
Signed-off-by: kse <kevin.sejourne@cloud-temple.com >
Co-authored-by: kse <kevin.sejourne@cloud-temple.com >
2026-04-20 13:19:57 +00:00
b82fc1364d
[Anthropic][Frontend] Added chat_template_kwargs to /v1/messages ( #40125 )
...
Signed-off-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
2026-04-20 06:10:45 -07:00
Yan Ma and GitHub
e06de7f005
[XPU] enable triton attention test on XPU by removing cuda device binding ( #39627 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-20 20:57:11 +08:00
zhanqiuhu and GitHub
cc3993b05d
nixl refactor [2/N]: unify TpKVTopology + HeteroTPTransferConfig into TransferTopology ( #39529 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-20 12:39:08 +02:00
Ilya Markov and GitHub
50dd4cb427
[EPLB] Add nixl-based eplb communicator ( #36276 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
2026-04-20 10:24:23 +00:00
f774ba028a
[kv_offload+HMA][4/N]: Support sliding window lookup ( #36645 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-04-20 12:53:51 +03:00
Fadi Arafeh and GitHub
2aab9acf48
[CPU][BugFix] Fix inter-node pipeline parallel ( #40150 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-04-20 17:21:12 +08:00
nemanjaudovic and GitHub
58631d7c3f
[Bugfix] Fix scaled_mm output narrowing for 3D input tensors ( #38093 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
2026-04-20 16:58:39 +08:00
Andreas Karatzas and GitHub
a943839e9a
[ROCm][CI] Introducing new MI300 nodes ( #39531 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-20 16:09:58 +08:00
milesial and GitHub
6d8b80802b
[Docs] Fix thinking_token_budget docs ( #40316 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
2026-04-20 08:09:44 +00:00
wuyingjun and GitHub
77fd2c8631
[Bugfix] Forward mm_processor_kwargs in offline generate APIs ( #40251 )
...
Signed-off-by: wuyingjun <wuyingjun_yewu@cmss.chinamobile.com >
2026-04-20 00:56:56 -07:00
San-Nguyen and GitHub
e729cc823d
[Fix] Add Spacing when Requesting Output Token > max_model_len ( #40324 )
...
Signed-off-by: San-Nguyen <san.nguyen@ibm.com >
2026-04-20 00:25:06 -07:00
velonica0 and GitHub
ec7aafc02a
[CPU][RISC-V] Support multiple RVV VLEN targets via compile-time dispatch ( #39478 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-04-20 14:36:59 +08:00
Julien Denize and GitHub
6097afb9bd
[BUGFIX] Fix Pixtral consolidated format vision weight loading ( #39916 )
...
Signed-off-by: Julien Denize <julien.denize@mistral.ai >
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-19 22:25:03 -07:00
4f4713f96e
[XPU] [torch.compile] Skipping CUDA graph memory estimation to avoid startup errors. ( #39977 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-20 13:04:39 +08:00
Tao He and GitHub
8936118134
[Qwen][Bugfix] Fixes sigmoid activation in torch impl of RMSNormGated. ( #40245 )
...
Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com >
2026-04-20 04:28:19 +00:00
Yuan Tang and GitHub
67ed01c353
fix: Do not make function calls when request has no tools for /v1/responses ( #40314 )
...
Signed-off-by: Yuan Tang <terrytangyuan@gmail.com >
2026-04-20 04:17:30 +00:00
6e10cb54f6
[Bugfix][Responses API] Fix streaming tool calls on /v1/responses ( #39892 )
...
Signed-off-by: Hoang Nguyen <118159510+hnt2601@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-20 11:24:52 +08:00
fcb31c1ac3
[Bugfix] Properly initialize PerTensorScaleParameter for fused-on-disk checkpoints ( #39765 )
...
Signed-off-by: Hemmi Shinichi <shemmi@preferred.jp >
Signed-off-by: Shinichi Hemmi <50256998+Alnusjaponica@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-20 02:53:04 +00:00
Lxx and GitHub
d886c26d4d
[Doc] Fix typos in token_embed pooling documentation ( #40266 )
...
Signed-off-by: YifanLi3 <lyfqlx3@gmail.com >
2026-04-19 19:27:32 -07:00
898beca5a8
[BugFix][XPU] fix lora ops bgmv_expand size not match ( #39989 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-20 08:24:50 +08:00
Kevin H. Luu and GitHub
629d45eacb
[ci] Make ecr authenticate non blocking ( #40305 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-04-19 15:37:53 -07:00
Andrew Barnes and GitHub
f150107efd
[ROCm] Fix cu_seqlens_q off-by-one in AITER FA speculative decode path ( #39120 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
2026-04-19 18:34:33 +00:00
danisereb and GitHub
d1135a5087
Fix MoE backend selection for LoRA (unquantized MoE) ( #40273 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-04-19 17:18:40 +00:00
982beae809
Optimize nemotron VL image/video preprocessing ( #40283 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: milesial <milesial@users.noreply.github.com >
2026-04-19 15:06:20 +00:00
TJian and GitHub
45232a454e
[FEAT] [Perf] [Gemma4] Fused Gemma4 Routing Function Triton ( #39083 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-04-19 09:57:39 +00:00
Flora Feng and GitHub
03ce1c6ed9
[Bugfix] Kimi-K2 tool parser streaming - fix token leakage, argument truncation, and content dropping ( #38579 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-19 01:30:27 -07:00
omerpaz95 and GitHub
4353c9cb4a
[KV Offload] Pass request context ( #39185 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
2026-04-19 08:54:59 +03:00
4b7f5ea1a0
[KV Connector] Allow metrics of multiple connectors of same types in multi connector. ( #40010 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-19 07:49:10 +03:00
38907e4391
[Frontend] Preserve structured output special tokens in offline LLM.chat ( #39352 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-04-18 19:46:07 -04:00
d0359f3e04
[Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100 ( #40191 )
...
Signed-off-by: ultranationalism <www913363043@gmail.com >
Signed-off-by: mgoin <mike.goin12@gmail.com >
Co-authored-by: mgoin <mike.goin12@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-04-18 13:58:46 -07:00
Dan Alistarh and GitHub
ed0622e3a8
[Attention] TurboQuant: remove redundant random signs, add prior art attribution ( #40194 )
...
Signed-off-by: Dan Alistarh <d.alistarh@gmail.com >
2026-04-18 14:31:59 -04:00
Yusuf Mohammad and GitHub
b5f6c5f834
Added general ND x ND matmul and unit test for it ( #39909 )
...
Signed-off-by: Yusuf <yusufmohammad@live.com >
2026-04-18 10:05:21 -04:00
Jee Jee Li and GitHub
bfde49e287
[DOC] Add fuse_minimax_qk_norm ( #39782 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-18 00:41:37 -07:00
153ba7f0f3
[Refactor] Drop direct dependency on librosa ( #39079 )
...
Signed-off-by: Nick Cao <ncao@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-18 06:55:38 +00:00
Chinmay-Kulkarni-AMD and GitHub
87518c3027
[ZenCPU] AMD Zen CPU Backend with supported dtypes via zentorch weekly ( #39967 )
...
Signed-off-by: Chinmay Kulkarni <Chinmay.Kulkarni@amd.com >
2026-04-18 06:22:37 +00:00
Rishapveer Singh and GitHub
aeee7ef939
[Bugfix] Fix k_proj's bias for GLM-ASR ( #40160 )
...
Signed-off-by: Rishapveer Singh <singhrishapveer@gmail.com >
2026-04-17 22:34:33 -07:00
z1ying and GitHub
cda19ecf4d
[Doc] Fix outdated source reference comment in anthropic/serving.py ( #40189 )
...
Signed-off-by: z1ying <tzzying@outlook.com >
2026-04-17 22:31:13 -07:00
80b18230e0
[Frontend] Add multimodal support to /inference/v1/generate endpoint ( #38405 )
...
Signed-off-by: Nithin Chalapathi <nithin.ch10@gmail.com >
Signed-off-by: Nithin Chalapathi <nithinc@berkeley.edu >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-17 20:31:56 -07:00
z1ying and GitHub
d0697cc7b6
[Doc] Add Realtime Transcription section to supported_models.md ( #39845 )
...
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-18 03:26:14 +00:00
b0755523dc
[Core] Reduce mm scheduler, get_num_embed overhead ( #40143 )
...
Signed-off-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-18 11:25:49 +08:00
993859ceb0
[XPU] fix all_reduce all-zero accuracy issue under torch.compile ( #39844 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-18 02:33:07 +00:00
Michael Goin and GitHub
48a65ccb02
[CI] Speed up test_fused_marlin_moe ( #40178 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-17 19:26:21 -07:00
55842a8d69
[XPU]fake impl for xpu fp8_gemm ( #39984 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-18 08:53:56 +08:00
Michael Goin and GitHub
1f45e83756
Remove outdated tests test_mixtral_moe and test_duplicated_ignored_sequence_group ( #40175 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-17 16:49:43 -07:00
Michael Goin and GitHub
a8bffaa133
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 ( #37463 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-17 16:42:32 -07:00
5cdddddd4a
[Kernel] [Helion] Force disable HOP path due to performance regression ( #40171 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
2026-04-17 17:36:49 -04:00
aditi-amd and GitHub
6ef1efd51f
[ROCm] Fix TurboQuant on ROCm: backend routing, flash-attn compat, int64 overflow ( #39953 )
...
Signed-off-by: aditi <aditi.rana@amd.com >
2026-04-17 13:08:37 -07:00
Ryan Rock and GitHub
58da4ee047
[AMD][CI] Update DeepEP branch ( #38396 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
2026-04-17 14:30:20 -05:00
Andreas Karatzas and GitHub
1ae11e2bfc
[ROCm][CI] Build fastsafetensors from source so it links against libamdhip64 ( #39978 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-17 14:30:08 -05:00
251c18d1f8
skip fp8e4b15 on xpu ( #39957 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-17 16:55:08 +00:00
512765d52d
[Misc][UX] Map mimo reasoning and tooling parsers ( #40089 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-04-17 16:49:21 +00:00
640cc9dd7d
feat: Add LoRA support for Gemma4ForConditionalGeneration ( #39291 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-17 09:39:19 -07:00
ceade1952c
[BugFix] Support custom tool parsers when tool_choice is required and named function ( #39870 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-04-17 16:38:10 +00:00
747256bb5d
[Bugfix][Core] Fix stuck chunked pipeline parallelism with async scheduling ( #38726 )
...
Signed-off-by: Jing Wang <jingwang96@qq.com >
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com >
2026-04-17 16:02:50 +00:00
Michael Goin and GitHub
1174723eba
Fix TURBOQUANT backend selection in cuda.py ( #40060 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-04-17 07:31:41 -07:00
sychen52 and GitHub
6b2b7bd0eb
Add nvfp4 support to reshape_and_cache_flash ( #37332 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-04-17 07:28:00 -07:00
Ben Browning and GitHub
70770268c3
Add @bbrowning to CODEOWNERS ( #40141 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-04-17 09:51:48 -04:00
Chauncey and GitHub
7a51b3e415
[Bugfix] Fix empty delta detection in Qwen3XMLToolParser streaming ( #40090 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-17 13:34:55 +00:00
Li, Jiang and GitHub
d02421a7db
[CPU] Refactor CPU affinity and memory management ( #39781 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-17 21:01:08 +08:00
Lukas Geiger and GitHub
b1dc87a098
[Models][Gemma4] Prevent GPU/CPU sync in embed_input_ids ( #39234 )
...
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com >
2026-04-17 12:37:21 +00:00
Or Ozeri and GitHub
79a5b63253
[kv_offload]: Fix num CPU blocks for UniformTypeKVCacheSpecs ( #39617 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-17 15:13:55 +03:00
Maral and GitHub
c0c98b8b9a
[Bugfix] Add Marlin kernel in block scaled mm kernel selection. ( #40105 )
...
Signed-off-by: maral <maralbahari.98@gmail.com >
2026-04-17 10:20:32 +00:00
wang.yuqi and GitHub
8d2cff8140
[Examples] Resettle Observability examples. ( #40123 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-17 03:13:31 -07:00
Cyrus Leung and GitHub
4f436782af
[Misc] Improve new PR bot trigger condition ( #40114 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-04-17 16:56:22 +08:00
z1ying and GitHub
bf45e6d0a5
[Doc] Add Gemma 4 to supported models list ( #39607 )
...
Signed-off-by: z1ying <tzzying@outlook.com >
Signed-off-by: Ziying Tao <tzzying@outlook.com >
2026-04-17 13:42:52 +08:00
978a4462bb
[CI Failure] Fix Plugin Tests (2 GPUs) Failure ( #40083 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Michele Gazzetti <michele.gazzetti1@ibm.com >
2026-04-17 04:17:39 +00:00
Michael Goin and GitHub
1948d0c467
[UX] Defer some imports on CLI paths to save ~2s ( #40056 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-04-16 19:48:37 -07:00
Shinichi Hemmi and GitHub
4c47710bf7
[CI/Build] Apply ruff formatter to pass pre-commit ( #40078 )
...
Signed-off-by: Hemmi Shinichi <shemmi@preferred.jp >
2026-04-17 08:54:32 +08:00
Giancarlo Delfin and GitHub
bf9a5ddb24
[MLA] Optimize mla indexer prepare uniform decode for MTP > 1 ( #39458 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-16 16:27:51 -07:00
bnellnm and GitHub
79e799ebbd
[Bugfix] Temporarily disable B200 fp4 MoE layer tests ( #40057 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-04-16 19:26:55 -04:00
Netanel Haber and GitHub
c4e601c73c
Bugfix: Parakeet: .conv.pointwise/depthwise_conv1/2.bias weigths can exist even if convolution_bias=False ( #40007 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-16 23:22:05 +00:00
BadrBasowid and GitHub
29057d3bee
[Compilation] Add Unit Tests for VllmFusionPatternMatcherPass ( #39692 )
...
Signed-off-by: BadrBasowid <badr.basowid@gmail.com >
2026-04-16 22:57:16 +00:00
Matthew Bonanni and GitHub
219bb5b8c0
[Misc] Update committers.md ( #40058 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-16 13:48:41 -07:00
Asaf Gardin and GitHub
ad2b1277f9
[Quantization] Consolidate experts_int8 with fp8 online quantization ( #38463 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-04-16 13:12:20 -07:00
roikoren755 and GitHub
b897f00c9c
Gate SSU dispatch setup ( #40039 )
...
Signed-off-by: Roi Koren <roik@nvidia.com >
2026-04-16 13:06:01 -07:00
adf9bb3c57
[CI] Add weight transfer tests to CI ( #39821 )
...
Signed-off-by: SumanthRH <sumanthrh99@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-16 15:51:45 -04:00
Flora Feng and GitHub
b16fda62b7
[Misc] Add @sfeng33 to CODEOWNERS ( #40048 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-04-16 12:25:29 -07:00
Yufeng He and GitHub
de111f3246
[Bugfix] Fix bench_serve UTF-8 decode crash on split multi-byte chars ( #38732 )
2026-04-16 12:01:25 -07:00
Jared Wen and GitHub
afabb5f45a
[bugfix] Normalize tool message content from array to string format ( #39899 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-04-16 11:54:39 -07:00
Roger Wang and GitHub
3abb7560c0
[Bugfix] Fix audioflamingo test ( #40052 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-04-16 11:53:58 -07:00
Isotr0py and GitHub
617d1c2ff1
[Misc] Move pyav and soundfile to common requirements ( #39997 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-16 08:52:37 -07:00
Nikita Shapovalov and GitHub
692db29cd4
[Bugfix] Fix Ray compiled-DAG SHM channel stalls by detaching zero-copy np.ndarray logprobs buffers ( #35736 )
...
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
2026-04-16 23:49:29 +08:00
Isotr0py and GitHub
82531edbfb
[Refactor] Remove resampy dependency ( #39524 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-04-16 08:48:17 -07:00
Nicolò Lucchesi and GitHub
3daca38e22
[Misc] toy_proxy_server handle min_tokens ( #39706 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-16 15:08:22 +00:00
daiyu1111 and GitHub
a302a8fd1b
[Bugfix] Fix LLM priority normalization for single-string prompts ( #40011 )
...
Signed-off-by: daiyu1111 <2356690121@qq.com >
2026-04-16 07:56:06 -07:00
4e8c3f1c19
[Frontend][last/5] Improve pooling entrypoints | clean up. ( #39675 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-16 07:53:23 -07:00
Vasiliy Kuznetsov and GitHub
5e5afafa21
[Doc] add docs for online quant frontend ( #39736 )
...
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com >
2026-04-16 07:52:58 -07:00
Li, Jiang and GitHub
324a3d2bd8
[CI/Build] Improve stability of CPU tests ( #39966 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-04-16 21:50:36 +08:00
4269b79409
[Model] Use mm_features to compute mrope positions for PaddleOCR-VL ( #39888 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
2026-04-16 06:14:00 -07:00
edc3648966
[Kernel][Helion] Fix inductor fusion of Helion HOP ( #39944 )
...
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-16 04:41:26 -07:00
Nicolò Lucchesi and GitHub
9965f501a8
[Nixl] Bump Nixl version to 0.10.1 ( #39922 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-16 11:53:21 +01:00
lalit10 and GitHub
17d87168d2
[Model] Use mm_features for Keye-VL and Keye-1.5-VL M-RoPE ( #39869 )
...
Signed-off-by: Lalit Laxminarayan Bangad <lalitbangad@gmail.com >
2026-04-16 02:16:06 -07:00
Netanel Haber and GitHub
98700c6105
Fix #33773 : Replace unconditional pandas import with PlaceholderModule ( #39990 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-04-16 02:06:51 -07:00
Simon Mo and GitHub
10e49d2638
[Docs] Update PR template to remove release notes google docs ( #39982 )
...
Signed-off-by: Simon Mo <simon.mo@hey.com >
2026-04-16 00:22:03 -07:00
Tim Messerschmidt and GitHub
8d7c962833
[Bugfix] Accept **kwargs in MiniMaxM2Parser.__init__() ( #39861 )
...
Signed-off-by: Tim Messerschmidt <timmesserschmidt@gmail.com >
2026-04-16 15:18:32 +08:00
f4ddaf8cf7
[XPU] use spawn multiproc method on xpu ( #39671 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-16 14:42:07 +08:00
2cdf86044d
Add Jina Embeddings v5 model support ( fixes #38633 ) ( #39575 )
...
Signed-off-by: Abhijit <abroy@redhat.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-16 06:37:10 +00:00
realliujiaxu and GitHub
7845379230
[Bugfix] add support for 'num_attention_groups' in ModelArchConfigConvertorBase for Step3p5 ( #39796 )
...
Signed-off-by: realliujiaxu <realliujiaxu@163.com >
2026-04-16 05:48:00 +00:00
R3hankhan and GitHub
4b7ca37bd4
[CPU][IBM Z][Dockefile][Docs] Fix s390x builds for torch 2.11 and update docs for s390x ( #39910 )
...
Signed-off-by: Rehan Khan <Rehan.Khan7@ibm.com >
2026-04-15 22:26:21 -07:00
445b7093fd
[perf][cpu] Accelerate BF16 GELU with LUT impl on Arm CPUs ( #37469 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-15 22:26:17 -07:00
18013df6ae
[Bugfix] Reject empty tools array with HTTP 400 ( #39780 )
...
Signed-off-by: Jigang Zhou <zjg0907008@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-16 12:08:04 +08:00
Julien Denize and GitHub
c0722f22de
[Mistral Grammar] Fix tool and reasoning parsing ( #39217 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-15 21:05:04 -07:00
Zhengxu Chen and GitHub
951dca8019
[compile] Invoke split FX graph by codegen. ( #38657 )
...
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
2026-04-15 21:03:41 -07:00
vllmellm and GitHub
5f7fab881a
[ROCm][FEAT] Integrate aiter gemm w8a8 ptpc ( #33773 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-04-16 09:55:29 +08:00
Giancarlo Delfin and GitHub
343f65234b
[Model Runner V2][BugFix] fix num_sampled dtype for probabilistic rej… ( #39951 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-04-15 18:09:11 -07:00
Asaf Gardin and GitHub
19fa90ed0d
[Quantization] - Layerwise reloading of Attention/KV quantized models ( #38995 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
2026-04-15 18:03:32 -07:00
03f8d3a548
Update to transformers v5 ( #30566 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Signed-off-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: khluu <khluu000@gmail.com >
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com >
Co-authored-by: jiang1.li <jiang1.li@intel.com >
2026-04-15 16:29:15 -07:00
6dc9491406
[Model] Fix Gemma 4 token repetition by dynamic BOS injection for PT models ( #39842 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-04-15 16:13:07 -07:00
Collin McCarthy and GitHub
27c0ca50a0
Update registry for Nemotron-v3 VL Nano/Super ( #39747 )
...
Signed-off-by: Collin McCarthy <cmccarthy@nvidia.com >
2026-04-15 16:09:11 -07:00
Wentao Ye and GitHub
7c636432c6
[CI Bug] fix flaky test ( #39938 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-15 17:20:06 -04:00
Matthew Bonanni and GitHub
c77e596e2e
[FlashAttention] Don't overwrite flash_attn_interface.py when installing precompiled ( #39932 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-04-15 16:43:15 -04:00
Benjamin Chislett and GitHub
ac3dac545b
[Bugfix][Perf] Indexer upcast WK to BF16 for fusion ( #38928 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-15 20:39:32 +00:00
Wentao Ye and GitHub
39ac640490
[Bug] Fix batch invariant test issue, bs=1 with max_seq_num = 1 ( #39320 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-04-15 16:28:43 -04:00
zhanqiuhu and GitHub
0b790a2501
[Speculative Decoding] Add DFlash speculators config parsing ( #38300 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-15 16:22:15 -04:00
zhanqiuhu and GitHub
41488f2acd
[Bugfix][NIXL] Fix _logical_to_kernel_block_ids conversion for non-mamba models ( #39724 )
...
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-04-15 20:08:58 +00:00
102d51c9f3
[CI] Only build release Docker images when NIGHTLY=1 ( #39882 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-15 19:01:13 +00:00
55e1a8e103
[Mooncake] Fix mixed MLA+Eagle block-size validation ( #39596 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-15 11:36:47 -07:00
Monishver and GitHub
21e5a9f48e
Bug/test eagle dp v2 ( #39838 )
...
Signed-off-by: Monishver Chandrasekaran <monishverchandrasekaran@gmail.com >
2026-04-15 17:48:12 +00:00
Mark McLoughlin and GitHub
8ad6ff0037
[Test] Fix @create_new_process_for_each_test("fork") in interactive shell pipeline ( #29130 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-04-15 12:22:20 -04:00
f2145efcb6
[BugFix] KeyError on scope["method"] for realtime api websocket in AuthenticationMiddleware ( #36934 )
...
Signed-off-by: daniebrill <50454544+daniebrill@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-15 16:15:01 +00:00
Roy Huang and GitHub
ed33310552
[KVConnector][LMCache] Propagate cache_salt through MP connector for per-user cache isolation ( #39837 )
...
Signed-off-by: royyhuang <royyhuang@gmail.com >
Signed-off-by: royyhuang <roy.y.huang@gmail.com >
2026-04-15 09:10:49 -07:00
3cc328a4be
[SpecDecode][Benchmark] Add SPEED-bench support to benchmarking CLI ( #36029 )
...
Signed-off-by: talora <talora@nvidia.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
2026-04-15 12:00:07 -04:00
3beb57a238
[XPU] properly handle q_descale on XPU as quant query input not supported ( #39676 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-04-15 21:52:58 +08:00
8b5531933a
FIX: support language_model.backbone naming in NemotronH Nano VL quantization config ( #39901 )
...
Signed-off-by: <>
Co-authored-by: root <root@lyris0144.lyris.clusters.nvidia.com >
2026-04-15 13:49:48 +00:00
Chauncey and GitHub
db8d4a4a06
[BugFix][Graph] fix: handle empty sym_shape_indices in PiecewiseBackend. ( #39395 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-15 09:28:09 -04:00
zofia and GitHub
fc701c8058
[XPU][MXFP4] add mxfp4 quant op for XPU ( #39857 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
2026-04-15 12:28:19 +00:00
Csrayz and GitHub
68be0f853e
[Metrics] Add request_id to FinishedRequestStats to enable correlation between metrics and requests ( #39710 )
...
Enables external `StatLogger` plugins to correlate per-request metrics
with request-level context. Also, this is a pre-requisite for Prometheus
exemplars in #30972 .
Signed-off-by: Csrayz <33659823+Csrayz@users.noreply.github.com >
2026-04-15 11:24:17 +00:00
Zhenzhong Xu and GitHub
60995c05b4
[Quantization][Autoround][CPU] Add W4A16 Support ( #38192 )
...
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
2026-04-15 18:38:31 +08:00
Yan Ma and GitHub
29e5d10205
fix online fp8 for MiniCPM models ( #39862 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-04-15 09:09:20 +00:00
Or Ozeri and GitHub
235e1f930a
[kv_offload+HMA][3/N]: Remove block_size from KVEvents ( #36644 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-04-15 11:53:19 +03:00
+86
431cea3eea
[Bugfix] Fix tool_calls Iterable consumed when debug logging is enabled ( #34844 )
...
Signed-off-by: Wojciech Wais <wojciech.wais@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Jaebok Lee <jaebok9541@naver.com >
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
Signed-off-by: yuwei <yuwei@dev.local >
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Signed-off-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Signed-off-by: Li <chuali@amd.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: lkm2835 <lkm2835@gmail.com >
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Signed-off-by: vnadathur <glvikramn@gmail.com >
Signed-off-by: WorldExplored <srreyansh.sethi@gmail.com >
Signed-off-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Elham Harirpoush <elham.harirpoush@arm.com >
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: jackcfwang <jackcfwang@tencent.com >
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
Signed-off-by: Richard Zou <zou3519@gmail.com >
Signed-off-by: milesial <milesial@users.noreply.github.com >
Signed-off-by: Elvir Crncevic <elvircrn@gmail.com >
Signed-off-by: whx-sjtu <2952154980@qq.com >
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: PatchouliTaisa <patchychen@tencent.com >
Signed-off-by: jatseng-ai <jatseng@amd.com >
Signed-off-by: jatseng-ai <janet.tseng@amd.com >
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
Signed-off-by: xaguilar-amd <xaguilar@amd.com >
Signed-off-by: rdondeti <ravitez.dondeti@gmail.com >
Signed-off-by: Ravitez Dondeti <ravitez.dondeti@gmail.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Signed-off-by: Peter Nguyen <petern0408@gmail.com >
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: Jesus Federico <jefp@amazon.com >
Signed-off-by: manu <fortin.emmanuel@gmail.com >
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: Yifan Zong <yzong@redhat.com >
Signed-off-by: Rahul-Tuli <rtuli@redhat.com >
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: leeyongjun <jqueen.astro@gmail.com >
Signed-off-by: Ziying Tao <tzzying@outlook.com >
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Signed-off-by: ShubyM <shubymishra20@gmail.com >
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: EdalatiAli <aliedalati@cohere.com >
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: r266-tech <r266.tech@gmail.com >
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Signed-off-by: Animesh Jain <anijain@umich.edu >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: zhxchen17 <zhxchen17@fb.com >
Signed-off-by: EricccYang <yangyang4991@gmail.com >
Signed-off-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Signed-off-by: sihao.li <sihao.li@intel.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com >
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Santino Ramos <santinor@inferact.ai >
Signed-off-by: haosdent <haosdent@gmail.com >
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: George-ao <yuyiao772@gmail.com >
Signed-off-by: Yuyi Ao <yuyiao772@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Mukesh Baphna <mukesh@hippocraticai.com >
Signed-off-by: Pedram Razavi <pedram.razavi@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: Rishi Puri <riship@nvidia.com >
Co-authored-by: zzaebok <44357534+zzaebok@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: yuwei <yuwei@dev.local >
Co-authored-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Ibrahim Arshad <38925737+ibrahim1023@users.noreply.github.com >
Co-authored-by: Chuan (Richard) Li <chuali@amd.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Ganesh R <ganesh.r@amd.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
Co-authored-by: Kyungmin Lee <30465912+lkm2835@users.noreply.github.com >
Co-authored-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Srreyansh Sethi <107075589+WorldExplored@users.noreply.github.com >
Co-authored-by: vnadathur <glvikramn@gmail.com >
Co-authored-by: vnadathur <236933696+vnadathur@users.noreply.github.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Elham <elham.harirpoush@arm.com >
Co-authored-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Chaofan Wang <jackcfwang@tencent.com >
Co-authored-by: Chendi.Xue <chendi.xue@intel.com >
Co-authored-by: Injae Ryou <injaeryou@gmail.com >
Co-authored-by: Richard Zou <zou3519@users.noreply.github.com >
Co-authored-by: milesial <milesial@users.noreply.github.com >
Co-authored-by: Elvir Crnčević <elvircrn@gmail.com >
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com >
Co-authored-by: Hexiang Wang <56632993+whx-sjtu@users.noreply.github.com >
Co-authored-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: PatchyTIS <58251192+PatchouliTIS@users.noreply.github.com >
Co-authored-by: PatchouliTaisa <patchychen@tencent.com >
Co-authored-by: jatseng-ai <janet.tseng@amd.com >
Co-authored-by: Matthias Gehre <matthias.gehre@amd.com >
Co-authored-by: xaguilar-amd <xavier.aguilarfruto@amd.com >
Co-authored-by: Ravitez Dondeti <dondetir@users.noreply.github.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Peter Nguyen <petern0408@gmail.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: zhrrr <43847754+izhuhaoran@users.noreply.github.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Jesus Federico <14651+jefp@users.noreply.github.com >
Co-authored-by: Manu <efortin@users.noreply.github.com >
Co-authored-by: zhanqiuhu <49648934+ZhanqiuHu@users.noreply.github.com >
Co-authored-by: yzong-rh <yzong@redhat.com >
Co-authored-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: Rahul-Tuli <rtuli@redhat.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <bchislett@nvidia.com >
Co-authored-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Co-authored-by: Lee Yongjun <35302114+elwhyjay@users.noreply.github.com >
Co-authored-by: z1ying <55220715+z1ying@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
Co-authored-by: Vibhav Agarwal <vibhavagarwal5@gmail.com >
Co-authored-by: vibhav-agarwal <vibhav.agarwal@glance.com >
Co-authored-by: ShubyM <shubymishra20@gmail.com >
Co-authored-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: EdalatiAli <aliedalati@cohere.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: r266-tech <r2668940489@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Le Yang <562593859@qq.com >
Co-authored-by: Animesh Jain <anijain@umich.edu >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Zhengxu Chen <zhxchen17@fb.com >
Co-authored-by: Kaicheng Yang <53411596+EricccYang@users.noreply.github.com >
Co-authored-by: maobaolong <baoloongmao@tencent.com >
Co-authored-by: sihao_li <165983188+1643661061leo@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
Co-authored-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: zofia <110436990+zufangzhu@users.noreply.github.com >
Co-authored-by: Tihomir Elek <tiho.elek@gmail.com >
Co-authored-by: Yi Liu <yi4.liu@intel.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Santino Ramos <51103228+santiramos27@users.noreply.github.com >
Co-authored-by: haosdent <haosdent@gmail.com >
Co-authored-by: JartX <sagformas@epdcenter.es >
Co-authored-by: Yuyi Ao <yuyiao772@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: mukesh-hai <mukesh@hippocraticai.com >
Co-authored-by: Pedram Razavi <pedram@sierra.ai >
2026-04-15 01:32:47 -07:00