Viktor Pus and GitHub
d5b31c954d
[Bugfix] Account for truncate_prompt_tokens when computing max_tokens ( #41800 )
...
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com >
2026-05-06 16:10:17 +00:00
David Zheng and GitHub
ee38750a75
[Bugfix] Fix spawn_new_process_for_each_test silently swallowing test failures ( #41423 )
...
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com >
2026-05-06 11:17:15 -04:00
27e0057aed
[Spec Decode] Add Gemma4 MTP speculative decoding support ( #41745 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-06 22:39:29 +08:00
Ronen Schaffer and GitHub
f39bcf1e30
[KV Offload] Return None from lookup() for in-flight blocks ( #41795 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-06 17:31:21 +03:00
6467213a9f
fix(openai): tolerate empty content in forced tool choice ( #40148 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-06 07:16:03 -07:00
df8e63f4ed
nixl refactor: new transfer design ( #40731 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
2026-05-06 06:16:25 -07:00
242afc6bf4
[MM][Gemma4] Respect max_soft_tokens in encoder budget ( #41799 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: gemini-code-assist <gemini-code-assist@google.com >
2026-05-06 05:54:42 -07:00
51c1ee9b7c
[Examples] Resettle Disaggregated examples. ( #40759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
Signed-off-by: wang.yuqi <noooop@126.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-06 01:20:38 -07:00
Lucas Kabela and GitHub
213f10bfdd
[Bugfix] Fix codegen for unqualified names ( #40726 )
...
Signed-off-by: Lucas Kabela <lucaskabela@meta.com >
2026-05-06 01:11:37 -07:00
Yuwen Zhou and GitHub
809b98e5b7
[CPU] Add FP8 W8A16 linear support ( #41186 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-05-06 07:05:27 +00:00
wi-adam and GitHub
b53c507bc9
[Bugfix] Skip PP sampled-token receive on last rank during async scheduling ( #40749 )
...
Signed-off-by: Adam Winstanley <adam@winstanley.industries >
2026-05-06 05:31:14 +00:00
16e336491e
[Mistral Tokenizer] allow more leniency in apply_chat_template ( #41658 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-05 19:56:15 -07:00
Matthew Bonanni and GitHub
01b9b5af67
[Attention] Minor refactor: layer takes ownership of the MLA prefill backend ( #41744 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-05 23:22:41 +00:00
Lanze Liu and GitHub
79246b5ea6
[Spec Decode] Fix max_model_len logging in speculative config for draft model ( #41571 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-05-05 21:56:06 +00:00
Julien Denize and GitHub
c6235ed180
[BUGFIX] Support streamed_args_for_tool in MistralToolParser ( #41730 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-05-05 17:48:53 +00:00
628c436301
[New Model][ROCm] Add AMD support for DeepSeek V4 ( #40871 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: tjtanaavllm <tunjian.tan@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaavllm <tunjian.tan@amd.com >
2026-05-05 08:55:37 -07:00
0a201b60cf
[Model] support Qianfan-OCR model ( #40136 )
...
Signed-off-by: bairongz <baiyuu.cs@gmail.com >
Signed-off-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-05 10:51:25 +00:00
8b9ea2f881
[Feature] Add Triton kernel JIT compilation monitor for inference ( #40137 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-05-05 14:08:57 +04:00
bee126165f
[P/D][Mooncake] Add KVConnectorStats for transfer observability ( #40414 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-05 02:17:38 -07:00
Bowen Bao and GitHub
1e9500410a
[ROCm][Quantization][2/N] Refactor quark_moe w4a8 w/ oracle ( #39136 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-04 19:50:38 -07:00
4f2af1a7c0
[Feature] TurboQuant: support hybrid models and uniform quantization ( #39931 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Sandermage <sandermage@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-04 20:14:01 -04:00
Wentao Ye and GitHub
577b9623e6
[Bug] Fix status update address for non-MOE model within external dp mode ( #40839 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-04 16:37:16 -07:00
Andreas Karatzas and GitHub
1cb0838721
[ROCm][CI] Fix MLA prefill scale for DeepSeek GSM8K ( #41569 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 16:32:55 -07:00
844df54269
feat: update xgrammar==0.2.0 to use structural tags for strict tool calling + reasoning for more models ( #40894 )
...
Signed-off-by: Yuchuan <yuchuan.7streams@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Ubospica <ubospica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Ubospica <ubospica@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-04 12:45:24 -07:00
422dd02598
[bugfix] Fix prompt logprobs on request eviction during chunked prefill ( #41411 )
...
Signed-off-by: Joachim Studnia <joachim@mistral.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-04 11:46:00 -07:00
8c780943b4
Fix Nano Nemotron text-only weight loading ( #41205 )
...
Signed-off-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Signed-off-by: Baekpica <35071468+Baekpica@users.noreply.github.com >
Signed-off-by: sunghoon.baek <seanbb93@gmail.com >
Co-authored-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-04 21:43:07 +03:00
712ad0286c
[Bugfix] KimiK2ReasoningParser: guard against buffered end-token in streaming ( #41068 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-04 17:42:05 +00:00
Ekagra Ranjan and GitHub
321fa2d6d1
Limit gpu utils and lower max BS on test_transcription_api_correctness.py ( #41649 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-05-04 10:30:02 -07:00
Netanel Haber and GitHub
8decbfa02c
Test nemotron nano-v2 and nemotron nano-v3 separately, disable super-omni redundant tests ( #41616 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-04 16:31:37 +03:00
Andreas Karatzas and GitHub
6ec9bbec38
[CI] Stabilize cpu offload compressed tensors test ( #41102 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 05:22:42 +00:00
Andreas Karatzas and GitHub
01d4d1ad37
[ROCm][CI] Align spec decode logprob test prefill settings ( #41335 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 04:33:29 +00:00
c103c02a1a
[Transformers v5] Vendor HCXVisionConfig for compatibility ( #38447 )
...
Signed-off-by: Fang Han <fhan0520@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-04 04:19:52 +00:00
Andreas Karatzas and GitHub
67058ca326
[CI] Clean up remote servers on pytest parent exit ( #41570 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 03:11:22 +00:00
Akim Tsvigun and GitHub
894a02500b
[Bench] Forward --seed to CustomDataset and CustomMMDataset shuffle ( #40788 )
...
Signed-off-by: akimtsvigun <akimtsvigun@gmail.com >
2026-05-04 00:39:10 +00:00
66dfee7121
[Bugfix] Fix degenerate KV cache stride causing TMA cudaErrorIllegalInstruction ( #40737 )
...
Signed-off-by: David Oy <david@baseten.co >
Signed-off-by: David Oy <58150256+the-david-oy@users.noreply.github.com >
Signed-off-by: David Oy <david.oy@baseten.co >
Co-authored-by: David Oy <david@baseten.co >
Co-authored-by: Claude <claude@anthropic.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-03 23:52:18 +00:00
Alex Brooks and GitHub
db9a84e0cd
[Bugfix] Fix FP8 Bias Loading ( #41424 )
...
Signed-off-by: Alex Brooks <albrooks@redhat.com >
2026-05-03 20:30:04 +00:00
c3ad791e1a
[Bugfix][Gemma 4] Clamp soft-token estimate to max_soft_tokens ( #40796 )
...
Signed-off-by: Hoang Nguyen <118159510+hnt2601@users.noreply.github.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-05-02 06:34:59 +00:00
Luka Govedič and GitHub
d58c42e19c
[vLLM IR] 2/N fused_add_rms_norm and maybe_inplace overload ( #36823 )
...
Signed-off-by: Luka Govedič <lgovedic@redhat.com >
Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
2026-05-01 23:41:15 -04:00
Ekagra Ranjan and GitHub
3e49479c4b
Limit concurrency on test_transcription_api_correctness.py ( #41478 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-05-02 03:19:07 +00:00
John Calderon and GitHub
964a4bc2a5
[MM][CG] Support ViT CG for Qwen2.5-VL ( #40830 )
...
Signed-off-by: John Calderon <jcalderon@nvidia.com >
2026-05-02 11:10:14 +08:00
f3fef12350
[Attention] Abstract the MLA prefill backends and eliminate cuDNN ( #32623 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 13:36:20 -04:00
Michael Goin and GitHub
3ccc1ff495
[Eval][CI] Add basic mrcr eval to tests/evals/ ( #40164 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-05-01 12:00:38 -04:00
529c671e80
[ROCm][FEAT] AITER Fused Allreduce + RMSNorm ( #37646 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com >
Signed-off-by: junkang1991 <junkangchow@gmail.com >
Co-authored-by: Rita Brugarolas <Rita.BrugarolasBrufau@amd.com >
Co-authored-by: junkang1991 <junkangchow@gmail.com >
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-01 23:07:18 +08:00
sungsoo ha and GitHub
4f7bde572a
[Kernel] Pack output and LSE in DCP A2A ( #41160 )
2026-05-01 09:01:17 -04:00
Or Ozeri and GitHub
2fa1f8ec00
[kv_offload+HMA][13/N]: Enable HMA support ( #41445 )
...
This is the final PR in a series to enables HMA support for the
offloading connector. The connector advertises `SupportsHMA`
and is validated with unit tests and e2e tests.
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-01 12:30:03 +01:00
a3ec4a35f5
[Bugfix][Metrics] Fix RayPrometheusMetric.labels() returning shared labeled child ( #40840 )
...
When vLLM runs with Ray Prometheus `vllm:request_success{finished_reason=...}`
only ever increments the repetition bucket regardless of the request's actual finish
reason; stop, length, abort, and error stay at zero. Root cause was `labels()` mutated
the wrapped Ray metric's default tags in place and returned self, so every `.labels(...)`
call on a given wrapper returned the same object.
Co-authored-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Signed-off-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
2026-05-01 08:43:39 +01:00
a07642667d
[Bugfix] Pass reasoning parser kwargs to structured output ( #41199 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-30 23:38:02 -07:00
baonudesifeizhai and GitHub
c3868bbbe4
[compile] Add FlashInfer FP8 async TP fusion and preserve allreduce fusion ordering #27893 ( #39505 )
...
Signed-off-by: baonudesifeizhai <baonudesifeizhai@gmail.com >
Signed-off-by: baonudesifeizhai <85092850+baonudesifeizhai@users.noreply.github.com >
Signed-off-by: roG0d <baonudesifeizhai@gmail.com >
2026-05-01 05:08:34 +00:00
sychen52 and GitHub
947138b6c2
Add nvfp4 kv cache support ( #40177 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-01 04:55:16 +00:00
6b6ac6c3c7
[Kernel][MoE] Support GELU on TRT-LLM NvFP4 fused MoE for Gemma4 ( #41050 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 03:37:43 +00:00