Or Ozeri and GitHub
2fa1f8ec00
[kv_offload+HMA][13/N]: Enable HMA support ( #41445 )
...
This is the final PR in a series to enables HMA support for the
offloading connector. The connector advertises `SupportsHMA`
and is validated with unit tests and e2e tests.
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-05-01 12:30:03 +01:00
a3ec4a35f5
[Bugfix][Metrics] Fix RayPrometheusMetric.labels() returning shared labeled child ( #40840 )
...
When vLLM runs with Ray Prometheus `vllm:request_success{finished_reason=...}`
only ever increments the repetition bucket regardless of the request's actual finish
reason; stop, length, abort, and error stay at zero. Root cause was `labels()` mutated
the wrapped Ray metric's default tags in place and returned self, so every `.labels(...)`
call on a given wrapper returned the same object.
Co-authored-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Signed-off-by: Marwan Sarieddine <sarieddine.marwan@gmail.com >
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
2026-05-01 08:43:39 +01:00
a07642667d
[Bugfix] Pass reasoning parser kwargs to structured output ( #41199 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-04-30 23:38:02 -07:00
baonudesifeizhai and GitHub
c3868bbbe4
[compile] Add FlashInfer FP8 async TP fusion and preserve allreduce fusion ordering #27893 ( #39505 )
...
Signed-off-by: baonudesifeizhai <baonudesifeizhai@gmail.com >
Signed-off-by: baonudesifeizhai <85092850+baonudesifeizhai@users.noreply.github.com >
Signed-off-by: roG0d <baonudesifeizhai@gmail.com >
2026-05-01 05:08:34 +00:00
sychen52 and GitHub
947138b6c2
Add nvfp4 kv cache support ( #40177 )
...
Signed-off-by: Shiyang Chen <shiychen@nvidia.com >
2026-05-01 04:55:16 +00:00
6b6ac6c3c7
[Kernel][MoE] Support GELU on TRT-LLM NvFP4 fused MoE for Gemma4 ( #41050 )
...
Signed-off-by: Juhi Mittal <juhim@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 03:37:43 +00:00
Dong W and GitHub
7198940b39
[Model] Add Moondream3 model support(only query and caption skills) ( #32325 )
...
Signed-off-by: Dong Wang <dongw2019@gmail.com >
2026-05-01 10:06:48 +08:00
14043dfecd
feat: Enable prompt_embeds Content Part Support in vLLM Chat Completions API ( #40720 )
...
Signed-off-by: Luis Robaina <luis@protopia.ai >
Signed-off-by: Luis Robaina 🚀 <luisfabian1545@gmail.com >
Signed-off-by: LuisRobaina <luis@protopia.ai >
Co-authored-by: Andrew Sansom <qthequartermasterman@gmail.com >
2026-05-01 10:05:55 +08:00
Andreas Karatzas and GitHub
1adaa5056b
[ROCm][CI] Add ROCm score absolute tolerance floor ( #41341 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-04-30 18:59:35 -07:00
Nick Hill and GitHub
dd5506a157
[Core] Simplify handling of scheduler_reserve_full_isl option ( #41064 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-30 18:10:00 -07:00
a3c83ff2fd
Faster per-token fp8 group quant packed kernel for blackwell ( #41326 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-30 18:09:55 -07:00
2917d6363a
[NVFP4][Hopper/AMD Instinct] Add Triton kernels for NVFP4 dequantization and QDQ emulation ( #40033 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-04-30 17:35:48 -04:00
Stefano Castagnetta and GitHub
efb4cdf2b8
[CI/Build] Skip Prithvi/Terratorch model-registry tests when terratorch is missing ( #41389 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-30 12:47:55 -07:00
92a7c121b6
[CI] Add MTP coverage: Qwen3.5 correctness + no-sync spec decode ( #40472 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-30 12:24:09 -07:00
Stefano Castagnetta and GitHub
10558f5f46
[CI/Build] Skip terratorch + torchgeo while PyPI has lightning quarantined ( #41377 )
...
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-30 07:59:07 -07:00
snadampal and GitHub
3179e53135
[P/D] Prefill compute optimizations with bi-directional KV cache transfers between P and D nodes ( #32553 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
2026-04-30 10:14:20 +00:00
Nicolò Lucchesi and GitHub
efdc95674d
[KVConnector] MultiConnector SupportsHMA ( #39571 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-04-30 02:10:50 -07:00
54146a9bf9
[Bugfix] correct h matrix layout in chunk_kda output kernel ( #40956 )
...
Signed-off-by: ChenxiQian <chenxi.qian.cq@outlook.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-30 16:22:41 +08:00
Ekagra Ranjan and GitHub
a04e0cf3b8
Fix Cohere ASR after HF upgrade ( #40582 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-04-29 23:39:04 -07:00
cb1b02d0e8
[Frontend] Add VLLM_SKIP_MODEL_NAME_VALIDATION environment variable ( #34676 )
...
Signed-off-by: Dhruv Singal <dhruvsingalabc@gmail.com >
Signed-off-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local >
Signed-off-by: Your Name <you@example.com >
Signed-off-by: vLLM Assistant <assistant@vllm.ai >
Signed-off-by: Simon Mo <simon.mo@hey.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local >
Co-authored-by: Your Name <you@example.com >
Co-authored-by: OpenCode <noreply@openai.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-04-29 23:19:09 -07:00
c42981d034
[Refactor][kv_offload] KV Offloading maintainability improvements ( #40538 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-30 05:55:31 +03:00
Wei Zhao and GitHub
0ff1bf9bb1
[Bugfix] Fix failure to allocate KV blocks error ( #41282 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-04-29 18:44:07 -07:00
Nick Hill and GitHub
18599bfdf2
[Ci][BugFix] Fix slow DP tests due to bad teardown logic ( #41166 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-29 19:31:00 -04:00
Thien Tran and GitHub
296741d025
[DSv4] Use cvt PTX for FP32->FP4 conversion ( #41015 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-04-29 16:16:40 -07:00
Hemanth Acharya and GitHub
6841f5dc77
[ROCm] Add env flags to disable dynamic MXFP4 quant and enable AITER tuned GEMMs for Attention Projection Layers ( #39987 )
...
Signed-off-by: Hemanth Acharya <heachary@amd.com >
2026-04-29 16:07:46 -07:00
ccfb620c62
Create tests/distributed/test_mnnvl_alltoall.py ( #35241 )
...
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Claude <claude@anthropic.com >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Claude <claude@anthropic.com >
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-04-29 21:56:56 +00:00
0335316a9b
[BUG] Two phase pause to prevent deadlock ( #39366 )
...
Signed-off-by: ahao-anyscale <ahao@anyscale.com >
Signed-off-by: Aaron Hao <ahao@anyscale.com >
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-04-29 17:51:03 -04:00
Laith Sakka and GitHub
6f20f81cbf
Replace shape_invariants with simpler apprach in dynamic_arg_dims utilizing shape_id property. ( #36194 )
...
Signed-off-by: Laith Sakka <lsakka@meta.com >
2026-04-29 18:32:15 +00:00
danisereb and GitHub
d1a75e303d
Fix timeout when using LoRA adapters with Nemotron Super ( #40916 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-04-30 01:39:49 +08:00
Terrence Zhao and GitHub
91a2d39014
[Models] Cohere MoE ( #40817 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-04-29 15:54:54 +00:00
Artem Perevedentsev and GitHub
b92ef9ec5a
[Perf] Enable FlashInfer top-k/top-p sampler by default ( #40376 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
2026-04-29 19:10:34 +04:00
22524f7a92
[Feat] CPU fp8 attn for AMX/AVX-512 ( #39445 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-04-29 20:43:21 +08:00
Bugen Zhao and GitHub
33f36d4260
[DSV4] Support max reasoning effort ( #40982 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-04-29 11:03:47 +00:00
3f1a4bb639
build: embed image provenance metadata in vLLM containers ( #40653 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-04-29 03:07:41 -07:00
Chauncey and GitHub
762022cafb
[Bugfix] DSV32/V4 add missing type conversion for non-streaming tool calls ( #41198 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 09:55:07 +00:00
Chauncey and GitHub
3885d340a4
[Frontend]Responses API supports Tool/Function calling with streaming with named tool/function ( #41110 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 09:11:27 +00:00
Chauncey and GitHub
92879e12ba
[CI] fix test_rotary_embedding_opcheck format error ( #41202 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-04-29 00:32:37 -07:00
68dd7db810
[Reasoning] Support for speculative decoding with thinking budget ( #34668 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-04-29 06:14:52 +00:00
8a8c9b564e
[KV Offload] Per-job store completion for CPU offloading connector ( #39186 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-04-29 08:52:55 +03:00
Jee Jee Li and GitHub
a269744e9f
[Bugfix] Fix rope ( #41113 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-04-28 22:42:35 -07:00
8b49cf3a37
[Bugfix] Fix max_num_batched_token not captured in cuda graph ( #40734 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com >
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-bia02.bia.clusters.nvidia.com >
2026-04-28 21:33:06 -07:00
liangel-02 and GitHub
7fd05e05ae
uncomment flex backend for batch invariant mode ( #40842 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-04-28 21:05:14 -07:00
haosdent and GitHub
75a7cf2c10
[CI] De-flake test_chat_completion_n_parameter_non_streaming ( #41147 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-29 03:23:59 +00:00
haosdent and GitHub
4b95e9cec4
[CI] Return HTTP 400 for unsupported chat content part type ( #41121 )
...
Signed-off-by: haosdent <haosdent@gmail.com >
2026-04-29 10:23:26 +08:00
rasmith and GitHub
856b15c62c
[CI][AMD][BugFix] Patch has_flashinfer decorator for test_select_rocm_aiter_backend ( #41072 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-29 02:12:17 +00:00
Nick Hill and GitHub
e68fa1b90a
[Core] Account for num_gpu_blocks_override in max_model_len checks ( #41069 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-04-28 15:44:09 -07:00
Julien Denize and GitHub
e9f8f31e9a
[FEATURE] Add EagleMistralForCausalLM ( #41024 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-04-28 12:22:20 -07:00
de3fe8dc62
[Bugfix] release KV blocks for skipped P-ranks to prevent invalid KV errors and timeouts when P_tp > D_tp and MLA ( #40449 )
...
Signed-off-by: yangruize <yangruize7@163.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-04-28 11:38:43 -07:00
0899f436aa
[New Model] Laguna XS.2 implementation ( #41129 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com >
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com >
2026-04-28 14:23:00 -04:00
rasmith and GitHub
358a755e43
[CI][AMD][BugFix] Update request URL in test_moriio_connector to match vllm-router compatibility changes ( #41076 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-04-28 13:14:59 -05:00