yewentao256
|
77230471c0
|
remove torch 2.9, 2.10 workround
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-30 21:10:50 +00:00 |
|
 Stefano CastagnettaandGitHub
|
efb4cdf2b8
|
[CI/Build] Skip Prithvi/Terratorch model-registry tests when terratorch is missing (#41389)
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
|
2026-04-30 12:47:55 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
92a7c121b6
|
[CI] Add MTP coverage: Qwen3.5 correctness + no-sync spec decode (#40472)
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-30 12:24:09 -07:00 |
|
 Stefano CastagnettaandGitHub
|
10558f5f46
|
[CI/Build] Skip terratorch + torchgeo while PyPI has lightning quarantined (#41377)
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
|
2026-04-30 07:59:07 -07:00 |
|
 snadampalandGitHub
|
3179e53135
|
[P/D] Prefill compute optimizations with bi-directional KV cache transfers between P and D nodes (#32553)
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
|
2026-04-30 10:14:20 +00:00 |
|
 Nicolò LucchesiandGitHub
|
efdc95674d
|
[KVConnector] MultiConnector SupportsHMA (#39571)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-04-30 02:10:50 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
54146a9bf9
|
[Bugfix] correct h matrix layout in chunk_kda output kernel (#40956)
Signed-off-by: ChenxiQian <chenxi.qian.cq@outlook.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-04-30 16:22:41 +08:00 |
|
 Ekagra RanjanandGitHub
|
a04e0cf3b8
|
Fix Cohere ASR after HF upgrade (#40582)
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com>
|
2026-04-29 23:39:04 -07:00 |
|
     
|
cb1b02d0e8
|
[Frontend] Add VLLM_SKIP_MODEL_NAME_VALIDATION environment variable (#34676)
Signed-off-by: Dhruv Singal <dhruvsingalabc@gmail.com>
Signed-off-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local>
Signed-off-by: Your Name <you@example.com>
Signed-off-by: vLLM Assistant <assistant@vllm.ai>
Signed-off-by: Simon Mo <simon.mo@hey.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Dhruv Singal <dsingal@Dhruvs-MacBook-Pro.local>
Co-authored-by: Your Name <you@example.com>
Co-authored-by: OpenCode <noreply@openai.com>
Co-authored-by: Simon Mo <simon.mo@hey.com>
|
2026-04-29 23:19:09 -07:00 |
|
 
|
c42981d034
|
[Refactor][kv_offload] KV Offloading maintainability improvements (#40538)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-04-30 05:55:31 +03:00 |
|
 Wei ZhaoandGitHub
|
0ff1bf9bb1
|
[Bugfix] Fix failure to allocate KV blocks error (#41282)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-04-29 18:44:07 -07:00 |
|
 Nick HillandGitHub
|
18599bfdf2
|
[Ci][BugFix] Fix slow DP tests due to bad teardown logic (#41166)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-29 19:31:00 -04:00 |
|
 Thien TranandGitHub
|
296741d025
|
[DSv4] Use cvt PTX for FP32->FP4 conversion (#41015)
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2026-04-29 16:16:40 -07:00 |
|
 Hemanth AcharyaandGitHub
|
6841f5dc77
|
[ROCm] Add env flags to disable dynamic MXFP4 quant and enable AITER tuned GEMMs for Attention Projection Layers (#39987)
Signed-off-by: Hemanth Acharya <heachary@amd.com>
|
2026-04-29 16:07:46 -07:00 |
|
  
|
ccfb620c62
|
Create tests/distributed/test_mnnvl_alltoall.py (#35241)
Signed-off-by: Rishi Puri <riship@nvidia.com>
Signed-off-by: Claude <claude@anthropic.com>
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
Co-authored-by: Claude <claude@anthropic.com>
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com>
|
2026-04-29 21:56:56 +00:00 |
|
  
|
0335316a9b
|
[BUG] Two phase pause to prevent deadlock (#39366)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
Signed-off-by: Aaron Hao <ahao@anyscale.com>
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-29 17:51:03 -04:00 |
|
 Laith SakkaandGitHub
|
6f20f81cbf
|
Replace shape_invariants with simpler apprach in dynamic_arg_dims utilizing shape_id property. (#36194)
Signed-off-by: Laith Sakka <lsakka@meta.com>
|
2026-04-29 18:32:15 +00:00 |
|
 daniserebandGitHub
|
d1a75e303d
|
Fix timeout when using LoRA adapters with Nemotron Super (#40916)
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
|
2026-04-30 01:39:49 +08:00 |
|
 Terrence ZhaoandGitHub
|
91a2d39014
|
[Models] Cohere MoE (#40817)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
|
2026-04-29 15:54:54 +00:00 |
|
 Artem PerevedentsevandGitHub
|
b92ef9ec5a
|
[Perf] Enable FlashInfer top-k/top-p sampler by default (#40376)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-29 19:10:34 +04:00 |
|
  
|
22524f7a92
|
[Feat] CPU fp8 attn for AMX/AVX-512 (#39445)
Signed-off-by: Li, Tianmu <tianmu.li@intel.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-04-29 20:43:21 +08:00 |
|
 Bugen ZhaoandGitHub
|
33f36d4260
|
[DSV4] Support max reasoning effort (#40982)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-04-29 11:03:47 +00:00 |
|
 
|
3f1a4bb639
|
build: embed image provenance metadata in vLLM containers (#40653)
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-04-29 03:07:41 -07:00 |
|
 ChaunceyandGitHub
|
762022cafb
|
[Bugfix] DSV32/V4 add missing type conversion for non-streaming tool calls (#41198)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-29 09:55:07 +00:00 |
|
 ChaunceyandGitHub
|
3885d340a4
|
[Frontend]Responses API supports Tool/Function calling with streaming with named tool/function (#41110)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-29 09:11:27 +00:00 |
|
 ChaunceyandGitHub
|
92879e12ba
|
[CI] fix test_rotary_embedding_opcheck format error (#41202)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-29 00:32:37 -07:00 |
|
  
|
68dd7db810
|
[Reasoning] Support for speculative decoding with thinking budget (#34668)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Signed-off-by: rishitdholakia13 <123388671+rishitdholakia13@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-04-29 06:14:52 +00:00 |
|
   
|
8a8c9b564e
|
[KV Offload] Per-job store completion for CPU offloading connector (#39186)
Signed-off-by: Itay Etelis <itay.etelis@ibm.com>
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com>
Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Or Ozeri <or@ozery.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-04-29 08:52:55 +03:00 |
|
 Jee Jee LiandGitHub
|
a269744e9f
|
[Bugfix] Fix rope (#41113)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-04-28 22:42:35 -07:00 |
|
 
|
8b49cf3a37
|
[Bugfix] Fix max_num_batched_token not captured in cuda graph (#40734)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com>
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-bia02.bia.clusters.nvidia.com>
|
2026-04-28 21:33:06 -07:00 |
|
 liangel-02andGitHub
|
7fd05e05ae
|
uncomment flex backend for batch invariant mode (#40842)
Signed-off-by: Angel Li <liangel@meta.com>
|
2026-04-28 21:05:14 -07:00 |
|
 haosdentandGitHub
|
75a7cf2c10
|
[CI] De-flake test_chat_completion_n_parameter_non_streaming (#41147)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-04-29 03:23:59 +00:00 |
|
 haosdentandGitHub
|
4b95e9cec4
|
[CI] Return HTTP 400 for unsupported chat content part type (#41121)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-04-29 10:23:26 +08:00 |
|
 rasmithandGitHub
|
856b15c62c
|
[CI][AMD][BugFix] Patch has_flashinfer decorator for test_select_rocm_aiter_backend (#41072)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-04-29 02:12:17 +00:00 |
|
 Nick HillandGitHub
|
e68fa1b90a
|
[Core] Account for num_gpu_blocks_override in max_model_len checks (#41069)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-04-28 15:44:09 -07:00 |
|
 Julien DenizeandGitHub
|
e9f8f31e9a
|
[FEATURE] Add EagleMistralForCausalLM (#41024)
Signed-off-by: juliendenize <julien.denize@mistral.ai>
|
2026-04-28 12:22:20 -07:00 |
|
 
|
de3fe8dc62
|
[Bugfix] release KV blocks for skipped P-ranks to prevent invalid KV errors and timeouts when P_tp > D_tp and MLA (#40449)
Signed-off-by: yangruize <yangruize7@163.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-04-28 11:38:43 -07:00 |
|
 
|
0899f436aa
|
[New Model] Laguna XS.2 implementation (#41129)
Signed-off-by: Joe Rowell <joerowell4@gmail.com>
Signed-off-by: Robert Shaw <robertgshaw2@gmail.com>
Co-authored-by: Robert Shaw <robertgshaw2@gmail.com>
|
2026-04-28 14:23:00 -04:00 |
|
 rasmithandGitHub
|
358a755e43
|
[CI][AMD][BugFix] Update request URL in test_moriio_connector to match vllm-router compatibility changes (#41076)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-04-28 13:14:59 -05:00 |
|
 wang.yuqiandGitHub
|
a8208e6a81
|
[Examples] Resettle features examples. (#40995)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-28 00:33:41 -07:00 |
|
  
|
7a1eb8ac2e
|
[Model] update for mimo v25 (#41029)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Isotr0py <Isotr0py@outlook.com>
Co-authored-by: Copilot <copilot@github.com>
|
2026-04-27 21:52:54 -07:00 |
|
 Matthew BonanniandGitHub
|
fd74c90d9c
|
[Attention][Spec Decode] Allow independent drafter attention backend selection (#39930)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-04-27 19:38:09 -07:00 |
|
 ChaunceyandGitHub
|
146f44b77d
|
[Frontend]Responses API supports Tool/Function calling with streaming with required (#40700)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2026-04-27 19:36:58 -07:00 |
|
 
|
0d4f714208
|
[Bugfix] Remove tokenizer encode/decode calls from Olmo3 reasoning parser (#40855)
Signed-off-by: Yifan <yzong@redhat.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-04-27 19:36:54 -07:00 |
|
 Angela YiandGitHub
|
03aeed802f
|
[Test] Fix test_dynamic_shapes_compilation for torch 2.12 (#40743)
Signed-off-by: Angela Yi <angelayi@meta.com>
|
2026-04-27 17:51:15 -07:00 |
|
 Andreas KaratzasandGitHub
|
5e2c37facd
|
[ROCm][CI] Add missing quantization methods and fix online quant test failures (#39801)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-27 15:08:57 -05:00 |
|
 Moritz SanftandGitHub
|
2c06cf3486
|
[Bugfix] use served_model_name for multimodal error message (#41003)
Signed-off-by: Moritz Sanft <58110325+msanft@users.noreply.github.com>
|
2026-04-27 08:22:35 -07:00 |
|
      
|
c245d35ff4
|
[Model] Add MiMo-V2.5 support (#40967)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Signed-off-by: Isotr0py <Isotr0py@outlook.com>
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>
Co-authored-by: zjy0516 <zhujiangyun@inferact.ai>
Co-authored-by: yasong <yasong.wang@inferact.ai>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Copilot <copilot@github.com>
|
2026-04-27 13:26:51 +00:00 |
|
 
|
ebf862c351
|
Add system_fingerprint field to OpenAI-compatible API responses (#40537)
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-04-27 16:17:52 +08:00 |
|
 Yongye ZhuandGitHub
|
706a04d34b
|
[DSV4] Add silu clamp limit to shared expert (#40950)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-04-27 00:37:43 -07:00 |
|