   
|
520a20ba4e
|
[Bugfix] MoRIIO toy P/D proxy: add /health (#45222)
Signed-off-by: Chaemin Lim <chaemin.lim@mangoboost.io>
Signed-off-by: Edwin Lim <edwinlim0919@gmail.com>
Co-authored-by: Edwin Lim <edwin.lim@mangoboost.io>
Co-authored-by: Jaeyoun Kim <jaeyoun.kim@mangoboost.io>
Co-authored-by: Edwin Lim <edwinlim0919@gmail.com>
|
2026-07-14 22:56:46 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9182e86971
|
Log fully resolved pooling config at startup (#48030)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-14 22:00:46 +00:00 |
|
 Matthew BonanniandGitHub
|
313d01f507
|
[CI][Bugfix] Fix FlashAttention reported MLA dimension support (#48631)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-07-14 21:33:02 +00:00 |
|
 Divakar VermaandGitHub
|
05d4f8bba3
|
[ROCm][CI] fix flashinfer import check (#48647)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-07-14 20:54:19 +00:00 |
|
 Michael GoinandGitHub
|
0b54201a04
|
[CI] Build macOS arm64 CPU wheel natively on the macmini queue (#48289)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-07-14 19:40:26 +00:00 |
|
 
|
32e632dfeb
|
[Reasoning] Optimize TPOT for thinking budget when used with speculative decoding (#46662)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-07-14 18:55:40 +00:00 |
|
 
|
7ffb98e248
|
[ROCm] Retune MI355 selective_state_update float32 config on the unified effective_batch grid (#48373)
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com>
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com>
|
2026-07-14 18:26:35 +00:00 |
|
  
|
cdaa40d2a8
|
[KV Offload] Split cpu_cache_usage_perc into write/read usage gauges (#47666)
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com>
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-14 20:13:41 +03:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
ca3618bc69
|
[Doc] Sync four function docstrings with their signatures (#45437)
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-14 13:10:13 -04:00 |
|
 Michael GoinandGitHub
|
b2f7d2560a
|
[Bugfix] Make MLA+SWA check the layer's backend, not the model config (#48520)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-07-14 09:53:34 -07:00 |
|
 Wentao YeandGitHub
|
1ff9429655
|
[CI Bug] Fully solve accuracy issue for DSv3.2 + MTP + Sequence Parallel (#48036)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-07-14 10:00:24 -04:00 |
|
 
|
af453e5647
|
[Bugfix] Gemma4 parser: classify channel-less output consistently in streaming and non-streaming (#48262)
Signed-off-by: Adhithya Balakrishnan <adhithya.b2004@gmail.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>
|
2026-07-14 09:30:16 -04:00 |
|
  
|
32aef44388
|
[Bugfix] Include inline per-token-head scales in offloaded page transfer width (#48411)
Signed-off-by: Itay Etelis <itay.etelis@ibm.com>
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com>
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com>
Co-authored-by: Itay Etelis <itay.etelis@ibm.com>
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com>
|
2026-07-14 16:07:26 +03:00 |
|
 
|
7a74a9662b
|
[NIXL] Avoid reading expired blocks in bidirectional turn-2 read (#47021)
Signed-off-by: Tomer Gilad <tgilad@nvidia.com>
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai>
|
2026-07-14 13:03:41 +00:00 |
|
 karthikandGitHub
|
b6754f536e
|
[Model] Enable LoRA support for tower and connector in LlavaNextVideo (#48594)
Signed-off-by: gangula-karthik <gkarthik923@gmail.com>
|
2026-07-14 20:09:38 +08:00 |
|
 Juan Pérez de AlgabaandGitHub
|
793cf79c89
|
[Bugfix][Security] Fix concurrent sparse invariant race bypassing CVE remediation (#48583)
Signed-off-by: jperezde <jperezde@redhat.com>
|
2026-07-14 11:08:24 +00:00 |
|
 
|
50ac1c7bab
|
[Misc] Rename VLLM_TRITON_ATTN_USE_TD to VLLM_TRITON_USE_TD (#45781)
Signed-off-by: Artur Fierka <artur.fierka@intel.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-14 10:32:57 +00:00 |
|
 
|
f04d3f640e
|
[Test] Enable KV cache events for HMA models in CPU offloading test (#47754)
Signed-off-by: Itay Etelis <92247226+Etelis@users.noreply.github.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-14 12:22:27 +03:00 |
|
 xiangdongandGitHub
|
0a9396a25e
|
[XPU][CI] Add tests/v1/e2e/general/test_correctness_sliding_window.py in Intel GPU CI (#47231)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Signed-off-by: xiangdong <40376367+zxd1997066@users.noreply.github.com>
|
2026-07-14 08:50:16 +00:00 |
|
 
|
038ec293b1
|
[Bugfix] Return 400 instead of 500 when multimodal data is sent to a text-only model (#48473)
Signed-off-by: Hoang Nguyen Tien <hoang.nguyentien.2601@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-14 08:15:43 +00:00 |
|
 
|
894ebb27f5
|
Add Cosmos3 Edge Reasoner model (#48291)
Signed-off-by: Bartosz Stefaniak <bstefaniak@nvidia.com>
Co-authored-by: Bartosz Stefaniak <bstefaniak@nvidia.com>
|
2026-07-14 08:14:50 +00:00 |
|
 Juan Pérez de AlgabaandGitHub
|
c9a788eedc
|
fix(security): guard lm-format-enforcer regex compile with timeout (#47595)
Signed-off-by: jperezde <jperezde@redhat.com>
|
2026-07-14 07:18:11 +00:00 |
|
 
|
0762f2afeb
|
[Perf][Feat] Add generic cuteDSL LL BF16 router (GEMM) (#42562)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
|
2026-07-13 23:01:21 -07:00 |
|
 
|
31be872f55
|
[ROCm] Retune MI355 selective_state_update float16 config on the unified effective_batch grid (#48372)
Signed-off-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com>
Co-authored-by: vanshbhatia-amd <210711135+vanshbhatia-amd@users.noreply.github.com>
|
2026-07-14 05:16:29 +00:00 |
|
 wangxiyuanandGitHub
|
94c0ef3001
|
[Misc] Clean up "swap_space" (#48549)
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
|
2026-07-14 04:43:45 +00:00 |
|
 Matt WoodsonandGitHub
|
af1f036a70
|
[Bugfix] Skip minimax_m3 tool parser tests when Rust extension is absent (#48523)
Signed-off-by: Matt Woodson <mwoodson@redhat.com>
|
2026-07-14 04:43:22 +00:00 |
|
    
|
95aab66e95
|
[ROCm][MiniMax-M3][Spec Decode] Support speculative decode with AITER sparse PA (#47984)
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com>
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com>
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com>
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-07-14 04:12:53 +00:00 |
|
 nemanjaudovicandGitHub
|
dcf4072da9
|
[Perf][ROCm] Fix GDN KKT warmup regression on RDNA by avoiding fp32 tl.dot (#45000)
Signed-off-by: Saeid Rostami <srostami@amd.com>
Signed-off-by: nemanjaudovic <nudovic@amd.com>
|
2026-07-13 20:48:54 -07:00 |
|
 
|
382bbd5144
|
[ROCm][Kernel] Add HybridW4A16LinearKernel: Triton prefill + HIP skinny decode (#40977)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-07-13 20:22:00 -07:00 |
|
 
|
b50ef9c6ed
|
[ROCm][MiniMax-M2] Dispatch fused QK-norm + AllReduce via AITER (#44849)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Pawel Kowalski <pawel.kowalski@amd.com>
|
2026-07-14 03:10:16 +00:00 |
|
 Dan BlanaruandGitHub
|
9e289c553c
|
up FI fp8 moe topk to 32 (#44462)
|
2026-07-14 02:58:16 +00:00 |
|
 
|
c4f5cd60da
|
[1/N] Add dense MHA path for sparse MLA short sequences (#47327)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-14 00:29:56 +00:00 |
|
 
|
0b0ef8d7eb
|
[Quantization][INC][ARK] Support INT2 XPU WOQ Linear (#47521)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-14 08:29:45 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
21472f32ea
|
add pad-aware swiglu limit kernel (#48287)
Signed-off-by: gnovack <novackgm@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-13 16:48:19 -07:00 |
|
 
|
fec64fea75
|
[BugFix] Correct OTEL span start time for Dynamo compilation (#40698)
Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai>
Co-authored-by: Simon Mo <simon.mo@hey.com>
|
2026-07-13 16:25:03 -07:00 |
|
 
|
8b8af2caf7
|
[Frontend] Expose logprob_token_ids on Python OpenAI endpoints (#43463)
Signed-off-by: Lang Zhao <lang.zhao@galileo.ai>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-13 14:40:21 -07:00 |
|
 SnehlataandGitHub
|
7738ef35b8
|
[Feat] Add Support for BertForMaskedLM to vLLM (#48463)
Signed-off-by: atalhens <sneh.lata@nutanix.com>
|
2026-07-13 20:56:25 +00:00 |
|
 
|
9a21f0d1a3
|
[BugFix] Initialize model_config for Qwen3-VL MoE (#44863)
Signed-off-by: wenpengw-nv <wenpengw@nvidia.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-07-13 13:43:53 -07:00 |
|
 Nick HillandGitHub
|
8ac8375270
|
[Core] Preserve Marconi caching with selective hybrid cache retention (#47782)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-13 21:24:20 +01:00 |
|
 shanjiazandGitHub
|
7dc447dda7
|
Added sliding window attention support for qwen-eagle3 architecture (#47568)
Signed-off-by: shanjiaz <zsjwpianpian@gmail.com>
|
2026-07-13 20:20:44 +00:00 |
|
 
|
7fc97042c3
|
Add DCP + Eagle support for Tokenspeed MLA backends (#48180)
Signed-off-by: Pavani Majety <pmajety@nvidia.com>
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com>
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com>
|
2026-07-13 11:46:02 -07:00 |
|
 Micah WilliamsonandGitHub
|
18c4067a54
|
[ROCm][CI] Unblock AMD: Language Models Test (Extended Pooling) (#48513)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-07-13 18:38:10 +00:00 |
|
 
|
550218b136
|
[Bugfix][Frontend] Flush engine reasoning parser at engine-reasoning → tool streaming boundary (#47606)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
Co-authored-by: Ben Browning <56071+bbrowning@users.noreply.github.com>
|
2026-07-13 14:06:10 -04:00 |
|
 Gavin MorrisandGitHub
|
5c342876a6
|
[Doc] Add DeepseekV32ForCausalLM to supported_models.md (#48293)
Signed-off-by: Gavin Morris <gmorriscs@gmail.com>
|
2026-07-13 17:43:59 +00:00 |
|
 
|
9427c45386
|
[ROCm][CI] Transformers: pass only one of input_ids/inputs_embeds (#48258)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-07-13 17:28:50 +00:00 |
|
 
|
43c8cbf79b
|
[EC Connector] CPU Offloading EC Connector (#47423)
Signed-off-by: omerpaz95 <omerpaz95@gmail.com>
Signed-off-by: Or Ozeri <oro@il.ibm.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-07-13 20:09:41 +03:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
62286308c9
|
[Misc] Improve Matryoshka pooling dimensions validation (#48057)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-07-13 12:57:36 -04:00 |
|
 Nick HillandGitHub
|
26587f9519
|
[BugFix][ModelRunner V2] Fix stale attn metadata in speculator prefill cudagraph capture (#48261)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-07-13 09:39:15 -07:00 |
|
 
|
93e3bc8f30
|
[XPU][CI]Adjust timeout_in_minutes in Intel GPU CI (#48418)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-07-13 23:11:16 +08:00 |
|
 Yan MaandGitHub
|
c2c9f7c5e2
|
remove force channels_last in Idefics3MultiModalProcessor (#48467)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-07-13 14:18:57 +00:00 |
|