Bugen Zhao
6e714a103c
stash
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-02 05:51:48 +00:00
Bugen Zhao
c9951fd5c7
separate vllm-model-files
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-07-01 14:24:43 +00:00
Harry Mellor and GitHub
a78c15616f
Migrate GPTBigCode and Starcoder2 to the Transformers modeling backend ( #30966 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 13:41:36 +00:00
5c4db60f01
docs(security): document gRPC interface as insecure for private use only ( #45903 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <russell.bryant@gmail.com >
Co-authored-by: Russell Bryant <rbryant@redhat.com >
2026-07-01 12:39:57 +00:00
4e5ca89cfe
[ROCm][MiniMax-M3] Cross-layer lightning-indexer top-k sharing ( #47269 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 10:50:09 +00:00
Harry Mellor and GitHub
a22e0dfc69
[Model] Remove AyaVision, MusicFlamingo ( #47263 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:39:33 +00:00
stevenkuang and GitHub
cc56379e28
[Model] Support Hy3 token suffix and JSON Schema array types ( #47192 )
...
Signed-off-by: stevenkuang-tencent <stevenkuang@tencent.com >
2026-07-01 10:16:07 +00:00
024b06b0dc
[Bugfix] Expose usage field in GenerateResponse for disaggregated serving ( #42748 )
...
Signed-off-by: AIvashov <ivashov.aleksey@proton.me >
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
Co-authored-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-07-01 10:00:19 +00:00
Harry Mellor and GitHub
e7d0fcbc09
[CI] Fix various failures on main ( #47197 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-07-01 10:35:34 +01:00
akii96 and GitHub
aa8bb5562e
[ROCm][Perf][Bugfix] DSv4 indexer: use platform FP8 dtype (fnuz) for Q-quant on gfx942 ( #46730 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-07-01 17:33:55 +08:00
Andy Lo and GitHub
fa4bec9056
[Bugfix] Fix pooled Whisper sliding-window KV sizing ( #47071 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-07-01 11:33:19 +02:00
dee5da1dec
[Test] Run SageMaker handler-override tests in-process via TestClient ( #47250 )
...
Signed-off-by: Jyothirmai Kottu <jkottu@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 09:14:00 +00:00
ed41aa270a
[ROCm][DSV4] Use aiter mHC pre/post as the default ROCm path ( #43950 )
...
Signed-off-by: Fangzhou Ai <fangzhou.ai@amd.com >
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 16:27:42 +08:00
77a9c5ae28
Weight sync refactor + move sparse nccl engine ( #44353 )
...
Signed-off-by: hao-aaron <ahao@anyscale.com >
Signed-off-by: haoaaron <ahao@anyscale.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
2026-07-01 01:25:19 -07:00
f651a8a9a4
[XPU][UT]Enable ut qk_norm_rope_fusion ( #42486 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-07-01 07:38:03 +00:00
Jee Jee Li and GitHub
8f82be5705
[CI/Build] Fix LoRA testing ( #47242 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-07-01 15:36:13 +08:00
Nils Matteson and GitHub
a461070d1c
[Core] Make sleep-mode backend capability flags communicator-agnostic ( #47243 )
2026-07-01 07:17:44 +00:00
4470ae84de
Remove mantis ( #46806 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-07-01 07:13:58 +00:00
Chauncey and GitHub
697c34b97b
[Bugfix] Fix beam search candidate indexing when logprobs count varies ( #47126 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-07-01 07:07:06 +00:00
Blas Rodriguez Irizar and GitHub
5b431b905c
[Rust Frontend] Coerce completion max_tokens: null to default ( #47166 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-07-01 06:41:33 +00:00
89e99202f2
[CPU][Perf]Added tanh AOR for faster gelu activations. ( #44639 )
...
Signed-off-by: Anna Mayne <anna.mayne@arm.com >
Signed-off-by: almayne <anna.mayne@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-30 23:24:40 -07:00
Micah Williamson and GitHub
b446792306
[ROCm][Bugfix] Fix Triton "out of resource: shared memory" Error In One-Shot LoRA MoE ( #47209 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:36 -07:00
Micah Williamson and GitHub
c3b1f9e827
[ROCm][CI] Enable LoRA TP Distributed Test Group In AMD CI ( #47193 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-30 23:24:32 -07:00
Jonathan Mamou and GitHub
df802a87b7
[CPU] Remove speculative decoding stream overrides from CPUModelRunner ( #47162 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
2026-07-01 06:12:49 +00:00
Nils Matteson and GitHub
93d8f834dd
[Core] Pluggable sleep-mode backend abstraction (RFC #34303 ) ( #44074 )
2026-06-30 22:00:53 -07:00
Maria Guevara and GitHub
aeb35b90f0
[Rust Frontend] Add error context in tool parser failures ( #46512 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-07-01 12:48:55 +08:00
Gabriel Wu and GitHub
9a08a5118e
fix: skip cooperative top-K on SM120 ( #47164 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
2026-06-30 21:32:54 -07:00
c5200d3565
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE ( #46076 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Signed-off-by: Jingyi Yang <girasoleyang@gmail.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: GirasoleY <girasoleyang@gmail.com >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-01 00:32:20 -04:00
Matt and GitHub
3c1396bab6
[Hardware][AMD][CI] Toggle test coredumps on ROCm debug agent ( #47222 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 23:30:10 -05:00
Benjamin Chislett and GitHub
9969466a59
[Spec Decode] Support SWA + DFlash for MiMo ( #46104 )
2026-06-30 20:34:47 -07:00
achyuthan.s and GitHub
3406e8f83d
[Bugfix][Frontend][gpt-oss] Return raw output when Harmony parser ends non-terminal ( #47062 )
...
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
2026-07-01 01:46:01 +00:00
a264e41975
[Distributed] Default FlashInfer allreduce to mnnvl on single node ( #47219 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-30 18:35:56 -07:00
Woosuk Kwon and GitHub
f098ee70c7
[GLM5] Support FlashMLA FP8 KV cache (Hopper & Blackwell) ( #47090 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-30 18:13:21 -07:00
9294dd27eb
fix(reasoning): guard rfind in ernie45 streaming </response> branch ( #46255 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-07-01 01:01:14 +00:00
yzong-rh and GitHub
b1190d03cc
[Refactor][GPT-OSS] Harmony Responses API Refactor to use HarmonyParser ( #47185 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-30 19:23:20 -04:00
92c7fac640
[Perf] Restore zero-init of swizzled NVFP4 scale buffer to recover Blackwell decode throughput ( #45739 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-30 22:56:56 +00:00
Ting SUN and GitHub
ac521f6237
[Bugfix][Structured Outputs] Reject degenerate structured_outputs that crash EngineCore ( #45346 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-30 22:41:33 +00:00
28242824e0
[Bugfix][Frontend] Normalize constrained Harmony recipients ( #45657 )
...
Signed-off-by: shaojunjie <626650687@qq.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-30 17:33:10 -04:00
VectorPeak and GitHub
68294739d1
[Bugfix] Align OpenCV video metadata timeline ( #47099 )
...
Signed-off-by: VectorPeak <73048950+VectorPeak@users.noreply.github.com >
2026-06-30 20:43:42 +00:00
c8d2f3cb14
[Bugfix] compressed-tensors: allow int8 grouped WNA16 MoE on Marlin ( #47154 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 12:50:46 -07:00
Matt and GitHub
345b28ff2f
[Hardware][AMD][CI] Bump timeouts of various test groups on AMD CI ( #47195 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-30 14:30:53 -05:00
248d1fbb71
[Feat][1/N] CuTeDSL warmup infrastructure, FA4 MLA ( #46182 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-30 12:17:34 -07:00
11b26c5528
[Bugfix][Tool Parser] PoolsideV1: fix logprobs AttributeError on Responses API ( #47138 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-30 19:14:09 +00:00
Roberto L. Castro and GitHub
20434c472e
[Feat] Improve Triton JIT diagnostics ( #46621 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-30 18:50:15 +00:00
Andreas Karatzas and GitHub
c8f9c156a5
[ROCm][V1][MLA] Clone prefill backend state per metadata builder ( #46993 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 11:43:54 -07:00
953bba488d
[PERF] Extend NCCL symmetric memory to AllGather and ReduceScatter ( #46703 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: snordmann <snordmann@nvidia.com >
2026-06-30 11:38:18 -07:00
Wentao Ye and GitHub
3a9784b82c
[Feature] DP supervisor using rust frontend ( #47076 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 14:34:05 -04:00
Giancarlo Delfin and GitHub
3cecee40f3
[Model Runner V2][Spec Decode] Fix stale values in idx_mapping from CG num reqs padding ( #47066 )
2026-06-30 11:25:32 -07:00
a7732537f4
[Bugfix] Restore part of bugfix #42650 after accidental deletion in #43241 ( #47039 )
...
Signed-off-by: zhanda <zhandazhu@gmail.com >
Signed-off-by: Nikita Shapovalov <nikita@poolside.ai >
Co-authored-by: Zhanda Zhu <49645678+zhandaz@users.noreply.github.com >
Co-authored-by: Shang Wang <shangw@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-30 11:07:59 -07:00
727971f1c1
Add Medusa speculative decoding e2e test ( #41396 )
...
Signed-off-by: Anshika Ojha <anshikao@nvidia.com >
Signed-off-by: Rishi Puri <riship@nvidia.com >
Signed-off-by: Rishi Puri <puririshi98@berkeley.edu >
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com >
Co-authored-by: Anshika Ojha <215760622+ojhaanshika@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Stefano Castagnetta <scastagnetta@nvidia.com >
2026-06-30 18:02:22 +00:00
25671cb520
[Parser][Bugfix] Ensure tool call or other special tokens don't leak in non-streaming tool parsing ( #46875 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-30 13:46:53 -04:00
27d5f78b63
[CI] Move distributed small LM eval to B200 ( #47048 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 13:34:25 -04:00
liuzhenwei and GitHub
7a341fa109
[XPU] Support ZE_AFFINITY_MASK passthrough in xpu_disagg_acc_test ( #47105 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-30 17:06:12 +00:00
Charlie Fu and GitHub
f41e8ddc97
[ROCm][CI] Move PyTorch Compilation Unit Tests to MI300(gfx942) ( #47065 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-30 11:32:58 -05:00
245888ff77
[Feature] Detect all2all peer fault with fault tolerance backend and prevent corrupted output ( #43637 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 09:00:25 -07:00
e840f0d3f5
[Platform] Replace torch.cuda.Event with torch.Event ( #47140 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 08:39:59 -07:00
fcaa84efa7
[BugFix] Gate MRV2 mixed sparse-MLA warmup on max_num_seqs > 1 ( #47050 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: ziminghuang <ziminghuang@inferact.ai >
2026-06-30 16:31:27 +01:00
Wentao Ye and GitHub
9e84ec8648
[Refactor] Remove dead minimax allreduce rms kernel ( #46842 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-30 08:29:21 -07:00
d8f483dc30
[Spec Decode] Fix hidden-state extraction block size for hybrid verifiers ( #46301 )
...
Signed-off-by: Igor Margulis <igor.margulis@intel.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: mgoin <mgoin64@gmail.com >
2026-06-30 08:19:51 -07:00
Nicolò Lucchesi and GitHub
dc148dc4d7
[CI][Bugfix] Fix Hybrid SSM NixlConnector PD prefix cache test (2 GPUs) ( #47157 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 23:14:13 +08:00
tc-mb and GitHub
7cf7cbcd95
[Bugfix] MiniCPM-V 4.6: fix grid rows/cols swap in placeholder generation ( #45918 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-30 08:12:44 -07:00
c231d1f290
fix(security): bound tokenizer work when explicit truncation_side is set ( #47007 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 23:08:51 +08:00
Giancarlo Delfin and GitHub
db808b3961
[Model Runner V2][Spec Decode] Implement block verification for rejection sampling ( #46781 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-30 08:07:24 -07:00
Arsalan Shakil and GitHub
00ebf19cca
[Bugfix][Quant] Raise actionable error instead of bare assert for group-size/TP mismatch ( #46230 ) ( #46236 )
...
Signed-off-by: Arsalan Shakil <shakil.arsalan@yahoo.com >
2026-06-30 14:57:14 +00:00
ded6676458
[Bugfix] Seed RayExecutorV2 TCPStore port by DP rank to avoid collisions ( #45960 )
...
Signed-off-by: Seiji Eicher <seiji@anyscale.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:37:34 -07:00
Bugen Zhao and GitHub
7a327f0b4f
[Rust Frontend] Simplify unit tests with shared TestTokenizer ( #47125 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 15:34:43 +01:00
Harry Mellor and GitHub
1ab9522935
Remove more unnecessary load_weights methods ( #47058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 15:22:16 +01:00
0fc2512094
[KV Offload] Pass ScheduleEndContext to on_schedule_end hook ( #46450 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 17:07:12 +03:00
Harry Mellor and GitHub
62c7d8009f
Forward fix nightly errors from #44589 ( #47151 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 14:02:34 +00:00
Isotr0py and GitHub
ab80b3dff4
[CI/Build] Bump PyNvVideoCodec version ( #47139 )
2026-06-30 06:38:46 -07:00
Qiming Zhang and GitHub
91055efd36
[XPU] C++ implementation for get_memory_info ( #47134 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-30 21:34:47 +08:00
Bugen Zhao and GitHub
3675bcff67
[Rust Frontend] Refactor TLS serve path with unified MaybeTlsListener ( #47101 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:31:58 +01:00
Bugen Zhao and GitHub
bdbd7278b6
[Rust Frontend] Extend renderer/parser roundtrip tests to support token ids ( #47110 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 14:27:45 +01:00
Harry Mellor and GitHub
5dc36a4fa5
[Model] Remove Tarsier, Tarsier2 ( #47143 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:20:33 +00:00
aab7af0bcb
[Bugfix][ROCm][MLA] Pass q/kv dtypes to get_mla_metadata_v1 in FP8 decode ( #46997 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-30 05:31:16 -07:00
536047755e
Bump actions/checkout from 6.0.1 to 7.0.0 ( #33057 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-30 13:16:20 +01:00
1907d3854a
[Bugfix] Reject negative values for max_logprobs and long_prefill_token_threshold ( #44002 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 13:01:03 +01:00
Chaojun Zhang and GitHub
ea9ddf59fc
[XPU][CI] Enable shared loader test ( #45977 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 11:20:33 +00:00
8cf7c4d8ad
[Attention Backend] add HPC-Ops Attention backend ( #46020 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 18:17:43 +08:00
8e9d70fdd5
[Kernel][XPU] Adjust kernel unit tests for XPU ( #45140 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-30 09:57:27 +00:00
Juan Pérez de Algaba and GitHub
364ee36af1
fix(security): prevent image decompression bomb OOM denial of service ( #47010 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-30 09:39:22 +00:00
Nicolò Lucchesi and GitHub
06fae69114
[Misc] Mistral label alert ( #47132 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-30 09:02:07 +00:00
14f8660a18
[CI/Build] Add CPU test dependency pre-commit hooks ( #47032 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-30 07:59:13 +00:00
aed541def4
[Bugfix][Responses] Set completed status for Harmony function calls ( #46945 )
...
Signed-off-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: amanambak <aman.paswan@ambak.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-30 07:55:14 +00:00
2bc20e8aba
[Frontend] Add Streaming Parser Engine and new Kimi k2.5/k2.6/k2.7 Parser ( #46610 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 07:53:17 +00:00
Chaojun Zhang and GitHub
8cc242335d
[XPU] Optimize XPU worker shutdown logic to prevent resource leak ( #46433 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-30 15:27:21 +08:00
Andreas Karatzas and GitHub
ba22cb6765
[ROCm][Ray][CI] Keep assigned GPU visible for weight transfer ( #47000 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:59:18 +08:00
Uros Markovic and GitHub
81bcced482
[Bugfix][ROCm] Preserve MoE weight padding for unquantized Triton path ( #46381 )
...
Signed-off-by: Uros Markovic <umarkovi@amd.com >
2026-06-30 14:47:57 +08:00
Kunshang Ji and GitHub
fb42e5219e
[Platform] Replace torch.cuda.mem_get_info with torch.accelerator.get_memory_info ( #44825 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-30 14:39:52 +08:00
Dakai An and GitHub
0feca7ffa8
PD disagg with Mooncake Connector: GDN support (Qwen3.5) and MLA support (Deepseek-V4-Flash) ( #46807 )
2026-06-29 23:29:04 -07:00
97b5ce5c39
[Bugfix] Raise VLLMValidationError for non-integer logit_bias keys ( #46612 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-30 06:18:59 +00:00
Andreas Karatzas and GitHub
4236514098
[ROCm][CI][Multimodal] Use ROCm-aware FA availability check for Unlimited-OCR ( #47004 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-30 14:03:13 +08:00
Blas Rodriguez Irizar and GitHub
e45c8a9f4b
[Rust Frontend] Start current wave for a stale DP FirstRequest ( #46833 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-30 05:13:09 +00:00
Wei Zhao and GitHub
b153dd3f28
[Bugfix] Use larger workspace size for Flashinfer MLA LSE ( #47074 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-29 22:11:03 -07:00
Reid and GitHub
930f8dc0a1
[Bugfix][Rust Frontend] Reject prompt_logprobs for streaming generate ( #46839 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 05:10:07 +00:00
Reid and GitHub
a16dbd5b85
[Rust Frontend] Avoid LoRA registry scans without active LoRA requests ( #47040 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-30 04:58:19 +00:00
bec232a914
Secondary tier implementation for PD disaggregation ( #42285 )
...
Signed-off-by: Liran Schour <lirans@il.ibm.com >
Signed-off-by: liranschour <liranschour@users.noreply.github.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-30 07:51:44 +03:00
b5c9e1ac33
[LoRA] Add language-backbone LoRA support for MiniCPM-V 4.6 ( #46740 )
...
Signed-off-by: linitra24 <Joy25810@foxmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-30 04:19:31 +00:00
ae2c4f3db7
[XPU][UT]Fix xpu pass_config.fuse_norm_quant assert issue ( #46804 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 21:13:44 -07:00
ganesh and GitHub
fca432e60a
[Bugfix] Propagate default stop_token_ids to per-request SamplingParams ( #35076 )
...
Signed-off-by: sriganesh123 <arjulasriganesh@gmail.com >
2026-06-30 12:10:09 +08:00
af1ee8c475
fix(config): reject negative max_logprobs (except -1) and long_prefill_token_threshold ( #44070 )
...
Signed-off-by: Chenglun Hu <chenglunhu@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-30 04:02:36 +00:00
5b4cb69523
[Bugfix][MLA] Fix LSE log-base mismatch in DCP + FlashInfer MLA decode ( #47079 )
...
Signed-off-by: girasoley <girasoleyang@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 19:15:02 -07:00
9fc0c08026
[ROCm][CI] Make tests/v1/shutdown an importable package ( #47085 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 21:01:27 -05:00
f2b5fabb23
[ROCm][CI] Move LM Eval Large Models (8 GPUs) to mi300 pool ( #47094 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 20:59:08 -05:00
b8cb75b149
[Rust Frontend] Add static HTTPS and mTLS support for HTTP and gRPC ( #45890 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-30 01:45:59 +00:00
Thien Tran and GitHub
43916891b2
[GDN] Improve kkt kernel of CuteDSL prefill backend ( #46346 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-29 18:34:18 -07:00
cda05ee8c4
[Bugfix][Reasoning] Fix thinking_token_budget not enforced on re-entry after forced end ( #43757 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Cursor Agent <cursor-agent@cursor.com >
Co-authored-by: Simon Mo <simon.mo@hey.com >
2026-06-30 01:04:25 +00:00
weishu and GitHub
77654d080c
[KVTransfer] MultiConnector: merge kv_transfer_params dicts across connectors ( #46777 )
...
Signed-off-by: deng451e <838677410@qq.com >
2026-06-30 00:25:05 +00:00
Wentao Ye and GitHub
75698e60b3
[Bug] Fix sparse attention issue for GLM5.2 non-torch compile path ( #47083 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-29 15:45:53 -07:00
Andreas Karatzas and GitHub
8632c884dc
[ROCm][CI] Use spawn around the threaded OTLP test ( #47003 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 16:34:05 -05:00
c3734e8334
[CI][Bugfix] Add cohere_melody to ROCm test requirements ( #47072 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-29 16:29:47 -05:00
53f7553f09
[ROCm][DeepEP] Stabilize high-throughput DBO for DP+EP ( #46990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-29 14:28:02 -07:00
4eb227992a
[ROCm][CI] Make memory sampling less racy in tests and sleep mode ( #45490 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-29 14:26:41 -07:00
Micah Williamson and GitHub
ebcf511ec3
[ROCm][CI] Soft Fail Spec Decode Ngram + Suffix and Entrypoints Integration (LLM) AMD Mirrors ( #47067 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-29 16:24:08 -05:00
Matthew Bonanni and GitHub
8fc1b2d046
Fix FA4 dynamic_causal for full attention layers ( #46659 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-29 14:23:34 -07:00
Harry Mellor and GitHub
5316638a5e
Fix transient dependency issues caused by requirements/common.txt ( #47015 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 14:20:33 -07:00
zhrrr and GitHub
61ab70ec3b
[Model Runner V2] support mamba hybrid models align prefix cache ( #42406 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-29 14:09:16 -07:00
Woosuk Kwon and GitHub
a309d4fe60
Support DCP with FlashInfer MLA ( #43729 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-29 13:24:29 -07:00
72f639927f
[XPU] [RMSNorm] revert weightless change on xpu ( #46987 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 19:03:06 +00:00
Nick Hill and GitHub
8ad4a01825
[ModelRunner V2] Simplify recent UnlimitedOCR-related changes ( #46975 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-29 09:56:17 -07:00
Jee Jee Li and GitHub
7be582697b
[Bugfix] Fix DeepseekV2Model hidden_size ( #46986 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-29 16:44:05 +00:00
030c9523bd
[Perf][1/N] Expand Triton kernel warmup coverage, DSv4 ( #46634 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-29 16:40:34 +00:00
4708292d48
Bump flashinfer version to 0.6.13 ( #46683 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-29 09:30:57 -07:00
debec6440b
Add MiniMax-M3 modelopt nvfp4 support ( #46756 )
...
Signed-off-by: Xin Li <xinli@nvidia.com >
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com >
Co-authored-by: Xin Li <xinli@nvidia.com >
2026-06-29 09:29:39 -07:00
c8fb2963bd
[FS-Offloading] Batch Lookup in C ( #46713 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 09:28:32 -07:00
HDCharles and GitHub
379acd4e4f
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression ( #46860 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
2026-06-29 15:55:42 +00:00
Martin Hickey and GitHub
07d33e575b
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager ( #44657 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 16:42:35 +01:00
36bbecd643
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" ( #46958 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-29 07:54:34 -07:00
Nicolò Lucchesi and GitHub
6149187a4c
[Kernel] Triton MLA logits workspace ( #46819 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-29 07:54:29 -07:00
Xiaohong (Sean) Chen and GitHub
49e28e8e91
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope ( #44010 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
2026-06-29 22:54:15 +08:00
0ca39c4f1f
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone ( #46973 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-29 10:00:31 -04:00
Blas Rodriguez Irizar and GitHub
6185d73882
[Rust Frontend] Keep literal "null" string for string-typed tool params ( #46827 )
...
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com >
2026-06-29 13:46:33 +00:00
bc8481af09
[MoE Refactor] Standardize Humming MoE experts + utilities ( #43373 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-29 06:19:29 -07:00
59575da46d
[XPU] exclude unsupported models for test_tensor_sechma.py ( #47008 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-29 12:30:28 +00:00
wang.yuqi and GitHub
3483240b7e
[Frontend] Consolidate scale out entrypoints ( #44512 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 03:18:53 -07:00
Roberto L. Castro and GitHub
eddfd4cf21
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen ( #46750 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-29 10:10:07 +00:00
Martin Hickey and GitHub
a4e3cb40d0
[mypy] Enable mypy for tests directory ( #47018 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-29 09:29:09 +00:00
soaringk and GitHub
ab132ee98b
Fix model info cache for package models ( #46567 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-29 09:17:54 +00:00
e186107870
[Bugfix] Use native SiLU activation in CPU fused MoE ( #45961 )
...
Signed-off-by: Alden Lobo <alden.lobo@arm.com >
Co-authored-by: Alden Lobo <alden.lobo@arm.com >
2026-06-29 09:12:20 +00:00
0e207dac78
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models ( #46835 )
...
Signed-off-by: John Langford <jl@hunch.net >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:59:15 +00:00
wang.yuqi and GitHub
9e86352c60
[CI Failure] Add transformers version check for openai/privacy-filter ( #47011 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-29 08:57:26 +00:00
Harry Mellor and GitHub
5051698e41
Remove unnecessary load_weights methods ( #44589 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 01:52:23 -07:00
Andreas Karatzas and GitHub
db28ae2d07
[ROCm][CI] Explicitly tear down multimodal offline LLMs ( #46999 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-29 07:59:24 +00:00
Harry Mellor and GitHub
f6bb8682ee
Fix docs on main ( #47009 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 15:50:57 +08:00
4559c43a95
[MM][CG] Gemma3 Encoder CUDA Graph ( #43591 )
...
Signed-off-by: JisoLya <523420504@qq.com >
Signed-off-by: Soyaazz <523420504@qq.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-29 04:52:00 +00:00
Bugen Zhao and GitHub
5274c1181d
[Rust Frontend] Add Harmony Renderer for GPT-OSS ( #46800 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-29 03:39:04 +00:00
Yuwen Zhou and GitHub
58d6a6e60a
[CPU] Support cpu compressed-tensor w8a8 int8 moe ( #42920 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com >
2026-06-29 03:04:05 +00:00
a2abce646f
[EPLB] Mask padding in EPLB load recording ( #38128 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-28 19:43:58 -07:00
Harry Mellor and GitHub
311ad689ad
Remove boilerplate missed by #46820 ( #46956 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-29 08:11:17 +08:00
Woosuk Kwon and GitHub
0472436541
[Spec Decode] Avoid redundant hidden-states gather in draft prefill ( #46968 )
2026-06-28 17:04:01 -07:00
4dfbf1503b
[Model] Add support for openai/privacy-filter ( #41026 )
...
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-28 16:18:22 -07:00
Wei Zhao and GitHub
95528527ea
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 ( #46855 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-28 14:36:23 -07:00
c2127a25c7
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm ( #46895 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:50:30 -05:00
03c6d01c30
[OCP MX ] Add back emulation to available OCP MX backends list ( #46629 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 12:43:19 -05:00
Woosuk Kwon and GitHub
4b643c463e
[GLM5] Fix minor typo ( #46961 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-28 08:37:00 -07:00
7544286b04
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE ( #46552 )
...
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-28 15:19:28 +00:00
Woosuk Kwon and GitHub
89876b0c54
[GLM5] Implement op fusion for GLM5/DSV3.2 ( #46876 )
2026-06-28 08:17:39 -07:00
Wentao Ye and GitHub
5c91039c41
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement ( #46635 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-28 14:55:54 +00:00
5ecae3266c
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) ( #45033 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-28 07:52:00 -07:00
6eb63a1da6
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers ( #46600 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-28 01:37:44 -07:00
09841ae705
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation ( #46846 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-28 00:07:33 -07:00
Matt and GitHub
a2a92cbbaa
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection ( #46930 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-28 00:07:14 -07:00
35e6c86caa
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL ( #46034 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-28 00:06:43 -07:00
c7ca0bccae
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 ( #44313 )
...
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com >
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com >
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com >
2026-06-28 00:04:08 -07:00
c6741b2ad4
[Model] Support Unlimited OCR ( #46564 )
...
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-27 23:09:18 -07:00
a65f93fb2e
[ROCm][CI] Add ci_base metadata for external cache orchestration ( #46886 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-28 12:51:19 +08:00
Chauncey and GitHub
11a12305c0
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models ( #46786 )
2026-06-27 18:38:07 -07:00
798185d438
[KV-Offloading] Fix tensors_per_block stride ( #46888 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-27 21:01:45 -04:00
Matt and GitHub
9036c89ee4
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now ( #46928 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-27 17:30:49 -05:00
Giancarlo Delfin and GitHub
b6caeb5a09
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance ( #46878 )
2026-06-27 14:09:25 -07:00
Taneem Ibrahim and GitHub
8bf064f8d3
Fixed chunked embedding aggregation with request-id metadata ( #46782 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-27 20:57:47 +00:00
ea2ead1db3
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod ( #46818 )
...
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com >
2026-06-27 20:23:59 +00:00
Wentao Ye and GitHub
56aa067bf0
[CI Bug] Fix h100 AssertionError: Cold-start child failed ( #46927 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-27 20:17:33 +00:00
xiaolinchen and GitHub
35e3850fa9
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) ( #46915 )
...
Signed-off-by: wentian-byte <2990624738@qq.com >
2026-06-27 14:30:10 -04:00
51a99565c3
[ROCm][Perf] Fused shared expert for Minimax M3 ( #46474 )
...
Signed-off-by: Fangzhou-Ai <fangzhouai@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-27 12:34:17 +00:00
867fd5e8ed
[ROCm][Perf] Use flydsl moe with Minimax-M3 mxfp8 weights on gfx950 and implemented moe-backend selection ( #46184 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
2026-06-27 10:22:57 +00:00
9fd00ee006
[ROCm][CI] Move remaining mi250_2 tests out of the MI250 queue ( #46905 )
...
Signed-off-by: Codex <codex@example.invalid >
Co-authored-by: Codex <codex@example.invalid >
2026-06-27 17:08:54 +08:00
091d13976c
[ROCm][CI] Add TRITON_ATTN score absolute tolerance floor ( #46891 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-27 06:35:50 +00:00
Wentao Ye and GitHub
b588f66dc2
[GLM5.2 Perf] fused_indexer_q_rope_quant triton kernel, 1.9% ~ 3.3% E2E Throughput improvement. ( #46862 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-26 22:16:20 -07:00
Benjamin Chislett and GitHub
455f25aa13
[CLI] Add flag to print TTFT and TPS in vllm chat ( #46775 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-26 22:15:10 -07:00
d706dec904
fix: Correct reasoning-end detection for prompt history ( #44551 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Signed-off-by: Jason Ozuzu <jasonozuzu@cohere.com >
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: Walter Beller-Morales <walterbm@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-26 22:15:06 -07:00
Divakar Verma and GitHub
68ee8300a0
[ROCm][CI]Fix test_concat_and_cache_mla_rope_fused on ROCm ( #46409 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 12:38:13 +08:00
ddd3855a28
[MoE Backend] add HPC-Ops MoE backend ( #45924 )
...
Signed-off-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: chengvjiang <chengvjiang@tencent.com >
Co-authored-by: youkaichao <youkaichao@gmail.com >
2026-06-27 11:18:07 +08:00
Divakar Verma and GitHub
00e045b7c7
[ROCm][CI TG] refactor and fix deepep_moe test group ( #46758 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:45:23 +08:00
Divakar Verma and GitHub
17a71d8702
[ROCm][CI] Relax fused layernorm quant test tolerances for one-ULP outliers ( #46658 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-27 10:44:29 +08:00
weizhoublue and GitHub
2e058851d3
fix(docker): eliminate race conditions in shared buildkit cache mounts ( #44984 )
2026-06-26 19:43:17 -07:00
Dāvis and GitHub
1a92dfcce4
[Build] Show error message when using ROCm with LTO and different compilers ( #35232 )
2026-06-26 19:43:00 -07:00
Chris Leonard and GitHub
d0f800811b
[Build] Update vllm to point to vllm-project/flash-attention commit that builds FA3 with torch stable API. ( #46644 )
2026-06-26 19:42:46 -07:00
Nick Hill and GitHub
c6dd32a810
[ModelRunner V2] Support realtime embeddings ( #46762 )
2026-06-26 19:42:27 -07:00
af16446bf3
Vram semaphore infra ( #44465 )
...
Signed-off-by: Brandon Pelfrey <bpelfrey@nvidia.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-26 17:32:51 -07:00
Harry Mellor and GitHub
3f67477497
[CI] Don't try and download files that we already know don't exist ( #46854 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-26 23:56:39 +00:00
Nick Hill and GitHub
1d41009e81
[ModelRunner V2] Fix cross-attention block table sizing ( #46753 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:34:21 -07:00
Nick Hill and GitHub
b94f212e37
[ModelRunner V2] Deduplicate ModelState init logic ( #46776 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 16:32:45 -07:00
Harry Mellor and GitHub
d8eb734d94
Fix Transformers backend FP8 MoE and remove some boilerplate ( #46820 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-27 00:16:05 +01:00
2ff76a5e85
[ROCm][Bugfix] Pass num_kv_splits to aiter mla_reduce_v1 ( #46760 )
...
Signed-off-by: Rohan Potdar <rohanpotdar138@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-26 21:58:40 +00:00
Yifan Qiao and GitHub
75fdcc82a5
[CI] Add @ivanium to CODEOWNERS for KV-cache/offload areas ( #46873 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-26 21:48:53 +00:00
yzong-rh and GitHub
77f8796d16
[Frontend][Gpt-oss] Use process_eos() to flush Harmony Parser outputs. ( #46437 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-26 17:18:47 -04:00
c40d307731
[Core] Remove FlashAttention block size restriction for hybrid models ( #36701 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-26 21:16:39 +00:00
Woosuk Kwon and GitHub
65e655d295
[GLM-5] Add DSV3.2/GLM5 to vllm/models/ ( #46808 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-26 14:09:05 -07:00
Charlie Fu and GitHub
6e2fb02fe5
[ROCm][CI] Fix rlhf_nccl.py on ROCm ( #46851 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-26 15:41:49 -05:00
Micah Williamson and GitHub
274325dd43
[ROCm][CI] Remove V1 Sample + Logits from mi250 Queue ( #46867 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-26 15:38:38 -05:00
Matt and GitHub
95e6442a6b
[Hardware][AMD][CI] Fix Kernels Quantization test timeout ( #46859 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-26 15:19:16 -05:00
701a23d99f
[Bugfix][Model] Support tensor parallelism for DiffusionGemma ( #45719 ) ( #46177 )
...
Signed-off-by: Carlos Alvarado <carlos-alvarado@outlook.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
2026-06-26 20:05:04 +00:00
Ben Browning and GitHub
dccb412e2c
[Bugfix][Parser] Pass token IDs to parser.parse() in Responses API and batch serving ( #46843 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-26 19:29:52 +00:00
c6554f321c
[CPU] Fix macOS/Apple Silicon hang by enabling OpenMP in the build ( #46769 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 14:32:21 -04:00
Julien Denize and GitHub
3d3b96488f
Migrate Voxtral to mistral-common 1.11.5 audio API ( #46705 )
...
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-06-26 11:06:31 -07:00
Nick Hill and GitHub
658b54efe4
[ModelRunner V2] Update scheduler tests to cover MRV2 paths ( #46771 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-26 09:36:31 -07:00
Li, Jiang and GitHub
abc71548ef
[CI/Build][CPU] Add test image cache clean-up ( #46831 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-26 23:28:49 +08:00
Nick Hill and GitHub
4e07ca2c92
[Core] Add VLLM_GPU_SYNC_CHECK env var ( #44800 )
2026-06-26 08:24:33 -07:00
Bugen Zhao and GitHub
e71bc6da85
[Rust Frontend] Use oss-harmony for Harmony output processing ( #46799 )
2026-06-26 08:24:13 -07:00
fxmarty-amd and GitHub
37ce34922f
[CI] Fix failing CUDA graph capture in Triton MOE ( #46735 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-26 07:21:20 -07:00
c2507fb293
[ROCm] [MoE] [Perf] Shared-expert fusion for bias-routed MoE; enable on MiniMax-M3 mxfp8 model ( #46545 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-26 07:05:20 -07:00
TJian and GitHub
8921c4be88
[ROCm] [Performance] Optimize aiter moe for DeepSeekV4 ( #46122 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 06:43:27 -07:00
8e394244a5
[ROCm]Enable AITER MoE backend for MiniMax-M3-MXFP4 ( #46419 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-26 06:35:35 -07:00
TJian and GitHub
302954e5f6
[ROCm] [CI] fix transcription flakiness AMD: Entrypoints Integration (API Server OpenAI - Part 1) (mi325_1) ( #46823 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-26 21:33:35 +08:00
Hyunkyun Moon and GitHub
950ee4c2e4
[API] Add token offsets to render endpoints (/v1/.../render) ( #44226 )
...
Signed-off-by: HyunKyun Moon <mhg5303@gmail.com >
2026-06-26 05:02:52 -07:00
d980a3cc6e
[ROCm] Fix AITER_UNIFIED_ATTN Dispatching After AITER Bump ( #46780 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-26 02:09:56 -07:00
bf292b5f6b
[Docs] Remove BambaForCausalLM from supported hybrid models list ( #46071 )
...
Signed-off-by: liejiang <jianglie2023@gmail.com >
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-06-26 08:02:50 +00:00
wang.yuqi and GitHub
5e3dad04b1
[Misc] Move the legacy api_server.py to the examples directory. ( #46783 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-26 07:43:29 +00:00
Joe Rowell and GitHub
63e161f296
[Bugfix][Tool Parser] PoolsideV1: fix string whitespace and required named tool choice ( #46486 )
...
Signed-off-by: Joe Rowell <joerowell4@gmail.com >
2026-06-26 06:05:16 +00:00
Tiezhen WANG and GitHub
c7645bce04
Remove grok model arch from vllm ( #46706 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
2026-06-25 23:02:10 -07:00
35a49fcfc2
[CI][Bugfix] Spawn engine in mm cache sleep test to fix ROCm HIP error ( #46749 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-26 00:38:26 -05:00
peizhang56 and GitHub
915e99ec67
[ROCm][Bugfix] Fix HIP fork re-init in multimodal offline examples ( #46741 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
2026-06-26 00:37:47 -05:00
Nick Hill and GitHub
5b33041746
[ModelRunner V2] Fix whisper test ( #46773 )
2026-06-25 22:10:36 -07:00
Matt and GitHub
1a4984520e
[Hardware][AMD][CI] Fix AMD CI image build ( #46792 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 22:05:12 -07:00
Reid and GitHub
e312c5cb25
[Rust Frontend] Make Granite4 string argument scanning incremental ( #46507 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-26 03:54:03 +00:00
Matti4 and GitHub
1502cf6274
Fix relative allowed local media paths ( #45263 )
2026-06-25 20:45:20 -07:00
d350fa8ddd
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens ( #46733 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-26 03:41:33 +00:00
dbc49b6b99
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 ( #45166 )
...
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com >
Signed-off-by: ovidiusm <ovidium@nvidia.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-25 19:33:42 -07:00
fxmarty-amd and GitHub
552a9dbe59
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears ( #44667 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-25 19:33:00 -07:00
02a1f23711
[DFlash] Fuse precompute kv per-layer rmsnorms ( #46761 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:32:07 -07:00
652d962bc9
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation ( #46448 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:30:07 -07:00
Giancarlo Delfin and GitHub
5314665bad
[Model Runner V2][DFlash] Enable dflash attention backend selection ( #46770 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 19:29:25 -07:00
Michael Goin and GitHub
3daea7ceb9
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models ( #46759 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 19:03:09 -07:00
Wentao Ye and GitHub
cc7981599e
[Refactor] Remove dead kernel code ( #46405 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 18:09:56 -07:00
Nick Hill and GitHub
32bb3195f0
[ModelRunner V2] Bound memory for large logprobs requests ( #46746 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-25 18:04:06 -07:00
ad28d605e6
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) ( #45544 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-25 17:46:28 -07:00
Bugen Zhao and GitHub
ae7c8ec223
[Rust Frontend] Switch rustls to native-tls/OpenSSL ( #46696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:19:44 -07:00
Bugen Zhao and GitHub
1d3f4cb3a4
[Rust Frontend] Extract renderer fixture test utilities ( #46719 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 17:12:38 -07:00
Bugen Zhao and GitHub
f9e684499f
[Rust Frontend] Migrate gemma4 to unified parser ( #46602 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 16:59:57 -07:00
Giancarlo Delfin and GitHub
c53994e134
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling ( #46665 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-25 23:46:10 +00:00
Matt and GitHub
27da2a2ac4
[Hardware][AMD][CI] Use Triton-based AITER MHA for LM Eval Qwen-3.5 Models Tests ( #46691 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 17:08:04 -05:00
Michael Goin and GitHub
a2e8ec3d52
[CI] Depend GPQA Eval DGX Spark job on arm64 image build ( #46736 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 17:07:04 -04:00
e8c24a7695
[Kernel] Vectorized fp32 moe_sum reduction and support any topk ( #46643 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 14:02:28 -07:00
Andreas Karatzas and GitHub
2a6f8f0c05
[ROCm][CI] Fine-tuning queues and test names ( #39238 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-25 13:24:09 -07:00
Robert Shaw and GitHub
c5e3c40877
Fix P/D with DP Supervisor ( #46628 )
...
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-25 13:13:08 -07:00
Wentao Ye and GitHub
8b4d93ba2b
[Perf] Remove redundant clone for GLM, Deepseek etc ( #46651 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-25 13:09:00 -07:00
Michael Goin and GitHub
e8e7b592d1
[Kernel][MoE] Tune block-FP8 fused MoE for low-batch decode ( #46642 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-25 12:38:28 -07:00
Rohan Potdar and GitHub
e53a17232c
[ROCm]: Bump aiter to 0.1.16.post2 ( #46692 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-25 11:53:26 -07:00
Flora Feng and GitHub
96eb8ddc41
[CI] Re-enable skipped glm and seedoss parser tests ( #46671 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-25 13:11:41 -04:00
Gabriel Wu and GitHub
8fa36fbbeb
[Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 ( #46506 )
2026-06-25 09:12:00 -07:00
Ranran and GitHub
e45b279928
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3Next ( #46316 )
...
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
2026-06-25 09:05:04 -07:00
d490b98162
[Core] Avoid mixed length specdec batches via padding ( #45237 )
...
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com >
2026-06-25 08:34:44 -07:00
haoyangli0109 and GitHub
1744adc256
[ROCM] [Communication] Add INT3 quantization method for quickreduce ( #45666 )
...
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com >
2026-06-25 15:14:15 +00:00
Divakar Verma and GitHub
cdfa2fd7e9
[ROCm][CI] rm duplicate Distributed Torchrun ci test ( #46729 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-25 09:58:17 -05:00
6f3da461d1
[Pooling] Fix Cohere embed billed image token accounting for mixed-content inputs ( #46093 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 10:44:29 -04:00
Russell Bryant and GitHub
d3130d878c
[CI] Pin GitHub Actions to commit hashes in macos-smoke-test.yml ( #38290 )
2026-06-25 13:48:44 +00:00
9bfd878a48
[MoE] [MoE Refactor] Add moe kernel oracle abc 37753 ( #43461 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
Signed-off-by: qyYue1389 <yueqiuyang1389@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 09:34:03 -04:00
Matt and GitHub
2365b7a8e7
[Hardware][AMD][CI] Mirror Basic Models (Others) and Weight Loading Multiple GPU test groups ( #46668 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 08:25:09 -05:00
15be78732b
[NIXL][Mamba] Add Mamba1 support to NIXL P/D disaggregation ( #45019 )
...
Signed-off-by: Josephasafg <ajgard7@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 05:50:41 -07:00
92221485aa
[CPU][CI/Build] Allow more CPU CI agents ( #46702 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 19:33:39 +08:00
xiangdong and GitHub
a6f41ab678
[XPU][CI]Refine .buildkite/ci_config_intel.yaml for Intel GPU CI ( #46674 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-25 08:58:26 +00:00
c63cd4906c
[ROCm][ [Perf] sparse attention optimization on minimax-m3 ( #46546 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: yueliu14 <yue.liu4@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-25 16:56:00 +08:00
638b1a99cc
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM ( #45269 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-25 08:18:10 +00:00
72adb20a6a
[Model] Remove AquilaForCausalLM, AquilaModel ( #46605 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-25 08:08:26 +00:00
2396d91e93
[CPU][Spec Decode] Enable DFlash SD for CPU ( #44029 )
...
Signed-off-by: guybd <guy.boudoukh@intel.com >
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 15:32:48 +08:00
9b215ae60b
[Rust Frontend] Forward VLLM_ENGINE_READY_TIMEOUT_S via --args-json ( #44610 )
...
Signed-off-by: kai <kai@example.com >
Co-authored-by: 图灵 <tuling.wk@alibaba-inc.com >
2026-06-25 07:25:08 +00:00
Bugen Zhao and GitHub
4d3b4b9b01
[Rust Frontend] Make ToolParserOutput a seq of ToolParserEvent to preserve order ( #46584 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 06:27:07 +00:00
Matthias Gehre and GitHub
77c1d9fe9b
[ROCm][Perf] Tune wvSplitK on gfx1151 ( #40784 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-06-25 14:17:46 +08:00
Jeff (Junze) Ma and GitHub
36fd7e8b86
[SimpleCPUOffloadConnector] Fix remaining global→block conversions under PCP/DCP ( #46394 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
2026-06-24 23:05:24 -07:00
fc61c6fc26
[Perf] Enable + tune FlashInfer fused allreduce at world_size=16 on SM 10.3 (GB300) ( #46392 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:04:17 -07:00
Matt and GitHub
e2af449c39
[Hardware][AMD][CI] Move Metrics, Tracing (2 GPUs) & make optional ( #46686 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-25 05:49:33 +00:00
3f5a1e1733
[ROCm][CI] Expand basic correctness target suites ( #46573 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matt <156021403+mawong-amd@users.noreply.github.com >
2026-06-25 12:18:57 +08:00
710ebaa189
[ROCm][Bugfix] Fix chunk alignment when using context parallelism with TRITON_MLA ( #46114 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:07:28 -04:00
1aad125815
[CPU] Enable chunked prefill and prefix caching for qwen3.5 ( #46202 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-25 03:49:21 +00:00
dc55936f64
[AMD][CI] Fix Pipeline + Context Parallelism test group ( #46650 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 22:23:42 -05:00
Bugen Zhao and GitHub
76c3c4ff63
[Rust Frontend] Introduce unified parser interface & combined parser ( #46583 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-25 03:17:31 +00:00
efb5acffd5
[Bugfix] fix: stream Mimimax m2 tool call string arguments ( #46382 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-25 03:12:45 +00:00
6e3a983cf3
[ROCm] Remove erroneous inclusion of gptq_marlin as supported quant scheme on ROCm ( #46655 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 21:27:19 -05:00
Xin Yang and GitHub
1273a8f05a
[Kernel] Add swap AB optimization to fused_moe_kernel ( #36559 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-06-25 01:44:30 +00:00
9e88e969c0
[Perf][KVConnector][Mooncake] Parallelize KV load with a receive-thread pool ( #45971 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 18:25:12 -07:00
dda3aca47f
[Speculative Decoding] Propagate norm_output and fc_norm config for Eagle3 speculators ( #46488 )
...
Signed-off-by: Orestis Zambounis <orestis.zambounis@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-25 00:51:33 +00:00
Jee Jee Li and GitHub
23aed9b0ee
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel ( #46508 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-25 08:42:51 +08:00
Maxwill Lin and GitHub
cd347298e8
[Frontend] Port seed_oss to the streaming parser engine as a Qwen3 subclass ( #46314 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 20:08:42 -04:00
Yifan Qiao and GitHub
b69816043a
[Bugfix][MooncakeStore] track resumed requests via scheduler's resumed_req_ids ( #46595 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 23:50:56 +00:00
Kaihang Jiang and GitHub
fc7fc421e9
[Kernel][MoE] Allow FlashInfer MXINT4 MoE for gated SiLU ( #46518 )
...
Signed-off-by: Kaihang Jiang <kaihangj@nvidia.com >
2026-06-24 18:32:50 -05:00
cyq and GitHub
e06a83445c
[Bugfix] Normalize slashes in Helion GPU names ( #46101 )
...
Signed-off-by: cyq <15000851237@163.com >
2026-06-24 18:22:49 -05:00
d7ab9be775
[Bugfix] Support -1 (invalid/non-local) slots in topk_ids for Triton MoE ( #46408 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:59:42 -07:00
6a1570711c
[Bugfix] Support non-power-of-2 top_k in legacy triton_kernels routing ( #46406 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 13:52:09 -07:00
Micah Williamson and GitHub
d6696e2385
[ROCm] Begin Deprecation Window for CUDA_VISIBLE_DEVICES on ROCm ( #46636 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-24 20:40:28 +00:00
Chauncey and GitHub
84c2f9f0fb
[Frontend] Fix Kimi K2 tool call IDs for required tool choice ( #46344 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-24 19:59:40 +00:00
49f2104c53
[Feature] Support DCP with FP8 KV cache in MLA decode path ( #44044 )
...
Signed-off-by: shivampr <shivampr.dev@gmail.com >
Signed-off-by: Shivam <shivamprasad91@gmail.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 19:28:17 +00:00
d511b5bae9
Chore: Fix minor doc sentence, grammar, quote errors ( #40469 )
...
Signed-off-by: Ashwin Phadke <23502062+ashwin-phadke@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-24 18:58:24 +00:00
3c43237233
[Bugfix][Model Runner V2][Spec Decode] Fix int32 offset overflow in sampler kernels ( #46560 )
...
Signed-off-by: xiaojun.wei <jessiewei747@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-24 11:00:57 -07:00
56ca5997ea
Humming support for 2/3/5/6/7-bit pack-quantized weight-only inference ( #46389 )
...
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-24 13:53:54 -04:00
Aarushi Jain and GitHub
cf57311187
Run DeepSeek-V2-Lite prefetch-offload eval eager on ROCm ( #46386 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-24 12:25:57 -05:00
Lucas Wilkinson and GitHub
e7df232288
[KV Offload] Gate packed HMA KV cache on cross-layer config ( #46252 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-06-24 11:55:30 -04:00
b3a688cb9e
[ROCm] Fix OOB During Model Warmup With ROCM_ATTN and MRV2 ( #46548 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-24 10:53:21 -05:00
Wentao Ye and GitHub
1cd3e0e945
[Bug] Fix IndentationError: expected an indented block after 'with' statement ( #46627 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-24 23:14:17 +08:00
Yiwei Hu and GitHub
f889325c51
[KV Offload] Use background thread for mmap / cpu_tensors pinning ( #45850 )
...
Signed-off-by: Sorryhorizon <arikara6666@gmail.com >
2026-06-24 18:13:27 +03:00
bb61177e49
[KV Offloading] Replace bool|None lookup return with LookupResult enum ( #46363 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 18:06:08 +03:00
7f99e80c3b
[Perf][ThinkingBudget] reduce search space for thinking tokens ( #46425 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 23:02:25 +08:00
2801b11156
[Test] Pin block_size in auto-fit max_model_len test ( #45914 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:56:21 +08:00
007b5a52ed
[Log] Update to log once ( #46511 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-24 14:45:16 +00:00
Cyrus Leung and GitHub
24d5186138
[Bugfix] Re-enable FP8 MoE on NVIDIA Thor ( #46339 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-24 07:35:46 -07:00
Nemani Harsha Vardhan and GitHub
7dc036058b
[Doc] Document Qwen3.6 (dense + MoE) ViT CUDA graph support ( #44720 )
...
Signed-off-by: harsha20032020 <nhvardhan2020@gmail.com >
2026-06-24 14:35:08 +00:00
61ee183d28
[ROCm] Fix AITER FP8 quantization schema tests ( #46414 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 22:29:19 +08:00
84c62e1cbd
[Model Runner V2][MM] Support EVS ( #46535 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:18:56 -04:00
Fadi Arafeh and GitHub
061043eaca
[CPU][Perf] Accelerate unquantized MoE for AArch64 ( #46353 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-24 14:14:35 +00:00
93ec645878
[Bugfix] Fix illegal memory access from a forward during a partial wake_up ( #44483 )
...
Signed-off-by: Meihan-chen <zr010426ztt@outlook.com >
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 22:12:23 +08:00
Kunshang Ji and GitHub
563c628968
[XPU] bump up vllm_xpu_kernels to v0.1.10.1 ( #46607 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 10:05:31 -04:00
0bc479e6eb
[Perf][LoRA] Replace O(n) list.index() with a dict in convert_mapping ( #46542 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 21:41:46 +08:00
Tae Jeong and GitHub
62890e204c
Fix duplicated logging when loading a corrupt or partial video ( #46467 )
...
Signed-off-by: hhhhhhhhhhhhhhhhho <man2719@naver.com >
2026-06-24 06:14:13 -07:00
Nicolò Lucchesi and GitHub
a2cb08b3d5
[Misc][PD] Disable bidirectional xfer mode for NixlPushConnector ( #46473 )
...
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai >
2026-06-24 21:14:05 +08:00
cf9fd6457e
Fix KV offload request-finished lifecycle contract ( #46284 )
...
Signed-off-by: test test <2260891073@qq.com >
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-24 15:42:40 +03:00
Kunshang Ji and GitHub
d4448b511d
[XPU][Docker] switch to ubuntu 24.04 as base image ( #45973 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 20:39:20 +08:00
f1a6703edd
[Bugfix][Config] Keep pydantic validation for fields with a TYPE_CHECKING Literal alias ( #46220 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 12:25:50 +00:00
Roy Wang and GitHub
160c80a34c
[Rust Frontend] Raise frontend JSON body limit ( #46582 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
2026-06-24 12:15:31 +00:00
Martin Hickey and GitHub
f237e16b41
[KV Offload] Replace OffloadingHandler with OffloadingWorker ( #45053 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-24 14:44:24 +03:00
70749fdcca
[Feature] Triton INT4 per-token-head KV cache quantization ( #40835 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 10:21:25 +00:00
d20dbf921b
[Mooncake] Only check and store new KV cache range ( #46412 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-24 03:10:50 -07:00
ede54b926e
set AttentionCGSupport.UNIFORM_BATCH for fa2 on xpu ( #46555 )
...
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-24 18:05:02 +08:00
52fbe12283
[Perf][Multimodal] Avoid building a full timestamps list in video frame sampling ( #46543 )
...
Signed-off-by: Lynn <lynnhe02@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-24 09:38:27 +00:00
Dakai An and GitHub
dc0d318177
[Attention] Add FLASH_ATTN_MLA_SPARSE backend for Hopper sparse MLA ( #46189 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-06-24 09:33:10 +00:00
soaringk and GitHub
d7c1821b5a
[Model][MiniMax-M3] Add pipeline parallelism support ( #45810 )
...
Signed-off-by: soaringk <k3vin.zhang@gmail.com >
2026-06-24 08:23:03 +00:00
4cd1a84c88
[Model] Remove BaiChuanForCausalLM and BaichuanForCausalLM ( #46362 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-24 16:13:57 +08:00
Mohammad Miadh Angkad and GitHub
191826ec61
[CI/Build] Fix topk histogram build on SM75 ( #46550 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-24 00:51:11 -07:00
Andreas Karatzas and GitHub
549c7074cd
[ROCm][CI] Skip the MoE Marlin tile-padding helper assertion ( #46580 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 07:31:33 +00:00
489abadfb8
feat: support to OpenMOSS-Team ( #44124 )
...
Signed-off-by: nagisa-kun <1434936049@qq.com >
Signed-off-by: nagisa19 <1434936049@qq.com >
Signed-off-by: nagisa <1434936049@qq.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 00:08:13 -07:00
Woosuk Kwon and GitHub
96de8bb389
[MoE] Free unused MXFP4 scales in OAI Triton Backend ( #46549 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-24 00:06:41 -07:00
Jee Jee Li and GitHub
9d6fdc2901
[Kernel] GLM5 Router GEMM ( #46385 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:54:50 -07:00
Benjamin Chislett and GitHub
4c5bc41ba6
[Bugfix][Spec Decode] Fix probabilistic sampling for parallel drafting ( #45956 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-24 05:36:23 +00:00
Michał Ganczarenko and GitHub
ac1fa74616
[Bugfix] Fix NemotronLayerNorm1P hardcoded cuda device type ( #46495 )
...
Signed-off-by: <Michal Ganczarenko> <michal.ganczarenko@intel.com >
2026-06-24 13:21:02 +08:00
Sting Lin and GitHub
556bc4e3a0
Upgrade tpu-inference to v0.23.0 ( #46568 )
2026-06-23 21:15:14 -07:00
Wei Zhao and GitHub
05a0caba91
[Mooncake] Optimize lookup pool key string construction ( #46188 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-24 11:51:53 +08:00
Nick Hill and GitHub
7ee4d22009
[Spec Decode] Reject placeholder (-1) draft tokens in rejection sampler ( #46533 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-24 03:32:32 +00:00
ce9f64020b
[Rust Frontend] Pass effective reasoning_parser_kwargs for structured output ( #46360 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 03:13:44 +00:00
4ed8eaafb0
[Rust Frontend] Integrate xgrammar-structural-tag for strict and required tool calling ( #46057 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-24 10:46:49 +08:00
6af0559ddb
[Core][DP] Throttle prefills based on local prefill work ( #46532 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 02:27:12 +00:00
e2bdc24612
[ROCm][Bugfix] Fix use_v2_model_runner inside Ray driver thread ( #45998 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-24 08:41:56 +08:00
Andreas Karatzas and GitHub
bcbeaac786
[ROCm][CI] Stage C-II of gating additional test groups ( #46537 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 17:36:40 -07:00
Maxwill Lin and GitHub
e48f2aa4ca
[Bugfix][Frontend] Emit a content block for empty Anthropic completions ( #46525 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-24 00:04:26 +00:00
Roberto L. Castro and GitHub
d86c66c981
[Feat] Add runtime monitor for post-warmup CuTeDSL compilation ( #46167 )
2026-06-23 23:33:17 +00:00
Nico Holmberg and GitHub
80e511772f
[ROCm][Bugfix][Perf] enable shared expert fusion for Qwen3.5 ( #44434 )
...
Signed-off-by: Nico Holmberg <nico.holmberg@amd.com >
2026-06-23 23:19:51 +00:00
Roberto L. Castro and GitHub
855cd4d787
[Perf][DSv4/DSv3.2] Add cluster-cooperative topK kernel for low-latency scenarios ( #43008 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-23 16:11:00 -07:00
3cc871aaf1
[Perf] Skip detokenization in online beam search ( #46422 )
...
Signed-off-by: Guy Stone <guys@spotify.com >
Signed-off-by: Guy Stone <guystone3@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 15:46:09 -07:00
0a3e2dbc09
[Optimization] Skip DP padding tokens in MoE ( #46428 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 14:54:46 -07:00
84f13374b3
[CI] Fix test_auto_gptq on ROCm CI ( #46164 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 16:38:06 -05:00
Micah Williamson and GitHub
b28103e1ca
[ROCm][CI] Shard LM Eval Qwen3-5 Models (B200-MI355) in AMD CI ( #46520 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 16:32:05 -05:00
Wentao Ye and GitHub
abc33134fa
[CI Test] Mark batch invariance test flaky ( #46530 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 21:01:34 +00:00
6617db1bfb
[Bugfix][Frontend] Emit non-ASCII tool-call arguments without \uXXXX escapes ( #46308 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-23 20:43:11 +00:00
899d72a58c
[Bugfix][ToolParser] Handle braces in required tool streaming strings ( #45389 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-23 20:29:34 +00:00
Yongye Zhu and GitHub
11b56b2ff2
[Kernel] Add FlashInferCutedslMxfp8LinearKernel (cute-dsl mm_mxfp8) ( #46393 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-23 12:45:49 -07:00
0d4d164488
[Bugfix] Allow flashinfer_cutlass as a clamped NVFP4 MoE backend ( #46492 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-23 12:43:36 -07:00
Mike G and GitHub
0775b882ba
[NVFP4 MoE/Deepseek V4] Marlin: wire SwiGLU clamp + allow it for clamped models on non-Blackwell ( #45836 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-23 12:21:19 -07:00
7c2e08451a
[Docker] Remove redundant flashinfer download-cubin step ( #46517 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 12:16:51 -07:00
Giancarlo Delfin and GitHub
ef361de916
[Model Runer V2][DFlash] Fix lm head sharing for dflash ( #46435 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-23 19:09:06 +00:00
Yan Ma and GitHub
acce57d8dd
Deprecate old FP8 online MoE quantization class ( #44514 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:53:38 -07:00
68afd78897
[Bugfix][ROCm] Fix cumem sleep and teardown ( #46203 )
...
Signed-off-by: pei.zhang <pei.zhang@amd.com >
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-24 02:45:31 +08:00
37a682d392
[Kernel] Extend Marlin thread-tile padding to MoE (WNA16 + FP8/MXFP8) ( #45703 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 11:45:10 -07:00
Rui Yin and GitHub
d8e422ccda
[Bugfix] Parse MiniMax M3 streaming reasoning by text markers ( #45718 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-23 14:43:58 -04:00
fxmarty-amd and GitHub
e368415daa
[AMD][OCP MX][CI] Fix tests to not dispatch on UNFUSED_TRITON backend on MI300, improve w_mxfp4_a_fp8 emulation support ( #46142 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-23 14:25:27 -04:00
Andreas Karatzas and GitHub
ceae5bcbda
[ROCm][CI] Fix nixl tests ( #45219 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 13:11:40 -05:00
6691f087a6
[Minimax-M3] BF16/FP8 Indexer using MSA ( #45892 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-23 10:28:49 -07:00
f4d5f73ffa
[Bugfix]: Fix unquantized gpt-oss weight loading broken by FusedMoE r… ( #45818 )
...
Signed-off-by: priyansh jain <priyansh.jain2@amd.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-23 16:56:17 +00:00
fd50a66015
[CI][ROCm] Skip unsupported test cases on ROCm ( #46160 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 11:35:49 -05:00
84586c9acc
[ROCm][CI] fix fp8 range in vit_fp8_quant ( #46410 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: Divakar Verma <137818590+divakar-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-23 11:34:21 -05:00
Taneem Ibrahim and GitHub
40e5522121
[Docs] Add Qwen3 forced alignment online example ( #46197 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-23 11:59:45 -04:00
Willow Lopez and GitHub
f3410b3bb1
fix(moe_wna16): access tp_size via moe_config for RoutedExperts compatibility ( #45404 )
...
Signed-off-by: Oxygen <1391083091@qq.com >
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-23 11:46:23 -04:00
568874fec2
[ROCm][CI] pass merge-base to container for python-only wheel metadata ( #45869 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-23 15:44:43 +00:00
275b43183c
[MyPy] Fix mypy for vllm/benchmarks ( #39896 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-23 15:22:29 +00:00
Yan Ma and GitHub
547d2c40d7
Add weights padding for fp8 per-block online quantization ( #44763 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-23 11:08:17 -04:00
2aaaf3febd
[ROCm][Test] Fix stale test_gfx950_moe MXFP4 oracle tests ( #46260 )
...
Signed-off-by: Spandan Tiwari <23646532+spandantiwari@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 23:07:46 +08:00
Micah Williamson and GitHub
156b12667c
[ROCm][CI] Skip Quark mxfp4 tests unless Quark version is compatible with Torch version ( #46431 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 22:26:39 +08:00
Jee Jee Li and GitHub
9f6f296428
[CI/Build] Remove BaiChuanForCausalLM from the LoRA test ( #46494 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-23 22:09:48 +08:00
e51e700470
[LoRA] Gate all_gather on fully_sharded_loras inside _mcp_apply; rewrite regression test ( #45715 )
...
Signed-off-by: lcheng <lcheng321@gatech.edu >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-23 07:08:33 -07:00
f59db63732
[Bugfix] GPT-OSS Autodrop reasoning in Response API and cleanup ( #45048 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 09:36:33 -04:00
Rukhaiya2004 and GitHub
9f5117820f
[HARDWARE][POWER] Enable fp16 support for PowerPC ( #46135 )
...
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
2026-06-23 13:24:49 +00:00
1bf149f334
Filter Pydantic-internal markers from validation error param ( #46457 )
...
Signed-off-by: muhammadfawaz1 <135441198+muhammadfawaz1@users.noreply.github.com >
Co-authored-by: Mahad Rehmann <114791389+mahadrehmann@users.noreply.github.com >
2026-06-23 13:20:50 +00:00
2a675a7b9f
[Bugfix] Responses API assistant EasyInputMessageParam input ( #44361 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-23 08:54:46 -04:00
7d47cff933
[Bugfix][KV Offload] Fix swap_blocks_batch on the default stream ( #46379 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-23 05:45:27 -07:00
091bc1026e
[KV Offloading] Add tiering metric plumbing ( #45959 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-23 15:10:36 +03:00
3554ada5d8
[CPU][Bugfix][Speculative Decoding] Accept USE_FP64_GUMBEL in CPU recovered-tokens sampler ( #46069 )
...
Signed-off-by: hillel.darshan <hillel.darshan@intel.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-23 11:54:07 +00:00
wang.yuqi and GitHub
31ca9504b1
[Frontend] Split ServingRender into renderer and entrypoint. ( #44285 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-23 11:19:09 +00:00
d32575a2d2
[ROCm][P/D] Support MoRIIO heterogeneous TP fan-in ( #46332 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-23 10:33:23 +00:00
Juan Pérez de Algaba and GitHub
83fa302ca4
fix(security): prevent infinite loop in split_audio with NaN audio sa… ( #46463 )
2026-06-23 10:24:51 +00:00
frida-andersson and GitHub
20b5af55c1
[ROCm][Perf] DSv3.2: fuse MLA Q concat+fp8-quant in forward_mqa ( #43673 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
2026-06-23 18:12:04 +08:00
Qiming Zhang and GitHub
901a3b091c
fix gpt_oss pp>1 with ep ( #46441 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
2026-06-23 16:59:11 +08:00
2d721ab5d8
[Rust Frontend] Align Rust allowed_token_ids validation with Python ( #46348 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: reidliu41 <reid201711@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 08:32:33 +00:00
accaa434f3
[Rust Frontend] Support echo for token-ID completion prompts ( #46219 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-23 08:04:41 +00:00
Sunny Yuan and GitHub
a04654da23
Doc: fix missing GLM-5.x in supported models ( #46452 )
...
Signed-off-by: Sunny Yuan <y.zichen@wustl.edu >
2026-06-23 07:42:27 +00:00
Bugen Zhao and GitHub
25bc3be49c
[Rust Frontend] Correct --reasoning-parser semantics ( #46359 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 15:38:39 +08:00
Ting SUN and GitHub
a46f3eb232
[Bugfix][Model Runner V2] Preserve all allowed_token_ids in the logit bias kernel ( #46245 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-23 07:01:13 +00:00
6c427dd401
[BugFix] Omit empty tool_calls from OpenAI chat responses ( #44105 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-23 13:43:53 +08:00
3ce5823762
[Refactor] Responses API parser state into conversation context ( #46030 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 13:42:58 +08:00
Woosuk Kwon and GitHub
04c2a8deac
[DeepEP V2] Fill invalid recv_topk_idx with -1 ( #46432 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-22 21:45:49 -07:00
7e47fb72b5
[ROCm][P/D] Fix MoRIIO WRITE mode for mixed KV layouts ( #46290 )
...
Signed-off-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Jun Kang Chow <junkangchow@gmail.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
2026-06-23 12:12:51 +08:00
a8481be7a9
[Rust Frontend][Perf] Use dedicated runtime for HTTP/request-processing/ZMQ ( #46051 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-23 04:03:20 +00:00
Kunshang Ji and GitHub
9d3317172c
[XPU][CI]fix xpu kv cache layout test ( #46429 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-23 03:43:29 +00:00
430a95ae3a
[v1][kvcache] Honor prefix-cache retention interval for Mamba/linear attention ( #45845 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 19:51:11 -07:00
Mike G and GitHub
56e5797511
[Quant] Enable modelopt_mixed on Turing (SM75) ( #45375 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-22 19:30:49 -07:00
8db12169a4
fix: stream Qwen3 tool call string arguments ( #46351 )
...
Signed-off-by: Rui Yin <2260891073@qq.com >
Co-authored-by: abinggo <107740309+abinggo@users.noreply.github.com >
2026-06-23 10:26:37 +08:00
33f50773cb
[Doc] Fix typos, grammar, and broken commands across docs ( #46398 )
...
Signed-off-by: MichaelCaoo <a992033227@163.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-23 02:01:22 +00:00
Micah Williamson and GitHub
fa36f86d77
[CI] Torch 2.11 flaky test_spec_decode_logprobs and gritlm tests ( #45772 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-23 01:26:54 +00:00
8207ce0850
[Bugfix] Fix humming lm_head crash and FusedMoE weight_shape coercion ( #46420 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 18:19:29 -07:00
e48592066e
[DeepEP V2] Bound num_max_tokens_per_rank in do_expand=False ( #46404 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Roy Wang <jasonailu87@gmail.com >
Co-authored-by: gnovack <novackgm@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 18:14:53 -07:00
91ba720b75
[ROCm][CI] Only require q_scale==1.0 for fp8 query in RocmAttention ( #46148 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 18:25:43 -05:00
fxmarty-amd and GitHub
6ead164e52
[CI] Add TP=4 requirement to test_mixed_precision_model_accuracies ( #46161 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-06-22 18:19:43 -05:00
c97e8f99d6
[ROCm][Quantization][4/N] refactor quark_moe fp8 w/ oracle ( #43721 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:58:03 -07:00
183b5f27ea
[Bugfix][V1][TurboQuant] Reserve workspace before CUDA graph capture ( #44053 )
...
Signed-off-by: Guipeng Zhang <zhangguipeng23z@ict.ac.cn >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 15:47:48 -07:00
ca5b24695b
Fix static actorder handling for compressed-tensors WNA16 MoE ( #41161 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:46:46 -07:00
Charlie Fu and GitHub
6f6bd3b8fe
[ROCm][CI] Increase the max wait time for server startup ( #46417 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-22 17:46:31 -05:00
Andreas Karatzas and GitHub
70ef4d3009
[ROCm][CI] Purging away redundant test group definitions ( #46418 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:42:47 -07:00
e2fe837572
[CI] Fix CPU-Multi-Modal Model Tests timeout by adding a 4th shard ( #46388 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 22:08:00 +00:00
Aarushi Jain and GitHub
fbf9ff7cf4
[CI][ROCm] Restrict MLA cross-layer KV cache test to supported backends on ROCm ( #46401 )
...
Signed-off-by: aarushjain29 <Aarushi.Jain2@amd.com >
2026-06-22 17:05:26 -05:00
6cc2c9ba3a
[CI] Add DGX Spark GPQA smoke test ( #39541 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 14:52:38 -07:00
c0b2d8f471
[Bugfix] FusedMoE: coerce shape-(1,) per-tensor scales to 0-D scalar … ( #43362 )
...
Signed-off-by: Varshith <kvarshithgowda@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-22 13:26:53 -07:00
Mohammad Miadh Angkad and GitHub
d1a38c2762
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend ( #42235 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-22 16:17:18 -04:00
2b4a7491ec
[ROCm][CI] Query total device memory via amdsmi to avoid HIP init ( #46141 )
...
Signed-off-by: stefankoncarevic <stefan.koncarevic@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-22 15:12:24 -05:00
Saddss and GitHub
82ede09a5a
[Bugfix][KVConnector] Fix SimpleCPUOffloadConnector GPU->CPU store race ( #46278 )
2026-06-22 13:08:47 -07:00
Nick Hill and GitHub
fbf520cf3a
[MRV2] Generalize use of WhisperModelState ( #46096 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-22 12:40:02 -07:00
44d95069e9
Enable DeepSeek V4 and GLM-5.1 on SM120 ( #43477 )
...
Signed-off-by: Zihua Wu <13583761+lucifer1004@users.noreply.github.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-22 11:54:14 -07:00
3ce15fd574
[v1][kvconnector] DecodeBenchConnector: fill list/tuple (Mamba/KDA) KV caches ( #45080 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-22 11:54:00 -07:00
e4b3da3feb
[Quantization][CI] add humming lm-eval test ( #43752 )
...
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:23:55 -07:00
3e6529cc0e
[Bugfix][Spec Decode] Fix EAGLE drafter multimodal encoder cache misses ( #46315 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 18:14:02 +00:00
ac614587f5
[EPLB] Enable nixl eplb communicator for elastic ep ( #45013 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
2026-06-22 10:54:08 -07:00
f2069b005b
[Pooling] Validate non-negative rerank top_n ( #46119 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 11:40:47 -04:00
Martin Hickey and GitHub
ccd49f6821
[MyPy] Fix mypy for vllm/lora ( #41722 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-22 10:57:09 -04:00
Li, Jiang and GitHub
1c7bc18318
[Bugfix][CPU] Fix CPU model runner v2 ( #46365 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-22 22:52:05 +08:00
AlexHuang and GitHub
9a938df64e
[Test][KV Offloading] Add unit tests for OffloadingSpecFactory and SecondaryTierFactory ( #46355 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
2026-06-22 17:45:04 +03:00
Liangliang Ma and GitHub
3da4a1b124
[XPU] add awq format for INCXPULinear ( #43404 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
2026-06-22 22:29:13 +08:00
6871738777
[Doc] Document pull request limit ( #46376 )
...
Signed-off-by: simon-mo <simon.mo@hey.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 14:04:56 +00:00
Yifan Qiao and GitHub
aa4990a9a2
[Attention] Re-enable cross-layer KV cache layout for MLA via stride-aware kernels ( #45111 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-22 06:57:02 -07:00
a4610da0c6
[docs] link security docs from AGENTS ( #46373 )
...
Add a security-review routing sentence to AGENTS.md that points agents to SECURITY.md, docs/usage/security.md, and docs/contributing/vulnerability_management.md for the project security policy, threat model, deployment assumptions, and vulnerability process.
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-22 06:28:25 -07:00
liuzhenwei and GitHub
09cdcf34aa
[XPU] update nixl to v1.2.0 ( #46327 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 20:55:06 +08:00
d2c671c29b
[CPU][RISC-V] Add RVV micro GEMM for WNA16 ( #44324 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-22 12:53:54 +00:00
xiangdong and GitHub
b5a2adec4b
[XPU][CI]Skip v1/spec_decode/test_speculators_correctness.py in intel GPU nightly ( #46356 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 19:30:41 +08:00
78739e3bda
[Bugfix] Reject matryoshka embedding dimensions above hidden size ( #46313 )
...
Signed-off-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
Co-authored-by: EazyReal <8047065+EazyReal@users.noreply.github.com >
2026-06-22 10:16:35 +00:00
Tuukka Sarvi and GitHub
89accad2cc
[ROCm][DSV4] Disable TileLang MHC dispatch on gfx942 ( #45931 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
2026-06-22 09:26:54 +00:00
3c8e49596c
[Model] ColQwen3.5: fix retrieval correctness (bias + bidirectional) ( #46108 )
...
Signed-off-by: Athrael Soju <athrael.soju@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-22 17:25:54 +08:00
cec2ec1176
[Bugfix] Avoid racy accepted counts in async spec decode ( #45100 )
...
Signed-off-by: Weiwei Sun <68775773+sunnweiwei@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-22 08:53:16 +00:00
liuzhenwei and GitHub
435f82d61a
[Bugfix] Fix Llama4ForCausalLM initialization test failure ( #46341 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-22 08:40:43 +00:00
Roger Wang and GitHub
1c4b51b990
Temporarily skip M3 on CI ( #46352 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:35:31 -07:00
2e2c47928b
[Doc] Update MiniMax-M3 ( #45940 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-22 01:23:27 -07:00
80abe0de7d
[Rust Frontend] Support thinking_token_budget for chat and completions ( #46137 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-22 16:00:02 +08:00
a9f7b2d41c
[feature][kv_offload] Self-describing KV events for OffloadingConnector ( #43468 )
...
Signed-off-by: Change72 <changg@nvidia.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-22 07:27:46 +00:00
d14e551a53
[Model] Remove MiniMaxText01, MiniMaxVL01, MiniMaxForCausalLM ( #45993 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-22 15:20:46 +08:00
68567ef2df
[CPUOffloadingManager] Maintain evictable list in LRUCachePolicy ( #46216 )
...
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-22 06:54:44 +00:00
6bc6f2d86d
[1/N][Core] add partial prefix cache primitives ( #45939 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 23:43:10 -07:00
wang.yuqi and GitHub
1eb2cc961e
[Frontend] Refactor ServingTokenization entrypoint. ( #46022 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-22 06:27:58 +00:00
31124749d1
[Bugfix] [Rust Frontend] Fix stop string truncation with repeated matches ( #46113 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-22 14:11:29 +08:00
Ma Jian and GitHub
9037498c22
[DSV4][XPU] Pass gemm1_clamp_limit to XpuFusedMoe ( #44517 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-22 12:57:10 +08:00
db32b53e30
[SpecDecode] Support DFlash with FlashInfer ( #43081 )
...
Signed-off-by: gss <2783977641@qq.com >
Co-authored-by: gss <2783977641@qq.com >
2026-06-22 04:55:30 +00:00
xiangdong and GitHub
b529bfd6c5
[XPU][CI] Add agent_tags for Intel GPU CI ( #45768 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-22 10:33:17 +08:00
Micah Williamson and GitHub
f3df7a7231
[ROCm][CI] Enable kv_connector unit tests on ROCm ( #45955 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-22 05:08:44 +03:00
485bbe1c6f
[CI] Fix missing tp_size attribute on RoutedExperts ( #46163 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 18:46:49 -06:00
Matt and GitHub
a19ff2218a
[Hardware][AMD][CI] Fix Spec Decode Eagle test group ( #46018 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:40:02 -05:00
Matt and GitHub
4f0d0049a0
[Hardware][AMD][CI] Fix Kernels Attention test groups ( #46080 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 17:10:51 -05:00
13b83d77ad
[ROCm][CI] skip test_double_aiter_rms_quant_fusion ( #45967 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-21 16:53:11 -05:00
Matt and GitHub
50241602fd
[Hardware][AMD][CI] Fix gfx942 Kernels MoE test group ( #46298 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-21 16:45:37 -05:00
Ting SUN and GitHub
12fe2a9aac
[Bugfix][Qwen3-VL] Fix multi-video crash with list-valued fps/num_frames ( #46305 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 14:31:23 -07:00
Benjamin Chislett and GitHub
89bd2c14d3
[Spec Decode] Add Qwen3 architecture support for EAGLE3 ( #43132 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
2026-06-21 13:55:26 -07:00
ZedongLiu and GitHub
9c450b1027
[Kernel][Bugfix] Fix INT8 per-token-head KV cache rounding in Triton reshape-and-cache ( #45361 )
...
Signed-off-by: ZedongLiu <113341356+Zedong-Liu@users.noreply.github.com >
2026-06-21 15:59:40 -04:00
635c38338a
[Multimodal] Add Qwen2-VL/Qwen2.5-VL processor-mapped video loader ( #45555 )
...
Signed-off-by: Ranran <hzz5361@psu.edu >
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-21 18:56:50 +00:00
c441ad1c07
[KV Offloading] Add labeled metrics support ( #45957 )
...
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-21 18:04:01 +00:00
Jee Jee Li and GitHub
745bba5ea8
[Model]Fix MiniMaxM2ForCausalLM perf regression ( #45935 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-22 00:28:52 +08:00
2cac89f9da
[Spec Decode] Support mixed KV page sizes for DFlash ( #45181 )
...
Signed-off-by: Alex Steiner <asteiner@nvidia.com >
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-21 22:45:14 +08:00
3e6e33526d
[Disagg] return routed_experts on streaming generate responses ( #44638 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-21 07:37:10 -07:00
b91b7726e0
[ROCm][P/D] Support MiniMax-M3 mixed KV layouts in MoRIIO READ mode ( #46039 )
...
Signed-off-by: Jun Kang Chow <junkangchow@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: Hongxia Yang <hongxia.yang@amd.com >
Co-authored-by: Tan Pin Siang <tanpinsiang@gmail.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
Co-authored-by: Chun Fang <chun.fang@amd.com >
Co-authored-by: TianDi101 <ditian12@amd.com >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-21 12:55:19 +00:00
Palaiologos1453 and GitHub
d3ad8e8bcd
[Bugfix] Defer offload reads while transfers are pending ( #46231 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-21 14:30:13 +03:00
b80ce9dd2f
[CI][test] Replace InternVL2-1B with InternVL3-1B in test_pipeline_parallel.py ( #46241 )
...
Signed-off-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-21 15:11:19 +08:00
b5495cc5f9
Fix memory pointer overflow in Mamba state buffers ( #44665 )
...
Signed-off-by: Shifani Rajabose <shifani.rajabose@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-21 14:00:50 +08:00
Ting SUN and GitHub
183a430c13
[Bugfix][Model Runner V2] Fix min_tokens off-by-one in the V2 GPU sampler ( #46243 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-21 05:06:49 +00:00
Matt and GitHub
a346d589f5
[Bugfix] Fix NVFP4/OCP MX MoE emulation ( #46254 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
2026-06-20 23:13:10 -05:00
Nick Hill and GitHub
7df3d7dada
[Core] Ensure memory is pinned prior to async h2d copy ( #45424 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 20:02:24 -07:00
8dd1b702f2
[Misc] Fix stale doc URL and docstring module path ( #35530 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-20 23:57:01 +00:00
f57ac274b2
[Render] Add reasoning/tool parsing to /derender + fix byte-fallback FFFD ( #45919 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-06-20 19:43:32 -04:00
6e919960af
[Perf] Skip/shrink all_token_ids copy in scheduler for non-async and V2 runner ( #45840 )
...
Signed-off-by: amanchugh89 <amanchugh.89@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-20 22:36:57 +00:00
Jonathan Chen and GitHub
c88d3d4775
[SimpleCPUOffloadConnector] PCP + DCP support ( #39831 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
2026-06-20 15:01:06 -07:00
Yifan Qiao and GitHub
ab7fcbdd5d
[Perf][KVConnector][Mooncake] Compact chunk-hash keys and zero-copy lookup wire format ( #45969 )
2026-06-20 15:00:11 -07:00
3b4a76b63f
[KV-Offloading] : Expose CPU cache usage metric ( #45737 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-20 21:21:55 +00:00
cc22621b51
[KV Offload] Support packed HMA KV cache layout ( #46205 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:19:40 +00:00
77148992cf
[Bugfix] Move extract_layer_index back inside is_v32 guard ( #46199 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 21:19:10 +00:00
891cc4b9c5
[Frontend] Report cache usage in Anthropic /v1/messages API ( #40912 )
...
Signed-off-by: mistral0105 <zhangshuoming17@mails.ucas.ac.cn >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-20 21:12:48 +00:00
TJian and GitHub
1bdf9810aa
[ROCm] [Bugfix] Bugfix ROCm Sparse Indexer ( #46222 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-20 13:38:42 -07:00
ebfbcfe46a
Stop setting CUDA_VISIBLE_DEVICES internally in vLLM, add device_ids arg ( #45026 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Codex <codex@openai.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: kourosh hakhamaneshi <kouroshHakha@users.noreply.github.com >
2026-06-20 13:38:10 -07:00
e9de72fe6c
[Bugfix] Guard model_config access in _log_compilation_config ( #46198 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-06-20 19:26:38 +00:00
d272418f45
[Perf] Optimize Qwen3-VL multi-video prompt processing ( #46026 )
...
Signed-off-by: Sirius29 <422058530@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 07:09:18 -07:00
Sumanth R Hegde and GitHub
7ff7f5c8eb
Revert "Fix Stale Encoder Cache After Weight Update" ( #46125 )
2026-06-20 07:09:09 -07:00
dced290769
[Hardware][AMD][CI] Fix e2e core test group ( #46024 )
...
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-20 02:04:35 -05:00
JasonLi314 and GitHub
93bad11912
[Bugfix] Fix gridDim.y overflow for large row counts ( #45255 )
...
Signed-off-by: Jason Li <li.jason.cs@gmail.com >
2026-06-19 23:27:45 -04:00
djramic and GitHub
0fbf42af84
[ROCm] Fix VRAM not freed in test_phi3v ( #46046 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 17:20:59 -05:00
Charlie Fu and GitHub
e6cd8913dd
[ROCm][CI] Skip Qwen3.5-35B-A3B-MXFP4-AITER-TP2 for non gfx950 ( #46109 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-19 17:20:10 -05:00
Ben Browning and GitHub
859e4d436b
[Bugfix][Parser] Fix U+FFFD leak at reasoning-to-content transition in engine parsers ( #46159 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 22:09:28 +00:00
Micah Williamson and GitHub
4a083cc858
[ROCm][CI] Pin test_rocm_compressed_tensors_w8a8 to TRITON_ATTN ( #46180 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-19 15:20:06 -05:00
Vadim Gimpelson and GitHub
ca7e1f2c43
Move CI failure diagnosis docs into ci-fails-buildkite skill ( #45975 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-06-19 20:12:40 +00:00
djramic and GitHub
dec860fb19
[ROCm] Use vLLM's fp8 quant max in AITER hipBLASLt accuracy test ( #46176 )
...
Signed-off-by: Djordje Ramic <djoramic@amd.com >
2026-06-19 13:24:02 -05:00
Harry Mellor and GitHub
0a49fb2b13
Fix dead link in docs ( #46181 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-19 18:16:09 +00:00
Ben Browning and GitHub
4a8abf37c7
[Test] Migrate test_openai_schema.py to schemathesis 4.x ( #46173 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-19 18:05:18 +00:00
01192139bf
[DSv4] Pack KV caches into contiguous per-block allocations for DeepSeek V4 ( #44577 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-19 12:55:42 -04:00
Chris Leonard and GitHub
b9a7cd464c
[12/n] final _C library kernel migration ( #45415 )
2026-06-19 06:57:26 -07:00
69bdd34542
[Bugfix] Fall back to Pydantic loc for param in validation errors ( #46038 )
...
Signed-off-by: professorsab <135441198+professorsab@users.noreply.github.com >
Co-authored-by: Mahad Durrani <114791389+mahadrehmann@users.noreply.github.com >
2026-06-19 19:11:11 +08:00
Kunshang Ji and GitHub
ec67d7ae61
[xpu] bump up vllm-xpu-kernels v0.1.10 and upgrade 2618 umd ( #40367 )
...
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-19 15:37:20 +08:00
ecf9d83520
[AMD][CI] Fix Language Models Test (Extended Generation) failures ( #45509 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-19 12:06:56 +08:00
Samuel Shen and GitHub
c9135db27c
[Docs] Update stale LMCache examples ( #45762 )
...
Signed-off-by: Samuel Shen <slshen@tensormesh.ai >
2026-06-19 03:21:36 +00:00
2a6c6b9429
[DeepSeek-V4] Support TEP=16 for the block-FP8 shared expert ( #46001 )
...
Signed-off-by: Jeff Ma <jeffjma@umich.edu >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:10:12 -07:00
Jared Wen and GitHub
ab66606993
[bugfix]Indexer init skip and MTP TopK share for iteration ( #45895 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-19 09:57:51 +08:00
9ea3a4015b
[Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded ( #42120 )
...
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 18:26:09 -07:00
Flora Feng and GitHub
560fb8b867
[Cohere] Remove dead prepare_structured_tag override in Cohere parser ( #46099 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-19 01:02:11 +00:00
Wentao Ye and GitHub
675cd5d228
[Model Runner V2] Fix MRv2 memory leak test ( #46095 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-19 00:36:40 +00:00
7f616c327d
[Bugfix] [Parser] Fix empty tool block silently dropping subsequent content ( #46091 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-18 23:17:18 +00:00
Ivy Xu and GitHub
c3c6d723fd
[Perf] Remove unused loggers in reasoning/ ( #45988 )
...
Signed-off-by: Ivy <fakeshadow1337@gmail.com >
2026-06-18 22:24:29 +00:00
41dcf49ca5
[Bugfix][KV Connector] Disable Mooncake TP put-striding when DCP > 1 ( #45371 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:13:44 -07:00
35e4dd4a69
[KV Connector][Mooncake] Async lookup to reduce scheduler overhead ( #45659 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-18 21:44:02 +00:00
4ce2d01453
fix(anthropic): auto-detect template support for mid-conversation system messages ( #46025 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:19:11 -04:00
Woosuk Kwon and GitHub
16908e132e
[MRV2] Make FP32 Gumbel sampling more accurate ( #45996 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-18 19:42:09 +00:00
Wentao Ye and GitHub
225936a1dd
[CI Bug] Revert #42379 to fix CI Multi-Modal Models (Extended Generation 1) ( #46070 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:37:39 -07:00
f6ba720963
(security) Upgrade Starlette to >= 1.0.1 to fix CVE-2026-48710 ( #45675 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-18 12:35:13 -07:00
Wentao Ye and GitHub
b53b1c7ffe
[Model Runner V2] Migration to support quantized model by default [5/N] ( #44446 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-18 12:20:44 -07:00
79ca54d221
[Bugfix][Quantization] Don't reject fp8_e5m2 KV cache for non-fp8 quantized checkpoints ( #45040 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:18:25 -04:00
Ben Browning and GitHub
09f3cd5c10
[Bugfix] [Parser] Fix Qwen3 latent bug in partial params dropping values containing < ( #46047 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 18:04:06 +00:00
ea6078fe6a
[KV Connector][Offloading] Disable parallel-agnostic fs-tier cache on V2 model runner ( #46044 )
...
Signed-off-by: Itay Etelis <etelis2019@gmail.com >
Co-authored-by: Itay Etelis <etelis2019@gmail.com >
2026-06-18 20:43:35 +03:00
Palaiologos1453 and GitHub
a0df04e477
[Tests] Add Qwen3 streaming parser delta boundary cases ( #45708 )
...
Signed-off-by: test test <2260891073@qq.com >
2026-06-18 17:37:39 +00:00
stefankoncarevic and GitHub
e2352c2974
[ROCm][Spec Decode] Fix probabilistic draft probs test attention backend ( #45706 )
...
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com >
2026-06-18 11:59:37 -05:00
qli88 and GitHub
25faa1f4cc
[CI]Enable mxfp4 lora test for ROCm platform ( #43802 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-18 16:59:09 +00:00
Humphrey and GitHub
4583630b56
[Bugfix][Kernel] Check output alignment in vectorize_with_alignment (fixes misaligned-address crash for non-multiple-of-8 head sizes) ( #45466 )
...
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com >
2026-06-18 16:58:22 +00:00
Divakar Verma and GitHub
21da47dabe
[ROCm][CI] move lora%N test to mi300 and gate ( #45970 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-19 00:50:32 +08:00
6c379b9e54
[Frontend] Add Streaming Parser Engine and new GLM4.7/GLM5.1/GLM5.2 Parser ( #45915 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-19 00:42:10 +08:00
Rohan Potdar and GitHub
5099474633
[Bugfix][ROCm] Fix rocm_aiter_per_tensor_quant custom op aliasing ( #45747 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
2026-06-18 11:30:21 -05:00
Yuwen Zhou and GitHub
058cc0a8b6
[Bugfix] Restore is_sym guard for zp in GPTQ/CT MoE to fix symmetric quant regression ( #45656 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-18 16:20:29 +00:00
837db7605e
[Bugfix][Tool Parser] Handle non-finite numbers in coerce_to_schema_type ( #43984 )
...
Signed-off-by: ashishpatel26 <shriganesh.patel@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-18 16:00:20 +00:00
Mark McLoughlin and GitHub
bf2a393034
Temporarily remove @markmc from CODEOWNERS ( #46053 )
...
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-18 14:15:43 +00:00
d682968aa9
[Model] Remove BambaForCausalLM ( #45990 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-18 06:51:00 -07:00
021cdf72bc
Fix _riscv_supports_rvv_vlen128() to detect RVV on hardware without zvl flags ( #43179 )
...
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn >
Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn >
2026-06-18 21:22:35 +08:00
Ashar and GitHub
4cb5e746b6
[Rust Frontend]: Add /get_world_size route with static parallel size ( #44801 )
2026-06-18 13:10:20 +00:00
Jee Jee Li and GitHub
22cc891108
[Kernel] Add PDL support for DeepGEMM kernel ( #46006 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-18 20:49:01 +08:00
afdcbd5d39
[ROCm][DSv4] Functional fixes for DeepSeek V4 on MI300X/MI325X ( #45681 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Jin Tao <jintao12@amd.com >
2026-06-18 12:21:14 +00:00
8d4f54966c
fix(quantization): Fix AWQ dequantize on Intel XPU and refactor AutoAWQ config ( #42727 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@tencent.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 20:12:28 +08:00
351c72d6e5
[CPU] Skip Triton kernel monkey-patches when Triton-CPU is available ( #44991 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-18 18:59:30 +08:00
Tahsin Tunan and GitHub
7299e6509e
[Rust Frontend] Return model metadata fields in /v1/models ( #45950 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-18 10:29:21 +00:00
littlecircle0730 and GitHub
08985351f3
Fix Stale Encoder Cache After Weight Update ( #45093 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
2026-06-18 09:32:10 +00:00
Wei Zhao and GitHub
5fd3b276f8
[Mooncake] Skip KV lookup for non-reachable SWA blocks ( #45444 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-18 02:23:20 -07:00
1e9f04da14
fix(anthropic): preserve inline system message position for prefix caching ( #44602 )
...
Signed-off-by: felix0080 <felix0080@users.noreply.github.com >
Co-authored-by: felix0080 <felix0080@users.noreply.github.com >
2026-06-18 15:58:11 +08:00
702214146c
[Bugfix][Frontend] Fix Anthropic count_tokens decorator order driving server load negative ( #44725 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 23:56:46 -07:00
a331589394
[XPU] Update nixl to v0.10.1 in Dockerfile ( #40287 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 14:01:26 +08:00
Micah Williamson and GitHub
e945169207
Revert "[Kernel] Add PDL support for DeepGEMM kernel" ( #45999 )
2026-06-17 22:59:48 -07:00
554352a311
[Test][KV Connector] Add request_finished fence population tests for offloading scheduler ( #45679 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@future.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:52 +03:00
421c1ec448
[KV Offloading] Remove dummy worker-side stats from OffloadingConnector ( #45905 )
...
Signed-off-by: Alex <alex.tech.lab@outlook.com >
Signed-off-by: AlexHuang <jihuihuang@alexai.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-18 08:13:28 +03:00
b4c80ec0fd
[Refactor] Remove dead cutlass mxfp8 code ( #44681 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-17 21:18:25 -07:00
Ronen Schaffer and GitHub
f428718ffe
[Fix][KV offload] Defer on_request_finished until in-flight transfers drain ( #45823 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-18 07:05:46 +03:00
Jee Jee Li and GitHub
4403af8fb5
[Kernel] Add PDL support for DeepGEMM kernel ( #42996 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-17 20:37:17 -07:00
d57888efa4
[SimpleCPUOffloadConnector]: Add support for reset_cache() ( #39726 )
...
Signed-off-by: Jonathan Chen <chenleejonathan@gmail.com >
Signed-off-by: Jonathan <chenleejonathan@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 19:47:12 -07:00
ed938ad7db
[CPUOffloading] Guard CPU eviction check ( #45757 )
...
Signed-off-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com >
2026-06-18 05:34:59 +03:00
Reid and GitHub
731fb3323d
[Rust Frontend] Validate tokenized bad_words vocabulary range ( #45876 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-18 02:28:45 +00:00
8dd8b6ed78
[XPU] Fix FP8 block-scaled scheme selection on non-CUDA platforms ( #43958 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-18 10:16:20 +08:00
e1a5fc406b
[Rust Frontend][Perf] O(n) argument scan in tool parser ( #45826 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-18 01:42:35 +00:00
Ace Eldeib and GitHub
b4092176b9
[Bugfix] Complete one-shot fused all-reduce PDL at end to avoid NaN ( #45448 )
2026-06-18 00:54:39 +00:00
Jee Jee Li and GitHub
ebbb2d55ac
[CI/Build][Bugfix] Fix SD LoRA ( #45941 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-18 00:34:15 +00:00
liuzhenwei and GitHub
2959a9273a
[XPU][CI] add model runner v2 into CI ( #44650 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-18 00:28:34 +00:00
1797576237
Revert "[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture" ( #45309 ) ( #45972 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-17 17:20:59 -07:00
0d339cf135
[Bugfix] Fix NixlConnector handshake block_len validation for GQA-replicated KV heads ( #45879 )
...
Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com >
Co-authored-by: waynehacking8 <waynehacking8@gmail.com >
2026-06-17 15:11:29 -07:00
5fd21eb0b2
[BUG] fix hidden states nan for hybrid attention models ( #45849 )
...
Signed-off-by: shanjiaz <hezhao@redhat.com >
Co-authored-by: shanjiaz <hezhao@redhat.com >
2026-06-17 18:02:24 -04:00
Ting SUN and GitHub
9d4b87f4f0
[Bugfix][Model] Validate DefaultModelLoader / LoadConfig and fail with clear errors ( #45196 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-17 21:46:33 +00:00
58b2e89642
[Bugfix][Gemma4] Render reasoning on assistant turns without tool_calls ( #45867 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-17 20:44:15 +00:00
Wentao Ye and GitHub
2659f60a1a
[Refactor] Remove dead quantization code and tests ( #45454 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 16:12:01 -04:00
091386a99b
[Bugfix] MiniMax-M3 (AMD): add packed_modules_mapping and pass swiglu… ( #45794 )
...
Signed-off-by: wangjiaxin99 <jiaxwang@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
2026-06-17 19:15:46 +00:00
qli88 and GitHub
d112eb1ac7
[feature] MiniMax-M3-MXFP4 support added ( #45896 )
...
Signed-off-by: Qiang Li <qiang.li2@amd.com >
2026-06-17 18:50:48 +00:00
Wentao Ye and GitHub
2a47a9ff0f
[DSV4 Perf] Optimize dsv4 cudagraph by reducing eager_break_during_capture, 26.8% ~ 27.9% E2E TTFT improvement ( #45309 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 09:34:53 -07:00
Wentao Ye and GitHub
9c7c74bf10
[Log] Update deepgemm log ( #45857 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 15:34:22 +00:00
danisereb and GitHub
5e27b2baf4
[Bugfix] Pass TP group to FlashInfer all-reduce fusion ( #45917 )
...
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com >
2026-06-17 15:24:26 +00:00
zhanqiuhu and GitHub
eb0fdeb1e8
[Bugfix][PD] Fix DSV4 disaggregated serving ( #45831 )
...
Signed-off-by: ZhanqiuHu <zhu@redhat.com >
2026-06-17 15:17:14 +00:00
46f74e144b
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_dynamic_per_token_quant ( #34432 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-17 23:03:54 +08:00
Wentao Ye and GitHub
0a7bacdcac
[DSv4 Perf] DSv4 flashinfer sparse index cache for metadata, 2%~4% TTFT improvement ( #45863 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-17 10:55:48 -04:00
amirkl94 and GitHub
8b2b566ea7
Feature: Enable Flashinfer non-gated MoE bf16 ( #43853 )
...
Signed-off-by: Amir Klein <203507526+amirkl94@users.noreply.github.com >
2026-06-17 14:32:49 +00:00
xaguilar-amd and GitHub
0b131b16c9
[ROCm][AITER][Quark] Tag per-channel FP8 weights as PER_CHANNEL so AITER pre-shuffled GEMM is selected ( #44626 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
2026-06-17 14:05:34 +00:00
bcb518ad7a
[quant][autoround]Refactor INC quantization into package with INCScheme orchestrator ( #40601 )
...
Signed-off-by: yiliu30 <yi4.liu@intel.com >
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com >
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com >
Co-authored-by: n1ck-guo <heng.guo@intel.com >
Co-authored-by: Zhenzhong1 <zhenzhong.xu@intel.com >
2026-06-17 21:51:32 +08:00
Chaojun Zhang and GitHub
06e1e0885c
[XPU] Fix test_logprobs_e2e import error: pin lm-eval[api]>=0.4.12 ( #44469 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-17 12:26:47 +00:00
Isotr0py and GitHub
1a59078c87
[CI/Build] Avoid duplicate ViT CG test introduced by accident ( #45654 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-17 12:23:44 +00:00
Oğuzhan KIR and GitHub
fa85ead2f3
[MM][Perf][CG] Support ViT full CUDA graph for Kimi-VL ( #41992 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
2026-06-17 12:14:01 +00:00
e28e8c8782
[ROCm][Quant] Minimax-M3: Enable fp8_per_channel for bf16 weights on mi300x ( #45854 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-17 12:02:40 +00:00
Angelo Ruocco and GitHub
ee0fd6984a
docs, kv_offloading: add docs for selective offload ( #45279 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
2026-06-17 14:58:00 +03:00
vllmellm and GitHub
d537122398
[ROCm][Bugfix]: Fallback GFX942 sparse MLA ops to Triton ( #45782 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-17 11:41:29 +00:00
Juan Pérez de Algaba and GitHub
3d20275bb4
fix(security): enforce audio decode duration limit in chat completions path ( #45908 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-17 11:07:13 +00:00
f694d43b33
[Bugfix][test] Use Salesforce/wikitext for ppl tests ( #45913 )
...
Co-authored-by: wentian-byte <192079369+wentian-byte@users.noreply.github.com >
2026-06-17 10:37:19 +00:00
Nikhilesh Chhetri and GitHub
3c6084bb0d
[Bugfix][Gemma4] Pre-initialise streaming reasoning state when prompt ends inside an open <|channel> ( fixes #45834 ) ( #45852 )
...
Signed-off-by: nikhilesh-csa <nchhetri@csa1.com >
2026-06-17 06:16:02 -04:00
Joel Smith and GitHub
68ff30d40e
[Bugfix] Fixes MiniCPM-O resampler device placement to avoid tensor device mismatch ( #42332 )
...
Signed-off-by: j9smith <j.smith9103@outlook.com >
2026-06-17 08:35:27 +00:00
6d8fff5698
[KV Connector][Offloading] Avoid blocking the engine to flush offloads on idle ( #45595 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Signed-off-by: Itay Etelis <Itay.etelis@gmail.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: Itay Etelis <Itay.etelis@gmail.com >
2026-06-17 11:35:07 +03:00
e2c58570ea
[Rust Frontend] Support hybrid/external DP LB in Python supervised bootstrap ( #45805 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-17 07:32:40 +00:00
Taneem Ibrahim and GitHub
43fa24e832
[Misc] Validate Cohere Embed Mixed Content Payloads ( #45873 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-17 06:57:12 +00:00
arghyadeep sarkar and GitHub
93bbe94d3a
[Kernel] Add weightless RMSNorm CUDA kernels for has_weight=False ( #41430 ) ( #44109 )
...
Signed-off-by: hello-args <args.sarkar@gmail.com >
2026-06-16 23:45:55 -07:00
Will Eaton and GitHub
17bc144556
[Rust Frontend] Add serde defaults for omit_defaults fields in EngineCoreSamplingParams ( #45848 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-17 06:40:49 +00:00
Sahil Singh and GitHub
295232a26a
[Rust Frontend] Add /abort_requests endpoint ( #44382 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-17 06:40:47 +00:00
Reid and GitHub
56e4345226
[Rust Frontend] Support prompt-only completions ( #44938 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-17 06:38:06 +00:00
Nick Hill and GitHub
e9993a52aa
[BugFix][CI] Fix scheduler plugin test ( #45897 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 06:30:49 +00:00
a46abb7ae6
[Bugfix][Quantization] Reject unsupported compressed tensors KV cache schemes ( #45312 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 05:08:41 +00:00
4c62663315
[M3] Enable FP8 sparse GQA ( #45744 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-16 21:38:03 -07:00
d78650cf97
[CI][NIXL] Pin NIXL to 1.2.0 ( #45843 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
Signed-off-by: Itay Alroy <75032521+itayalroy@users.noreply.github.com >
Co-authored-by: ovidiusm <ovidium@nvidia.com >
2026-06-16 21:29:34 -07:00
5bdc01bcc3
[M3] Tune Triton indexer score decode for spec-decode ( #45743 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 21:07:34 -07:00
liangel-02 and GitHub
20a5f8b43b
[FlexAttention] make custom mask mods fully cudagraphable ( #45232 )
...
Signed-off-by: Angel Li <liangel@meta.com >
2026-06-17 11:53:12 +08:00
7b5d60cc37
[Bugfix][V1] Clean up compiled-model bytecode hooks on VllmRunner exit ( #45195 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 20:31:17 -07:00
Nick Hill and GitHub
14b438a98b
[ModelRunnerV2] Various model/config compatibility fixes ( #45868 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-17 03:23:01 +00:00
2785a5e0e6
[Bugfix][ROCm] Fix FP8 per-tensor scale rank mismatch causing Inductor assertion failure ( #44912 )
...
Signed-off-by: nehmathe2 <nehmathe2@gmail.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Signed-off-by: nehmathe <nehmathe@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:17:42 -07:00
efd15e192a
[Bugfix][ROCm] Fix MiniMax-M3 FP8 KV cache dtype ( #45720 )
...
Signed-off-by: Cam Quilici <cjquilici@gmail.com >
Signed-off-by: Cameron Quilici <cjquilici@gmail.com >
Co-authored-by: Hongxia Yang <62075498+hongxiayang@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-17 03:14:45 +00:00
556b063e45
[XPU] Fix test_spec_decode_logprobs: use FLASH_ATTN for XPU in GPU_DETERMINISM_KWARGS ( #44468 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 11:07:04 +08:00
aa0ac8a661
[CI] Run pre-commit on self-hosted vllm-runners ( #45865 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 19:49:22 -07:00
Federico and GitHub
b831374cf1
[Bugfix][Gemma4] Fix parsing when thinking is disabled ( #45832 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-17 02:41:36 +00:00
71bc19dbdd
[Bugfix] Fix MoE model load OOM in FlashInfer_TRTLLM backend with sleep mode ( #45589 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 19:36:51 -07:00
Kunshang Ji and GitHub
ef2c40dc00
[XPU][CI] fix server test file path ( #45870 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-17 09:06:25 +08:00
4bf699d310
[Kernel] Support DS Mamba tail copy for MTP align mode ( #45473 )
...
Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com >
Co-authored-by: Thomas Parnell <tom.parnell@gmail.com >
2026-06-16 22:50:30 +00:00
Stan Wozniak and GitHub
520828789c
Apply LRU policy only to proper cache entries ( #42656 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-16 21:49:15 +00:00
9d4dc4ca2f
[Kernel] Support GLM-5 dimensions for TRT-LLM ragged MLA prefill ( #43525 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 20:49:47 +00:00
Federico and GitHub
b9684d99e9
[Bugfix] Gemma4: skip forced JSON for required/named tool choice ( #45795 )
...
Signed-off-by: Federico Iezzi <fiezzi@google.com >
2026-06-16 20:38:29 +00:00
Divakar Verma and GitHub
4fadf9c92c
[ROCm][CI] fix multimodel run cmds ( #45858 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-16 15:31:52 -05:00
Nick Hill and GitHub
d8d95998dc
[Core] Add prefill step cadence for better non-PD DP balancing ( #44558 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 13:17:18 -07:00
Flora Feng and GitHub
475a6ad18a
[Misc] Update Mergify tool-calling label ( #45853 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-16 19:08:00 +00:00
Hongxia Yang and GitHub
f2beaa80c8
[ROCm][Quant] mxfp8 moe/linear gfx950 tuning for MiniMax-M3 ( #45725 )
...
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com >
2026-06-16 18:50:40 +00:00
8e27a9c215
[PERF] Fuse multi-group block table staged writes ( #44944 )
...
Signed-off-by: jesse <szxfml@gmail.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-16 10:53:27 -07:00
7d567172fc
[Bugfix] Fix Qwen3 prompt tool-call reasoning false positive ( #45763 )
...
Signed-off-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
Co-authored-by: Alex Bilichenko <alexbi29@users.noreply.github.com >
2026-06-16 17:48:01 +00:00
Chauncey and GitHub
f00e163f35
[Frontend] Add Streaming Parser Engine and new MinimaxM2 Parser ( #45701 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-16 13:38:17 -04:00
44b2512767
[KV Connector][Mooncake] Add cache_prefix to namespace store keys ( #45767 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 10:24:20 -07:00
188c68798e
[KVConnector][MoRIIO] Allow overriding the advertised host IP ( #45488 )
...
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-16 17:18:37 +00:00
c45f681932
[Bugfix][Core] Fall back when numactl --membind is blocked in constrained containers ( #45438 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-16 09:49:56 -07:00
89e8645a9e
[Model] Remove Dots1ForCausalLM ( #45637 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-17 00:32:18 +08:00
Wentao Ye and GitHub
88a9cdd439
[Model Runner V2] Enable GraniteMOE for MRv2 by default ( #45461 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-16 09:31:32 -07:00
Micah Williamson and GitHub
6f612fbedf
[ROCm][CI] Patch conftest to resolve occasional OOMs ( #45722 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-16 10:00:15 -05:00
Sting Lin and GitHub
506ec6d656
Upgrade tpu-inference to v0.22.1 ( #45793 )
2026-06-16 07:54:57 -07:00
a52205bccf
[Model] Add HrmTextForCausalLM (Hierarchical Reasoning Model — Text) ( #43098 )
...
Signed-off-by: Wuyifei <wuyifei@me.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-16 22:41:41 +08:00
3d34f8cbdc
[ROCm][Cleanup] Remove stale AITER FA hybrid KV-cache TODO ( #44178 )
...
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-16 07:28:06 -07:00
Carl Y and GitHub
eb04c769d3
feat: MLA prefill enable FA4 fp8 output ( #43050 )
...
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com >
2026-06-16 07:10:59 -07:00
ce3ef17bec
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_per_block_quant ( #36895 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-16 22:09:52 +08:00
bf5149b516
[Bugfix] Fix FlashMLA sparse accuracy with topk_length and zero-init padding ( #36616 )
...
Signed-off-by: AjAnubolu <anuboluajay@gmail.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-16 07:09:00 -07:00
Tahsin Tunan and GitHub
cca3365b73
[Rust Frontend] Add CORS support ( #45753 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
2026-06-16 13:47:11 +00:00
040df8f2ea
[CI] Fix attention benchmark smoke test ( #45728 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 13:43:37 +00:00
ced32bb474
[Perf] Add VLLM_TRITON_FORCE_FIRST_CONFIG to skip Triton autotuning ( #42425 )
...
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-16 15:16:45 +02:00
c5e5c33fcd
[Bugfix][MoE] Restore routed output unpadding before shared expert add ( #45707 )
...
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 16:06:28 +03:00
Mike G and GitHub
a8c86eeb16
[Quant] Support modelopt_mixed on Ampere (SM80/SM86) ( #45306 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
2026-06-16 08:43:44 -04:00
Andreas Karatzas and GitHub
7e179e4bc0
[ROCm][CI] Gate incompatible HF references on Transformers v5 ( #41532 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-16 20:34:11 +08:00
405c7cf283
[ZenCPU] Add zencpu Platform Runtime Logging and Docs ( #42726 )
...
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-16 08:23:12 -04:00
3f53e2138f
[Refactor] Remove Fp8OnlineLinearMethod as scheduled ( #45463 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-16 04:35:58 -07:00
Hank Han and GitHub
d53f4593ce
[KV Connector][Mooncake] Pipeline-parallel support for PD-disaggregated serving with Mooncake connector ( #44528 )
...
Signed-off-by: hanhan.hank <hanhan.hank@bytedance.com >
Signed-off-by: Hank Han <hanhan7630@outlook.com >
2026-06-16 04:35:38 -07:00
ad32608e24
[MM][Perf][CG] Support dual-path ViT full CUDA graph for DeepSeek-OCR ( #43586 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 04:35:20 -07:00
Thien Tran and GitHub
b2cfae777d
Add Triton recompile detection ( #45631 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-06-16 18:25:28 +08:00
wangxiyuan and GitHub
3f1ff1ff14
[Misc]Clean up useless test ( #45792 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
2026-06-16 09:53:08 +00:00
c69c73418a
[XPU][CI] add intel xpu cases for nightly CI ( #44372 )
...
Signed-off-by: wenjun.liu <wenjun.liu@intel.com >
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: zengxian <xiangdong.zeng@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-16 16:35:08 +08:00
Thomas Parnell and GitHub
ebf3a6d705
[Bugfix] Fix trtllm fused allreduce+rms_norm for transformers backend ( #45307 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-16 08:34:27 +00:00
wang.yuqi and GitHub
c4fd9794e9
[Frontend] Remove AsyncMicrobatchTokenizer. ( #45759 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-16 08:02:11 +00:00
7ad894c86a
[Bugfix] Prevent cuMemcpyBatchAsync segfault with MTP and KV offloading ( #44784 )
...
Signed-off-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: joshua <joshua.abraham@multicorewareinc.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-16 07:58:39 +00:00
Li, Jiang and GitHub
a7fdfeef72
[CPU] Support Gemma Diffusion ( #45690 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-16 14:39:56 +08:00
Jimmy Lee and GitHub
8bf374955f
[Bug Fix] Allow pinned memory for WSL2 ( #41496 )
...
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com >
2026-06-16 05:56:26 +00:00
Cyrus Leung and GitHub
9096659edb
[Cleanup] Remove dead env ( #45777 )
...
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk >
2026-06-15 22:56:23 -07:00
Taneem Ibrahim and GitHub
81d8f4ebac
[Misc] Added validation for Cohere /v2/embed input field exclusivity ( #45640 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-16 05:42:43 +00:00
a9a8a32dcd
Register parsed config classes before tokenizer init ( #40299 )
...
Signed-off-by: Bortlesboat <bortstheboat@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-16 05:33:08 +00:00
9d808e2309
[Core] Use fastsafetensors ParallelLoader for weight loading ( #40183 )
...
Signed-off-by: Git Bisector <gitbisector@gmail.com >
Signed-off-by: gitbisector <gitbisector@gmail.com >
Signed-off-by: git bisector <gitbisector@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-15 22:32:05 -07:00
Ben Browning and GitHub
f3858d5422
[Frontend] [Parser] Migrate Nemotron V3 to streaming parser engine ( #45755 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-16 05:31:21 +00:00
Bugen Zhao and GitHub
259ff891be
[Rust Frontend] Require ModelConfig.vocab_size to be present ( #45696 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 05:30:25 +00:00
6607a80dab
[Bugfix][Gemma4] Fix offline parser truncation, adjust_request token leak, and chat template sync ( #45553 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-16 04:31:53 +00:00
liuzhenwei and GitHub
b8bd773fe4
[XPU] Fix Triton attn fp8/bf16 check failing ( #45758 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
2026-06-16 12:31:20 +08:00
Ruinan Ma and GitHub
2addbb9cc9
[BugFix] Support async scheduling with prompt embeds for multimodal models ( #45673 )
...
Signed-off-by: Ruinan Ma <r7ma3088@gmail.com >
2026-06-16 04:12:54 +00:00
Isotr0py and GitHub
e3cfea2e1b
[Multimodal] Add Qwen3-VL video loader ( #44412 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-16 03:45:34 +00:00
Bugen Zhao and GitHub
f99260d2aa
[Rust Frontend] Lower out-of-vocab validation to text layer ( #45685 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 03:37:58 +00:00
Bugen Zhao and GitHub
3f65e21e32
[Rust Frontend] Support max_logprobs validation ( #45674 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-16 10:57:56 +08:00
xx-thomas and GitHub
b00e76ff72
[Misc][Model] add io processor for query/document embeddings from ColBERT (jinaai/jina-colbert-v2) ( #45210 )
...
Signed-off-by: thomas <thomas.varghese@columbia.edu >
2026-06-16 01:32:32 +00:00
Woosuk Kwon and GitHub
f4359a70f9
[DSV4][Minor] Fix supported KV cache dtypes ( #44892 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-16 00:14:51 +00:00
Itay Alroy and GitHub
3afe659b6b
[EP] Enable DBO with NIXL EP ( #45275 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 23:37:22 +00:00
Itay Alroy and GitHub
16e91176cf
[EP] Query NIXL EP top-k index dtype ( #45298 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:50:18 +00:00
Itay Alroy and GitHub
ab8b0fe338
nixl_ep: Skip post-receive quantization for NVFP4 ( #45606 )
...
Signed-off-by: Itay Alroy <ialroy@nvidia.com >
2026-06-15 22:42:05 +00:00
d467a2a7f2
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer ( #45357 )
...
Signed-off-by: llx-08 <2596671364@qq.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-15 21:36:09 +00:00
76a373eff4
[Frontend] Replace legacy Gemma4 parsers with engine-based implementation ( #45588 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 21:34:07 +00:00
Zang Peiyu and GitHub
25ee659db0
Fix parallel_tool_calls: null treated as false instead of default true ( #44955 )
...
Signed-off-by: factnn <166481866+factnn@users.noreply.github.com >
2026-06-15 21:14:10 +00:00
eacff17c8d
[Model Runner V2][Bugfix] Fix MRV2 LoRA warmup ( #35536 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-15 13:17:23 -07:00
Flora Feng and GitHub
cd9078fe59
[Frontend] Skip structural tags for auto tool_choice without strict mode ( #45600 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 19:55:31 +00:00
Wentao Ye and GitHub
e18fe932ca
[Perf] Optimize DSv4 prefill chunk planning, 4.0% E2E Throughput Improvement ( #45061 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-15 19:50:21 +00:00
51ec5cf08f
[Bugfix] Chat Completions Harmony Refactor Clean up ( #45464 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-15 14:45:19 -04:00
7e612a0f06
[KV Offloading] Implement reset_cache for TieringOffloadingManager ( #44541 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 18:42:53 +00:00
+1
0a1c5034f5
[Model] Add MiniMax M3 support ( #45381 )
...
Signed-off-by: youkaichao <youkaichao@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-06-16 01:01:25 +08:00
RoyWang and GitHub
a3195fab7b
[AMD][Bugfix][Quantization] Honor fused-name match in is_layer_skipped ( #43981 )
2026-06-15 09:37:52 -07:00
Flora Feng and GitHub
0d80979644
[Chore] Consolidate reasoning/tool parser attributes into unified Parser in chat serving ( #45548 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 11:16:45 -04:00
Saddss and GitHub
588db18362
[Bugfix] Two-phase KV allocation for cross-group prefix cache hits (supersedes #33775 ) ( #44409 )
...
Signed-off-by: Saddss <2872669061@qq.com >
2026-06-15 22:39:59 +08:00
fa63bb9db6
Remove redundant Triton KV cache dtype asserts and enforce architectural support (fp8 >= sm89) ( #43914 )
...
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-15 06:49:57 -07:00
5ed15f42b9
Fix the E8M0 scale computation in the MXFP4 (W4A4) MOE CUTLASS kernel ( #43557 )
...
Signed-off-by: Xin He <xin3.he@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-15 06:04:54 -07:00
Juan Pérez de Algaba and GitHub
b997071ec4
(security) Enforce audio upload size limit before full file materialization ( #45510 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-15 10:25:24 +00:00
Martin Kukla and GitHub
6c5872efc5
[Bugfix] Unset HF's default max_new_tokens for DiffusionGemma ( #45417 )
...
Signed-off-by: Martin Kukla <martin.kukla@cantab.net >
2026-06-15 17:31:57 +08:00
wang.yuqi and GitHub
1d88c4dadd
[Docs] Update the online serving docs. ( #45676 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-15 17:23:36 +08:00
vllmellm and GitHub
25c53d1293
[ROCm][Doc] Add installation notes about python version requirement ( #45671 )
...
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-06-15 17:22:55 +08:00
Yejing Lai and GitHub
9872921c5f
[XPU] skip UT test_with_ngram_gpu_spec_decoding ( #44423 )
...
Signed-off-by: Lai, Yejing <yejing.lai@intel.com >
2026-06-15 08:46:30 +00:00
Reid and GitHub
c17e2f7c84
[Bugfix][Rust Frontend] Make metrics respect --served-model-name ( #45465 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-15 08:05:10 +00:00
FAUST and GitHub
40eac9a9d9
[Rust Frontend] Support parallel_tool_calls = false ( #44760 )
...
Signed-off-by: zhoujinyu <2319109590@qq.com >
2026-06-15 07:50:48 +00:00
b5adb027ad
[Models] Fix MiMo v2.x QKV TP sharding + FP4 support ( #45200 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-15 15:13:34 +08:00
Sahil Singh and GitHub
64833f8158
[Rust Frontend] Add external→internal request-id map for abort() ( #45137 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
2026-06-15 06:51:24 +00:00
ddad5dbda2
[Bugfix][Rust] Sync EngineCoreReadyResponse with the Python dataclass ( #45557 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-15 06:49:42 +00:00
Peter Pan and GitHub
ebb0a71ad0
[Bugfix] Reject out-of-range temperature values in SamplingParams ( #44965 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
2026-06-14 23:12:44 -07:00
Ting SUN and GitHub
48df95c43e
[Feature][Frontend] Report multimodal token counts in usage.prompt_tokens_details ( #45458 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-15 05:20:58 +00:00
7df4fe1bd7
[Model] Remove XverseForCausalLM ( #45638 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-14 22:09:00 -07:00
b8336c3c7c
[Bugfix][V1] Split V2 model-runner attention groups on num_heads_q ( #45564 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-14 21:49:46 -07:00
e8d3e22c88
Fix included router missing path for FastAPI >=0.137 ( #45629 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com >
2026-06-15 04:28:52 +00:00
c4a3f9d137
[Frontend] Add Streaming Parser Engine and new Qwen3 Parser ( #45413 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-15 11:59:05 +08:00
Flora Feng and GitHub
e3e3cd5458
[Bugfix][CI] Update Dockerfile dependency graph PNG ( #45602 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-15 10:35:24 +08:00
Li, Jiang and GitHub
8760f972ca
[CPU] Refine CPU attention frontend ( #45391 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-14 19:26:54 -07:00
b675cb7d0f
[Bugfix][CPU] Honor cgroup memory limit when computing KV cache size ( #45086 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-14 19:26:50 -07:00
Chaojun Zhang and GitHub
2725c84aae
[XPU] Enable sequence parallel support for XPU ( #38608 )
...
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-06-14 19:26:46 -07:00
Noa Neria and GitHub
1801fad0ba
[Bugfix] Stream Llama4 weight loading to avoid host-OOM with copy-returning loaders ( #44645 )
...
Signed-off-by: Noa Neria <nneria@nvidia.com >
2026-06-14 19:23:44 -07:00
Ting SUN and GitHub
3d6ce816f0
[Bugfix][Model] Validate runai_streamer model_loader_extra_config ( #45291 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-14 19:23:30 -07:00
Taneem Ibrahim and GitHub
2c764c089a
Added real /v1/embeddings support for messages + chat_template_kw ( #45173 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-15 09:08:10 +08:00
Michael Ma and GitHub
c621af1690
[BugFix] Fix prompt_embeds for multimodal models ( #45383 )
...
Signed-off-by: ruinan ma <r7ma3088@gmail.com >
2026-06-14 01:44:56 -07:00
Roger Wang and GitHub
e2bf2b3d84
[Perf] Use bisect for mm feature lookup in model runner v2 ( #45566 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
2026-06-14 00:22:53 -07:00
Amanzhol Salykov and GitHub
725c3bc808
[ROCm][Perf] Enable W4A16 FlyDSL MoE ( #44400 )
...
Signed-off-by: amd-asalykov <asalykov@amd.com >
Signed-off-by: Amanzhol Salykov <asalykov@amd.com >
2026-06-14 00:14:39 -07:00
9548a1887f
[XPU] Support int4 group_size=32 W4A16 MoE ( #45136 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-14 00:14:35 -07:00
Jeff (Junze) Ma and GitHub
9fd737badc
[Bugfix][DCP] Fix illegal memory access in DCP a2a decode under full CUDA graphs ( #45487 )
2026-06-14 00:14:31 -07:00
4ef4492e9b
[V1][Spec Decode] Add Dynamic SD ( #32374 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com >
2026-06-14 00:14:27 -07:00
78e7293bb1
[Build] Fix CUDA arch build coverage gaps ( #45277 )
...
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Xin Li <xinli-sw@users.noreply.github.com >
Co-authored-by: ShawRong <ShawRong@users.noreply.github.com >
Co-authored-by: Change72 <Change72@users.noreply.github.com >
2026-06-13 22:09:20 -07:00
54bbf51668
[Bugfix] nightly Docker images crash with ImportError: AnthropicOutputConfig since May 28 ( #44795 )
...
Signed-off-by: achyuthan.s <113010327+Achyuthan-S@users.noreply.github.com >
Signed-off-by: Achyuthan S <achyuthan.sivasankar@gmail.com >
Signed-off-by: Achyuthan Sivasankar <achyuthan.sivasankar@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-13 21:45:29 -07:00
Nick Hill and GitHub
cf027b86af
[Core] Simplify MRV2 async output handling ( #45442 )
2026-06-13 18:15:36 -07:00
71b961dd35
[Perf] SM90 cutlass fp8 mm supports odd M by swap_ab, 180~290% kernel performance improvement ( #44572 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 12:05:45 -07:00
521b88c29e
[Bugfix] Reject structured outputs for diffusion decoders with a clear error ( #45468 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-13 12:04:01 -07:00
Harry Mellor and GitHub
b3f0a0a0df
Fix docs build on main ( #45536 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 08:53:23 -07:00
Juan Pérez de Algaba and GitHub
470229c37e
[Security] Fix DoS via prompt_embeds on M-RoPE models ( #45252 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-13 10:17:38 +00:00
2b3006076c
[Security] Add timeout guard for regex compilation in structured outp… ( #45118 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-13 09:52:56 +00:00
Wentao Ye and GitHub
96fa5cdd9e
[CI Bug] Fix ValueError: There is no module or parameter named 'model.vision_tower.vision_model' ( #45478 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 02:38:37 -07:00
Andreas Karatzas and GitHub
9261dbbc55
Treat null completion max_tokens like the default ( #45491 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 09:34:09 +00:00
Wentao Ye and GitHub
2ecf7d0eb4
[Model Runner V2] Fix openai.InternalServerError: Error code: 500 - 'list index out of range' ( #45467 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-13 01:44:16 -07:00
midas and GitHub
0d29612292
[Doc] Fix uv dependency resolution failure for setuptools during CPU source builds (x86 & ARM) ( #45412 )
...
Signed-off-by: midas <the.anon.github@gmail.com >
2026-06-13 06:18:58 +00:00
WEI CHENG CHIU and GitHub
5b2943f5a6
[Bugfix] Return the tokenizer from maybe_make_thread_pool so it survives pickling ( #45460 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 06:01:35 +00:00
43f0e024bc
[Render] Add /derender endpoints for disaggregated postprocessing ( #43606 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-13 13:55:33 +08:00
Andreas Karatzas and GitHub
1033ffac2e
[CI] Wait for SSL cert refresher events in the test ( #45489 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-13 04:57:18 +00:00
ff5a30cfac
[Bugfix] Replace deprecated Qwen2VLImageProcessorFast with Qwen2VLImageProcessor ( #42700 )
...
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-12 21:04:31 -07:00
WEI CHENG CHIU and GitHub
17ee5b1ac5
[Bugfix] Set type/role explicitly in streaming message_start event ( #45376 )
...
Signed-off-by: Wayne Chiu <waynehacking8@gmail.com >
2026-06-13 01:40:50 +00:00
Nick Hill and GitHub
1a369783e9
[BugFix] Avoid prematurely freeing cached mm encoder outputs ( #45347 )
...
Signed-off-by: Roger Wang <hey@rogerw.io >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 15:39:40 -07:00
Kevin H. Luu and GitHub
e3e31e54b0
[Bugfix][CPU] Don't build triton-cpu on arm64 release image ( #45401 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-06-12 14:51:45 -07:00
badddd254f
[ROCm][DSV4][Perf] Fuse inverse-RoPE and cache bf16 wo_a in o-projection ( #45103 )
...
Signed-off-by: Fangzhou Ai <fangzhouai@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 15:57:09 -05:00
c90650088d
Add the QuantizedActivation linear-kernel contract ( #44260 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-12 13:48:15 -07:00
Michael Goin and GitHub
9eaacb23ec
[Kernel] Consolidate Marlin thread-tile padding across all dense Marlin paths ( #45295 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-12 13:46:21 -07:00
78739c1946
[Model Runner v2] Migration from v1 to v2, with Qwen and DSv2 MOE models [3/N] ( #42667 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 20:44:52 +00:00
Matthew Bonanni and GitHub
cf567cbc71
[Attention] Improve attention benchmarks: configs and profiling ( #39336 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-12 16:24:25 -04:00
Micah Williamson and GitHub
39cb9bf292
[ROCm] Bump Torch to 2.11 ( #45362 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-12 15:22:26 -05:00
Flora Feng and GitHub
6e4a547176
[Refactor] Deprecate ResponsesParser wrapper, inline parsing into ParsableContext ( #45431 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-12 16:15:41 -04:00
aab639c705
[Core][AMD] Propagate shutdown timeout to MultiprocExecutor ( #43154 )
...
Signed-off-by: Ryan Rock <ryan.rock@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-12 15:13:31 -05:00
efe7adb5e1
[Perf] Use native DSA indexer decode path for next_n > 2 on SM100 ( #45322 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-12 12:54:00 -07:00
Isotr0py and GitHub
6635279d8a
[Migration] Migrate GGUF quantization support to plugin ( #39612 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 12:02:21 -07:00
Jonas I. Liechti and GitHub
d6fd7ce8da
[Model][Dflash] Enable Dflash support for Qwen3NextForCausalLM targets ( #45319 )
...
Signed-off-by: Jonas I. Liechti <j-i-l@t4d.ch >
2026-06-12 10:30:09 -07:00
272c16953e
[Kernel][Helion][1/N] Add Helion kernel for dynamic_per_token_scaled_fp8_quant ( #33790 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
2026-06-12 12:50:06 -04:00
Yi Zhong and GitHub
053e7daa79
[Model] Add encoder CUDA graph support to Lfm2VL ( #44930 )
...
Signed-off-by: vincentzed <207368749+vincentzed@users.noreply.github.com >
2026-06-12 09:17:26 -07:00
5af4aec141
[Rust Frontend] Add standalone granite4 tool parser ( #45216 )
...
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-13 00:16:36 +08:00
Sai Sridhar Tarra and GitHub
a30addc754
[Docs][KV Connector][NIXL] document KV Transfer stat logging and Prometheus metrics ( #44055 )
...
Signed-off-by: Sai Sridhar <tarrasridhar1154@gmail.com >
2026-06-12 15:39:11 +00:00
Chauncey and GitHub
3b8fc3fe6d
[Frontend] Support strict mode for tool calling with ResponsesAPI ( #45396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-06-12 10:59:59 -04:00
9ff278b1d2
[Core][KV Connector] fix scheduler KV connector stats aggregation ( #43877 )
...
Fixes scheduler-side KV connector stats collection so that:
1. update_connector_output() runs before scheduler-side stats are collected.
2. worker-side and scheduler-side KV connector stats are aggregated when both are present.
3. scheduler-only KV connector stats are still emitted when no worker-side stats exist.
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
2026-06-12 14:51:55 +00:00
Guan-Ming (Wesley) Chiu and GitHub
c7aa3d2630
[Core] Support structured outputs for beam search ( #35022 )
...
Signed-off-by: Guan-Ming (Wesley) Chiu <guanmingchiu@gmail.com >
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com >
2026-06-12 06:56:25 -07:00
fbc3a1907a
[Bug] Migrate Reset cache for both v2 and v1 model runner ( #42759 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 09:38:12 -04:00
4171ae406c
[V1][Metrics] Add MLA attention metrics for DeepSeek MFU estimation ( #39457 )
...
Signed-off-by: Thillai Chithambaram <thillaichithambaram.a@gmail.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-06-12 14:28:40 +01:00
Ethan Feng and GitHub
b7f9b6ab27
[Metrics] Add group-aware KV cache capacity to vllm:cache_config_info ( #42206 )
...
The startup log already reports the correct group-aware KV cache capacity for
hybrid models, but Prometheus did not expose matching info in 'vllm:cache_config_info`.
This PR adds kv_cache_size_tokens and kv_cache_max_concurrency.
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-06-12 11:49:44 +00:00
8af550b399
[BUGFIX][XPU] Update fa interface for compatibility ( #45394 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-12 11:45:01 +00:00
f1e13f7df9
[Model] Remove Mono-InternVL (InternLM2VEForCausalLM) ( #45129 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 10:41:09 +00:00
88ed636218
[KV Connector]: Support KV push from Prefill to Decode node using Nixl KV Connector ( #35264 )
...
Signed-off-by: Sunita Nadampalli <nadampal@amazon.com >
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-06-12 10:38:41 +00:00
a014dddbaa
[11b/n] Migrate Machete kernels to torch stable ABI ( #45304 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-12 10:36:49 +00:00
Thomas Parnell and GitHub
a37b4a940e
[Doc] AGENTS.md: add section about coding style ( #45301 )
...
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com >
2026-06-12 06:23:04 -04:00
Juan Pérez de Algaba and GitHub
f715f25f29
Fix misleading error for audio duration limit rejection ( #45113 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-12 09:58:08 +00:00
Fynn Schmitt-Ulms and GitHub
462ef83d58
Update hidden states extraction integration test triggers ( #45294 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
2026-06-12 01:05:19 -07:00
1ae1051b4b
[Bugfix][Rust Frontend] Return 400 for prompt-validation submit errors ( #45286 )
...
Signed-off-by: xiaguan <751080330@qq.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 07:53:11 +00:00
2043258dec
[Frontend] Support strict mode for tool calling ( #45003 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: cjackal <44624812+cjackal@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 07:51:48 +00:00
bd59c913bc
[CI] ci-fetch-log.sh: fetch all failed jobs from a build URL or PR number ( #45274 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude Fable 5 <noreply@anthropic.com >
2026-06-12 00:42:18 -07:00
04cec9e4d8
[XPU][DeepSeek-V4] Fix MTP: sync with upstream fixes #44821 and #43746 ( #45240 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 15:41:36 +08:00
Will Eaton and GitHub
87b98d6d6c
[Rust Frontend][Bugfix] Forward --shutdown-timeout and --disable-log-stats to the managed Python engine ( #45300 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
2026-06-12 07:39:27 +00:00
Yuwen Zhou and GitHub
0cd9b7af25
[CPU] Support CPU W4A16 INT4 MoE ( #43409 )
...
Signed-off-by: yuwenzho <yuwen.zhou@intel.com >
2026-06-12 07:12:37 +00:00
Isotr0py and GitHub
a2c72d4388
[Bugfix] Fix Dockerfile dependency graph pre-commit error ( #45374 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-12 07:10:18 +00:00
Rohan Potdar and GitHub
fe04238292
[ROCm][gpt-oss] Pass GateMode.INTERLEAVE for MXFP4 W4A16 fused MoE ( #44893 )
...
Signed-off-by: Rohan Potdar <rohan.potdar@amd.com >
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-12 01:02:04 -05:00
39dee1114a
[MM][Perf][CG] Support ViT full cudagraphs for mllama4 ( #40660 )
...
Signed-off-by: allgather <all2allops@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-11 22:17:55 -07:00
+1
eb28452b10
[Model] Add DiffusionGemma Support ( #45163 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Martin Kukla <martin.kukla@cantab.net >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Dipika Sikka <dsikka@redhat.com >
Co-authored-by: NickLucche <nlucches@redhat.com >
Co-authored-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: Alec Kohlhoff <134344302+aleckohlhoff@users.noreply.github.com >
Co-authored-by: Porras Huang <20535584+porrashuang@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: scoootscooob <167050519+scoootscooob@users.noreply.github.com >
2026-06-11 22:17:35 -07:00
Divakar Verma and GitHub
1ce3cdc5c1
[ROCm][CI] fix fp8 support for test_deepep_moe ( #45302 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-12 00:16:14 -05:00
Dao007forever and GitHub
6fbfdd1831
[NIXL] Per-region KV transfer classification for mixed full-attn + MLA groups ( #44583 )
2026-06-11 21:42:41 -07:00
Chris Leonard and GitHub
7021be66e8
[11a/n] Migrate Marlin kernels to torch stable ABI ( #45176 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
2026-06-11 21:22:37 -07:00
Ekagra Ranjan and GitHub
226ba9fc9e
[ASR] Add Long Audio benchmark and correctness test ( #44587 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-06-12 04:11:16 +00:00
b927004c44
[Bugfix] Mamba CPU Offloading ( #44599 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-11 21:07:35 -07:00
e0b9fb1290
[ASR] Optimize CPU preproc to get 2.5x RTFx via multi-threading ( #44612 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 21:05:11 -07:00
42ae5e7ac6
[Bugfix] Fix --enable-prompt-tokens-details omitting zero cached tokens ( #44383 )
...
Signed-off-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Sasindharan Sankar <sasindharansankar@email.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-11 20:37:42 -07:00
Nick Hill and GitHub
2263f8a3de
[CI][BugFix] Fix broken test_mamba_prefix_cache.py due to stale mock ( #45345 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-12 03:26:17 +00:00
Ting SUN and GitHub
c1076839c9
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads ( #45308 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-11 20:21:46 -07:00
fcf5115c45
[ROCm][DSv4][Perf] Flash-decode split-K decode attention kernel ( #44899 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-12 03:17:52 +00:00
4bc83323f2
[Bugfix] OffloadingConnector: respect skip_reading_prefix_cache flag ( #44592 )
...
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: littlecircle0730 <43994952+littlecircle0730@users.noreply.github.com >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-06-12 02:20:39 +00:00
yzong-rh and GitHub
e0871ad225
[Refactor] Chat Completions Streaming Harmony Refactor and Bugfixes ( #45104 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-12 01:09:47 +00:00
6f573f486b
[Bugfix] Initialize missing attributes in mistral eagle ( #45217 )
...
Signed-off-by: jpwang <jpwang@smail.nju.edu.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 08:21:01 +08:00
Neil Schemenauer and GitHub
9bbf42be26
Make mistral_common optional by deferring MistralToolCall import ( #45305 )
...
Signed-off-by: Neil Schemenauer <nas@arctrix.com >
2026-06-11 22:59:11 +00:00
8a91228dbe
[Bugfix][KVConnector][Mooncake] Close MooncakeDistributedStore on connector teardown ( #45206 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 14:33:48 -07:00
yzong-rh and GitHub
f712fd0d7d
[Refactor] Chat Completions Harmony Refactor, non-streaming path. ( #45171 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-11 21:18:30 +00:00
Wentao Ye and GitHub
5a6c7b7ab5
[Bug] Fix test flashmla for DSv4 ( #45052 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 16:22:26 -04:00
c9340e6f35
[Model] Remove InternLMForCausalLM registry alias ( #45128 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 20:02:51 +00:00
Ben Browning and GitHub
235b63c004
[Bugfix] Fix Anthropic tool_use content handling dropping args ( #45287 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-11 20:01:29 +00:00
3b03a2cf47
[Rust Frontend] Support continuous_usage_stats stream option ( #43965 )
...
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 17:50:59 +00:00
wentian-byte and GitHub
b8142294b7
[Bugfix] Restrict FlashInfer cuDNN FP8 ViT attention gate to Blackwell (SM 100) ( #45251 )
...
Signed-off-by: Wentian Byte <3400259131@qq.com >
2026-06-11 16:39:24 +00:00
2ec6594db9
[Kernel][Helion][1/N] Add Helion kernel for per_token_group_fp8_quant ( #36902 )
...
Signed-off-by: Sean Chen <seachen@redhat.com >
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 08:59:08 -07:00
vraiti and GitHub
79f8c5bd8c
[Metrics] Scope unregister_vllm_metrics() to strictly "vllm:" metrics ( #42331 )
...
`unregister_vllm_metrics()` currently uses "vllm" in `collector._name` to decide
which collectors to remove from the Prometheus registry, removing every even
metrics registered by other subsystems or downstream extensions like "vllm_omni:"
Signed-off-by: vraiti <vraiti@redhat.com >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
2026-06-11 15:43:14 +00:00
Jiangyun Zhu and GitHub
f81daf8880
[Attention] add triton diff-kv backend for mimo ( #41797 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-11 11:36:31 -04:00
4085ff7cb4
[Core] Add kvcache watermark to reduce preemptions ( #44594 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-11 08:27:31 -07:00
23eb7c8fbb
[Bugfix] Fix NixlEPAll2AllManager's dependency on --enable-elastic-ep to function ( #44422 )
...
Signed-off-by: fangyuchu <fangyuchu@qq.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-06-11 08:14:49 -07:00
wineandchord and GitHub
c2b4cd39ac
[Doc][Attention] Fix MLA top-of-file comments ( #37047 )
...
Signed-off-by: wineandchord <guoqizhou19@gmail.com >
2026-06-11 08:14:45 -07:00
Kai K. and GitHub
f1d8d99717
[Bugfix] CohereModel.load_weights: skip modelopt _quantizer.* keys ( #43495 )
...
Signed-off-by: Kai Köhler <kai.koehler@web.de >
2026-06-11 08:14:21 -07:00
Nicolò Lucchesi and GitHub
750aab5b8e
[Bugfix] Fix CPU memory leak related to not cleaning up old remotes data ( #44424 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-11 07:54:52 -07:00
5edf7ff489
[Core] Release cached device memory under pressure on UMA GPUs during weight loading ( #45179 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-11 17:49:50 +03:00
b78fc47f05
[Docs] Add redirect for moved lmcache examples page ( #45218 )
...
Signed-off-by: nataliepjlin <nataliepjlin@gmail.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:41:08 -04:00
Harry Mellor and GitHub
03878d1c22
Deprecations for v0.23 and v0.24 ( #44992 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 14:35:38 +00:00
55911db580
[PD][Core] Fix Mamba prefix cache hit rate in PD disaggregation ( #44243 )
...
Co-authored-by: lHrHenry233 <2381623149@qq.com >
Co-authored-by: underfituu <hzhucong@163.com >
Signed-off-by: Zhanqiu Hu <zhu@redhat.com >
2026-06-11 14:10:25 +00:00
cc640ee8bc
[Rust Frontend][Metrics] Export vllm:lora_requests_info from frontend ( #45030 )
...
Signed-off-by: Will Eaton <weaton@redhat.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 06:45:03 -07:00
ebc6ef971a
Hidden states extraction improvements ( #43805 )
...
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 09:44:45 -04:00
tc-mb and GitHub
ab3a1fd2e6
minicpmv4_6: fix ImageSize (W,H) order for placeholder token calculation ( #45244 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-11 13:43:56 +00:00
c3662b36ea
[KV offload] Parallel-agnostic fs-tier cache for single full-attention group ( #44733 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-06-11 15:48:37 +03:00
Juan Pérez de Algaba and GitHub
e62d00ab73
docs: add fix disclosure policy to SECURITY.md ( #45253 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 12:48:00 +00:00
1f60771c74
fix: guard flash-attn rotary import ( #42679 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-11 08:43:31 -04:00
05d9848267
[Build] Upgrade CUDA Dockerfiles from GCC 10 to GCC 12 for C++20 compatibility ( #44923 )
...
Signed-off-by: Richard Barnes <rbarnes@meta.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:26:52 +00:00
jasen and GitHub
ef67071b21
[Build] Skip spinloop extension on Python < 3.11 ( #44783 )
...
Signed-off-by: Jasen2201 <yajizhan@amd.com >
2026-06-11 11:23:21 +00:00
x41lakazam and GitHub
3508cb78d4
[Bugfix] Fix broken profile_modular_kernel.py ( #43300 )
2026-06-11 12:17:23 +01:00
Harry Mellor and GitHub
432905d5d6
Only enable PR docs builds manually ( #45262 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 03:14:29 -07:00
1f9dd7900d
[Bugfix][Rust Frontend] Validate out-of-vocab token ids in request params ( #44680 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 03:14:11 -07:00
9492362972
[Security] Apply sanitize_message to Anthropic and STT error paths ( #45119 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-11 10:05:34 +00:00
7852e50e4d
[docs] Document --scheduler-cls base class requirement (extend AsyncScheduler, not Scheduler) ( #43724 )
...
Signed-off-by: Georgii Kliukovkin <kliukovkin@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 10:49:51 +01:00
Reid and GitHub
0d657e44dc
[Rust Frontend] Fix DeepSeek V3.2 continue_final_message rendering ( #45155 )
...
Signed-off-by: reidliu41 <reid201711@gmail.com >
2026-06-11 09:34:19 +00:00
aa1df36c53
Fix/minicpmv46 missing version ( #44980 )
...
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-11 09:20:45 +00:00
f06aefb4e3
[CPU] Add missing scalar fallback for CPU W4A8 INT4 GEMM ( #44523 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn >
2026-06-11 08:52:01 +00:00
Julien Denize and GitHub
1c3a72b8b2
[Bugfix] Add fetch_images to MistralCommonImageProcessor ( #45180 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-11 16:13:01 +08:00
Juan Pérez de Algaba and GitHub
d598d23973
[Security] Reject non-finite temperature and repetition_penalty values ( #45116 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 01:12:14 -07:00
Juan Pérez de Algaba and GitHub
f219788f91
[Security] Fix info disclosure via int32 truncation in GGUF dequantize kernels ( #44971 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-11 08:05:14 +00:00
6e64c1bab1
[10c/n] Migrate MoE kernels to torch stable ABI ( #44565 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-10 23:02:26 -07:00
Kevin H. Luu and GitHub
2f2c5cf4f1
[release] Always block release images to dockerhub ( #45236 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-10 22:53:04 -07:00
Mohammad Miadh Angkad and GitHub
40e065e86a
[Docker] Fix CUTLASS DSL cu13 install order in Dockerfile ( #45204 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-11 05:19:36 +00:00
0b995f8609
Use std::bit_cast for type punning in CPU kernels ( #45089 )
...
Signed-off-by: Yuanyuan Chen <cyyever@outlook.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-10 22:07:44 -07:00
Bugen Zhao and GitHub
43914dd743
[Rust Frontend] Add Python bridge for Rust tool parsers ( #44624 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-11 04:51:06 +00:00
3501324957
[Build] fix self-contradictory precompiled-flag orthogonality test ( #44942 )
...
Signed-off-by: pjdurden <prajjwalchittori1@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-11 12:49:08 +08:00
Flora Feng and GitHub
3a04061701
[Refactor][Parser] Unify Response API to use parser.parse() like Chat Completion API ( #45190 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-11 04:37:51 +00:00
Yifan Qiao and GitHub
f272dfdce1
[KV Connector] Mooncake store: prefix-cache retention interval for sparse attention ( #44774 )
2026-06-10 21:36:34 -07:00
velonica0 and GitHub
f31bc2ea60
[CPU][RISC-V] Enable oneDNN W8A8 INT8 to run on RISC-V ( #44478 )
...
Signed-off-by: velonica0 <like@mail.nankai.edu.cn >
2026-06-11 04:09:05 +00:00
248e33c40d
[Bugfix][Responses API] Set id on function_call item in streaming done event ( #44608 )
...
Signed-off-by: Aniruddh Krovvidi <aniruddh.krovvidi@oracle.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-11 03:52:42 +00:00
Bugen Zhao and GitHub
5d5591d99b
[Rust Frontend] Populate cached_token_count in responses ( #44887 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 20:50:05 -07:00
Wentao Ye and GitHub
85a0ffae42
[CI Bug] Remove qwen test ValueError: No example model defined for Qwen/Qwen-7B-Chat ( #45194 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-10 20:11:00 -07:00
Harry Mellor and GitHub
18d87a87dc
Deprecate Transformers v4 support ( #45161 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-11 11:04:01 +08:00
Flora Feng and GitHub
b038a2f73b
[CI][Bugfix] Update Dockerfile dependency graph PNG ( #45209 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 19:40:25 -07:00
Ting SUN and GitHub
2d481f8a94
[Bugfix][Rust Frontend] Stop unescaping XML-style tool-call parameter values ( #45025 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-10 19:05:23 -07:00
7920ccb97c
[Bugfix]: Fix Quark gpt-oss weight loading broken by FusedMoe refactor ( #45067 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 18:17:46 -07:00
Wentao Ye and GitHub
86111c00c7
[Chore] Add Github notification for MRv2 for @yewentao256 ( #45191 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-11 09:01:49 +08:00
qizixi and GitHub
e2db0222e9
[Perf][Attention] Pin MLA chunked-context metadata tensors so H2D copies are truly non-blocking ( #45074 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-10 15:56:49 -07:00
82d6b59f04
[CI/Build] Skip test_use_trtllm_attention on non-CUDA platforms ( #44687 )
...
Signed-off-by: Dan Blanaru <48605845+DanBlanaru@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 18:18:42 -04:00
Andreas Karatzas and GitHub
16282a9c4e
[ROCm][CI] Moving MI300 tests to MI325 until cluster is stabilized ( #45170 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 20:26:17 +00:00
5b6b536fdc
[ROCm][Bugfix] Make intermediate_pad TP-aware in rocm_aiter_fused_experts ( #44679 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 15:10:50 -05:00
12f3f19c19
feat(qwen3-asr): support prompt parameter in v1/audio/transcriptions ( #35415 )
...
Signed-off-by: Nathan Price <nathan@abridge.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 19:54:59 +00:00
Ilya Markov and GitHub
6471ec75bd
[EPLB] Reject NCCL-based EPLB communicators with async EPLB ( #44978 )
...
Signed-off-by: Markov Ilya <markovilya197@gmail.com >
2026-06-10 19:51:27 +00:00
3d300aecb1
[Doc] Switch K8S examples to default MP mode ( #39400 )
...
Signed-off-by: Peter Pan <Peter.Pan@daocloud.io >
Signed-off-by: Peter Pan <peter.pan@daocloud.io >
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com >
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-06-10 18:17:11 +00:00
ffce72c041
[Model Runner V2] Fix v2 AttributeError: 'CohereASRDecoder' object has no attribute 'embed_input_ids' ( #44568 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 11:06:01 -07:00
TJian and GitHub
bfe1001ab6
[Bugfix] [DSV4] [ROCm] Pin apache-tvm-ffi version to 0.1.10 ( #45169 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-10 17:41:15 +00:00
fa8c868a3c
[Bugfix] Fix Llama4 weight loading ( #45047 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-10 13:40:45 -04:00
Ben Browning and GitHub
d1bcb4b44c
[Bugfix] Fix tool parsing crash with non-function tool types (e.g. WebSearchTool) ( #45147 )
...
Signed-off-by: Ben Browning <bbrownin@redhat.com >
2026-06-10 17:17:16 +00:00
bnellnm and GitHub
29026682cb
[Bugfix] Fix nemotron accuracy drop introduced by #41184 ( #45037 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 13:16:25 -04:00
Stan Wozniak and GitHub
dc66e01a70
[Hybrid] Marconi-style admission policy for hybrid cache ( #37898 )
...
Signed-off-by: Stanislaw Wozniak <stw@zurich.ibm.com >
2026-06-10 10:03:13 -07:00
Yongye Zhu and GitHub
2ba68d9bf7
[Test] Fix one-sided MNNVL alltoall test workspace under-reservation ( #44946 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-11 00:43:12 +08:00
Julien Denize and GitHub
2131b597b1
[CI] Ping Mistral team for ministral/voxtral/mixtral/pixtral changes ( #45153 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-06-10 08:48:00 -07:00
0bae1d3848
[MRV2][Spec Decode] DFlash ( #44586 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-10 08:47:46 -07:00
Yufeng He and GitHub
4673ca1d78
fix: prefix DeepSeek V4 MTP projections ( #44821 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
2026-06-10 08:47:04 -07:00
Angela Yi and GitHub
de900fa7e5
fix: AOT compile cache collision for dataclass-based HF configs ( #45059 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-06-10 08:05:29 -07:00
Divakar Verma and GitHub
166d14e9bf
[bugfix] skip conch kernel for g_idx reordering ( #45072 )
...
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
2026-06-10 23:04:19 +08:00
af65e08fc5
KV-Cache multi-tier offloading async batched lookup ( #44193 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 14:59:30 +00:00
Harry Mellor and GitHub
3cc9fecd58
Deprecated 1st generation Qwen and QwenVL models ( #45131 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 14:55:33 +00:00
ccc05de038
[Bugfix] Fix missing sequence_lengths in EXAONE-4.5 vision encoder ( #45073 )
...
Signed-off-by: Jongsu Liam Kim <jongsukim8@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 15:44:34 +01:00
6ec7dcd641
[Frontend][Metrics] Add vllm:tool_call_parser_invocations_total Prometheus metric ( #44448 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 10:29:11 -04:00
c9e5bf8135
[Bugfix] Fix layerwise reload dropping params after a composed weight loader ( #44814 )
...
Signed-off-by: hallerite <git@hallerite.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com >
2026-06-10 06:42:05 -07:00
Roberto L. Castro and GitHub
6850839c6f
[Perf] Fix dsv3_router_gemm heuristic ( #44217 )
...
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com >
2026-06-10 06:08:41 -07:00
87c15d46e3
[Bugfix] Lazily import the humming quantization backend ( #44921 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 06:06:17 -07:00
4882fd7632
[Bugfix][Reasoning] Nemotron V3: surface reasoning as content when thinking is unterminated ( #39091 )
...
Signed-off-by: Andrii Skliar <askliar@nvidia.com >
Co-authored-by: Andrii Skliar <askliar@nvidia.com >
2026-06-10 05:58:19 -07:00
77f42d9725
[Model] Remove obsolete ERNIE models ( #45127 )
...
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 20:54:30 +08:00
9dfc313bdc
Feature/offloading manager stats ( #35669 )
...
Signed-off-by: Sriusa4414@gmail.com
Signed-off-by: srinivas_oo7 <Sriusa4414@gmail.com >
Signed-off-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Signed-off-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Signed-off-by: Or Ozeri <oro@il.ibm.com >
Co-authored-by: srinivas_oo7 <sklinkedin0120@gmail.com >
Co-authored-by: Srinivasoo7 <158864704+Srinivasoo7@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com >
2026-06-10 12:44:55 +00:00
9ad08c4d15
[Bugfix][Rust Frontend] Fix missing added tokens in hf/fastokens tokenizer ( #44683 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 03:52:41 -07:00
Shantipriya Parida and GitHub
a1ec011a83
[Bugfix] Add deepseek_v32 to Quark dynamic MXFP4 model type check ( #39498 )
...
Signed-off-by: Shantipriya Parida <shantipriya.parida@amd.com >
2026-06-10 02:52:33 -07:00
Bugen Zhao and GitHub
fdfb2566c0
[Rust Frontend] [CI] Unify Rust artifact builds with setuptools-rust ( #44981 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-10 17:48:34 +08:00
Juan Pérez de Algaba and GitHub
8a5cf1ccd6
[Security] Fix remote DoS via invalid recovered token reinjection ( #44744 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 02:31:43 -07:00
Kunshang Ji and GitHub
fe1d923afc
[BUGFIX][XPU] fix xpu flash_attn_varlen_func interface ( #45110 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-10 17:07:40 +08:00
32daf56b42
[Refactor] Rename rocm_moe.py to rocm_moe_rdna.py ( #45011 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-10 17:02:09 +08:00
Andreas Karatzas and GitHub
82a42234be
[ROCm][CI] Defer AITER sampler import and isolate server test PYTHONPATH ( #44823 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-10 08:56:11 +00:00
Harry Mellor and GitHub
af9f583344
Revert "[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures" ( #45029 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 01:37:03 -07:00
yiheng and GitHub
bd2d83ff31
[SpecDecode] Reduce TP communication for large-vocab draft models speculative decoding ( #39419 )
...
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn >
2026-06-10 07:59:24 +00:00
xiaohuguo2023 and GitHub
bb78168b21
[ROCm][gpt-oss] Hybrid CDNA4 swizzle gate for A8W4 MoE ( #44804 )
...
Signed-off-by: Xiaohu Guo <Xiaohu.Guo@amd.com >
2026-06-09 23:59:44 -07:00
89c6a41001
[Bench] Add BFCL dataset for vllm bench serve tool-calling workloads ( #42457 )
...
Signed-off-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Li Zhang <lzhanga@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-09 23:59:18 -07:00
7fdfa6441d
Model/colbert autoweightsloader ( #44999 )
...
Signed-off-by: Furkan Fidan <dev@yufufi.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 23:58:50 -07:00
Harry Mellor and GitHub
e9b728de8a
Change from owning configs to owning config utils ( #45058 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-10 06:40:25 +00:00
5828a205ef
Fix Harmony tool descriptions for optional fields ( #44686 )
...
Signed-off-by: Varun Shenoy <varun.vinayak.shenoy@oracle.com >
Co-authored-by: Codex <codex@openai.com >
2026-06-09 23:29:22 -07:00
Yaoming Zhan and GitHub
7a74f31d2e
[Rust Frontend] Add seed_oss and step3p5 reasoning parsers ( #44552 )
...
Signed-off-by: yzhan1 <zhanyaoming2014@gmail.com >
2026-06-09 23:01:33 -07:00
47930b59ca
[Bugfix] Handle HWC images in ImageProcessorItems.get_image_size ( #45057 )
...
Signed-off-by: YellowFoxH4XOR <yellowfoxh4xor@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-10 05:35:50 +00:00
Flora Feng and GitHub
6aec99f030
[Refactor] Remove dead states from chat completion serving ( #45081 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-09 22:20:15 -07:00
bnellnm and GitHub
f4966f8b3d
[Bugfix] Fix weight loading issues caused by #41184 ( #45054 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-10 01:20:13 -04:00
Mohammad Miadh Angkad and GitHub
2c9c07c85e
[Bugfix][CI/Build] Fix Rust frontend build after chat conversion refactor ( #45085 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 20:04:41 -07:00
Change72 and GitHub
320c52b134
[Bench] benchmark_serving_multi_turn: make non-standard conversation_id payload opt-in ( #43756 )
...
Signed-off-by: Change72 <cguo51@asu.edu >
2026-06-09 19:41:56 -07:00
6deb05e0e4
[Core][Model] Gemma4: Unified FA4 for all layers + FlashAttention mm_prefix support ( #42175 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-09 17:45:39 -07:00
Flora Feng and GitHub
d82ac00923
[Refactor][Mistral] Extract parsing logic into MistralParser ( #44596 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-10 00:12:23 +00:00
Bugen Zhao and GitHub
dac9e9a640
[Rust Frontend] Extract shared options in route helper params ( #44884 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 17:02:35 -07:00
Wentao Ye and GitHub
d7607ad273
[Bug] Fix deepseek v4 OOM issue ( #44914 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-09 15:47:06 -07:00
d955745d58
[ROCm][CI] fix test_rope_kvcache_fusion.py ( #44678 )
...
Signed-off-by: charlifu <charlifu@amd.com >
Co-authored-by: Rohan Potdar <66227218+Rohan138@users.noreply.github.com >
2026-06-09 21:53:46 +00:00
Micah Williamson and GitHub
e1ed89dbee
Revert "[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuf… ( #45066 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 14:12:06 -07:00
1c2ffc6f88
feat(multi-turn-bench): add api_key and custom headers for multi turn benchmark ( #44516 )
...
Signed-off-by: Jimmy <jinmingyi1998@sina.cn >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: simon-mo <simon.mo@hey.com >
2026-06-09 14:00:07 -07:00
Jiangyun Zhu and GitHub
ca4cfd8731
[Bugfix] fix qwen3.5 ep weight loading ( #45002 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 13:55:30 -07:00
Micah Williamson and GitHub
c9c1540e61
[ROCm][V2] Fix failed assertion in Llama models when using EAGLE with ROCM_AITER_FA ( #44936 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-09 13:30:52 -05:00
c1d754d681
[Mooncake] Use all HCAs on multi-NIC hosts instead of GPU-indexed RNIC selection ( #43799 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-06-09 11:05:36 -07:00
01d8cd92dd
[ROCm][Perf] Use fused softplus-sqrt-topk router under AITER fused-MoE ( #44945 )
...
Co-authored-by: vLLM Contributor <contributor@vllm.ai >
2026-06-09 17:53:05 +00:00
a4b14b98c6
[Kernel] Speed up silu_and_mul_per_block_quant with warp-shuffle reduction + vectorized I/O ( #44173 )
...
Signed-off-by: SII-yangdian <yangdian@sii.edu.cn >
Co-authored-by: SII-yangdian <yangdian@sii.edu.cn >
2026-06-09 10:41:26 -07:00
Juan Pérez de Algaba and GitHub
cf1c906724
[Security] Fix image EXIF orientation and tRNS transparency handling ( #44974 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-09 09:34:44 -07:00
766ce2bb6b
Fix MiDashengLM TP>1 crash in audio encoder attention ( #44408 )
...
Signed-off-by: Michał Ganczarenko <michal.ganczarenko@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-09 09:29:14 -07:00
3d119f78f7
[Docs] Add KV offloading usage guide (single- and multi-tier) ( #44415 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 19:20:23 +03:00
Juan Pérez de Algaba and GitHub
1b1359c332
[Security] Fix DoS via audio decompression bomb in speech-to-text endpoint ( #44970 )
...
Signed-off-by: jperezde <jperezde@redhat.com >
2026-06-10 00:18:53 +08:00
Tyko Niemi and GitHub
cad4ca12b8
[Bugfix] Add X-Session-ID from conversation_id in multi-turn benchmark ( #44663 )
...
Signed-off-by: Tyko Niemi <tyko.niemi@amd.com >
2026-06-09 08:57:00 -07:00
Andreas Karatzas and GitHub
b697119800
[ROCm][CI] Stabilize ModernBERT token-classification parity against Hugging Face ( #44040 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:52:36 +01:00
Kunshang Ji and GitHub
b4c6dc6454
[WIP][XPU] upgrade torch-xpu to 2.12 ( #42262 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-09 15:51:39 +00:00
Raushan Turganbay and GitHub
2ee5106372
Remove raw_inputs from transformers backend ( #39425 )
...
Signed-off-by: raushan <raushan@huggingface.co >
2026-06-09 15:01:04 +00:00
Jiangyun Zhu and GitHub
7a89b72564
[Perf] fuse qk rmsnorm rope gate for qwen3.5 ( #44176 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-09 22:12:17 +08:00
Jee Jee Li and GitHub
dc10e467a9
[Bugfix] Fix minimax_qk_norm_fusion ( #44983 )
2026-06-09 06:43:46 -07:00
Terrence Zhao and GitHub
ee4d7df2b5
[Cohere] Cohere2 moe parser fix ( #44907 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:32:18 -07:00
Terrence Zhao and GitHub
3e8afdf785
[Cohere] Fix Cohere2MoE weight loading when using Transformers ≥5.10 ( #44747 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-09 06:27:40 -07:00
Nicolò Lucchesi and GitHub
6690a0c4de
[PD][Bugfix] Fix KV Cache sharing with HMA ( #44629 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 06:10:06 -07:00
Maria Guevara and GitHub
1c23c42030
[Rust Frontend] Support Kimi K2 tool call IDs ( #44901 )
2026-06-09 05:31:26 -07:00
xiangdong and GitHub
b12e42d132
[XPU][CI] Refine docker image build and pull/create lock mechanism in Intel GPU CI ( #44481 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-09 20:20:32 +08:00
69fdaffbcd
[Rust Frontend] Add /tokenize and /detokenize endpoints ( #44222 )
...
Signed-off-by: Tan Ngoc Do <darkknightkhtn2008@gmail.com >
Signed-off-by: TanNgocDo <darkknightkhtn2008@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-09 05:11:37 -07:00
80e2c4462d
[ROCm][Compile] Fuse AR + RMSNorm + per-group FP8 quant (+ DSv3.2 indexer fan-out) ( #42864 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
2026-06-09 12:06:56 +00:00
Sage and GitHub
5b3807e862
[KV Events] Switch event structs from array to map encoding ( #42892 )
...
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com >
2026-06-09 11:39:52 +00:00
Qiuyang Yue and GitHub
59401ac9f1
[Kernel][Perf] Tune fused_moe FP8 config for Qwen3-Next-80B tp=4 on H100 (+25% at batch 96-512) ( #44830 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-09 04:15:51 -07:00
d841386d27
[Rust Frontend] Support API key authentication ( #44321 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 10:15:20 +00:00
Mohammad Miadh Angkad and GitHub
fff9210b2a
[CI/Docs] Remove stale disagg prefill links ( #44918 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-09 03:05:53 -07:00
Ma Jian and GitHub
70db1488c5
[DSV4][XPU] Add MHC fused_post_pre support ( #44144 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-09 17:23:17 +08:00
Andreas Karatzas and GitHub
2385e140d6
[ROCm][CI] Stabilize sleep-mode memory release ( #43022 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 16:51:12 +08:00
Nicolò Lucchesi and GitHub
dab60fc658
[Bugfix][CI] Fix test_offloading_connector.py::test_fs_tiering_offloading ( #44903 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-09 00:57:34 -07:00
wang.yuqi and GitHub
996222f4bf
[CI] Reorganize entrypoints CI ( #44947 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-09 00:46:11 -07:00
e6fc848d4f
[Bugfix][MiniCPM-o] Fix cuda/cpu device mismatch in Resampler2_5 pos_embed ( #43844 )
...
Signed-off-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Parth Ashwin Jain <parthash@amd.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 23:28:26 -07:00
Andreas Karatzas and GitHub
f843ac1a1c
[Bugfix][CI] Gemma3 Transformers multimodal encoder profiling and build prompt-embedding fixtures ( #44952 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-09 05:49:30 +00:00
7c2aa3108a
fix: prevent MM cache hang from stale LRU order keys ( #43595 )
...
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:48:31 -07:00
ebf53ba373
[Bugfix][Rust Frontend] Set a structured-output backend so requests do not 500 ( #44729 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 22:30:54 -07:00
baacbfcebf
[ROCm][MLA][Bugfix] Reserve FP8 prefill workspace before lock for Kimi-K2.5 ( #42978 )
...
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 22:25:52 -07:00
d8218b1ee7
[Bugfix] Propagate ImportError from load_audio_pyav when vllm[audio] … ( #44750 )
...
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com >
Signed-off-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
Co-authored-by: Hsiao-Yuan Chen <hy.c@Hsiao-YuandeMacBook-Pro.local >
2026-06-09 04:24:52 +00:00
9f153aa781
[MM][Perf][CG] Support ViT full CUDA graph for glm4_1v image and video inference ( #40576 )
...
Signed-off-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: grYe99 <guorongye99@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-09 11:13:56 +08:00
Kunshang Ji and GitHub
d3de61502f
[XPU][CI] fix test case path ( #44940 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-08 20:02:31 -07:00
Lanze Liu and GitHub
540aaf2140
[Bugfix][Model] Qwen3-Omni: move cu_seqlens to GPU before VIT attention ( #44264 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-08 20:02:27 -07:00
Lanze Liu and GitHub
4128605ad4
[Docs] Remove broken link to deleted disaggregated_prefill.sh ( #44929 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-06-09 01:40:06 +00:00
e2f993dc41
[WideEP] Integrate DeepEP v2 ( #41183 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-06-08 18:07:29 -07:00
Andreas Karatzas and GitHub
05cb606cad
[ROCm][CI] Re-route NixlConnector jobs ( #44809 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 18:57:11 -05:00
3f627ebef7
[Misc] usage_stats: report more engine, spec-decode, and EP config ( #44595 )
...
Signed-off-by: Zach Xi <zachary.xi@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-08 15:20:00 -07:00
Bugen Zhao and GitHub
bc941f375d
[Rust Frontend] [Refactor] Refine utility call interfaces ( #44856 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 15:13:08 -07:00
Michael Goin and GitHub
6afa25000c
[Bugfix] Canonicalize FP8 weight layout to (K, N) at the source ( #44735 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-08 14:37:36 -06:00
Mohammad Miadh Angkad and GitHub
823a0ab754
[Bugfix][MoE] Fix fused MoE expert mapping helper call sites ( #44897 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-08 13:35:04 -07:00
Wentao Ye and GitHub
2c27c294c0
[Model Runner V2] Fix mrv2 mm lora issue ( #44450 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-08 14:30:09 -04:00
ba94a3b998
[Attention] Extract KV-cache update from CPU attention backend ( #40470 )
...
Signed-off-by: Diego Maniloff <diego.maniloff@gmail.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-08 15:43:05 +00:00
bnellnm and GitHub
dc68bd8c41
[MoE Refactor] FusedMoE/MoERunner inversion refactor ( #41184 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-08 10:42:58 -04:00
753e9d55e6
[Quantization] add online fp8 ptpc ( #44132 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-08 22:42:11 +08:00
akii96 and GitHub
ac3409d162
[Benchmark] Auto-detect and correct client/server tokenizer mismatch for random dataset ( #44708 )
2026-06-08 06:10:20 -07:00
wang.yuqi and GitHub
93ee4cd47f
[CI] Consolidate multimodal entrypoint tests. ( #44819 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-08 04:48:08 -07:00
Li, Jiang and GitHub
980796cd07
[CI/Build][CPU] Fix flaky CI image build failure and unexpected warnings ( #44852 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-08 11:10:06 +00:00
Nicolò Lucchesi and GitHub
5add018beb
[Connector] Remove P2pNcclConnector ( #44854 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-08 18:58:29 +08:00
d5fe994e79
[CPU][Spec Decode] Warn about throughput loss when libiomp5 is not preloaded ( #44419 )
...
Signed-off-by: jmamou <jonathan.mamou@intel.com >
Signed-off-by: Jonathan Mamou <jonathan.mamou@intel.com >
Co-authored-by: Li, Jiang <bigpyj64@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-08 03:45:08 -07:00
Chaojun Zhang and GitHub
fa662b1a8b
[XPU] Cap topk/topp Triton BLOCK_SIZE to 4096 to fix Top-p mask difference failures ( #44470 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-08 09:36:51 +00:00
3c0b4432be
[Rust Frontend] Add /pause, /resume, /is_paused endpoints ( #44499 )
...
Signed-off-by: Sahil Singh <sahiilsiingh37@gmail.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-08 17:28:37 +08:00
Sungjae Lee and GitHub
469f3dcf1d
[BugFix] Use served model name in gemma4 audio-tower error message ( #44828 )
...
Signed-off-by: Sungjae Lee <33976427+llsj14@users.noreply.github.com >
Signed-off-by: Sungjae Lee <sung-jae.lee@navercorp.com >
2026-06-08 06:58:31 +00:00
xiangdong and GitHub
94fcdd007f
[XPU][CI] Add more test cases in Intel GPU CI ( #43663 )
...
Signed-off-by: zengxian <xiangdong.zeng@intel.com >
2026-06-08 06:21:24 +00:00
Andreas Karatzas and GitHub
d9ff7e4e9a
[ROCm][CI] Stabilizing teardown and timeout of flaky tests to prevent rare OOMs ( #44761 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-08 14:11:17 +08:00
Andreas Karatzas and GitHub
967c5c3bc3
[ROCm][CI] Stage C mirrors ( #42793 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 23:00:59 -07:00
Yan Ma and GitHub
54c660c3a6
[XPU][Minor] format moe kernel name and add in kernel list ( #44771 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-08 13:58:16 +08:00
Shanshan Shen and GitHub
8fb0274415
[MM][CG] Simplify ViT CUDA graph interfaces ( #44484 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-08 05:57:06 +00:00
Ma Jian and GitHub
eebce65756
[XPU]feat: add DeepSeek-V4 XPU attention decode path ( #42953 )
...
Signed-off-by: Ma Jian <jian1.ma@intel.com >
2026-06-08 13:27:12 +08:00
303916e93d
[Bugfix]: Fix assertion in MambaManager.allocate_slots() ( #39562 )
...
Signed-off-by: Holworth <kangqihan17@mails.ucas.ac.cn >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-08 00:34:37 -04:00
Taneem Ibrahim and GitHub
5633405964
Added extra_repr() to pooler classes to improve debuggability ( #44805 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-08 03:19:31 +00:00
6124a98a9b
[Bugfix] Fix FunASR-Nano crash during initialization ( #44215 )
...
Signed-off-by: SunskyXH <sunskyxh@gmail.com >
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-06-07 20:00:02 -07:00
Agata Dobrzyniewicz and GitHub
2ed0a9627b
[Kernel][Test] Make kernel tests for mamba dual-HW (CUDA + XPU) ( #42736 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-08 08:22:47 +08:00
4dcd10eb0d
[1/N][KV-Cache Layout Refactor] Refactor DSV4 KV cache config construction ( #44454 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-07 14:53:37 +00:00
Charlie Fu and GitHub
228bcc436b
[ROCm][Kernel] Enable permute_cols for ROCm ( #44674 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-07 09:50:03 +00:00
3d3ba460a2
Modify torch dependency in xpu.txt ( #43087 )
...
Signed-off-by: Bram Vanroy <2779410+BramVanroy@users.noreply.github.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-07 16:33:50 +08:00
Mohammad Miadh Angkad and GitHub
66ecfd0568
[Dependency] Remove stale cuDNN frontend upper bound ( #42599 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-06-07 16:09:25 +08:00
Andreas Karatzas and GitHub
f0f6805d8a
[CI] Stabilize the multi-audio OpenAI server path ( #44051 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-07 15:54:32 +08:00
15652a6b70
[Doc] Fix multimodal torch.compile troubleshooting to not use removed VLLM_TORCH_COMPILE_LEVEL ( #44378 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-07 00:34:07 -07:00
Yifan Qiao and GitHub
51ef688831
[Bugfix][Mooncake] Fix per-group block_size/block_hash and group_idx in MooncakeStoreConnector KV events ( #44103 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-07 07:12:43 +00:00
Jared Wen and GitHub
6ac69203e8
[videoloader] implement glm46v video loader ( #44417 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
2026-06-07 06:27:20 +00:00
1505b3d8a1
[Cohere] Enable Cohere Mini Code model and update Command A-plus test registry ( #44707 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:44:16 -07:00
32f34d3935
[feature] add index share feature for DSA MTP ( #44420 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:04:14 -07:00
Qiuyang Yue and GitHub
9c7f7741d4
[Bugfix] Fix benchmark_moe.py after inplace mechanism removal ( #44041 )
...
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com >
2026-06-07 00:32:00 -04:00
6181e80fe0
[XPU] add xpu branch in compressed_tensors_moe_w4a4_mxfp4 ( #44540 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com >
Co-authored-by: Kunshang Ji <jikunshang95@gmail.com >
2026-06-07 12:27:34 +08:00
Yan Ma and GitHub
3bb46975bd
[XPU][Feature] transparent sleep mode support for XPU platform ( #37149 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-07 10:45:31 +08:00
Chaojun Zhang and GitHub
810966453a
[XPU] Support cpu kv offloading and tiering offloading on XPU platform ( #36423 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-07 09:59:28 +08:00
Woosuk Kwon and GitHub
2a983c79ac
[DSV4] Decouple DS V4 Sparse MLA Metadata from DS V3.2 ( #44699 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-06 20:37:56 -04:00
bc5745a00f
[ROCm][MLA] Replace torch.cat in sparse-MLA forward_mqa with fused concat_mla_q ( #42838 )
...
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com >
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 18:20:50 -05:00
Nick Hill and GitHub
3b3d5287fa
[BugFix] Resolve multiple async kv load deadlock ( #44560 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-06 23:05:47 +00:00
062b05ff3a
[ROCm][Perf] Fused MoE W4A16 HIP kernel for AMD RDNA3 (gfx1100) ( #44075 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 15:30:39 -05:00
fa27d4e9cf
[PERF] [Qwen3.5] Split mixed prefill+decode batches: route decodes to the recurrent kernel ( #44700 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-06 22:13:50 +08:00
Vadim Gimpelson and GitHub
67d3792d99
[Bugfix] Fix Qwen3.5-FP8 nightly fail. Guard fused_add_rms_norm input/weight dtype mismatch in RMSNorm + quant fusion ( #44694 )
2026-06-06 08:46:14 -04:00
00d1fb7747
[Bugfix][ROCm] ApplyRotaryEmb: fall back to native when flash_attn rotary grid would exceed the HIP per-dim limit ( #43684 )
...
Signed-off-by: vLLM ROCm fix <noreply@example.com >
Signed-off-by: amd-fuweiy <fuweiy@amd.com >
Co-authored-by: vLLM ROCm fix <noreply@example.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-06 02:29:16 -07:00
c9b4b184b4
[Bugfix][Voxtral] Add fetch_audio to MistralCommonFeatureExtractor (transformers>=5.10 compat) ( #44559 )
...
Signed-off-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Yadan Wei <weiyadan@amazon.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 07:58:09 +00:00
f87df1df9e
[Bugfix][MoE] Snapshot max_cudagraph_capture_size into FusedMoEConfig ( #44613 )
...
Signed-off-by: aoshen02 <aoshen@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-06-05 23:14:41 -07:00
Taneem Ibrahim and GitHub
eafbb06331
[Misc] Replaced asserts with proper exceptions to improve UX for pooling ( #44593 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-06 05:57:26 +00:00
ec0a31d4aa
[Bugfix][Kernel] Fix mHC fused-RMSNorm big-fuse miscompile for hidden_size != 4096 ( #44692 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-06 10:44:21 +08:00
Devin Lai and GitHub
c8beda4cc3
[Rust Frontend] Add Phi-4 mini JSON tool parser ( #44213 )
2026-06-06 10:40:00 +08:00
2f27c9a150
Preserve layout-changing clones ( #44574 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-06-05 20:45:24 -04:00
4765f0f189
[Bugfix] Fix sequence_parallel_chunk_impl custom op aliasing its input ( #44130 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-05 23:56:36 +00:00
Terrence Zhao and GitHub
a50e675b0d
[Cohere] fix RoutingMethodType ( #44021 )
...
Signed-off-by: Terrencezzj <terrence@cohere.ai >
2026-06-05 16:25:53 -07:00
Daoyuan Li and GitHub
f6a708ab2b
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models ( #44435 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-05 16:04:32 -07:00
4200f62147
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path ( #42832 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 16:22:19 -05:00
Walter Beller-Morales and GitHub
c73b0d0db9
[Core][Engine] allow DP ray placement groups to be set on specific nodes ( #44669 )
...
Signed-off-by: walterbm <walter.beller.morales@gmail.com >
2026-06-05 20:07:47 +00:00
Harry Mellor and GitHub
e28e369f78
Male Mergify comment less spammy ( #44666 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 10:56:52 -07:00
yzong-rh and GitHub
703fb17b13
[Bugfix] GPT-OSS instruction rendering ( #44330 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-06-05 13:52:32 -04:00
Sting Lin and GitHub
b593396c7a
Upgrade tpu-inference to v0.21.0 ( #44621 )
...
Signed-off-by: StingLin <sting.lin@cienet.com >
2026-06-05 16:12:49 +00:00
Flame and GitHub
91e17d4315
Fix sarvam forward compatibility with transformers v5 ( #38804 )
...
Signed-off-by: vikrantpalle <vikrantpalle@gmail.com >
2026-06-05 11:51:44 -04:00
TJian and GitHub
aa6fb8a329
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm ( #44648 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-06-05 15:51:17 +00:00
Effi Ofer and GitHub
6a894574bf
Add objectstore as a secondary tier to multi-tier kv cache offloading ( #41968 )
...
Signed-off-by: Effi Ofer <effi.ofer@gmail.com >
2026-06-05 18:05:41 +03:00
Yan Ma and GitHub
7f003a1285
Support MiniCPMV batched preprocessing ( #44609 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
2026-06-05 15:05:31 +00:00
Harry Mellor and GitHub
ef0df7dbd6
[CI] Bump mypy version 1.19.1 -> 1.20.2 ( #44647 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:56:27 +00:00
Harry Mellor and GitHub
a80af24356
Speed up docs build ( #44635 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:51:44 +00:00
Harry Mellor and GitHub
c66b19800b
[CI] Bump mistral-common ( #44649 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-05 14:18:50 +00:00
6a11d72df7
[Reasoning][Structured Outputs] Add Command A plus tags for structural tags ( #44588 )
...
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-06-05 06:51:20 -07:00
Woosuk Kwon and GitHub
02d2da0748
[DSV4] Move more ops out of eager breakpoint ( #44561 )
2026-06-05 06:42:41 -07:00
adhithyamulticoreware and GitHub
bbb6c274c8
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call ( #44615 )
...
Signed-off-by: ADHITHYA BALAKRISHNAN <adhithya.balakrishnan@multicorewareinc.com >
2026-06-05 12:47:56 +00:00
62215e72c6
Remove KV cache scale boilerplate from model weight loading methods ( #43167 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 05:19:04 -07:00
7fe7800fa4
[BUG] Fix FP64 Gumbel precision coverage ( #43150 )
...
Signed-off-by: tianyu-z <zhangtianyupro@gmail.com >
Signed-off-by: Tianyu Zhang <53099276+tianyu-z@users.noreply.github.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-05 19:04:14 +08:00
8a83e6f2d7
[Rust Frontend] Batch auto-abort requests by engine ( #44591 )
...
Signed-off-by: Hugh Ryan <197298026+HueCodes@users.noreply.github.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-05 02:59:09 -07:00
Chunyang Wen and GitHub
efc347f1b2
docs: fix tokenizer optimization typo ( #44066 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-06-05 02:12:49 -07:00
Nicolò Lucchesi and GitHub
d98b8f371c
[NixlConnector] Initiate deprecation cycle for kv_both role ( #43874 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-05 11:08:17 +02:00
Chao-Ju Chen and GitHub
e64237ae82
[Rust Frontend] Support include_reasoning=false ( #44391 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-06-05 16:47:50 +08:00
d61d8566ec
[Bugfix] Update mistral tokenizer test for continue_final_message fix ( #44622 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 16:13:26 +08:00
Uranus and GitHub
d2f70da116
fix: pad dummy run query_start_loc ( #44603 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-05 00:43:04 -07:00
6542d48964
[Bugfix] Fix test_invocations flaky failure with newer openai SDK ( #44618 )
...
Signed-off-by: Xu Zhou <xuzhou9417@163.com >
Co-authored-by: Xu Zhou <xuzhou9417@163.com >
2026-06-05 07:36:20 +00:00
Ting SUN and GitHub
ca73293fa6
[Bugfix][Rust Frontend] Fix UTF-8 char-boundary panic in incremental detokenizer ( #44620 )
...
Signed-off-by: Ting Sun <suntcrick@gmail.com >
2026-06-05 07:36:17 +00:00
Vic Wen and GitHub
ef3af56a97
Fix LLM.wait_for_completion output type docstring ( #44617 )
...
Signed-off-by: viiccwen <viiccwen@gmail.com >
2026-06-05 00:16:38 -07:00
b4a6f26c90
[ROCm][perf] Use workspace manager for sparse indexer allocations ( #41002 )
...
Signed-off-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com >
Co-authored-by: Stig-Arne Grönroos <stig-arne.gronroos@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:46:29 -07:00
165b7864d0
[ROCM] [FEAT] Integrate Aiter hipBLASLt GEMM online tuning ( #40426 )
...
Signed-off-by: hanlin12 <hanlin12@amd.com >
Signed-off-by: Han Lin <hanlin12@amd.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-06-04 23:45:36 -07:00
Li, Jiang and GitHub
c505cd93ef
[CI/Build] Disable CPU-Compatibility Tests ( #44605 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-05 13:14:26 +08:00
qizixi and GitHub
96229fa99e
[KVConnector][1/N] PP-aware handshake aggregation and intermediate-PP output plumbing ( #43720 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
2026-06-04 22:04:19 -07:00
da1daf40bf
[Bugfix] Exclude vision embedder from quantization in Gemma4 Unified ( #44571 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-04 20:47:38 -07:00
Woosuk Kwon and GitHub
4efd6ffde0
[DSV4] Refactor DeepseekV4Attention ( #44569 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-04 20:23:07 -07:00
Chris Leonard and GitHub
56aff0dd15
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to torch stale ABI. ( #44334 )
2026-06-04 20:14:43 -07:00
zofia and GitHub
063ce98fb7
[XPU][MoE] support block_fp8_moe on xpu ( #42139 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com >
2026-06-05 08:36:58 +08:00
Bugen Zhao and GitHub
62d6f06e3d
[Rust Frontend] Skip loading multimodal processor if --language-model-only is specified ( #44500 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-04 17:02:54 -07:00
Schwinn Saereesitthipitak and GitHub
b7c5baf63d
fix: keep DeepSeek V4 RoPE cache on inv_freq device ( #43926 )
...
Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com >
Signed-off-by: Schwinn Saereesitthipitak <17022745+galletas1712@users.noreply.github.com >
2026-06-05 02:30:29 +04:00
Jiangyun Zhu and GitHub
a55fccfc7c
[mamba] unify KDA conv states into one cache to match 2-state SSM layout ( #44539 )
2026-06-04 20:38:05 +02:00
Wentao Ye and GitHub
41a4829f22
[Logs Refactor] Optimize shutdown logs, easier to follow and consistent ( #43707 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-04 14:36:32 -04:00
38fd2405f3
use split_group for pytorch process group creation ( #41980 )
...
Signed-off-by: Tushar Jain <tushar00jain@users.noreply.github.com >
Co-authored-by: Tushar Jain <tushar00jain@users.noreply.github.com >
2026-06-04 14:36:07 -04:00
Agata Dobrzyniewicz and GitHub
a947f7a420
[Kernel][Test] Extend lightning_attn and awq_triton kernel tests to XPU ( #43307 )
...
Signed-off-by: Dobrzyniewicz, Agata <agata.dobrzyniewicz@intel.com >
2026-06-04 14:25:59 -04:00
bnellnm and GitHub
439203d32c
[Bugfix] Fix test_cutlass_moe.py ( #44380 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-06-04 14:18:52 -04:00
Taneem Ibrahim and GitHub
8d9536a775
[Misc] Add unit tests for pooler head classes ( #44471 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
2026-06-04 17:59:25 +00:00
Fadi Arafeh and GitHub
3da29aa4a5
[DOC] Add INT8 W4A8 docs and Arm's supported quantization schemes ( #34894 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
2026-06-04 16:27:17 +00:00
06f94633e7
[ROCm][CI] Add test for Aiter unified attn kernel ( #44436 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Divakar Verma <divakar.verma@amd.com >
Co-authored-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 16:15:05 +00:00
99ef652907
[Bugfix] Reject non-positive values for ParallelConfig int knobs ( #44057 )
...
Signed-off-by: jwzheng96 <jianweizheng@pku.edu.cn >
Signed-off-by: JianweiZheng <32029023+jwzheng96@users.noreply.github.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-06-04 11:46:50 -04:00
Tyler Michael Smith and GitHub
4cc78c9d5d
[Core] Freeze garbage collector in workers after model initialization ( #44363 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 08:39:04 -07:00
tc-mb and GitHub
3dbb4e0ace
[Bugfix] MiniCPM-V-4.6 video inference crash: placeholder count mismatches visual embedding count ( #44509 )
...
Signed-off-by: tc-mb <tianchi_cai@icloud.com >
2026-06-04 08:22:30 -07:00
b21443e23c
Add model support for granite speech plus ( #43519 )
...
Signed-off-by: Zvi Kons[WSL] <zvi@il.ibm.com >
Signed-off-by: Zvi Kons (BlueVela) <zvi@il.ibm.com >
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com >
2026-06-04 14:47:48 +00:00
Michael Goin and GitHub
06ee2d8433
[Quant] Support compressed-tensors WNA8O8Int linears and WNInt embeddings ( #44340 )
...
Signed-off-by: mgoin <mgoin64@gmail.com >
2026-06-04 07:40:33 -07:00
Yongye Zhu and GitHub
b5235fca2e
[DSv4] Adding TRTLLM gen attention kernel ( #43827 )
2026-06-04 07:35:09 -07:00
Andreas Karatzas and GitHub
3e77036768
[ROCm][CI] Specifying time outs for the lm eval models ( #44255 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:35:00 +08:00
Andreas Karatzas and GitHub
6f68ca3e91
[ROCm][CI] Stabilize memory-release in the Hybrid model generation tests ( #44046 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 22:34:24 +08:00
Turner Jabbour and GitHub
0c96dd64fb
[ROCm] Bump fastsafetensors to v0.3.2 from PyPI, remove git source build ( #43625 )
...
Signed-off-by: Turner Jabbour <doubleujabbour@gmail.com >
2026-06-04 07:30:57 -07:00
Nicolò Lucchesi and GitHub
68f5e565c9
[PD][Nixl] Mamba prefix caching mode support ( #42554 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 06:41:46 -07:00
QiliangCui2023 and GitHub
9354fb1ba5
[Bugfix][Compile] Guard per_token_group_fp8_quant lookup on non-CUDA platforms ( #44476 )
2026-06-04 09:31:50 -04:00
Harry Mellor and GitHub
f35b557239
Add GH token to docs build pre run check ( #44534 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-06-04 05:43:49 -07:00
Dipika Sikka and GitHub
e68988a248
Refactor CT NVFP4 linear to use a single class ( #42443 )
2026-06-04 08:25:08 -04:00
4b87b3e845
[Bugfix] fix EVS for qwen3-vl ( #44205 )
...
Signed-off-by: Rui "Garry" Gao <garrygaogg@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-06-04 11:06:51 +00:00
90619351e3
[Attention] Mamba attention module refactor - LINEAR ( #43556 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 18:45:29 +08:00
d0975a4b50
[perf] Add gemma RMS AR fusion ( #42646 )
...
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-04 01:33:59 -07:00
Kevin_Xiong and GitHub
1bdc60ed53
Fix Kimi-K2.5 FlashInfer ViT metadata ( #44493 )
...
Signed-off-by: Kevin-XiongC <kevin_xiong1997@outlook.com >
2026-06-04 08:14:35 +00:00
a6183563b6
[Prefix Caching] DeepSeekv4 - Support selective prefix-cache retention for sliding-window KV cache ( #43447 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
Co-authored-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-04 00:48:31 -07:00
Andreas Karatzas and GitHub
22c2e87555
[CI] Reverted gitignore changes ( #44497 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-04 00:37:44 -07:00
wang.yuqi and GitHub
d01d0b4646
[Frontend] Consolidate online serving utils. ( #44479 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-04 06:49:31 +00:00
b4b4aaa70e
[Inductor] Fast-path Inductor fallback for vllm::*/vllm_aiter::* custom ops ( #42129 )
...
Signed-off-by: Oxana Korzh <okorzh@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-04 00:03:52 -05:00
Andreas Karatzas and GitHub
5e2af28838
[CI] Resolve release V2 docker build after ROCm CI wheels change ( #44463 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-03 21:35:40 -07:00
4f423bd5bc
[EPLB] Nixl communicator optimization. Zero-copy transfers ( #41633 )
...
Signed-off-by: ilmarkov <markovilya197@gmail.com >
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
2026-06-04 03:40:34 +00:00
f0cd590d62
optimize the compressor 128 split cutedsl kernel ( #44230 )
...
Signed-off-by: Jie Fang <jief@nvidia.com >
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-03 20:22:57 -07:00
e6018c644a
[Refactor] Remove dead code in tests and parallel_state ( #41471 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 19:32:39 -07:00
f25952e59b
[MM][Perf][CG] Support ViT full CUDA graph for InternVL ( #41759 )
...
Signed-off-by: oguz <oguzhankir17@gmail.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-04 10:24:25 +08:00
maobaolong and GitHub
b58e082d95
[KV Connector] Update lmcache kv_offloading_backend to use LMCacheMPConnector ( #42865 )
...
Signed-off-by: baoloongmao <baoloongmao@tencent.com >
2026-06-03 19:23:55 -07:00
Ted Mostly and GitHub
0c1e6f63f5
[Bugfix] Fix VLLMNotFoundError when using LoRA adapter name in poolin… ( #44410 )
...
Signed-off-by: Ted Mostly <wanghenshui@qq.com >
2026-06-04 02:22:03 +00:00
Giancarlo Delfin and GitHub
ceb0111a90
[Model Runner V2][Spec Decode] Add Gemma4 MTP support ( #43241 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-06-04 00:51:06 +00:00
0414d75410
[XPU] skip unapplied UT in test_gpu_model_runner.py ( #44289 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-04 08:48:17 +08:00
128adabfe0
[Bugfix] Fix Gemma4 MTP block_table batch_size mismatch under concurrent load ( #43982 )
...
Signed-off-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
Co-authored-by: Dmytro Kuntso <dkuntso@amazon.co.uk >
2026-06-03 17:11:10 -07:00
bdbf08fc02
Bump actions/stale from 10.1.1 to 10.2.0 ( #35078 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-03 14:14:41 -07:00
Woosuk Kwon and GitHub
6bad553f4e
[Minor] Remove FlashInfer version check in topk_topp_sampler ( #44442 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 21:06:00 +00:00
91945b6e4a
[Bug Fix][Model Runner V2][Spec Decode] Warmup & capture with different attention states for speculator prefill ( #44253 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-03 13:32:40 -07:00
2b237c7a41
[Bugfix] Honor tool_choice="none" in Chat Completions streaming ( #42752 )
...
Signed-off-by: hoobnn <111053672+hoobnn@users.noreply.github.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 13:27:45 -07:00
Wentao Ye and GitHub
dad95e34d8
[Feature] Support batch invariant rms norm with residual ( #42453 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 15:22:01 -04:00
a248b45d05
[Model] Add Gemma4 Unified (encoder-free) support ( #44429 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-03 12:01:39 -07:00
linitra24 and GitHub
271328e256
[LoRA] Fix dedup for post-replacement module aliases ( #44413 )
...
Signed-off-by: bk-201 <joy25810@foxmail.com >
2026-06-03 18:23:23 +00:00
Wentao Ye and GitHub
2b91012650
[Refactor] Remove dead code fp quant ( #44122 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-06-03 14:22:23 -04:00
JartX and GitHub
5b2a2beade
[ROCm][CI] Move Model Executor test step from MI250 to MI300 (gfx942) ( #44370 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-03 12:23:51 -05:00
59d0236193
[10b/n] Migrate custom all-reduce, DeepSeek V4 fused MLA, MiniMax reduce-RMS, and MXFP8 MoE to libtorch stable ABI ( #44365 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Shengqi Chen <harry-chen@outlook.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-04 00:29:46 +08:00
0a5cbf633e
Handle spinloop ext load failure gracefully ( #43659 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-06-03 16:09:52 +00:00
Willow Lopez and GitHub
51e0c579b0
fix(config): validate max_num_scheduled_tokens >= 0 on all paths ( #44207 )
...
Signed-off-by: Oxygen56 <1391083091@qq.com >
2026-06-03 16:06:45 +00:00
0c6631f02a
[KVCache] Support Pluggable KVCacheSpec ( #37505 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: zjy0516 <riverclouds.zhu@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 09:05:16 -07:00
Nicolò Lucchesi and GitHub
df7252c343
[CI] Align PD tests to HMA on by default ( #44174 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-06-04 00:04:30 +08:00
Jee Jee Li and GitHub
4d1fd13613
[CI/Build] Fix LoRA testing ( #44425 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-03 08:58:06 -07:00
Nick Hill and GitHub
ec8d60bea8
[Model Runner V2] Use FlashInfer sampler ( #42472 )
2026-06-03 07:59:31 -07:00
27f1d34a23
[Frontend][Responses API] Move developer-to-system conversion into HF renderer ( #43590 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: kdcyberdude <kdsingh.cyberdude@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-06-03 14:52:24 +00:00
Flora Feng and GitHub
e3e132d2dd
[Refactor] Suppress SyntaxWarning from ast.literal_eval in tool parsers ( #44346 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 10:42:19 -04:00
e5232679a3
[XPU] Add XPU block-scaled W8A8 fp8 path ( #39968 )
...
Signed-off-by: Wu, Xiaochang <xiaochang.wu@intel.com >
Signed-off-by: Xiaochang Wu <xiaochang.wu@intel.com >
Co-authored-by: Yuxiang <yuxiang.liang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 20:16:19 +08:00
309385a359
[Rust Frontend] Add /server_info to Rust frontend ( #43942 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 04:30:47 -07:00
3d76f395e3
[SharedOffloadRegion] Align blocks to page-size ( #43689 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 14:25:57 +03:00
Li, Jiang and GitHub
823d271c0d
[Attention][CPU] Standardize kv layout to blocks first ( #44393 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-06-03 19:03:09 +08:00
Andy Lo and GitHub
95b1615ec9
[Perf] Improve multimodal item handling from O(n) to O(log n) per step ( #44212 )
...
Signed-off-by: Andy Lo <andy@mistral.ai >
2026-06-03 11:00:26 +00:00
1fa9ea09f6
[Perf] Triton fast path for small CPU→GPU swap_blocks_batch in the offloading connector ( #42212 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 13:38:17 +03:00
02564b4de0
[XPU]fallback to TRITON_ATTN for vit attn on xpu when use float32 dtype ( #43759 )
...
Signed-off-by: Yan Ma <yan.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-03 03:20:21 -07:00
Flora Feng and GitHub
209709a8c1
[Bugfix] Fix unstreamed tool call args dropped in Responses API streaming ( #44348 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-03 03:19:08 -07:00
ace95c9cf8
[Bugfix] Update TrtLLM MoE routing methods ( #44347 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 02:56:43 -07:00
Shanshan Shen and GitHub
0e2b13103b
[Doc] Update ViT CUDA graph interfaces ( #44388 )
...
Signed-off-by: shen-shanshan <467638484@qq.com >
2026-06-03 01:20:59 -07:00
Bugen Zhao and GitHub
449be4f934
[Rust Frontend] Fix several hf chat template rendering issues ( #44311 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 01:04:43 -07:00
6550ff12f2
[Rust Frontend] Add dynamic LoRA endpoints ( #43778 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:55:29 +00:00
4aaed4ca22
[Rust Frontend] Add server router extension hook ( #43774 )
...
Signed-off-by: NolanHo <kujyo.eia.serias@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-03 07:45:31 +00:00
7268457999
[KV Offloading] Enable HMA models for Tiering Offloading ( #44287 )
...
Signed-off-by: varun sundar rabindranath <vsundarr@redhat.com >
Co-authored-by: varun sundar rabindranath <vsundarr@redhat.com >
2026-06-03 10:03:00 +03:00
9af53a3c13
[Perf] Add tuned selective_state_update configs for H200 and RTX PRO … ( #44251 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-06-02 23:59:01 -07:00
Andreas Karatzas and GitHub
87954eb50e
[ROCm][CI] Optimize ROCm Docker build: registry cache, DeepEP, and ci-bake script ( #36949 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 23:43:07 -07:00
Charlie Fu and GitHub
71df063c49
Enable perf_token_group_quant/_C_stable_libtorch for ROCm ( #42758 )
...
Signed-off-by: charlifu <charlifu@amd.com >
2026-06-02 23:23:28 -07:00
Albert Cheng and GitHub
e0081ef8cf
[Benchmark] Enable reasoning-model (thinking) benchmarking via --chat-template-kwargs for client-rendered datasets ( #44244 )
...
Signed-off-by: Albert Cheng <albertching0112@gmail.com >
2026-06-02 22:49:51 -07:00
f0204358d9
[Bugfix] fix crash in postprocess for null tool args ( #43862 )
...
Signed-off-by: William-Rom <william.rom@intility.no >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 22:17:26 -07:00
Willow Lopez and GitHub
597bc15936
fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init ( #44236 )
...
Signed-off-by: Willow Lopez <100782273+Oxygen56@users.noreply.github.com >
2026-06-03 01:07:10 -04:00
Rotem Shavitt and GitHub
3f0a91bb96
Nit Changes in Tiered KV Offload ( #44293 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-06-02 21:53:21 -07:00
Flora Feng and GitHub
e67063826b
[CI] Add missing vllm/parser/ CI trigger and fix test_parse.py ( #44352 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 21:05:19 -07:00
Andreas Karatzas and GitHub
53b88d1dfc
[CI] Reject out-of-vocabulary before they reach the GPU logprob path ( #44042 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-02 22:27:52 -05:00
JartX and GitHub
7b476c8f14
[ROCm][CI] Skip fp8 reload tests on gfx90a (MI250) ( #44369 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:27:14 -05:00
JartX and GitHub
4454a18695
[ROCm][CI] Fix stale wvSplitK GEMM fallback test for N=5 ( #44368 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-02 22:00:25 -05:00
02a01496fc
[Platform] Add is_cumem_allocator_available ( #43838 )
...
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-03 10:54:50 +08:00
Kevin H. Luu and GitHub
27a93cd426
[docker] Stop using extra-index-url for flashinfer-jit-cache ( #44366 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-06-02 18:58:22 -07:00
Wei Zhao and GitHub
969aec4bc8
[Bugfix] Fix Deepseek v4 non-mega-moe model init error ( #44356 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-06-02 18:26:30 -07:00
ca17b6b17d
[Perf] Apply single-pass min_larger finding and binary search in Triton Top-p path. ( #42191 )
...
Signed-off-by: js_park <cakeng@naver.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 17:57:26 -07:00
Woosuk Kwon and GitHub
b254e0456c
[DSV4] Minor cleanup for DeepseekV4MegaMoEExperts ( #44367 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 17:54:27 -07:00
Daoyuan Li and GitHub
bd98e97557
[Misc] Remove dead VLLM_RPC_TIMEOUT env var and fix profiling doc that references it ( #44128 )
...
Signed-off-by: Daoyuan Li <94409450+DaoyuanLi2816@users.noreply.github.com >
2026-06-03 00:22:10 +00:00
a4ac746405
[MoE/b12x] Accept W4A16 (kNvfp4Static, None) in FlashInferB12xExperts supports check ( #43332 )
...
Signed-off-by: Junhao Shen <junshen@nvidia.com >
Co-authored-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-06-02 15:20:37 -07:00
8b3b71ee9d
[CI/Build] Bump flashinfer to v0.6.12 ( #44036 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-06-02 15:19:05 -07:00
Siddharth Bedekar and GitHub
0917a009d3
Fix sparse NCCL weight transfer test construction ( #44345 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
2026-06-02 21:51:21 +00:00
3099de3617
[Kernel][MoE] Add GELU_TANH to CPU, CUTLASS, and WNA16 MoE backends ( #42027 )
...
Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com >
Co-authored-by: lesj0610 <lesj0610@users.noreply.github.com >
2026-06-02 17:12:08 -04:00
Nick Hill and GitHub
e15f20258b
[ModelRunnerV2] Avoid pipeline parallel bubbles ( #42187 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 14:02:01 -07:00
557781131a
[Misc] Remove stray empty file ( #44350 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-02 12:53:03 -07:00
Yifan Qiao and GitHub
e9e08c49b9
[Bugfix] Cache the EAGLE/MTP lookahead block in the SWA prefix-cache mask ( #44082 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 12:21:07 -07:00
Woosuk Kwon and GitHub
e4a2e584e5
[MRV2] Remove assignment of graph_pool in cudagraph_utils ( #44338 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 11:50:27 -07:00
b8b49e2395
Bump actions/github-script from 8.0.0 to 9.0.0 ( #39667 )
...
Signed-off-by: dependabot[bot] <support@github.com >
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-02 11:26:57 -07:00
da107a59e5
[MRV2] Also enable MRV2 for Llama and Mistral dense models ( #43458 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: yewentao256 <zhyanwentao@126.com >
2026-06-02 11:18:46 -07:00
ed9a7526b6
[Anthropic] Support system role messages inside messages array ( #44283 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
Co-authored-by: Aleksandar Yanakiev <alexander.yanakiev@discretestack.com >
Co-authored-by: Ang Kah Min, Kelvin <syraxius@hotmail.com >
2026-06-02 18:13:54 +00:00
2427094152
[Feature] Support EPLB for DeepSeek v4 Mega Moe ( #43339 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
Co-authored-by: Wei Zhao (Engrg-Hardware 1) <weizha@login-lyris01.lyris.clusters.nvidia.com >
2026-06-02 10:56:44 -07:00
Kartavya sonar and GitHub
fe32e7830b
[Bugfix] flashinfer: fail fast when --kv-cache-dtype nvfp4 used on unsupported arch ( #43669 )
...
Signed-off-by: Kartavya Sonar <sonarkartavya@gmail.com >
2026-06-02 10:50:00 -07:00
afcb580715
[BugFix] Fix Humming MoE deploy error ( #43100 )
...
Signed-off-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Signed-off-by: Alireza Dadgarnia <49554709+adotdad@users.noreply.github.com >
Co-authored-by: Alireza Dadgarnia <dadgarnia@Alirezas-MacBook-Pro-2.local >
Co-authored-by: Jinzhen Lin <linjinzhen@hotmail.com >
2026-06-02 09:32:50 -07:00
3f3e2702c2
[XPU] Enable rms_norm/act quant fusions ( #43963 )
...
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 16:14:41 +00:00
Flora Feng and GitHub
478b49ddec
[Refactor] Remove dead code from parser infrastructure ( #44279 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 12:08:27 -04:00
Nick Hill and GitHub
cab5c9a2a9
[Core] Move max_concurrent_batches to VllmConfig ( #44274 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-02 08:57:25 -07:00
Brian Dellabetta and GitHub
774e552397
[compressed-tensors] Asymmetric support for MoE WNA16 marlin ( #44025 )
...
Signed-off-by: Brian Dellabetta <bdellabe@redhat.com >
2026-06-02 08:51:45 -07:00
XiaoZ and GitHub
53fa09d085
[Misc] Support local image encoding in benchmarks ( #43843 )
...
Signed-off-by: xiaoz <Sukra1@outlook.com >
2026-06-02 15:15:06 +00:00
Chris Leonard and GitHub
4d93bc35c9
Migrate header files to torch stable abi ( #44013 )
2026-06-02 08:09:52 -07:00
Bugen Zhao and GitHub
586201ebdc
[Rust Frontend] Cover different thinking modes in roundtrip tests ( #44320 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:51:25 -07:00
pschlan-amd and GitHub
88f172188b
[ROCm] Fix AITER RMSNormQuantFusion for Kimi-Linear ( #44308 )
...
Signed-off-by: Patrick Schlangen <pschlan@amd.com >
2026-06-02 14:50:21 +00:00
Bugen Zhao and GitHub
880fc032f4
[Rust Frontend] Support recursive tool parameter conversion ( #44299 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-06-02 07:45:35 -07:00
6314de8bad
[XPU] [Bug] remove xpuw4a16 output size check ( #44168 )
...
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 22:26:20 +08:00
IdoAtadTD and GitHub
c91a87f01a
[BugFix] [GDN] Read linear_key_head_dim from hf_text_config for multimodal models ( #43978 )
...
Signed-off-by: IdoAtadTD <ido.atad@twodelta.com >
2026-06-02 17:17:55 +03:00
Matthew Bonanni and GitHub
ea0d045a05
[FlashAttention] Sync FA with upstream ( #44065 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-06-02 07:15:37 -07:00
0bdfd5eb84
[Bugfix] Vendor MiniCPMV/MiniCPMO processors to unblock Transformers v5 ( #44282 )
...
Signed-off-by: guanwei-wu <b08901019@ntu.edu.tw >
Signed-off-by: wjinxu <1299461899@qq.com >
Co-authored-by: guanwei-wu <b08901019@ntu.edu.tw >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-02 07:14:38 -07:00
0cbc48c4f9
Support ModelOpt MXFP8 non-gated MoE ( #42958 )
...
Signed-off-by: tbarnatan <tbarnatan@nvidia.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-06-02 13:56:03 +00:00
2fd0e52252
[Bugfix] Fix Gemma4 startup crash with recent transformers multimodal processor ( #44232 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-06-02 13:42:40 +00:00
654bd2bca4
[Bugfix] Sync block_size from EngineCore to frontend for hybrid Mamba… ( #42967 )
...
Signed-off-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Amit Gruner <agruner@crusoe.ai >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-02 13:41:00 +00:00
wang.yuqi and GitHub
b623f7ea95
[Frontend] Consolidate dev entrypoints. ( #44170 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-02 06:30:21 -07:00
Shreyas Kulkarni and GitHub
0eeba5eec1
Fix DFlash prefix cache corruption due to missing lookahead block ( #42971 )
...
Signed-off-by: Shreyas Kulkarni <shreyas.gp269@gmail.com >
2026-06-02 12:06:33 +00:00
f69ede495b
[XPU][Mamba] Triton-based selective scan forward op for XPU ( #43421 )
...
Signed-off-by: Marceli Fylcek <marceli.fylcek@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:50:26 -07:00
Ronen Schaffer and GitHub
2a2b5ca791
[KV Offload] Add on_schedule_end() hook to separate step lifecycle from event draining ( #44206 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 13:42:52 +03:00
689b0eeb9e
[HARDWARE][POWER] Enable SHM communicator support for PowerPC ( #43754 )
...
Signed-off-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Signed-off-by: Rukhaiya <bibirukhaiya123@gmail.com >
Co-authored-by: Rukhaiya <rukhaiya@c643n08aix1-lp1.pok.stglabs.ibm.com >
Co-authored-by: Akash kaothalkar <61960177+Akashcodes732@users.noreply.github.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 18:06:32 +08:00
Isotr0py and GitHub
f8e9c56d15
[Multimodal] Automatically select registered video loader for VLM ( #44126 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-02 09:09:47 +00:00
alberto and GitHub
e30313220c
[Parser] Migrate ResponsesParser to unified Parser interface ( #42977 )
...
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com >
2026-06-02 08:50:05 +00:00
d247a9dc13
[EC Connector] Non blocking EC Connector lookup ( #41627 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-02 08:48:25 +00:00
Yifan Qiao and GitHub
7c37096620
[Core][Refactor]: thread scheduler_block_size into KVCacheManager and KVCacheCoordinator ( #44165 )
...
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai >
2026-06-02 01:14:44 -07:00
Maria Guevara and GitHub
b817b23f7b
[Rust Frontend] add --enable-request-id-headers flag support. ( #43883 )
...
Signed-off-by: Maria Guevara <kawaiiplush14@gmail.com >
2026-06-02 16:08:37 +08:00
Ronen Schaffer and GitHub
93da882e73
[kv_offload] Add @override decorators to subclass method implementations ( #44177 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-06-02 08:07:47 +00:00
0b25cf4419
[CPU][Perf] Enable fused kernels for GDN's gated delta rules ( #43534 )
...
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-02 08:00:48 +00:00
Jiangyun Zhu and GitHub
dcdfe66bfa
[Perf] use triton moe backend on hopper by default ( #44220 )
...
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com >
2026-06-02 15:52:30 +08:00
Flora Feng and GitHub
68dafcca75
[Refactor] Unify reasoning + tool-call parsing behind Parser.parse() ( #44267 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 15:11:42 +08:00
zhrrr and GitHub
1edfd09ffd
[Model Runner V2] Use actual batch max_seq_len for attn metadata ( #43991 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 06:07:56 +00:00
zhrrr and GitHub
8a9eb40808
[Model Runner V2] Support zeroing freshly allocated KV blocks for hybrid + fp8 KVCache ( #43990 )
...
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com >
2026-06-02 05:56:53 +00:00
f91fb2fcf3
[Bugfix] Convert Gemma4-MM ViT linear layers to vllm native impl ( #43798 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: ZiTian Zhao <zitian.zhao@tencentmusic.com >
Co-authored-by: B-201 <Joy25810@foxmail.com >
2026-06-01 21:41:16 -07:00
JooHo Lee and GitHub
a045c7425f
[MM][CG] Profile encoder CUDA graph pool memory ( #41714 )
...
Signed-off-by: JooHo Lee <jooho414@gmail.com >
2026-06-02 12:27:34 +08:00
a3a5a5ece5
[XPU][Bugfix] Fix per_token_group_fp8_quant missing dummy args on XPU ( #43930 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-06-02 03:09:21 +00:00
Or Ozeri and GitHub
480fadab1b
[BugFix][kv_offload]: Prevent offloading stale sliding window blocks ( #42959 )
...
Signed-off-by: Or Ozeri <oro@il.ibm.com >
2026-06-02 05:59:48 +03:00
279d25f5cb
[BugFix] Fix TypeError in MiniCPM-O audio feature unpadding ( #38053 )
...
Signed-off-by: Krishna Chaitanya Balusu <krishnabkc15@gmail.com >
Signed-off-by: wjinxu <1299461899@qq.com >
Signed-off-by: Kc Balusu <kcbalusu@users.noreply.github.com >
Co-authored-by: wjinxu <1299461899@qq.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
Co-authored-by: Kc Balusu <kcbalusu@users.noreply.github.com >
2026-06-01 19:57:28 -07:00
Andreas Karatzas and GitHub
54d0c36fff
[CI] Stabilize OpenAI schema fuzzing for malformed structural tags ( #44131 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-06-01 19:56:15 -07:00
Flora Feng and GitHub
9affc17a05
[Refactor] Move unstreamed tool-arg flush from serving layer to parser ( #44017 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-06-02 10:37:43 +08:00
Alec and GitHub
816cc73a9b
[Bugfix][CI] Normalize NIXL connector CUDA wheel installs ( #44266 )
...
Signed-off-by: Alec Flowers <aflowers@nvidia.com >
2026-06-01 19:34:05 -07:00
Micah Williamson and GitHub
2588ec4f0a
[ROCm] Upgrade AITER to v0.1.13.post1 ( #44265 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-06-02 01:48:59 +00:00
d68f0b220e
[Bugfix][Mooncake] Release GPU pin on failed store in MooncakeStoreConnector ( #43742 )
...
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-06-01 18:29:18 -07:00
Woosuk Kwon and GitHub
517e74a964
[DSV4] Refactor RoPE initialization ( #44262 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-02 01:26:58 +00:00
JartX and GitHub
48c0d13e65
[ROCm][CI] Skip unbacked dynamic shapes tests on PyTorch < 2.11 ( #44256 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-06-01 19:09:01 -05:00
Woosuk Kwon and GitHub
8c3cc98cff
[DSV4] Remove unncessary classes & functions ( #44246 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-06-01 14:43:00 -07:00
Nick Hill and GitHub
e4cbc4385d
[Test][BugFix] Fix double-BOS in PD+specdec acceptance test ( #44234 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:31:12 -07:00
Nick Hill and GitHub
6f8b40a23f
[BugFix][CI] Fix added _has_module tests ( #44248 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-06-01 14:23:12 -07:00
266b9d9c64
[Frontend][Core] Add sparse NCCL weight transfer support for in-place updates ( #40096 )
...
Signed-off-by: Siddharth Bedekar <bedeksid@gmail.com >
Co-authored-by: OpenAI Codex <codex@openai.com >
2026-06-01 15:37:30 -04:00
182c67daf1
[Rust Frontend] Support streaming generate endpoint ( #43779 )
...
Signed-off-by: xunzhuo <xunzhuo@vllm-semantic-router.ai >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 19:30:55 +00:00
fd9e91d7e4
[ROCm][CI] Fix and stabilize EAGLE3 acceptance tests ( #41294 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
Co-authored-by: Micah Williamson <micah.williamson@amd.com >
2026-06-01 12:40:01 -05:00
Yongye Zhu and GitHub
035733515f
[Kernel][DSv4] Optimize sparse FP8 compressor kernels ( #44161 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-06-02 00:18:32 +08:00
023808c23d
[Feature] Add support for JetBrains' Mellum v2 code generation model ( #43992 )
...
Signed-off-by: Madeesh Kannan <madeeswaran.kannan@jetbrains.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-06-01 10:11:35 -04:00
985c97a6a8
[Perf] Optimize cutlass fp8 scaled mm bypassing padding, 20% kernel performance improvement ( #43706 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-01 09:05:21 -04:00
Chaojun Zhang and GitHub
bd0aecdc08
[XPU][CI] Fix test_audio_in_video flake by using module-scoped server fixture ( #44146 )
...
Signed-off-by: Chaojun Zhang <chaojun.zhang@intel.com >
2026-06-01 11:21:36 +00:00
8796838910
[Bugfix] fix wrong partial_rotary_factor calculation for bailing_moe model. ( #43770 )
...
Signed-off-by: zzt <zengzetang.zzt@antgroup.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-06-01 02:42:49 -07:00
de21863419
[Rust Frontend] Add InternLM2 tool parser ( #43481 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Bugen Zhao <i@bugenzhao.com >
2026-06-01 08:58:46 +00:00
wang.yuqi and GitHub
0910f7e0e1
[Frontend] Resettle generative scoring entrypoint. ( #44153 )
...
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-06-01 07:54:59 +00:00
Uranus and GitHub
1f6048abe5
fix: glm5.1 pp model loading ( #42944 )
...
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com >
2026-06-01 15:14:47 +08:00
98f1279815
[CPU][RISC-V] Add missing RVV cpu_types helpers for WNA16 ( #42730 )
...
Signed-off-by: wcy <233313160abc@gmail.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-06-01 14:56:41 +08:00
Isotr0py and GitHub
1fd8bd02a4
[Docs] Replace broken video url in examples ( #44159 )
...
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
2026-06-01 06:01:10 +00:00
29d69332aa
[BugFix] Fix _has_module to verify native deps via trial import ( #44035 )
...
Signed-off-by: esmeetu <jasonailu87@gmail.com >
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: esmeetu <jasonailu87@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-31 22:06:33 -07:00
Lucas Wilkinson and GitHub
4721bb3aa4
[MRV2] Remove Eagle's dedicated CUDA graph pool ( #44078 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
2026-05-31 22:00:33 -07:00
Umut Polat and GitHub
f46e6be169
[Misc] Use VLLMValidationError consistently in chat completion and completion protocol validators ( #36254 )
...
Signed-off-by: umut-polat <52835619+umut-polat@users.noreply.github.com >
2026-06-01 04:04:11 +00:00
8b8546da1c
docs: fix MLA attention docstring examples ( #44118 )
...
Co-authored-by: nightcityblade <nightcityblade@gmail.com >
2026-05-31 12:28:38 -07:00
Jee Jee Li and GitHub
6bdabbad5b
[CI/Build] Enable Step3p7ForConditionalGeneration testing ( #43956 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-31 05:16:12 +00:00
3fd9d2d357
[CPU][Zen] Route W8A8 and W4A16 linear inference through zentorch on AMD Zen CPUs ( #41813 )
...
Signed-off-by: R <Ganesh.R@amd.com >
Signed-off-by: Harshal Adhav <harshal.adhav@amd.com >
Signed-off-by: Aakar Dwivedi <aadwived@amd.com >
Co-authored-by: R <Ganesh.R@amd.com >
Co-authored-by: Harshal Adhav <harshal.adhav@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
2026-05-30 14:17:21 -05:00
Woosuk Kwon and GitHub
27fa5aa3b9
[MRV2] Support breakable CUDA graph ( #44050 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-30 09:40:52 -07:00
e1105064b2
[Bug] Fix gemma4 MTP IMA issue when TP>1, CUDA error: an illegal memory access was encountered ( #43909 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-30 10:34:33 -04:00
Bugen Zhao and GitHub
50c80d7923
[Governance] Add @BugenZhao as Rust frontend code owner ( #44047 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-30 22:23:54 +08:00
3becc5db40
[ROCm] Add attention sink support to AITer flash attention backend ( #43817 )
...
Signed-off-by: Xiaoran Chen <xiaoran@fb.com >
Co-authored-by: Xiaoran Chen <xiaoran@fb.com >
2026-05-30 18:13:18 +08:00
124fac10cb
[Bugfix] Fix RMSNorm kernels to multiply in weight's native dtype ( #42379 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 23:16:53 -07:00
e9499996df
[BugFix][Platform] Fix import vllm.platforms.rocm error on non-CUDA test_gpt_oss.py ( #43571 )
...
Signed-off-by: Ma, Liangliang <liangliang.ma@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 23:16:49 -07:00
c0056b19bf
[ROCm] cmake: support PYTORCH_FOUND_HIP for torch 2.13 native HIP language support ( #43881 )
...
Signed-off-by: nemanjaudovic <nudovic@amd.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 22:16:57 -07:00
Andreas Karatzas and GitHub
ef8840adc7
[ROCm][CI] Fix failure in the Phi3V pooling test ( #44028 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-30 12:14:37 +08:00
Flora Feng and GitHub
1a096d8208
[Refactor] Remove dead current_tool_name_sent assignments from tool parsers ( #43997 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 21:45:15 -04:00
Gagan Dhakrey and GitHub
1e2ce5d11a
offload prompt_embeds decode in render_prompts_async to avoid blocking ( #43792 )
...
Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com >
2026-05-30 01:36:34 +00:00
559d6710bf
[PERF]MiniMax-M2 gate kernel ( #38445 )
...
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com >
Signed-off-by: qianlihuang <91178480+qianlihuang@users.noreply.github.com >
Co-authored-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com >
2026-05-29 18:28:34 -07:00
bnellnm and GitHub
187457a952
Revert "[MoE Refactor] Migrate MoeWNA16Method quantization to MK orac… ( #44033 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 16:45:29 -07:00
8fad266507
[CI] Fix smoke test step key to bypass block gate ( #43974 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 16:28:32 -07:00
Flora Feng and GitHub
8c6daf6e2f
[CI] Remove duplicate Harmony test coverage ( #44023 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 22:52:46 +00:00
bnellnm and GitHub
7b98f498cd
[MoE Refactor] Remove supports_expert_map ( #43108 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
2026-05-29 17:26:56 -04:00
106aa92f04
[MoE Refactor] Migrate MoeWNA16Method quantization to MK oracle ( #42647 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 17:19:31 -04:00
yzong-rh and GitHub
46409fd2a1
[Fronten] Clean up stop_token_ids override for Harmony ( #44009 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-29 13:28:06 -07:00
38b864d81d
[Metrics] Exclude KV transfer tokens from iteration_tokens_total ( #43346 )
...
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-29 19:56:44 +00:00
Wentao Ye and GitHub
5dbf1605a0
[Feature] SSL support for dp supervisor ( #43688 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 19:28:12 +00:00
Kevin H. Luu and GitHub
acbc203340
Add @khluu to CODEOWNERS ( #44019 )
...
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
2026-05-29 12:24:29 -07:00
Flora Feng and GitHub
6de08e8b46
[CI] Remove redundant test_chat_with_tool_reasoning.py ( #44011 )
...
Signed-off-by: sfeng33 <4florafeng@gmail.com >
2026-05-29 19:23:56 +00:00
6aabe221a5
[CI] Make Model Executor test hangs fail fast with a traceback ( #43971 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 11:58:25 -07:00
Wentao Ye and GitHub
739096a028
[Bug] Fix torch device issue for MOE permute ( #44005 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-29 18:55:00 +00:00
czhu-cohere and GitHub
8b9deeec4b
[Bugfix] Fix Ray placement group allocation with grouped nodes ( #43998 )
...
Signed-off-by: <conway.zhu@cohere.com >
Signed-off-by: root <conway.zhu@cohere.com >
2026-05-29 12:51:05 -06:00
d07ad0693b
[Bugfix] Use storage_block_size in KV cache reshape for compressed specs (DeepSeek V4) ( #43988 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-05-29 11:14:25 -07:00
4aaba00f92
[EPLB] Make async EPLB default ( #43219 )
...
Signed-off-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Markov Ilya <markovilya19@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-29 18:07:16 +00:00
84b2a8a7e7
[MoE Refactor] WNA16 MoE backend selection into oracle module ( #42553 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 13:11:17 -04:00
4ff865c38e
[Bugfix] Disable allreduce_rms_fusion when pipeline_parallel_size > 1 ( #43616 )
...
Signed-off-by: zixi-qi <zixi@inferact.ai >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-29 22:57:43 +08:00
5502c3b52d
[Misc] added unit tests for the core pooling methods ( #43818 )
...
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-29 14:40:31 +00:00
Chunyang Wen and GitHub
f191d5630e
docs: clarify ITL acronym in optimization docs ( #43922 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-29 07:40:05 -07:00
11dfa3169d
Add vLLM library info to Hugging Face Hub requests ( #43857 )
...
Signed-off-by: Wauplin <lucainp@gmail.com >
Signed-off-by: Lucain Pouget <lucain@huggingface.co >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-05-29 14:04:58 +00:00
Li, Jiang and GitHub
3f6f508e14
[Bugfix][CPU] Remove invalid extra deps ( #43977 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-29 22:02:09 +08:00
Harry Mellor and GitHub
0585b5ba2e
Skip docs build if PR doesn't affect docs ( #43972 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-29 12:09:52 +00:00
Thien Tran and GitHub
d2889722ff
[Bugfix] Corrupted MLA + linear attention ( #43961 )
...
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg >
2026-05-29 05:00:51 -07:00
0b56815a24
[ROCm][Perf] DSv3.2 MI355X TP4 decode-step orchestration cleanup (3 micro-opts) ( #42982 )
...
Signed-off-by: Frida Andersson <fanderss@amd.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-29 04:26:57 -07:00
ab12aab127
[Bugfix] [ROCm] [DSV4] Fix AITER MXFP4 MoE weight loading and shuffle… ( #42595 )
...
Co-authored-by: MHYangAMD <MHYangAMD@users.noreply.github.com >
2026-05-29 04:08:33 -07:00
JartX and GitHub
0cff0741ff
[Kernel][ROCm] Native W4A16 kernel for AMD RDNA3 (gfx1100) — fp16 + bf16 ( #41394 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
2026-05-29 11:04:40 +00:00
60a7a2214f
[Bugfix] Fix Step3 pipeline parallel KeyError for residual tensor ( #37622 )
...
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com >
2026-05-29 03:04:02 -07:00
Nicolò Lucchesi and GitHub
7ebc0ec104
[CI] Nixl+SimpleCPUOffloadingConnector unit tests ( #43871 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-29 02:40:42 -07:00
e8b5199973
[XPU] support MTP of gdn attention ( #43565 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 17:10:24 +08:00
Simon Danielsson and GitHub
b7fb747d8d
[CI][ROCm] Don't skip MoRI-IO Connector tests ( #43703 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-29 17:06:23 +08:00
Kunshang Ji and GitHub
30c6289b8e
[XPU] fix xpu install document triton-xpu version ( #43947 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 02:05:12 -07:00
Andreas Karatzas and GitHub
ff990d0d32
[ROCm][CI] Fix AITER unified attention for encoder-decoder cross-attention ( #43945 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-29 16:43:39 +08:00
Chauncey and GitHub
87f12e5c7c
[Frontend]Responses API supports chat_template_kwargs ( #43761 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-29 07:58:19 +00:00
kliuae and GitHub
ab7521d77c
[ROCm][DSv4] Remove device pipeline stall in sparse attention ( #43898 )
...
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
2026-05-29 15:42:40 +08:00
94d3f4d205
[CPU Backend] CPU top-k and top-p sampling kernels using Triton ( #43633 )
...
Signed-off-by: Li, Tianmu <tianmu.li@intel.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 15:02:39 +08:00
04516eabc8
[XPU] add gelu_tanh to xpu moe backend supported activations ( #42822 )
...
Signed-off-by: yintong-lu <yintong.lu@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 14:37:20 +08:00
648c3ebee6
[CI] Separate non-root smoke tests from image build step ( #43712 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 23:34:16 -07:00
22a58640b4
[9/n] Migrate attention and cache kernels to torch stable ABI (continued) ( #43717 )
...
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-29 04:44:45 +00:00
710f077617
[Refactor] Remove dead code ( #43234 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 00:29:56 -04:00
d63108fb18
[kv_offload] Skip decode-phase blocks in CPU offload ( #43797 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-05-29 06:39:43 +03:00
9636709372
[XPU] add scale transpose to prepare_fp8_moe_layer_for_xpu and bump up kernels ( #43277 )
...
Signed-off-by: mayuyuace <qiming1.zhang@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-29 03:22:51 +00:00
Weida Hong and GitHub
dfe8ba7c80
Adjust design around encoder_cudagraph_forward ( #42288 )
...
Signed-off-by: Weida Hong <wdhongtw@google.com >
2026-05-29 03:02:52 +00:00
212deff2ec
[feat] add GlmgaProcessor specific logits in glm4_1v.py ( #43575 )
...
Signed-off-by: JaredforReal <w13431838023@gmail.com >
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
2026-05-29 02:56:02 +00:00
Woosuk Kwon and GitHub
7bd45da585
[DSv4] Move mHC tilelang kernels & Don't use CustomOP in dsv4/nvidia ( #43905 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-29 10:25:02 +08:00
bf18d7e0b4
[Misc][NUMA] Auto-bind to PCT priority cores on DGX B300 + widen EngineCore across shard NUMA nodes ( #43270 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Cursor <noreply@cursor.com >
2026-05-29 10:07:44 +08:00
Bugen Zhao and GitHub
1521173c17
[Rust Frontend] Add /version endpoint using engine-reported value ( #43854 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-29 00:32:27 +00:00
b690b2bb67
[Model]Support Step-3.7-Flash ( #43859 )
...
Signed-off-by: luotingdan <luotingdan@stepfun.com >
Signed-off-by: Isotr0py <Isotr0py@outlook.com >
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
Co-authored-by: luotingdan <luotingdan@stepfun.com >
Co-authored-by: Isotr0py <Isotr0py@outlook.com >
Co-authored-by: Yu Huang <yuhuang@nvidia.com >
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-28 17:01:48 -07:00
yzong-rh and GitHub
325a1ec4fb
[CI] Enable prefix caching in BFCL benchmark ( #43925 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
2026-05-28 23:36:31 +00:00
69c9f19957
fix(frontend): Add multimodal placeholders to Gemma4 tool message template ( #41459 )
...
Signed-off-by: Harshal Janjani <harshaljanjani@gmail.com >
Co-authored-by: Ben Browning <bbrownin@redhat.com >
2026-05-28 14:48:12 -07:00
rasmith and GitHub
9769e2df2a
[AMD][CI][BugFix] Fix Distributed Compile Unit Tests (2xH100-2xMI300) group ( #43120 )
...
Signed-off-by: Randall Smith <Randall.Smith@amd.com >
2026-05-28 14:39:01 -07:00
Michael Goin and GitHub
03f03f9630
Refactor output filename handling in ci-fetch-log.sh ( #43901 )
...
Signed-off-by: Michael Goin <mgoin64@gmail.com >
2026-05-28 14:20:12 -07:00
Benjamin Chislett and GitHub
9202ea6fda
[Spec Decode] Allow causal DFlash ( #43445 )
2026-05-28 21:18:44 +00:00
Woosuk Kwon and GitHub
69b8956dcd
[Model Refactoring] Remove unncessary torch op registration for DSv4 ( #43891 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:04:55 -07:00
a3ed5ab10c
[KV Offload] Add per-request offloading policy via on_new_request lifecycle hook ( #43205 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 20:45:18 +00:00
7e53283b1c
[Core] Cleanup KVConnector handling with PP + fix MRV2 ( #43732 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 13:12:03 -07:00
9090368b65
[Feat] Add support for per GPU worker RDMA NIC selection ( #42083 )
...
Signed-off-by: Raj Joshi <rajjoshi@redhat.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-05-28 12:45:23 -07:00
Harry Mellor and GitHub
085ac221a3
Deprecate JAISLMHeadModel ( #43784 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 18:29:12 +00:00
Hua Huang and GitHub
9006204e90
[MM][CG] Avoid over-padding Qwen2.5-VL encoder cudagraph window metadata ( #42796 )
...
Signed-off-by: Hua Huang <huah@nvidia.com >
2026-05-28 11:22:56 -07:00
ed7fe831da
[ROCm] Enable the aiter top-k/top-p sampler by default ( #43331 )
...
Signed-off-by: John Qin <yanyuan.qin@amd.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-28 13:19:59 -05:00
Nicolò Lucchesi and GitHub
5b115bb8a3
[Attention][AMD] Standardize kv layout to blocks first for AMD ( #43660 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
2026-05-28 12:28:50 -05:00
53a2088675
Allow native KV cache dtype in Triton cache update ( #43330 )
...
Signed-off-by: Michael Gschwind <mgschwind@nvidia.com >
Co-authored-by: Michael Gschwind <mgschwind@nvidia.com >
2026-05-28 16:51:40 +00:00
Chao-Ju Chen and GitHub
099024762c
[Rust Frontend] Optimize multimodal prompt expansion ( #43670 )
...
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com >
2026-05-28 09:46:18 -07:00
9aa131f944
Add Cosmos3 Reasoner model ( #43356 )
...
Signed-off-by: Maciej Bala <mbala@nvidia.com >
Signed-off-by: MaciejBalaNV <mbala@nvidia.com >
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn >
Co-authored-by: Isotr0py <2037008807@qq.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io >
2026-05-28 09:43:55 -07:00
Micah Williamson and GitHub
1b5437cec8
[ROCm] Bump ROCm to 7.2.3 ( #43136 )
...
Signed-off-by: Micah Williamson <micah.williamson@amd.com >
2026-05-28 09:42:43 -07:00
3207e7680e
[XPU][MoE] Add WNA16 oracle backend for GPTQ sym-int4 (xpu_fused_moe) ( #41426 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-28 16:30:48 +00:00
Matthias Gehre and GitHub
a9ec46d4b7
[ROCm][Perf] Support N=5 in wvSplitK skinny GEMM kernels for speculative decoding ( #40687 )
...
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com >
2026-05-28 16:28:21 +00:00
Ronen Schaffer and GitHub
4bfa0f2b14
[KV Offload] Rename SecondaryTierManager.get_finished() to get_finished_jobs() ( #43870 )
...
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com >
2026-05-28 16:00:18 +00:00
Vadim Gimpelson and GitHub
5d126dd155
[Bugfix] Exclude Ray DP from #42585 's deferred port allocation ( #43864 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
2026-05-28 15:55:14 +00:00
c08ebebf30
[Perf] Add do_not_specialize to Mamba SSD chunk kernels ( #43803 )
...
Signed-off-by: Majid Taheri Andani <tahemaji@amazon.com >
Co-authored-by: Majid Taheri Andani <tahemaji@amazon.com >
2026-05-28 15:40:02 +00:00
Wentao Ye and GitHub
be4062fd6c
[Bug] Fix tests/distributed/test_elastic_ep.py - assert False ( #43813 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 11:00:56 -04:00
577d693838
[rust] fix: aggregate is_sleeping and reset_prefix_cache across DP engines ( #43429 )
...
Signed-off-by: Will.hou <1205157517@qq.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-28 07:56:56 -07:00
Bugen Zhao and GitHub
61a1e30473
[Rust Frontend] Reduce Gemma4 tool parser args scan complexity ( #43850 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:52:29 +00:00
Bugen Zhao and GitHub
3a282230ee
[Rust Frontend] Add hy_v3 tool parser ( #43872 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 14:42:47 +00:00
Li, Jiang and GitHub
20d69d100a
[CPU] Migrate cpu_awq into awq_marlin ( #43841 )
...
Signed-off-by: jiang1.li <jiang1.li@intel.com >
2026-05-28 22:36:31 +08:00
Simon Danielsson and GitHub
552eb81918
[Bugfix][ROCm] Resolve MoRI connector hangs at high concurrency ( #40344 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
2026-05-28 14:30:21 +00:00
Woosuk Kwon and GitHub
9957e4d240
[Model Refactoring] Remove torch compile dependency in DSv4 ( #43746 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 14:26:25 +00:00
864990e8d9
Add token-offset based selective offload in OffloadConnector ( #39983 )
...
Signed-off-by: Angelo Ruocco <ang@zurich.ibm.com >
Co-authored-by: Or Ozeri <or@ozery.com >
2026-05-28 14:11:02 +00:00
f3b2a819f7
[Perf][KDA] Fuse gate softplus, chunk-local cumsum, and RCP_LN2 scaling ( #43667 )
...
Signed-off-by: haojiangzheng <justineric096@gmail.com >
Co-authored-by: haojiangzheng <justineric096@gmail.com >
2026-05-28 13:47:08 +00:00
Wentao Ye and GitHub
64e1218673
[Perf] Optimize moe permute by pre-allocate buffer, 9~14% kernel performance improvement ( #43014 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-28 06:18:26 -07:00
Julien Denize and GitHub
02606b0b09
[BUGFIX] Multimodal benchmark with MistralTokenizer ( #42965 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
Signed-off-by: Julien Denize <40604584+juliendenize@users.noreply.github.com >
2026-05-28 05:36:24 -07:00
19af4e6dd4
Fix OlmoHybridForCausalLM not initialising ( #43846 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-28 05:33:31 -07:00
omerpaz95 and GitHub
811d805195
[EC Connector] Add shutdown API to EC Connector. ( #42423 )
...
Signed-off-by: omerpaz95 <omerpaz95@gmail.com >
2026-05-28 12:28:01 +00:00
Vadim Gimpelson and GitHub
c1c4db8b4b
Log dummy DP step in iteration details ( #41406 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
2026-05-28 12:18:39 +00:00
Chauncey and GitHub
d692b89c2c
[Feature] Add structured output and effort support to Anthropic Messages API ( #42396 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-28 12:06:48 +00:00
Bugen Zhao and GitHub
8e0580f4ee
[CI] Auto-apply rust label to relevant PRs ( #43866 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 11:57:22 +00:00
61288b5458
[Bugfix] Fix HyperCLOVAX CI failure after upstream removed remote code ( #43860 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-28 03:37:36 -07:00
a583c84e2b
[Bugfix][ROCm] Fix Accuracy Drop in Sparse Indexer on gfx950 ( #43781 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com >
Signed-off-by: kliuae <kuanfu.liu@embeddedllm.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: vllmellm <vllm.ellm@embeddedllm.com >
2026-05-28 03:37:15 -07:00
4ec2817313
[Model][Bugfix] Rename weight_mapper to hf_to_vllm_mapper in LlamaNemotronVL pooling models ( #43581 )
...
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com >
Co-authored-by: opencode <noreply@opencode.ai >
Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com >
2026-05-28 03:32:22 -07:00
Wei Zhao and GitHub
f2caefe226
[UX] Increase DP Coordinator startup timeout from 30s to 120s ( #42343 )
...
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com >
2026-05-28 03:31:45 -07:00
Animesh Trivedi and GitHub
bfb9ebc211
[Feature] Add support for timed trace replay in vllm bench serve to replay Moonshot and Alibaba workload traces ( #39795 )
...
Signed-off-by: Animesh Trivedi <Animesh.Trivedi@ibm.com >
2026-05-28 03:31:34 -07:00
Andreas Karatzas and GitHub
a9bc0ad8e4
[ROCm][CI] Move workload from MI300 to MI325 ( #43824 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 03:31:29 -07:00
b372ad3e90
[Bugfix] Stream DeepSeek DSML tool-call argument deltas incrementally ( #42879 )
...
Signed-off-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: QwertyJack <7554089+QwertyJack@users.noreply.github.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-28 17:50:23 +08:00
Harry Mellor and GitHub
2a781756a1
Restore Literal for WeightTransferConfig.backend ( #43183 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-28 09:39:41 +00:00
Woosuk Kwon and GitHub
a04afd76aa
[DSV4] Remove AMD/XPU path in deepseek_v4/nvidia ( #43829 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-28 08:00:52 +00:00
6cc8577421
[Kernel] Marlin MoE: include SM 12.x in default arch list ( #40923 )
...
Signed-off-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: Tony Liu <tonyliu0512@gmail.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-28 15:30:26 +08:00
d6b48f928f
[BugFix] Fix hard-coded timeout for multi-API-server startup ( #43768 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:09:13 -07:00
Rotem Shavitt and GitHub
1b16f2ddc9
change name of fs_python secondary tier to fs. ( #43600 )
...
Signed-off-by: Rotem Shavitt <rshavitt@gmail.com >
2026-05-28 07:05:48 +00:00
TJian and GitHub
0ba46d4b11
[ROCm][DSV4] Enable Tilelang MHC replacing torch/triton mhc ( #43679 )
...
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
2026-05-28 07:05:28 +00:00
JINO ROHIT and GitHub
e1814f822d
minor docs: fix incorrect example path ( #43830 )
...
Signed-off-by: JINO-ROHIT <find.jinorohit@gmail.com >
2026-05-27 22:58:43 -07:00
7909f82a45
[Bugfix][Frontend] streaming tool-call serializer drops first args chunk when name and args share a DeltaMessage ( #42683 )
...
Signed-off-by: ignaciosica <mignacio.sica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-28 05:20:55 +00:00
Nick Hill and GitHub
626fa9bba5
[BugFix] Fix blocked reasoning parsing with MRV2 ( #43808 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 04:59:34 +00:00
Thien Tran and GitHub
e54eff769d
[Bugfix] Pass routed_scaling_factor to FlashInfer TRTLLM BF16 MoE ( #43769 )
2026-05-27 21:29:14 -07:00
05ac829629
fix: parse Qwen3 XML JSON arguments first ( #43243 )
...
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-28 03:35:59 +00:00
Andreas Karatzas and GitHub
33e94fc3ad
[ROCm][CI] Stabilize Cargo cache and pre-test image checks ( #43815 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-28 11:24:44 +08:00
413ac5c070
[Misc][Rocm] Remove redundant AiterUnifiedAttentionBackend block size log ( #43664 )
...
Signed-off-by: NickLucche <nlucches@redhat.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 22:19:11 -05:00
Yongye Zhu and GitHub
2d2c660104
[MoE] Remove inplace fused experts mechanism ( #43727 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 20:00:19 -07:00
Benjamin Bartels and GitHub
05eec7120e
Fix RunAI streamer tensor buffer reuse during weight loading ( #43464 )
...
Signed-off-by: bbartels <benjamin@bartels.dev >
2026-05-27 19:16:52 -07:00
Bugen Zhao and GitHub
c87f62ccf8
[Rust Frontend] Introduce mock engine for benchmark baseline ( #43469 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-28 01:40:35 +00:00
1223732dda
[ModelRunnerV2][Hybrid model] Support kernel block size in hybrid model ( #38831 )
...
Signed-off-by: MengqingCao <cmq0113@163.com >
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Signed-off-by: Mengqing Cao <cmq0113@163.com >
Co-authored-by: Nick Hill <nickhill123@gmail.com >
2026-05-28 00:55:55 +00:00
amitz-nv and GitHub
381edde1b9
[Bugfix][Kernel] TRTLLM NVFP4 MoE chunking ( #43599 )
...
Signed-off-by: amitz-nv <203509407+amitz-nv@users.noreply.github.com >
2026-05-28 00:36:21 +00:00
Andreas Karatzas and GitHub
094124af15
Add @AndreasKaratzas to CODEOWNERS ( #43740 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-27 16:14:50 -07:00
Dakai An and GitHub
5963c19478
Fix Qwen3-VL and Qwen3-omni-thinker accuracy degradation from deepstack inputs under torch.compile ( #43617 )
...
Signed-off-by: Dakai An <dakaian108@gmail.com >
2026-05-27 15:34:08 -07:00
7fb9c0197a
[Bugfix][DFlash]allocate the proper number of lookahead slots ( #43733 )
...
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com >
Signed-off-by: Benjamin Chislett <chislett.ben@gmail.com >
Co-authored-by: Nicolò Lucchesi <nicolo.lucchesi@gmail.com >
2026-05-27 21:45:34 +00:00
Harry Mellor and GitHub
2c2c966669
Validate against some config fields being set to 0 ( #43794 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 21:14:49 +00:00
Harry Mellor and GitHub
2616f67faa
Remove Transformers forward/backward compatibility tests ( #43785 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 12:46:36 -07:00
206b72c982
[Quantization] Fix Humming RoutedExperts import ( #43540 )
...
Signed-off-by: Minh Vu <vuhoangminh97@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 10:51:56 -07:00
284e6f543d
[8/n] Migrate merge_attn_states, mamba, sampler to torch stable ABI (continued) ( #43361 )
...
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Signed-off-by: Chris Leonard <chleonar@redhat.com >
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com >
Co-authored-by: Shengqi Chen <harry-chen@outlook.com >
2026-05-27 09:35:24 -07:00
jatseng-ai and GitHub
05c50c721e
[ROCm] mori: add InterNodeV1LL inter-node kernel selection via VLLM_MORI_INTERNODE_KERNEL ( #41751 )
...
Signed-off-by: jatseng-ai <jatseng@amd.com >
2026-05-28 00:33:32 +08:00
Harry Mellor and GitHub
41688e2dc7
Fix early CUDA init ( #43791 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-27 09:30:11 -07:00
Chunyang Wen and GitHub
49a3510266
[Docs] Fix the duplicate doc icon issue ( #43546 )
...
Signed-off-by: chunyang.wen <chunyang.wen@gmail.com >
2026-05-27 16:09:58 +00:00
Injae Ryou and GitHub
165460941f
[BugFix] HFValidationError with cloud storage URIs when HF_HUB_OFFLINE=1 ( #39155 )
...
Signed-off-by: Injae Ryou <injaeryou@gmail.com >
2026-05-27 10:53:32 -05:00
Yongye Zhu and GitHub
03d9cc2fe2
[misc] Bump cutedsl version to 4.5.2 ( #43745 )
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-27 08:25:36 -07:00
52a31ccecc
[Bugfix] Map reasoning_effort to enable_thinking in chat template kwargs ( #43401 )
...
Signed-off-by: Ashwin Giridharan <girida@amazon.com >
Signed-off-by: Chauncey <chaunceyjiang@gmail.com >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 05:39:49 -07:00
2272062471
[Kernel] Enable TritonW4A16LinearKernel as CUDA fallback for non-Marlin-aligned W4A16 shapes ( #43731 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 18:36:27 +08:00
Mohammad Miadh Angkad and GitHub
158289e0fc
[Docs] Fix MLA prefill backend default docs ( #43697 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-27 10:13:22 +00:00
Bugen Zhao and GitHub
396c8fee50
[Rust Frontend] Align tool parser fallback behavior between streaming & non-streaming paths ( #43662 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 10:13:12 +00:00
ad464e16c0
[Doc] Add Ascend NPU tab to the quickstart installation guide ( #43550 )
...
Signed-off-by: Aditya Singh <adisin650@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-27 08:41:29 +00:00
akii96 and GitHub
de12f5ca0b
[ROCm][GPT-OSS] Avoid repeated compile-time cos_sin_cache.to(bf16) casts in rotary path ( #42833 )
...
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com >
2026-05-27 16:22:27 +08:00
683033d4ba
[Frontend] Add MiniCPM5 XML tool call parser ( #43175 )
...
Signed-off-by: zhangtao <zhangtao2@modelbest.cn >
Signed-off-by: zhangtao2 <zhangtao2@modelbest.cn >
Co-authored-by: zhangtao <zhangtao2@modelbest.cn >
Co-authored-by: Chauncey <chaunceyjiang@gmail.com >
2026-05-27 00:39:35 -07:00
8c94938cfb
[MRV2][BugFix] Fix KV connector handling in spec decode case ( #43719 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com >
2026-05-27 06:37:56 +00:00
Nico Holmberg and GitHub
7b54690244
[ROCm][Perf] Expose AITER MoE sorting dispatch policy via env var ( #39177 )
...
Signed-off-by: nholmber <nholmber@users.noreply.github.com >
2026-05-27 13:11:02 +08:00
1fc2cee50a
[KVConnector][Mooncake] Wire reset_cache cascade end-to-end ( #42694 )
...
Signed-off-by: aoshen524 <aoshen524@gmail.com >
Signed-off-by: Ao Shen <aoshen@inferact.ai >
Co-authored-by: aoshen524 <aoshen524@gmail.com >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 20:52:35 -07:00
Angela Yi and GitHub
0fa3114ae1
Fix test_aot_compile for torch 2.12 ( #43695 )
...
Signed-off-by: Angela Yi <yiangela7@gmail.com >
2026-05-26 23:12:49 -04:00
Woosuk Kwon and GitHub
adaa5e455a
[DSv4] Refactor compressor & Fix ROCm compatibility ( #43710 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 19:56:46 -07:00
c02c758ea4
[Deprecation] Deprecate functions as scheduled for v0.21.0 ( #43358 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 19:56:21 -07:00
Matthew Bonanni and GitHub
aa6138169f
[MLA][Attention] Add OOT MLA prefill backend registration mechanism ( #43325 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-26 19:56:09 -07:00
7e33081cee
[Attention] Make FlexAttention and FlashAttention use num-blocks first layouts ( #42095 )
...
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com >
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com >
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com >
2026-05-26 19:55:56 -07:00
Xin Yang and GitHub
d8eebe6d97
[Perf] Optimize Fp8BlockScaledMMLinearKernel input_scale tensor using new_empty() ( #43677 )
...
Signed-off-by: Xin Yang <xyangx@amazon.com >
2026-05-26 19:55:52 -07:00
Andreas Karatzas and GitHub
5bdb181df5
[ROCm][CI] Fix ROCm multimodal Qwen2.5-VL activation compile and Phi4MM ragged image mask handling ( #43647 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 19:53:34 -07:00
Bugen Zhao and GitHub
0b68f21e7c
[Rust Frontend] Add reasoning/tool parser & renderer roundtrip tests ( #43582 )
...
Signed-off-by: Bugen Zhao <i@bugenzhao.com >
2026-05-27 00:49:30 +00:00
dede691c95
[Bugfix] Split attention groups by num_heads_q for spec-decode drafts ( #43543 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-27 00:11:01 +00:00
e19b9b1045
[ci] Add arm64 ci image ( #41303 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Signed-off-by: Kevin H. Luu <khluu000@gmail.com >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 14:38:09 -07:00
812e7e7364
[Bugfix][V1] Fix TOCTOU race causing intermittent EADDRINUSE on multi-API-server DP startup ( #42585 )
...
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com >
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 14:06:00 -07:00
d98cbf472b
[KV Connector] MooncakeStore: drop dead discard_partial_chunks parameter ( #43627 )
...
Signed-off-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Zhewen Li <zhewen@inferact.ai >
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
2026-05-26 13:40:21 -07:00
Jee Jee Li and GitHub
6e503868ca
[Kernel] Porting fuse_minimax_qk_norm to manual fusion ( #43410 )
...
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai >
2026-05-26 13:16:03 -07:00
49b4882779
[CI] Soft-fail AMD entrypoints mirror tests ( #43709 )
...
Signed-off-by: Kevin Luu <kevin@inferact.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-26 13:08:48 -07:00
Woosuk Kwon and GitHub
193ce8812e
[DSv4] Drop _get_compressed_kv_buffer in DeepseekCompressor ( #43690 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 10:11:25 -07:00
3aea37d28e
[Doc] Add line limit to AGENTS.md ( #43635 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
Signed-off-by: Mark McLoughlin <markmc@redhat.com >
Co-authored-by: Mark McLoughlin <markmc@redhat.com >
2026-05-26 09:31:23 -07:00
Wei-Ming Chen and GitHub
6f5b533241
Add LM head quantization support for ModelOpt ( #42124 )
...
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com >
2026-05-26 09:21:05 -07:00
Woosuk Kwon and GitHub
c8414a8271
[ROCm] Remove MegaMoE integration in deepseek v4 ( #43629 )
...
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai >
2026-05-26 08:56:04 -07:00
f51bbc694d
[MoE Refactor] W4a8 int8 oracle ( #42789 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:15:42 -04:00
b226ddacfd
[MoE Refactor] Migrate ModelOptMxFp8FusedMoE to oracle ( #42768 )
...
Signed-off-by: Bill Nell <bnell@redhat.com >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-26 11:14:14 -04:00
Yongye Zhu and GitHub
6ab6ffb428
[Feat][DSV4] Fuse q pad into deepseek v4 fused kernel ( #43162 )
2026-05-26 05:12:54 -10:00
Andreas Karatzas and GitHub
445ded18c1
[ROCm][CI] Extend ROCm quick reduce coverage ( #40990 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-26 21:57:13 +08:00
d565357a90
[Docs][ROCm] MoRI-IO Connector Usage Guide ( #43603 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Signed-off-by: Simon Danielsson <70206058+simondanielsson@users.noreply.github.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-26 21:52:30 +08:00
Mohammad Miadh Angkad and GitHub
a970fb5a1a
Fix CuPy runtime deps and restore humming ( #43530 )
...
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com >
2026-05-26 05:59:40 -07:00
Chaojun Zhang and GitHub
861b97765d
[XPU] Fix fused MoE LoRA kernel crash on XPU by using platform-agnos num_compute_units ( #43646 )
...
Signed-off-by: Chaojun,Zhang <chaojun.zhang@intel.com >
2026-05-26 03:40:32 -07:00
ebd0692f80
[Model] Use AutoWeightsLoader for InternLM2 ( #38278 )
...
Signed-off-by: Jesus De Jesus <dejesus.9297@gmail.com >
Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com >
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk >
2026-05-26 03:39:26 -07:00
739af5c7e1
[Reasoning] [Bugfix] Reject invalid thinking_token_budget values ( #43402 )
...
Signed-off-by: linzm1007 <linzm1007@126.com >
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-26 03:37:30 -07:00
Thibault Castells and GitHub
5d09f471f4
[Misc] Support interleaved custom image benchmark datasets ( #43636 )
...
Signed-off-by: ThibaultCastells <thib.castells@icloud.com >
2026-05-26 03:37:25 -07:00
681d7dd38b
[Misc][Refactor][ROCm] Convert MoRI-related envvars to extra config args ( #43303 )
...
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com >
Co-authored-by: TJian <tunjian.tan@embeddedllm.com >
2026-05-26 03:33:35 -07:00
Ethan Feng and GitHub
755043cf3c
[KV Transfer] Enable HMA by default for connectors that support it ( #41847 )
...
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com >
2026-05-26 12:28:51 +02:00