 Divakar VermaandGitHub
|
ca7e4546da
|
[CI] set max transformers version for skywork model (#42104)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-05-13 16:53:49 -07:00 |
|
 
|
b2198670b1
|
[Bugfix] V1: support tuple model outputs in ubatch wrapper (dbo + spec decode) (#40789)
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-13 15:47:51 -07:00 |
|
 Mohammad Miadh AngkadandGitHub
|
f1cc7aad3c
|
[Bugfix] Fix DeepSeek V4 MTP HC state handling (#42320)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-13 15:44:52 -07:00 |
|
 Lukas GeigerandGitHub
|
597ed13803
|
[Core][MM] Do not use urllib3 to parse data URLs (#42535)
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com>
|
2026-05-13 22:21:01 +00:00 |
|
 liangel-02andGitHub
|
6b5c389ee3
|
expose flex block size for batch invariant mode (#41252)
Signed-off-by: Angel Li <liangel@meta.com>
|
2026-05-13 14:11:57 -07:00 |
|
 Michael GoinandGitHub
|
8efd508204
|
[Quantization] Rework quantization_config to use QuantKey and allow for activation override (#41566)
|
2026-05-13 16:58:32 -04:00 |
|
 ovidiusmandGitHub
|
cca32d55a2
|
[PD] Fix broken NIXL EP installation (#42542)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
|
2026-05-13 13:55:51 -07:00 |
|
 Walter Beller-MoralesandGitHub
|
873910d608
|
[Frontend] add support for thinking_token_budget in completions (#42116)
|
2026-05-13 16:01:52 -04:00 |
|
 Wentao YeandGitHub
|
3f611f6106
|
[CI] Fix pre-commit issue (#42563)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-13 12:37:26 -07:00 |
|
 Nick HillandGitHub
|
a505cf807e
|
[ModelRunner V2] Share identical MTP weights (#42538)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-13 18:57:04 +00:00 |
|
 
|
40330967ab
|
[Quark] Support loading Quark NVFP4 checkpoints in vLLM (#35859)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: fxmarty-amd <felmarty@amd.com>
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-05-13 11:17:36 -07:00 |
|
 Fynn Schmitt-UlmsandGitHub
|
ab1ad0d7a9
|
Remove verifier model type check in speculative config (#42536)
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com>
|
2026-05-13 18:14:39 +00:00 |
|
 Ben BrowningandGitHub
|
0f69128a37
|
[Bugfix] Handle real-world gpt-oss tool call output in Harmony parsing (#42454)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
|
2026-05-13 17:54:46 +00:00 |
|
 
|
b3c69595a6
|
[MM][CG] Support ViT CG for Qwen2-VL (#41736)
Signed-off-by: John Calderon <jcalderon@nvidia.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-14 01:52:35 +08:00 |
|
 
|
2f821faeae
|
[Spec Decode] Support hybrid attention models in extract_hidden_states (#39949)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-13 10:45:53 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
5794c65f8c
|
[Bugfix][Model] Gemma4 MoE routing closure captures per_expert_scale, breaking functional_call substitution (#42250)
Signed-off-by: Noelia <noeliabentancor1@gmail.com>
Signed-off-by: Noelia Bentancor <71080743+NoeliaBentancor@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-13 17:43:12 +00:00 |
|
 CynicDoraandGitHub
|
256dbcaabf
|
[Feature] Support custom callable proposer backend for speculative decoding (#39487)
Signed-off-by: 524031910363 <hyzhyzsh@sjtu.edu.cn>
Signed-off-by: CynicDora <hyzhyzsh@sjtu.edu.cn>
|
2026-05-13 16:53:01 +00:00 |
|
 Wentao YeandGitHub
|
e35c0d4c63
|
[Feature] Support compile mode for batch invariance on SM80 (#42456)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-13 11:02:39 -04:00 |
|
 Ronen SchafferandGitHub
|
11f6b545d4
|
[kv_offload] Add multi-tier KV cache offloading framework (#40020)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-13 17:21:43 +03:00 |
|
 
|
a8887c208f
|
[Bugfix] [ROCm] [DSV4] [Perf] Add aiter mhc support (#41946)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-13 21:43:15 +08:00 |
|
 
|
0ddaf6dffa
|
[XPU] [CT] Enable CT W4A4MxFp4 path and add xpu kernel (#38896)
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
Signed-off-by: zofia <110436990+zufangzhu@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-13 06:43:00 -07:00 |
|
 Marek WawrzosandGitHub
|
67671692ac
|
[CI] Re-enable Nemotron Parse parity test and switch testing to nemotron-parse v1.2 (#42498)
Signed-off-by: <mwawrzos@nvidia.com>
|
2026-05-13 21:05:27 +08:00 |
|
 hissu-hyvarinenandGitHub
|
0a62f5eec9
|
[AMD] skip machete tests for rocm (#42326)
Signed-off-by: Hissu Hyvarinen <hissu.hyvarinen@amd.com>
|
2026-05-13 12:11:03 +00:00 |
|
 PikaPikachuandGitHub
|
3b1ef03be4
|
[Bugfix][Quark] Fix W8A8 INT8 garbage outputs on Step-3.5-Flash (and other 3-key fused-MoE Quark exports) (#41892)
Signed-off-by: kangletian <kangletian@hotmail.com>
|
2026-05-13 11:59:49 +00:00 |
|
 
|
3c413a5481
|
Triton attention: add USE_TD constexpr for tensor descriptor Q/K/V load/store (#40327)
Signed-off-by: Artur Fierka <artur.fierka@intel.com>
Co-authored-by: quinnlp <quinnlp@users.noreply.github.com>
|
2026-05-13 13:57:41 +02:00 |
|
 Ronen SchafferandGitHub
|
79fd1bc7ed
|
[kv_offload] Add req_id to ReqContext for per-request tracking (#42507)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-13 11:11:10 +00:00 |
|
 SILONG ZENGandGitHub
|
cee6751e54
|
[Bugfix][Qwen3-VL] Fix pipeline-parallel deepstack initialization (#42394)
Signed-off-by: MrZ20 <2609716663@qq.com>
|
2026-05-13 10:58:42 +00:00 |
|
 
|
16863072ca
|
[Bugfix] Fix scipy audio resampling ratio (#42233)
Signed-off-by: JooHo Lee <BWAAEEEK@users.noreply.github.com>
Co-authored-by: JooHo Lee <BWAAEEEK@users.noreply.github.com>
|
2026-05-13 18:52:41 +08:00 |
|
 Andreas KaratzasandGitHub
|
d628a3c5cb
|
[ROCm][CI] Skip ROCm batch invalid-input test pending torch fix (#41572)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-05-13 18:50:47 +08:00 |
|
 akii96andGitHub
|
74dffae666
|
[ROCm] Run AITER RMSNorm pad fusion before AR RMS fusion (#42411)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
|
2026-05-13 18:35:12 +08:00 |
|
 
|
97c4317bf5
|
[Bugfix][Frontend] Default max_tokens server-side on /inference/v1/generate (#42329)
Signed-off-by: hallerite <git@hallerite.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-13 11:16:46 +02:00 |
|
 
|
f6e868fbdf
|
[CI] Use uv with Python 3.12 for PyPI wheel upload (#42470)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-13 02:12:06 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
13bf242100
|
[Feat][KVConnector] Add bind_gpu_block_pool() to KVConnectorBase_V1 (#39654)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-13 02:10:29 -07:00 |
|
 Jiangyun ZhuandGitHub
|
140dc2ec30
|
[Bugfix] Install nvidia-cutlass-dsl[cu13] extra on CUDA 13 platforms (#42438)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
|
2026-05-13 01:57:21 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9ce74042d3
|
[Bugfix][SimpleCPUOffloadBackend] Dedup in-flight CPU offload stores across scheduler steps (#41289)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-13 01:53:32 -07:00 |
|
 sychen52andGitHub
|
a8c13d2837
|
Patch SlidingWindowSpec.real_page_size_bytes for nvfp4 kv (#42464)
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
2026-05-13 01:46:30 -07:00 |
|
 Shanshan ShenandGitHub
|
92def124bc
|
[MM][Perf][CG] Support ViT full CUDA graph for Qwen3.5 (#42151)
Signed-off-by: shen-shanshan <467638484@qq.com>
|
2026-05-13 16:00:32 +08:00 |
|
 
|
85b2fecab7
|
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (#42339)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
|
2026-05-13 07:24:39 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
503697c9ce
|
[chore] Refactor pooling metadata token ID accessors (#42368)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-13 06:08:01 +00:00 |
|
 Nicolò LucchesiandGitHub
|
71bcd02ef3
|
[Bugfix][PD] Fix multi-node TP (TP>8) (#39907)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-12 22:20:57 -07:00 |
|
 Matthew BonanniandGitHub
|
dcacdf9a88
|
[Attention] Sync FA with upstream (#41052)
|
2026-05-12 23:34:18 -04:00 |
|
 bnellnmandGitHub
|
18f6bf5a21
|
[MoE Refactor] Add sequence parallel tests to test_moe_layer.py (#41299)
Signed-off-by: Bill Nell <bnell@redhat.com>
|
2026-05-12 21:52:19 -04:00 |
|
 AlecandGitHub
|
07534b8782
|
[PD] Bump NIXL connector dependency to 1.x (#42364)
Signed-off-by: Alec Flowers <aflowers@nvidia.com>
|
2026-05-12 18:05:01 -07:00 |
|
 Wentao YeandGitHub
|
3d635c58c0
|
[Perf] Optimize MLA compute_prefill_context memory allocation (#42460)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-12 16:23:46 -07:00 |
|
+4        
|
ebeb09d822
|
[KV Transfer] Add MooncakeStoreConnector for KV cache offloading via Mooncake distributed store (#40900)
Signed-off-by: leichao.lc <leichao.lc@antgroup.com>
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: leichao.lc <leichao.lc@antgroup.com>
Co-authored-by: ivanium <yifanqiao@inferact.ai>
Co-authored-by: aoshen524 <aoshen@inferact.ai>
Co-authored-by: Dao007forever <daole@inferact.ai>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
Co-authored-by: Pz1116 <zpbzpb123123@gmail.com>
Co-authored-by: foraxe <1055696449@qq.com>
Co-authored-by: Skywalker-EP <173423846@qq.com>
Co-authored-by: fems14 <1804143737@qq.com>
Co-authored-by: jianzs <zheng.shoujian@outlook.com>
Co-authored-by: baxingpiaochong <771405853@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-12 16:09:10 -07:00 |
|
 Michael GoinandGitHub
|
184577ae46
|
[Build] DeepGEMM: trim comments, add integration notes + TODOs (#42429)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-05-12 15:57:58 -07:00 |
|
 Kevin H. LuuandGitHub
|
8c4fc4202a
|
[CI] Inline build artifact annotations in release pipeline (#42357)
Signed-off-by: khluu <khluu000@gmail.com>
|
2026-05-12 15:57:43 -07:00 |
|
 Nick HillandGitHub
|
fe8b42e80c
|
[CI] Fix test_async_scheduling.py flakiness (#42455)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-12 21:38:32 +00:00 |
|
 Giancarlo DelfinandGitHub
|
fe5b4e0fe7
|
[Model Runner V2] Apply synthetic mode to probabilistic rejection sampler (#41035)
|
2026-05-12 13:37:03 -07:00 |
|
  
|
0ce6613b9c
|
platforms: add uses_cpu_device() hook to Platform for DeviceConfig (#42313)
Signed-off-by: Viktor Pus <viktorpus@tenstorrent.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-05-12 12:39:17 -07:00 |
|