![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
0d4d334eaa
|
Bump llguidance to 1.7 (#42150)
Signed-off-by: RickyChen / 陳昭儒 <ricky.chen@infinirc.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-14 20:35:27 -04:00 |
|
  
|
fa2a33b893
|
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py (#38288)
Signed-off-by: Chengyi Nie <cnie@roblox.com>
Co-authored-by: Chengyi Nie <cnie@roblox.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-15 08:25:52 +08:00 |
|
 Giancarlo DelfinandGitHub
|
3b6a204789
|
[Model Runner V2][Bug Fix][DSV4] Ensure lazy attention state initializations happen during cudagraph capture (#42444)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-05-14 16:16:17 -07:00 |
|
 
|
f8848b2f2d
|
[Bugfix] Add swiglu limits to deepgemm fp8 methods (#41986)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-14 15:43:13 -07:00 |
|
 Charlie FuandGitHub
|
4cfcc0866f
|
[CI][ROCm] Remove unsupported cases in test_fusion.py (#38680)
Signed-off-by: charlifu <charlifu@amd.com>
|
2026-05-14 17:37:18 -04:00 |
|
 
|
f887aa1a53
|
[Aiter][ROCm] RMSNormGated+GroupedQuantFP8 fusion (#40710)
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Tres Popp <trespopp@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-14 15:37:09 -04:00 |
|
 Matthew BonanniandGitHub
|
9898f94abe
|
[Attention] Remove deprecated MLA prefill arguments (#42555)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-14 10:34:06 -07:00 |
|
 
|
ae4f59f0ec
|
[Model Runner v2] Oracle for model runner v2 - qwen3 dense model by default [1/N] (#39337)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-14 10:02:33 -07:00 |
|
 RanranandGitHub
|
f3d5360591
|
[Bugfix][Multimodal] PyAV video backend returns keyframes labeled as targets (#42586)
Signed-off-by: Ranran <hzz5361@psu.edu>
|
2026-05-14 08:56:59 -07:00 |
|
 Baorun (Lauren) MuandGitHub
|
a7737cb4f3
|
[Fix] Misc Fixes in ViT CUDA Graph (#38040)
Signed-off-by: Baorun Mu <bmu@nvidia.com>
|
2026-05-14 23:49:06 +08:00 |
|
 Cyrus LeungandGitHub
|
b8a25d0e12
|
[Bugfix] Fix LM detection for Nemotron Parse (#42641)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-05-14 23:42:10 +08:00 |
|
 frida-anderssonandGitHub
|
f07b1da797
|
[ROCm] Enable gluon paged MQA logits on gfx950 (MI355X) (#42062)
Signed-off-by: Frida Andersson <fanderss@amd.com>
|
2026-05-14 15:39:26 +00:00 |
|
 
|
f60c6b33a5
|
[V1][DP][LB] Publish request counts at the start of each engine step (#41626)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
Signed-off-by: Vadim Gimpelson <156319763+vadiklyutiy@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-14 15:39:24 +00:00 |
|
 
|
24337fb860
|
PD disagg with NIXL Connector: GDN support (Qwen3.5) (#41869)
Signed-off-by: Zhanqiu Hu <zhu@redhat.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-14 16:33:01 +02:00 |
|
   
|
c7560af424
|
[RFC] Replace shared-memory routed experts with ModelRunnerOutput transfer and HTTP support (#39568)
Signed-off-by: xhx1022 <1737006628@qq.com>
Signed-off-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: arlenxu <arlenxu@tencent.com>
Co-authored-by: Junjie Zhang <junj.jay.zhang@gmail.com>
|
2026-05-14 14:12:30 +00:00 |
|
 Mohammad Miadh AngkadandGitHub
|
2317682f95
|
[Bugfix] Fix TRTLLM ragged MLA prefill workspace warmup (#42112)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-14 09:48:56 -04:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
5bd8c71e79
|
[kv_offload] Implement reset_cache() for the offloading connector (#41956)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Or Ozeri <or@ozery.com>
|
2026-05-14 16:00:10 +03:00 |
|
 Wentao YeandGitHub
|
6548560496
|
[Compile] Fix compile warning with topk softplus sqrt (#41261)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-14 05:12:50 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
0a65d46628
|
[DSV4] Fuse norm and router for low latency scenario (#41263)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
Signed-off-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io>
Co-authored-by: jeejeelee <jeejeelee@verda-b300-05.datacrunch.io>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-14 05:11:02 -07:00 |
|
 Zhenzhong XuandGitHub
|
1ea9401364
|
[Quantization][Autoround][Toolkit] Add W4A16 Support (#39778)
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong Xu <zhenzhong.xu@intel.com>
|
2026-05-14 19:18:49 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9946c38b7f
|
[XPU] Fix double-transpose in XPUFP8ScaledMMLinearKernel for W8A8 quant method (#41689)
Signed-off-by: Libin Tang <libin.tang@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-14 17:17:39 +08:00 |
|
 
|
23c85343fb
|
[Bug] Fix DeepSeek V4 AttributeError: module 'cutlass.cute.nvgpu' has no attribute 'LoadCacheMode' (#42342)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-14 02:00:20 -07:00 |
|
 rasmithandGitHub
|
768f4a6f26
|
[CI][AMD][BugFix] Prevent triton compiler error when running test_moe_layer with use_ep = True on ROCm (#40857)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-05-14 08:44:22 +00:00 |
|
 rasmithandGitHub
|
addef3299c
|
[CI][AMD] Skip tests where models have problems or fails on both HW types (#42126)
Signed-off-by: Randall Smith <Randall.Smith@amd.com>
|
2026-05-14 08:21:06 +00:00 |
|
   
|
ce29c26b31
|
Update Dockerfile.rocm for AINIC & Thor NIC (#40453)
Signed-off-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu>
Signed-off-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Co-authored-by: root <root@gbt350-odcdh5-wbb3.png-odc.dcgpu>
Co-authored-by: Jhao-Ting Chen <jhaotingc@nvidia.com>
Co-authored-by: simondanielsson <simon.danielsson99@hotmail.com>
|
2026-05-14 15:24:27 +08:00 |
|
 aoshen02andGitHub
|
8c79ad6580
|
Revert "[Core] Replace routing replay with device cache and async D2H pipeline" (#39917) (#42434)
Signed-off-by: aoshen02 <aoshen@inferact.ai>
|
2026-05-13 23:49:01 -07:00 |
|
 
|
0d2732dd91
|
[MLA Attention Backend] Add TOKENSPEED_MLA backend for DSR1/Kimi K25 prefill + decode on Blackwell (#41778)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Signed-off-by: Roger Wang <hey@rogerw.io>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-13 23:48:02 -07:00 |
|
 Rebecca LeeandGitHub
|
fd7d858c8a
|
Use hidden_pad and intermediate_pad from vLLM #34301 (#42098)
Signed-off-by: Rebecca Lee <Rebecca.Lee@amd.com>
|
2026-05-14 14:21:04 +08:00 |
|
 liuzhenweiandGitHub
|
b26558d4a3
|
[CI][XPU] skip ut of offload connector (#42598)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-05-14 13:13:53 +08:00 |
|
 Sarah SalahandGitHub
|
bf0d2dc6d7
|
[Misc] Fix mypy error in parser_manager type narrowing (#42441)
Signed-off-by: Sarah-Salah <11881117+Sarah-Salah@users.noreply.github.com>
|
2026-05-14 02:48:59 +00:00 |
|
 
|
ca60a4e84f
|
[Fix] Weight loading for qwen3_5 using runai_streamer (#42521)
Signed-off-by: Harsh Shah <iharsh@google.com>
Co-authored-by: Harsh Shah <iharsh@google.com>
|
2026-05-14 10:36:20 +08:00 |
|
 Roy WangandGitHub
|
77e1421a68
|
[Bugfix] Fix EPLB initialization for VLM wrapper models (#39805)
Signed-off-by: esmeetu <jasonailu87@gmail.com>
|
2026-05-14 02:26:15 +00:00 |
|
 Kunshang JiandGitHub
|
751b9f14bd
|
[XPU][CT] Support mxfp8 moe model (#41918)
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-14 09:47:10 +08:00 |
|
 Krish GuptaandGitHub
|
70c00163ff
|
[Feature] Add instruction support for score/rerank chat templates (#42412)
Signed-off-by: KrxGu <krishom70@gmail.com>
|
2026-05-14 09:41:22 +08:00 |
|
 Siddharth BedekarandGitHub
|
f51f6844f9
|
[Bugfix][Spec Decode] Wire draft_probs into probabilistic draft_model rejection (#40269)
|
2026-05-13 21:04:03 -04:00 |
|
 longguoandGitHub
|
665f9c4253
|
[Bugfix] Fix Gemma4ToolParser streaming float corruption (#42128)
Signed-off-by: abinggo <107740309+abinggo@users.noreply.github.com>
|
2026-05-13 18:03:30 -07:00 |
|
 Flora FengandGitHub
|
1087676a90
|
[Refactor] Use shared utils in hermes tool parser (#42570)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-13 20:35:45 -04:00 |
|
   
|
63cc8a55a9
|
fix(tool-parser): preserve "none"/"nil" strings as valid enum values in minimax_m2 (#39599)
Signed-off-by: Yiyang Liu <yiyangliu@microsoft.com>
Signed-off-by: Yiyang Liu <37043548+ianliuy@users.noreply.github.com>
Signed-off-by: sfeng33 <4florafeng@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: sfeng33 <4florafeng@gmail.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-13 20:35:34 -04:00 |
|
 Divakar VermaandGitHub
|
ca7e4546da
|
[CI] set max transformers version for skywork model (#42104)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-05-13 16:53:49 -07:00 |
|
 
|
b2198670b1
|
[Bugfix] V1: support tuple model outputs in ubatch wrapper (dbo + spec decode) (#40789)
Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-13 15:47:51 -07:00 |
|
 Mohammad Miadh AngkadandGitHub
|
f1cc7aad3c
|
[Bugfix] Fix DeepSeek V4 MTP HC state handling (#42320)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-05-13 15:44:52 -07:00 |
|
 Lukas GeigerandGitHub
|
597ed13803
|
[Core][MM] Do not use urllib3 to parse data URLs (#42535)
Signed-off-by: Lukas Geiger <lukas.geiger94@gmail.com>
|
2026-05-13 22:21:01 +00:00 |
|
 liangel-02andGitHub
|
6b5c389ee3
|
expose flex block size for batch invariant mode (#41252)
Signed-off-by: Angel Li <liangel@meta.com>
|
2026-05-13 14:11:57 -07:00 |
|
 Michael GoinandGitHub
|
8efd508204
|
[Quantization] Rework quantization_config to use QuantKey and allow for activation override (#41566)
|
2026-05-13 16:58:32 -04:00 |
|
 ovidiusmandGitHub
|
cca32d55a2
|
[PD] Fix broken NIXL EP installation (#42542)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
|
2026-05-13 13:55:51 -07:00 |
|
 Walter Beller-MoralesandGitHub
|
873910d608
|
[Frontend] add support for thinking_token_budget in completions (#42116)
|
2026-05-13 16:01:52 -04:00 |
|
 Wentao YeandGitHub
|
3f611f6106
|
[CI] Fix pre-commit issue (#42563)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-13 12:37:26 -07:00 |
|
 Nick HillandGitHub
|
a505cf807e
|
[ModelRunner V2] Share identical MTP weights (#42538)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-13 18:57:04 +00:00 |
|
 
|
40330967ab
|
[Quark] Support loading Quark NVFP4 checkpoints in vLLM (#35859)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: fxmarty-amd <felmarty@amd.com>
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-05-13 11:17:36 -07:00 |
|
 Fynn Schmitt-UlmsandGitHub
|
ab1ad0d7a9
|
Remove verifier model type check in speculative config (#42536)
Signed-off-by: Fynn Schmitt-Ulms <fschmitt@redhat.com>
|
2026-05-13 18:14:39 +00:00 |
|