 Jee Jee LiandGitHub
|
ee05e8137e
|
[Minor] Bigger overlap for FI AR (#43103)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-05-20 21:20:57 -07:00 |
|
 Louie TsaiandGitHub
|
5d041cc1fe
|
update GPU json file based on h200 recipes (#43262)
Signed-off-by: louie-tsai <louie.tsai@intel.com>
|
2026-05-21 03:57:48 +00:00 |
|
 
|
9640970de2
|
[Model Runner V2] Fix lora Triton Error [CUDA]: device-side assert triggered (#43139)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-21 01:00:30 +00:00 |
|
 
|
63ea11709b
|
[CI] Add composed-schema regression tests for DeepSeek V3.2/V4 parsers (#43255)
Signed-off-by: Ace Eldeib <aeldeib@coreweave.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-05-21 00:36:16 +00:00 |
|
 akii96andGitHub
|
bde560ed6e
|
[ROCm] Add QuickReduce min-size override and codec threshold (#41675)
Signed-off-by: <>
|
2026-05-20 17:46:51 -05:00 |
|
 Jiangyun ZhuandGitHub
|
6dc0a71843
|
[Misc] downgrade nvidia-cutlass-dsl to 4.5.0 (#43230)
Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
|
2026-05-20 14:19:50 -07:00 |
|
 Michael GoinandGitHub
|
5774aad9c5
|
[Perf][gpt-oss] Downgrade triton_kernels to v3.5.1 (#43135)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-05-20 14:13:12 -07:00 |
|
 Douglas LehrandGitHub
|
452baa860b
|
Add dllehr-amd to CODEOWNERS and committers list (#42772)
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
|
2026-05-20 16:10:44 -05:00 |
|
 Flora FengandGitHub
|
2a43b407c5
|
[Bugfix][CI] Add missing import of pad_nvfp4_activation_for_cutlass in flashinfer (#43237)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-20 11:59:12 -07:00 |
|
 
|
53ff50fcd3
|
[Perf] Optimize CutlassFP8ScaledMMLinearKernel when padding needed by pre-weight processing, 13.5% TTFT improvement (#42651)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-05-20 11:57:42 -07:00 |
|
 
|
363fc84407
|
Integrate flashinfer b12x MoE and FP4 GEMM kernels for SM120/121 (#40082)
Signed-off-by: Meenakshi Venkataraman <meenakshiv@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-20 17:21:11 +00:00 |
|
 
|
f2d5e3d3ae
|
[CI] Lower granite-4.0-h-tiny gsm8k threshold for Hybrid SSM NixlConnector PD accuracy tests (4 GPUs) (#43186)
Signed-off-by: haosdent <haosdent@gmail.com>
Signed-off-by: NickLucche <nlucches@redhat.com>
Co-authored-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 17:00:24 +00:00 |
|
 
|
2d6b3489b9
|
[R3] Add routed experts to openai entrypoint (#38939)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-05-20 09:07:59 -07:00 |
|
 Vadim GimpelsonandGitHub
|
9c78c99995
|
[MISC] Fix symm_mem cap-equal gate; log AR backend selection (#42993)
Signed-off-by: Vadim Gimpelson <vadim.gimpelson@gmail.com>
|
2026-05-20 08:50:24 -07:00 |
|
 Flora FengandGitHub
|
a10d69116c
|
[Bugfix] Use shared coerce_to_schema_type in DeepSeekV32 tool parser (#43019)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-20 10:21:00 -04:00 |
|
 
|
644b2a28e7
|
[Bugfix] Use enable_sm120_family for per-tensor FP8 CUTLASS kernels on SM12.1 (#41215)
Signed-off-by: j9smith <j.smith9103@outlook.com>
Signed-off-by: Joel Smith <j.smith9103@outlook.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-20 14:10:01 +00:00 |
|
 
|
ded871201a
|
[Bug][Structured Outputs] Fix bug that leads to unconstrained generations with structural tags (#42452)
Signed-off-by: rishitdholakia13 <rishit+github@cohere.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-20 07:08:58 -07:00 |
|
 Dipika SikkaandGitHub
|
df84fb07a6
|
Remove additional dead code as a follow-up to #42889 (#43144)
Signed-off-by: Dipika Sikka <dipikasikka1@gmail.com>
|
2026-05-20 10:01:45 -04:00 |
|
 Benjamin ChislettandGitHub
|
0a508743d4
|
[Spec Decode] Support non-MTP speculation for NemotronH (#43130)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-05-20 09:15:52 -04:00 |
|
 KebeandGitHub
|
19cf334207
|
[Feature] Support manually enabling the cumem allocator (#33648)
Signed-off-by: Kebe <mail@kebe7jun.com>
|
2026-05-20 08:58:30 -04:00 |
|
 
|
87e31455b0
|
[Doc] Sync CLI guide with actual help modes and launch subcommand (#40326)
Signed-off-by: Rui Wang <raygorous@gmail.com>
Co-authored-by: Rui Wang <raygorous@gmail.com>
|
2026-05-20 02:32:03 -07:00 |
|
 
|
cb600d1cdb
|
[Frontend] Forward X-data-parallel-rank header on /inference/v1/generate (#42330)
Signed-off-by: hallerite <git@hallerite.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-20 08:58:46 +00:00 |
|
 xiangdongandGitHub
|
6f21558da1
|
[XPU][CI] Add 2 server model test files in Intel GPU CI (#42499)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
|
2026-05-20 16:54:58 +08:00 |
|
 Artem PerevedentsevandGitHub
|
1cb224430b
|
[GDN] Enable FI Blackwell GDN prefill kernel (#40717)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-05-20 01:46:55 -07:00 |
|
 Harry MellorandGitHub
|
9b343dd4f5
|
Enable mermaid diagrams in the docs (#43192)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-20 08:10:00 +00:00 |
|
  
|
07aeaf9d4d
|
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) (#42663)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-20 00:18:12 -07:00 |
|
 Nicolò LucchesiandGitHub
|
40651c0207
|
[Docs][PD][NIXL] Bidirectional kv-cache transfer (#43097)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 09:02:36 +02:00 |
|
 Nicolò LucchesiandGitHub
|
7e4bc2cecb
|
[Docs][PD][NIXL] Lease extension mechanism for blocks on P (#43099)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-20 08:58:25 +02:00 |
|
 Kevin H. LuuandGitHub
|
85959567c3
|
[ci] Revert model executor test back to L4 (#43188)
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
|
2026-05-19 23:01:41 -07:00 |
|
 Ronen SchafferandGitHub
|
4f940896a3
|
[KV Offload] Pass OffloadingSpec instead of VllmConfig to secondary tiers (#43076)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-05-20 03:32:08 +00:00 |
|
 Michael GoinandGitHub
|
cd0ff26e7a
|
[CI] Add DSV4-Flash to gsm8k moe-refactor/config-b200.txt (#42111)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-05-19 20:21:01 -07:00 |
|
 Izik GolanandGitHub
|
2ae910ed88
|
[Perf] Avoid forward scan for async output placeholders (#42938)
|
2026-05-19 20:16:07 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
fadf5d332c
|
add enqueue all option to throughput benchmark (#42975)
Signed-off-by: Philip Maybank <pmaybank@amd.com>
Signed-off-by: pmaybank <113125070+pmaybank@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-19 20:16:02 -07:00 |
|
 Benjamin ChislettandGitHub
|
c628a93a64
|
[Perf][Bugfix] Update dflash aux layer indexing (#40727)
Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
|
2026-05-19 20:15:57 -07:00 |
|
 Terrence ZhaoandGitHub
|
5774aaed0c
|
[Cohere] Enable Cohere MoE (#43143)
Signed-off-by: Terrencezzj <terrence@cohere.ai>
|
2026-05-19 19:32:06 -07:00 |
|
 Nick HillandGitHub
|
39bba710be
|
[MRV2][BugFix] Fix default-stream CG capture in P/W LoRA case (#43160)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-19 19:19:05 -07:00 |
|
 Aaron HaoandGitHub
|
73dd2f33b7
|
[bug] fix WeightTransferConfig.backend to allow for all strings (#43121)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
|
2026-05-19 21:01:29 -04:00 |
|
 Fadi ArafehandGitHub
|
be16785998
|
[CPU][DOC] Fix installation commands for Arm CPUs (#43115)
Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
|
2026-05-19 23:31:15 +00:00 |
|
 ![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
117afeea46
|
Fix error in Dynamic NTK scaling (#41277)
Signed-off-by: Max de Bayser <mbayser@br.ibm.com>
Signed-off-by: Max de Bayser <maxdebayser@gmail.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-19 17:27:54 -04:00 |
|
 Doğaç EldenkandGitHub
|
1242196295
|
[Model] Support post-norm architecture for EAGLE-3 supeculators (#42764)
Signed-off-by: Doğaç Eldenk <dogacel@gmail.com>
|
2026-05-19 13:39:00 -07:00 |
|
 Kevin H. LuuandGitHub
|
a65093c1a3
|
[ci] Move language models tests (hybrid) back to L4 (#43129)
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
|
2026-05-19 11:51:34 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9aaf83ef50
|
[CI failure] Temporarily disable using persistent cache for flashinfer autotune (#43119)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Signed-off-by: Wei Zhao <51183510+wzhao18@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-19 11:44:32 -07:00 |
|
 tomeras91andGitHub
|
f54721bcc3
|
[Bugfix][MoE] FlashInfer one-sided: workspace union across heterogeneous layers (#42976)
Signed-off-by: Tomer Asida <57313761+tomeras91@users.noreply.github.com>
|
2026-05-19 14:43:04 -04:00 |
|
 
|
aed2eb355a
|
[Docs] Fix MooncakeStoreConnector role in disaggregated example (#42994)
Signed-off-by: Dao Le <Dao007forever@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-19 11:14:43 -07:00 |
|
 Dom BrownandGitHub
|
d247a931cc
|
[feat] Add FP8 per-tensor Q scale support to Triton attention backend (#42080)
Signed-off-by: Dom Brown <3886319+DomBrown@users.noreply.github.com>
|
2026-05-19 09:02:05 -07:00 |
|
 Jinzhen LinandGitHub
|
8200fbe1ac
|
[Misc] add humming to dependencies (#42540)
Signed-off-by: Jinzhen Lin <jinzhen.ljz@antgroup.com>
|
2026-05-19 08:36:47 -07:00 |
|
 Flora FengandGitHub
|
42b4f1fdf7
|
[Refactor] Extract extract_types_from_schema utility from Minimax M2 tool parser (#43025)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-19 11:21:12 -04:00 |
|
 Wang YiwenandGitHub
|
1c6158083a
|
[Model] Openvla support (#42654)
Signed-off-by: Wang Yiwen <121547057+yiwen101@users.noreply.github.com>
|
2026-05-19 08:17:42 -07:00 |
|
 Xinyu ChenandGitHub
|
d740e2c029
|
[XPU] update xpu graph usage (#43043)
Signed-off-by: Xinyu Chen <xinyu1.chen@intel.com>
|
2026-05-19 23:09:07 +08:00 |
|
 Nick HillandGitHub
|
b82e908b4c
|
[Perf][4/n] Eliminate various GPU<->CPU syncs (#42347)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-19 10:35:54 -04:00 |
|