 Ethan FengandGitHub
|
a43bc34baf
|
[Docs] Update server entrypoint examples (#42077)
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com>
|
2026-05-09 02:03:52 +00:00 |
|
 Ethan FengandGitHub
|
236bf9d152
|
[Docs] Fix RLHF example links (#42073)
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com>
|
2026-05-09 02:03:42 +00:00 |
|
 Lucas WilkinsonandGitHub
|
b1728c1e66
|
[Attention][Cleanup] Remove tree attention (#42121)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
|
2026-05-08 18:36:19 -07:00 |
|
 
|
be0dcc29dc
|
[XPU] remove q/k/v force contiguous for flash_attn (#40356)
Signed-off-by: Yan Ma <yan.ma@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-05-09 01:19:05 +00:00 |
|
 Sumanth R HegdeandGitHub
|
e3b65a5ba0
|
[feat] Add explicit /start_weight_update and /finish_weight_update APIs for weight transfer (#39212)
|
2026-05-08 18:03:33 -07:00 |
|
 Harry MellorandGitHub
|
30f519e947
|
Use pre-commit / pre-run-check to gate docs build too (#42053)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-09 00:02:51 +00:00 |
|
 Roy WangandGitHub
|
60851b1d22
|
[Bugfix][KV Transfer] Reject NixlConnector + expandable_segments:True (#41237)
|
2026-05-08 16:47:33 -07:00 |
|
 Michael GoinandGitHub
|
8bcd8a260c
|
[Bugfix] Fix FlashInfer CUTLASS MXFP4-MXFP8 MoE by restoring swizzled scale (#42089)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-05-08 15:59:06 -07:00 |
|
 John CalderonandGitHub
|
8a2fc80b84
|
[CUDA][CUTLASS] Enable cutlass scaled mm for non-compatible sizes (#41868)
Signed-off-by: John Calderon <jcalderon@nvidia.com>
|
2026-05-08 15:58:05 -07:00 |
|
 
|
6881c754e1
|
use HIP_VERSION variables to guard against duplicate atomicAdd definitions (#41802)
Signed-off-by: Philip Maybank <pmaybank@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-05-08 18:44:37 -04:00 |
|
 Kevin H. LuuandGitHub
|
0c2e9d4892
|
[CI] Narrow misc.yaml source dependencies (#42059)
Signed-off-by: khluu <khluu000@gmail.com>
|
2026-05-08 15:10:12 -07:00 |
|
 Kevin H. LuuandGitHub
|
d2f22dfc9f
|
[CI] Narrow engine.yaml source dependencies (#42055)
Signed-off-by: khluu <khluu000@gmail.com>
|
2026-05-08 14:55:33 -07:00 |
|
 Kevin H. LuuandGitHub
|
f4dd5c116c
|
[CI] Narrow Platform Tests (CUDA) source dependencies (#42054)
Signed-off-by: khluu <khluu000@gmail.com>
|
2026-05-08 14:54:06 -07:00 |
|
 Kevin H. LuuandGitHub
|
f47ccc8b1c
|
[CI] Narrow pytorch.yaml compile job source dependencies (#42057)
Signed-off-by: khluu <khluu000@gmail.com>
|
2026-05-08 14:43:17 -07:00 |
|
 
|
dbd86a67e3
|
[Bugfix][Gemma4] Fix infinite loop and array boundary issues in tool parser (#41991)
Signed-off-by: David Oy <david.oy@baseten.co>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-05-08 17:24:37 -04:00 |
|
      
|
2c6b59b807
|
[ROCm][Perf] Add Fused Shared Expert (FSE) support for Qwen3-Next (#39280)
Signed-off-by: nholmber <nholmber@users.noreply.github.com>
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Doug Lehr <douglehr@amd.com>
Co-authored-by: nholmber <nholmber@users.noreply.github.com>
Co-authored-by: Tres <tpopp@users.noreply.github.com>
Co-authored-by: Tres Popp <tres.popp@amd.com>
Co-authored-by: Doug Lehr <douglehr@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Douglas Lehr <91553416+dllehr-amd@users.noreply.github.com>
|
2026-05-08 15:38:00 -04:00 |
|
 
|
44e6b44a21
|
[CI][Elastic EP] Fix Elastic EP Scaling Test Failure (#41792)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
|
2026-05-08 15:17:44 -04:00 |
|
 Hiroaki MikamiandGitHub
|
90f145aaf7
|
[Models][Gemma3/Gemma4] Support hidden_act variants in gated MLP (#40588)
Signed-off-by: Hiroaki Mikami <hiroaki8270+github@gmail.com>
|
2026-05-08 11:29:11 -07:00 |
|
 Ethan FengandGitHub
|
4140faa4a5
|
[Docs] Fix OpenAI batch model argument examples (#42066)
Signed-off-by: Ethan Feng <ethan.fengch@gmail.com>
|
2026-05-08 14:02:46 +00:00 |
|
 liuzhenweiandGitHub
|
f2bbd575e2
|
[CI][XPU] Skip fork-dependent logits processor test (#42013)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-05-08 06:10:19 -07:00 |
|
 haosdentandGitHub
|
52458b60a8
|
[CI][Examples][RLHF] Disable async scheduling in rlhf_async_new_apis (#42042)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-08 04:58:48 -07:00 |
|
 Harry MellorandGitHub
|
630820a59b
|
Make docs environment deterministic (#41926)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-05-08 10:13:03 +00:00 |
|
 Chaojun ZhangandGitHub
|
19df11f5d1
|
[CI][XPU]Ignore some lora tests from LoRA Intel CI pipeline (#42010)
Signed-off-by: chaojun-zhang <chaojun.zhang@intel.com>
|
2026-05-08 17:34:27 +08:00 |
|
 haosdentandGitHub
|
36b2c79d4b
|
[CI][Bugfix] Drop duplicated examples/ prefix in tensorize_vllm_model command (#42039)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-08 02:23:22 -07:00 |
|
 haosdentandGitHub
|
160858cba4
|
[CI][Bugfix] Surface subprocess output in spawn_new_process_for_each_test (#41943)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-08 10:39:37 +02:00 |
|
 Simon DanielssonandGitHub
|
f9b9bf3bbb
|
[CI][ROCm] Ship RIXL with vllm/vllm-openai-rocm (#41634)
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
2026-05-08 07:05:17 +00:00 |
|
 
|
445d747434
|
[Bugifx] Missing Renderer for fastokens mode (#41984)
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-05-07 23:45:14 -07:00 |
|
 wang.yuqiandGitHub
|
77b13b9602
|
[Docs] Reorganize examples docs. (#41082)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-05-07 23:23:44 -07:00 |
|
  
|
ed582b6a4c
|
[Aiter][ROCm] gdn_linear_attn kernel fusion (#40711)
Signed-off-by: Tres Popp <tres.popp@amd.com>
Signed-off-by: Chuan Li <chuali@amd.com>
Co-authored-by: hellozhuo <zhuo.su@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-07 23:11:37 -07:00 |
|
 David ZhengandGitHub
|
1acd67a795
|
[Bugfix] Fix XPU/ROCm compatibility in spawn_new_process_for_each_test (#41895)
Signed-off-by: dqzhengAP <dqzheng1996@gmail.com>
|
2026-05-08 00:47:22 -04:00 |
|
 
|
0b99971352
|
[Kernel][Helion] Optimize Helion config parsing latency (#40850)
Signed-off-by: Yanan Cao <gmagogsfm@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
|
2026-05-07 20:27:34 -07:00 |
|
 
|
baf068d8be
|
enable persistent mla for sparse mla backend (#41990)
Signed-off-by: ganyi <ygan@amd.com>
Signed-off-by: Douglas Lehr <Doug.Lehr@amd.com>
Co-authored-by: ganyi <ygan@amd.com>
|
2026-05-07 20:10:50 -07:00 |
|
 SamareshSinghandGitHub
|
01b0f3adab
|
fix: default TILELANG_CLEANUP_TEMP_FILES=1 to avoid shared /tmp conflicts (#41486)
Signed-off-by: Samaresh Kumar Singh <ssam3003@gmail.com>
|
2026-05-07 19:59:00 -07:00 |
|
 Nick HillandGitHub
|
989c176c0a
|
[Perf][3/n] Eliminate GPU<->CPU syncs in attention impls (#41434)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-07 19:44:24 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
cd58e30872
|
[Perf] Use numpy zero-copy path for embedding float response serialization (#41681)
Signed-off-by: Shrinav Loka <lokashrinav@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-05-07 19:42:21 -07:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1d694e78c9
|
[Examples][last/6] Resettle examples. (#41084)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-07 19:42:12 -07:00 |
|
 haosdentandGitHub
|
57c2f724c1
|
[CI][Bugfix] Fix CI failures for "PyTorch Compilation Unit Tests" (#41940)
Signed-off-by: haosdent <haosdent@gmail.com>
|
2026-05-07 19:42:00 -07:00 |
|
 
|
5f6a02812a
|
[CI][Bugfix] Fix failure CI step "PyTorch Fullgraph Smoke Test" (#41953)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-07 19:41:56 -07:00 |
|
 
|
50f2db2555
|
add: LFM2/2.5 Tool Parser (#39243)
Signed-off-by: Jonathan Buchanan <jonathan.buchanan@liquid.ai>
Co-authored-by: Chauncey <chaunceyjiang@gmail.com>
|
2026-05-08 09:58:17 +08:00 |
|
 
|
09a7cc5ba9
|
[KV Connector] Opt DecodeBenchConnector into SupportsHMA (#41770)
Signed-off-by: Zijing Liu <liuzijing2014@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-07 16:10:00 -07:00 |
|
 Nick HillandGitHub
|
10ebb40d62
|
[Core] Avoid using extra thread in UniProcExecutor (#40891)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-05-07 15:33:00 -07:00 |
|
 
|
54f548e9e5
|
[Bugfix] Restore moe_forward output shape invariant on TRTLLM MXFP4 path (#41646)
Signed-off-by: Stefano Castagnetta <scastagnetta@nvidia.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-07 15:26:06 -07:00 |
|
 Kyle SayersandGitHub
|
c1819ca283
|
[Compressed Tensors] Allow configs with non-explicit ignores (#41965)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-05-07 14:03:45 -07:00 |
|
 
|
969fbfb4a9
|
Laguna xs dflash support (#41880)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-05-07 13:31:16 -07:00 |
|
 akii96andGitHub
|
3af561ec0a
|
[ROCm] Fix AITER AR+RMSNorm no-residual fusion (#41972)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
|
2026-05-07 13:14:58 -07:00 |
|
 akii96andGitHub
|
c936548ce6
|
[ROCm][DeepSeek] Enable V3.2 TP4 AITER MLA (#41835)
Signed-off-by: Aakif Nawaz <aakif.nawaz@amd.com>
|
2026-05-07 15:10:57 -05:00 |
|
 TomerBN-NvidiaandGitHub
|
8189a15914
|
[Core] Replace routing replay with device cache and async D2H pipeline (#39917)
Signed-off-by: Tomer Barnatan <tbarnatan@nvidia.com>
|
2026-05-07 11:24:56 -07:00 |
|
 Flora FengandGitHub
|
8eb401134e
|
[Refactor] Consolidate required/named tool_choice streaming into DelegatingParser (#41876)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-05-07 09:50:59 -07:00 |
|
 Nicolò LucchesiandGitHub
|
9d6500b89d
|
[Misc] Delay EPLB Nixl import until needed (#41805)
Signed-off-by: NickLucche <nlucches@redhat.com>
|
2026-05-07 09:43:07 -07:00 |
|
 zhrrrandGitHub
|
7a08b34fbf
|
[Model Runner V2] support qwen35 / mamba hybrid model (#35520)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
|
2026-05-07 09:31:05 -07:00 |
|