 MattandGitHub
|
1a4984520e
|
[Hardware][AMD][CI] Fix AMD CI image build (#46792)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 22:05:12 -07:00 |
|
 ReidandGitHub
|
e312c5cb25
|
[Rust Frontend] Make Granite4 string argument scanning incremental (#46507)
Signed-off-by: reidliu41 <reid201711@gmail.com>
|
2026-06-26 03:54:03 +00:00 |
|
 Matti4andGitHub
|
1502cf6274
|
Fix relative allowed local media paths (#45263)
|
2026-06-25 20:45:20 -07:00 |
|
 
|
d350fa8ddd
|
[Bugfix][Rust Frontend] Reject min_tokens above max_tokens (#46733)
Co-authored-by: Bugen Zhao <i@bugenzhao.com>
Signed-off-by: reidliu41 <reid201711@gmail.com>
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-26 03:41:33 +00:00 |
|
 
|
dbc49b6b99
|
[CI][NIXL] Fix NIXL EP import canary for the nixl 1.3.0 wheel and pin nixl==1.3.0 (#45166)
Signed-off-by: Ovidiu Mara <ovidium@nvidia.com>
Signed-off-by: ovidiusm <ovidium@nvidia.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
|
2026-06-25 19:33:42 -07:00 |
|
 fxmarty-amdandGitHub
|
552a9dbe59
|
[NVFP4][Emulation] Fuse NVFP4 weight dequantization with compute in triton kernel for w13/w2 MOE MLP linears (#44667)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
|
2026-06-25 19:33:00 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
02a1f23711
|
[DFlash] Fuse precompute kv per-layer rmsnorms (#46761)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:32:07 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
652d962bc9
|
[Model Runner V2][Spec Decode] Reduce TP communication for draft token generation (#46448)
Signed-off-by: EanWang211123 <wangyiheng@sangfor.com.cn>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:30:07 -07:00 |
|
 Giancarlo DelfinandGitHub
|
5314665bad
|
[Model Runner V2][DFlash] Enable dflash attention backend selection (#46770)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-06-25 19:29:25 -07:00 |
|
 Michael GoinandGitHub
|
3daea7ceb9
|
[Bugfix][MRV2] Forward seq_lens_cpu_upper_bound for mamba hybrid models (#46759)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-06-25 19:03:09 -07:00 |
|
 Wentao YeandGitHub
|
cc7981599e
|
[Refactor] Remove dead kernel code (#46405)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-25 18:09:56 -07:00 |
|
 Nick HillandGitHub
|
32bb3195f0
|
[ModelRunner V2] Bound memory for large logprobs requests (#46746)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-25 18:04:06 -07:00 |
|
 
|
ad28d605e6
|
[Bugfix] Default tie_weights to sharing the weight (fix tied quantized embeddings, e.g. ModelOpt Gemma4) (#45544)
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
|
2026-06-25 17:46:28 -07:00 |
|
 Bugen ZhaoandGitHub
|
ae7c8ec223
|
[Rust Frontend] Switch rustls to native-tls/OpenSSL (#46696)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 17:19:44 -07:00 |
|
 Bugen ZhaoandGitHub
|
1d3f4cb3a4
|
[Rust Frontend] Extract renderer fixture test utilities (#46719)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 17:12:38 -07:00 |
|
 Bugen ZhaoandGitHub
|
f9e684499f
|
[Rust Frontend] Migrate gemma4 to unified parser (#46602)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 16:59:57 -07:00 |
|
 Giancarlo DelfinandGitHub
|
c53994e134
|
[Model Runner V2][Spec Decode] Use log1p to compute residual during rejection sampling (#46665)
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
|
2026-06-25 23:46:10 +00:00 |
|
 MattandGitHub
|
27da2a2ac4
|
[Hardware][AMD][CI] Use Triton-based AITER MHA for LM Eval Qwen-3.5 Models Tests (#46691)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 17:08:04 -05:00 |
|
 Michael GoinandGitHub
|
a2e8ec3d52
|
[CI] Depend GPQA Eval DGX Spark job on arm64 image build (#46736)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-06-25 17:07:04 -04:00 |
|
 
|
e8c24a7695
|
[Kernel] Vectorized fp32 moe_sum reduction and support any topk (#46643)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-25 14:02:28 -07:00 |
|
 Andreas KaratzasandGitHub
|
2a6f8f0c05
|
[ROCm][CI] Fine-tuning queues and test names (#39238)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-25 13:24:09 -07:00 |
|
 Robert ShawandGitHub
|
c5e3c40877
|
Fix P/D with DP Supervisor (#46628)
Signed-off-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-06-25 13:13:08 -07:00 |
|
 Wentao YeandGitHub
|
8b4d93ba2b
|
[Perf] Remove redundant clone for GLM, Deepseek etc (#46651)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-25 13:09:00 -07:00 |
|
 Michael GoinandGitHub
|
e8e7b592d1
|
[Kernel][MoE] Tune block-FP8 fused MoE for low-batch decode (#46642)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-06-25 12:38:28 -07:00 |
|
 Rohan PotdarandGitHub
|
e53a17232c
|
[ROCm]: Bump aiter to 0.1.16.post2 (#46692)
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
|
2026-06-25 11:53:26 -07:00 |
|
 Flora FengandGitHub
|
96eb8ddc41
|
[CI] Re-enable skipped glm and seedoss parser tests (#46671)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-06-25 13:11:41 -04:00 |
|
 Gabriel WuandGitHub
|
8fa36fbbeb
|
[Bugfix] FLASHINFER_MLA_SPARSE_SM120 compatibility with GLM-5 NVFP4 (#46506)
|
2026-06-25 09:12:00 -07:00 |
|
 RanranandGitHub
|
e45b279928
|
[Bugfix] Fix NVFP4+MTP crash: force unquantized mtp.fc for Qwen3Next (#46316)
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com>
|
2026-06-25 09:05:04 -07:00 |
|
    
|
d490b98162
|
[Core] Avoid mixed length specdec batches via padding (#45237)
Signed-off-by: Yiliu Dong <91178480+qianlihuang@users.noreply.github.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Jade Zheng <zheng.shoujian@outlook.com>
Co-authored-by: Giancarlo Delfin <gdelfin@inferact.ai>
Co-authored-by: Zijing Liu <liuzijing2014@gmail.com>
|
2026-06-25 08:34:44 -07:00 |
|
 haoyangli0109andGitHub
|
1744adc256
|
[ROCM] [Communication] Add INT3 quantization method for quickreduce (#45666)
Signed-off-by: Haoyang Li <lihaoyang0109@gmail.com>
|
2026-06-25 15:14:15 +00:00 |
|
 Divakar VermaandGitHub
|
cdfa2fd7e9
|
[ROCm][CI] rm duplicate Distributed Torchrun ci test (#46729)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-06-25 09:58:17 -05:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
6f3da461d1
|
[Pooling] Fix Cohere embed billed image token accounting for mixed-content inputs (#46093)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 10:44:29 -04:00 |
|
 Russell BryantandGitHub
|
d3130d878c
|
[CI] Pin GitHub Actions to commit hashes in macos-smoke-test.yml (#38290)
|
2026-06-25 13:48:44 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
9bfd878a48
|
[MoE] [MoE Refactor] Add moe kernel oracle abc 37753 (#43461)
Signed-off-by: Qiuyang Yue <yueqiuyang1389@gmail.com>
Signed-off-by: qyYue1389 <yueqiuyang1389@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 09:34:03 -04:00 |
|
 MattandGitHub
|
2365b7a8e7
|
[Hardware][AMD][CI] Mirror Basic Models (Others) and Weight Loading Multiple GPU test groups (#46668)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 08:25:09 -05:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
15be78732b
|
[NIXL][Mamba] Add Mamba1 support to NIXL P/D disaggregation (#45019)
Signed-off-by: Josephasafg <ajgard7@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 05:50:41 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
92221485aa
|
[CPU][CI/Build] Allow more CPU CI agents (#46702)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 19:33:39 +08:00 |
|
 xiangdongandGitHub
|
a6f41ab678
|
[XPU][CI]Refine .buildkite/ci_config_intel.yaml for Intel GPU CI (#46674)
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
|
2026-06-25 08:58:26 +00:00 |
|
   
|
c63cd4906c
|
[ROCm][ [Perf] sparse attention optimization on minimax-m3 (#46546)
Signed-off-by: Hongxia Yang <hongxia.yang@amd.com>
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: yueliu14 <yue.liu4@amd.com>
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com>
|
2026-06-25 16:56:00 +08:00 |
|
  
|
638b1a99cc
|
[CPU][RISC-V] Add RVV path for W4A8 INT4 GEMM (#45269)
Signed-off-by: wcy <233313160abc@gmail.com>
Co-authored-by: lyd1992 <liuyudong@iscas.ac.cn>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-06-25 08:18:10 +00:00 |
|
 
|
72adb20a6a
|
[Model] Remove AquilaForCausalLM, AquilaModel (#46605)
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-25 08:08:26 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
2396d91e93
|
[CPU][Spec Decode] Enable DFlash SD for CPU (#44029)
Signed-off-by: guybd <guy.boudoukh@intel.com>
Signed-off-by: Guy Boudoukh <guy.boudoukh@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 15:32:48 +08:00 |
|
 
|
9b215ae60b
|
[Rust Frontend] Forward VLLM_ENGINE_READY_TIMEOUT_S via --args-json (#44610)
Signed-off-by: kai <kai@example.com>
Co-authored-by: 图灵 <tuling.wk@alibaba-inc.com>
|
2026-06-25 07:25:08 +00:00 |
|
 Bugen ZhaoandGitHub
|
4d3b4b9b01
|
[Rust Frontend] Make ToolParserOutput a seq of ToolParserEvent to preserve order (#46584)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-25 06:27:07 +00:00 |
|
 Matthias GehreandGitHub
|
77c1d9fe9b
|
[ROCm][Perf] Tune wvSplitK on gfx1151 (#40784)
Signed-off-by: Matthias Gehre <matthias.gehre@amd.com>
|
2026-06-25 14:17:46 +08:00 |
|
 Jeff (Junze) MaandGitHub
|
36fd7e8b86
|
[SimpleCPUOffloadConnector] Fix remaining global→block conversions under PCP/DCP (#46394)
Signed-off-by: Jeff Ma <jeffjma@umich.edu>
|
2026-06-24 23:05:24 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
fc61c6fc26
|
[Perf] Enable + tune FlashInfer fused allreduce at world_size=16 on SM 10.3 (GB300) (#46392)
Signed-off-by: Jeff Ma <jeffjma@umich.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-24 23:04:17 -07:00 |
|
 MattandGitHub
|
e2af449c39
|
[Hardware][AMD][CI] Move Metrics, Tracing (2 GPUs) & make optional (#46686)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-25 05:49:33 +00:00 |
|
  
|
3f5a1e1733
|
[ROCm][CI] Expand basic correctness target suites (#46573)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matt <156021403+mawong-amd@users.noreply.github.com>
|
2026-06-25 12:18:57 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
710ebaa189
|
[ROCm][Bugfix] Fix chunk alignment when using context parallelism with TRITON_MLA (#46114)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-25 00:07:28 -04:00 |
|