 Wentao YeandGitHub
|
b41f7d8ffd
|
Merge branch 'main' into wentao-optimize-per-token-group-quant
|
2026-06-29 17:37:43 -04:00 |
|
 Andreas KaratzasandGitHub
|
8632c884dc
|
[ROCm][CI] Use spawn around the threaded OTLP test (#47003)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-29 16:34:05 -05:00 |
|
 
|
c3734e8334
|
[CI][Bugfix] Add cohere_melody to ROCm test requirements (#47072)
Signed-off-by: pei.zhang <pei.zhang@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-29 16:29:47 -05:00 |
|
 
|
53f7553f09
|
[ROCm][DeepEP] Stabilize high-throughput DBO for DP+EP (#46990)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Andreas Karatzas <Andreas.Karatzas@amd.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-06-29 14:28:02 -07:00 |
|
 
|
4eb227992a
|
[ROCm][CI] Make memory sampling less racy in tests and sleep mode (#45490)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Codex <codex@example.invalid>
Co-authored-by: Codex <codex@example.invalid>
|
2026-06-29 14:26:41 -07:00 |
|
 Micah WilliamsonandGitHub
|
ebcf511ec3
|
[ROCm][CI] Soft Fail Spec Decode Ngram + Suffix and Entrypoints Integration (LLM) AMD Mirrors (#47067)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-06-29 16:24:08 -05:00 |
|
 Matthew BonanniandGitHub
|
8fc1b2d046
|
Fix FA4 dynamic_causal for full attention layers (#46659)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-06-29 14:23:34 -07:00 |
|
 Harry MellorandGitHub
|
5316638a5e
|
Fix transient dependency issues caused by requirements/common.txt (#47015)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 14:20:33 -07:00 |
|
 zhrrrandGitHub
|
61ab70ec3b
|
[Model Runner V2] support mamba hybrid models align prefix cache (#42406)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
|
2026-06-29 14:09:16 -07:00 |
|
 Woosuk KwonandGitHub
|
a309d4fe60
|
Support DCP with FlashInfer MLA (#43729)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-29 13:24:29 -07:00 |
|
 
|
72f639927f
|
[XPU] [RMSNorm] revert weightless change on xpu (#46987)
Signed-off-by: Zhu, Zufang <zufang.zhu@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-29 19:03:06 +00:00 |
|
 Nick HillandGitHub
|
8ad4a01825
|
[ModelRunner V2] Simplify recent UnlimitedOCR-related changes (#46975)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-29 09:56:17 -07:00 |
|
 Jee Jee LiandGitHub
|
7be582697b
|
[Bugfix] Fix DeepseekV2Model hidden_size (#46986)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-29 16:44:05 +00:00 |
|
 
|
030c9523bd
|
[Perf][1/N] Expand Triton kernel warmup coverage, DSv4 (#46634)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
Signed-off-by: Roberto L. Castro <38211239+LopezCastroRoberto@users.noreply.github.com>
Co-authored-by: Lucas Wilkinson <LucasWilkinson@users.noreply.github.com>
|
2026-06-29 16:40:34 +00:00 |
|
 
|
4708292d48
|
Bump flashinfer version to 0.6.13 (#46683)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-06-29 09:30:57 -07:00 |
|
 
|
debec6440b
|
Add MiniMax-M3 modelopt nvfp4 support (#46756)
Signed-off-by: Xin Li <xinli@nvidia.com>
Signed-off-by: jasonlizhengjian <jasonlizhengjian@gmail.com>
Co-authored-by: Xin Li <xinli@nvidia.com>
|
2026-06-29 09:29:39 -07:00 |
|
 
|
c8fb2963bd
|
[FS-Offloading] Batch Lookup in C (#46713)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-29 09:28:32 -07:00 |
|
 HDCharlesandGitHub
|
379acd4e4f
|
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression (#46860)
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com>
|
2026-06-29 15:55:42 +00:00 |
|
 Martin HickeyandGitHub
|
07d33e575b
|
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager (#44657)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
2026-06-29 16:42:35 +01:00 |
|
 
|
36bbecd643
|
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" (#46958)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-29 07:54:34 -07:00 |
|
 Nicolò LucchesiandGitHub
|
6149187a4c
|
[Kernel] Triton MLA logits workspace (#46819)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
|
2026-06-29 07:54:29 -07:00 |
|
 Xiaohong (Sean) ChenandGitHub
|
49e28e8e91
|
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope (#44010)
Signed-off-by: Sean Chen <seachen@redhat.com>
|
2026-06-29 22:54:15 +08:00 |
|
 
|
0ca39c4f1f
|
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone (#46973)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-29 10:00:31 -04:00 |
|
 Blas Rodriguez IrizarandGitHub
|
6185d73882
|
[Rust Frontend] Keep literal "null" string for string-typed tool params (#46827)
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com>
|
2026-06-29 13:46:33 +00:00 |
|
 
|
bc8481af09
|
[MoE Refactor] Standardize Humming MoE experts + utilities (#43373)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-06-29 06:19:29 -07:00 |
|
 
|
59575da46d
|
[XPU] exclude unsupported models for test_tensor_sechma.py (#47008)
Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-29 12:30:28 +00:00 |
|
 wang.yuqiandGitHub
|
3483240b7e
|
[Frontend] Consolidate scale out entrypoints (#44512)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-29 03:18:53 -07:00 |
|
 Roberto L. CastroandGitHub
|
eddfd4cf21
|
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen (#46750)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
|
2026-06-29 10:10:07 +00:00 |
|
 Martin HickeyandGitHub
|
a4e3cb40d0
|
[mypy] Enable mypy for tests directory (#47018)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
2026-06-29 09:29:09 +00:00 |
|
 soaringkandGitHub
|
ab132ee98b
|
Fix model info cache for package models (#46567)
Signed-off-by: soaringk <k3vin.zhang@gmail.com>
|
2026-06-29 09:17:54 +00:00 |
|
 
|
e186107870
|
[Bugfix] Use native SiLU activation in CPU fused MoE (#45961)
Signed-off-by: Alden Lobo <alden.lobo@arm.com>
Co-authored-by: Alden Lobo <alden.lobo@arm.com>
|
2026-06-29 09:12:20 +00:00 |
|
 
|
0e207dac78
|
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models (#46835)
Signed-off-by: John Langford <jl@hunch.net>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 08:59:15 +00:00 |
|
 wang.yuqiandGitHub
|
9e86352c60
|
[CI Failure] Add transformers version check for openai/privacy-filter (#47011)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-29 08:57:26 +00:00 |
|
 Harry MellorandGitHub
|
5051698e41
|
Remove unnecessary load_weights methods (#44589)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 01:52:23 -07:00 |
|
 Andreas KaratzasandGitHub
|
db28ae2d07
|
[ROCm][CI] Explicitly tear down multimodal offline LLMs (#46999)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-29 07:59:24 +00:00 |
|
 Harry MellorandGitHub
|
f6bb8682ee
|
Fix docs on main (#47009)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 15:50:57 +08:00 |
|
 
|
4559c43a95
|
[MM][CG] Gemma3 Encoder CUDA Graph (#43591)
Signed-off-by: JisoLya <523420504@qq.com>
Signed-off-by: Soyaazz <523420504@qq.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-29 04:52:00 +00:00 |
|
 Bugen ZhaoandGitHub
|
5274c1181d
|
[Rust Frontend] Add Harmony Renderer for GPT-OSS (#46800)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-29 03:39:04 +00:00 |
|
 Yuwen ZhouandGitHub
|
58d6a6e60a
|
[CPU] Support cpu compressed-tensor w8a8 int8 moe (#42920)
Signed-off-by: yuwenzho <yuwen.zhou@intel.com>
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com>
|
2026-06-29 03:04:05 +00:00 |
|
 
|
a2abce646f
|
[EPLB] Mask padding in EPLB load recording (#38128)
Signed-off-by: ilmarkov <markovilya197@gmail.com>
Signed-off-by: Markov Ilya <markovilya19@gmail.com>
Co-authored-by: Markov Ilya <markovilya19@gmail.com>
|
2026-06-28 19:43:58 -07:00 |
|
 Harry MellorandGitHub
|
311ad689ad
|
Remove boilerplate missed by #46820 (#46956)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 08:11:17 +08:00 |
|
 Woosuk KwonandGitHub
|
0472436541
|
[Spec Decode] Avoid redundant hidden-states gather in draft prefill (#46968)
|
2026-06-28 17:04:01 -07:00 |
|
  
|
4dfbf1503b
|
[Model] Add support for openai/privacy-filter (#41026)
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-06-28 16:18:22 -07:00 |
|
 Wei ZhaoandGitHub
|
95528527ea
|
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 (#46855)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-06-28 14:36:23 -07:00 |
|
  
|
c2127a25c7
|
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm (#46895)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:50:30 -05:00 |
|
 
|
03c6d01c30
|
[OCP MX ] Add back emulation to available OCP MX backends list (#46629)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:43:19 -05:00 |
|
 Woosuk KwonandGitHub
|
4b643c463e
|
[GLM5] Fix minor typo (#46961)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-28 08:37:00 -07:00 |
|
 
|
7544286b04
|
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE (#46552)
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-28 15:19:28 +00:00 |
|
 Woosuk KwonandGitHub
|
89876b0c54
|
[GLM5] Implement op fusion for GLM5/DSV3.2 (#46876)
|
2026-06-28 08:17:39 -07:00 |
|
 Wentao YeandGitHub
|
5c91039c41
|
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement (#46635)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-28 14:55:54 +00:00 |
|