 
|
c8fb2963bd
|
[FS-Offloading] Batch Lookup in C (#46713)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-29 09:28:32 -07:00 |
|
 HDCharlesandGitHub
|
379acd4e4f
|
[Bugfix][Quantization] Fix W8A8 int-quantized scheme selection regression (#46860)
Signed-off-by: HDCharles <charlesdavidhernandez@gmail.com>
|
2026-06-29 15:55:42 +00:00 |
|
 Martin HickeyandGitHub
|
07d33e575b
|
[MyPy] Fix mypy incompatible assignment errors in LRUCacheLoRAModelManager (#44657)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
2026-06-29 16:42:35 +01:00 |
|
 
|
36bbecd643
|
[BugFix] Revert "[KV Offload] Use background thread for mmap / cpu_tensors pinning" (#46958)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-29 07:54:34 -07:00 |
|
 Nicolò LucchesiandGitHub
|
6149187a4c
|
[Kernel] Triton MLA logits workspace (#46819)
Signed-off-by: NickLucche <nicolo.lucchesi@mistral.ai>
|
2026-06-29 07:54:29 -07:00 |
|
 Xiaohong (Sean) ChenandGitHub
|
49e28e8e91
|
[Kernel][Helion][1/N] Add Helion kernel for fused_qk_norm_rope (#44010)
Signed-off-by: Sean Chen <seachen@redhat.com>
|
2026-06-29 22:54:15 +08:00 |
|
 
|
0ca39c4f1f
|
[Bugfix] Capture final-layer aux hidden state in deepseek_v2 backbone (#46973)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-29 10:00:31 -04:00 |
|
 Blas Rodriguez IrizarandGitHub
|
6185d73882
|
[Rust Frontend] Keep literal "null" string for string-typed tool params (#46827)
Signed-off-by: Blas Rodriguez Irizar <rodrigblas@gmail.com>
|
2026-06-29 13:46:33 +00:00 |
|
 
|
bc8481af09
|
[MoE Refactor] Standardize Humming MoE experts + utilities (#43373)
Signed-off-by: Bill Nell <bnell@redhat.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-06-29 06:19:29 -07:00 |
|
 
|
59575da46d
|
[XPU] exclude unsupported models for test_tensor_sechma.py (#47008)
Signed-off-by: Yan Ma <yan.ma@intel.com>
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-29 12:30:28 +00:00 |
|
 wang.yuqiandGitHub
|
3483240b7e
|
[Frontend] Consolidate scale out entrypoints (#44512)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-29 03:18:53 -07:00 |
|
 Roberto L. CastroandGitHub
|
eddfd4cf21
|
[Perf][2/N] Expand Triton kernel warmup coverage, Qwen (#46750)
Signed-off-by: LopezCastroRoberto <rocastro@redhat.com>
|
2026-06-29 10:10:07 +00:00 |
|
 Martin HickeyandGitHub
|
a4e3cb40d0
|
[mypy] Enable mypy for tests directory (#47018)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
2026-06-29 09:29:09 +00:00 |
|
 soaringkandGitHub
|
ab132ee98b
|
Fix model info cache for package models (#46567)
Signed-off-by: soaringk <k3vin.zhang@gmail.com>
|
2026-06-29 09:17:54 +00:00 |
|
 
|
e186107870
|
[Bugfix] Use native SiLU activation in CPU fused MoE (#45961)
Signed-off-by: Alden Lobo <alden.lobo@arm.com>
Co-authored-by: Alden Lobo <alden.lobo@arm.com>
|
2026-06-29 09:12:20 +00:00 |
|
 
|
0e207dac78
|
[Bugfix] Transformers backend: apply learned lm_head.bias for tied-embedding models (#46835)
Signed-off-by: John Langford <jl@hunch.net>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 08:59:15 +00:00 |
|
 wang.yuqiandGitHub
|
9e86352c60
|
[CI Failure] Add transformers version check for openai/privacy-filter (#47011)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-29 08:57:26 +00:00 |
|
 Harry MellorandGitHub
|
5051698e41
|
Remove unnecessary load_weights methods (#44589)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 01:52:23 -07:00 |
|
 Andreas KaratzasandGitHub
|
db28ae2d07
|
[ROCm][CI] Explicitly tear down multimodal offline LLMs (#46999)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-29 07:59:24 +00:00 |
|
 Harry MellorandGitHub
|
f6bb8682ee
|
Fix docs on main (#47009)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 15:50:57 +08:00 |
|
 
|
4559c43a95
|
[MM][CG] Gemma3 Encoder CUDA Graph (#43591)
Signed-off-by: JisoLya <523420504@qq.com>
Signed-off-by: Soyaazz <523420504@qq.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-29 04:52:00 +00:00 |
|
 Bugen ZhaoandGitHub
|
5274c1181d
|
[Rust Frontend] Add Harmony Renderer for GPT-OSS (#46800)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-29 03:39:04 +00:00 |
|
 Yuwen ZhouandGitHub
|
58d6a6e60a
|
[CPU] Support cpu compressed-tensor w8a8 int8 moe (#42920)
Signed-off-by: yuwenzho <yuwen.zhou@intel.com>
Signed-off-by: Yuwen Zhou <yuwen.zhou@intel.com>
|
2026-06-29 03:04:05 +00:00 |
|
 
|
a2abce646f
|
[EPLB] Mask padding in EPLB load recording (#38128)
Signed-off-by: ilmarkov <markovilya197@gmail.com>
Signed-off-by: Markov Ilya <markovilya19@gmail.com>
Co-authored-by: Markov Ilya <markovilya19@gmail.com>
|
2026-06-28 19:43:58 -07:00 |
|
 Harry MellorandGitHub
|
311ad689ad
|
Remove boilerplate missed by #46820 (#46956)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 08:11:17 +08:00 |
|
 Woosuk KwonandGitHub
|
0472436541
|
[Spec Decode] Avoid redundant hidden-states gather in draft prefill (#46968)
|
2026-06-28 17:04:01 -07:00 |
|
  
|
4dfbf1503b
|
[Model] Add support for openai/privacy-filter (#41026)
Signed-off-by: Fabian Joswig <fjosw@users.noreply.github.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
|
2026-06-28 16:18:22 -07:00 |
|
 Wei ZhaoandGitHub
|
95528527ea
|
[Bugfix][Mooncake] Fix Mooncake lookup prefixes with DCP > 1 (#46855)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-06-28 14:36:23 -07:00 |
|
  
|
c2127a25c7
|
[ROCm][CI] Fix rlhf_async_new_apis Example On ROCm (#46895)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Matthew Wong <Matthew.Wong2@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:50:30 -05:00 |
|
 
|
03c6d01c30
|
[OCP MX ] Add back emulation to available OCP MX backends list (#46629)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 12:43:19 -05:00 |
|
 Woosuk KwonandGitHub
|
4b643c463e
|
[GLM5] Fix minor typo (#46961)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-28 08:37:00 -07:00 |
|
 
|
7544286b04
|
[Bugfix] Transformers backend: recompute mm_token_type_ids per request for M-RoPE (#46552)
Signed-off-by: Gonzague de Carpentier <decarpentierg@gmail.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-28 15:19:28 +00:00 |
|
 Woosuk KwonandGitHub
|
89876b0c54
|
[GLM5] Implement op fusion for GLM5/DSV3.2 (#46876)
|
2026-06-28 08:17:39 -07:00 |
|
 Wentao YeandGitHub
|
5c91039c41
|
[GLM5.2 Perf] Replace MOE all-reduce with reduce-scatter, 3.1%~3.2 E2E Throughput improvement (#46635)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-28 14:55:54 +00:00 |
|
 
|
5ecae3266c
|
[ROCm][Perf][MLA] Add AITER FlashAttention MLA prefill backend (ROCM_AITER_FA) (#45033)
Signed-off-by: Xavier Aguilar <xavier.aguilarfruto@amd.com>
Signed-off-by: Xavier Aguilar <Xavier.AguilarFruto@amd.com>
Co-authored-by: TJian <tunjian.tan@embeddedllm.com>
|
2026-06-28 07:52:00 -07:00 |
|
 
|
6eb63a1da6
|
[Bugfix][DSv3.2] Skip indexer weights for index-cache-skipped layers (#46600)
Signed-off-by: Frida Andersson <fanderss@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-28 01:37:44 -07:00 |
|
 
|
09841ae705
|
[Render][Speculator] Add return_loss_mask to render endpoint for training data generation (#46846)
Signed-off-by: Ranran Haoran Zhang <ranzhang@redhat.com>
Co-authored-by: Benjamin Chislett <chislett.ben@gmail.com>
|
2026-06-28 00:07:33 -07:00 |
|
 MattandGitHub
|
a2a92cbbaa
|
[Hardware][AMD][CI] Tweak mirrored tests; improve CI base dependency change detection (#46930)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-28 00:07:14 -07:00 |
|
 
|
35e6c86caa
|
[Bugfix][MM][CG] Enable dual-path ViT CUDA graph for Step3-VL (#46034)
Signed-off-by: shen-shanshan <467638484@qq.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-28 00:06:43 -07:00 |
|
  
|
c7ca0bccae
|
[ROCm][Perf] Add Fused Shared Expert (FSE) support for GLM-4.5/6/7 (#44313)
Signed-off-by: Olga Miroshnichenko <olga.miroshnichenko@amd.com>
Signed-off-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com>
Co-authored-by: Mehdi Ghanimifard <mghanimi@amd.com>
Co-authored-by: Mehdi Ghanimifard <mehdi.ghanimifard@amd.com>
|
2026-06-28 00:04:08 -07:00 |
|
  
|
c6741b2ad4
|
[Model] Support Unlimited OCR (#46564)
Signed-off-by: Tianyu Guo <guoty9@mail2.sysu.edu.cn>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-06-27 23:09:18 -07:00 |
|
 
|
a65f93fb2e
|
[ROCm][CI] Add ci_base metadata for external cache orchestration (#46886)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
Signed-off-by: Codex <codex@example.invalid>
Co-authored-by: Codex <codex@example.invalid>
|
2026-06-28 12:51:19 +08:00 |
|
 ChaunceyandGitHub
|
11a12305c0
|
[Model Runner V2][Spec Decode] Handle tuple hidden states from MTP draft models (#46786)
|
2026-06-27 18:38:07 -07:00 |
|
 
|
798185d438
|
[KV-Offloading] Fix tensors_per_block stride (#46888)
Signed-off-by: <>
Co-authored-by: Varun Sundar Rabindranath <varun-sundar-rabindranath@h100-01.nemg-001.lab.rdu2.dc.redhat.com>
|
2026-06-27 21:01:45 -04:00 |
|
 MattandGitHub
|
9036c89ee4
|
[Hardware][AMD][CI] Patch Whisper multi LoRA test to use TRITON_ATTN for now (#46928)
Signed-off-by: Matthew Wong <Matthew.Wong2@amd.com>
|
2026-06-27 17:30:49 -05:00 |
|
 Giancarlo DelfinandGitHub
|
b6caeb5a09
|
[Model Runner V2][Spec Decode] Use fp32 uniform threshold for acceptance (#46878)
|
2026-06-27 14:09:25 -07:00 |
|
 Taneem IbrahimandGitHub
|
8bf064f8d3
|
Fixed chunked embedding aggregation with request-id metadata (#46782)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-06-27 20:57:47 +00:00 |
|
 
|
ea2ead1db3
|
[Misc] Fix incorrect layer type annotation in Fp8LinearMethod (#46818)
Signed-off-by: shaojinjie.sjj <shaojinjiesjj@gmail.com>
Co-authored-by: shaojinjie.sjj <shaojinjiesjj@gmail.com>
|
2026-06-27 20:23:59 +00:00 |
|
 Wentao YeandGitHub
|
56aa067bf0
|
[CI Bug] Fix h100 AssertionError: Cold-start child failed (#46927)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-27 20:17:33 +00:00 |
|
 xiaolinchenandGitHub
|
35e3850fa9
|
[Bugfix][Test] Fix test_flashinfer_cutlass_mxfp4_fused_moe on sm90 (stale weight/scale interleave) (#46915)
Signed-off-by: wentian-byte <2990624738@qq.com>
|
2026-06-27 14:30:10 -04:00 |
|