 Chris LeonardandGitHub
|
fbc9ba6d30
|
New stable abi cleanup (#46656)
Signed-off-by: Chris Leonard <chleonar@redhat.com>
|
2026-07-03 14:02:26 +08:00 |
|
 Jee Jee LiandGitHub
|
23aed9b0ee
|
[Kernel] Enable PDL for per_token_group_quant_8bit_kernel (#46508)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-25 08:42:51 +08:00 |
|
 JasonLi314andGitHub
|
93bad11912
|
[Bugfix] Fix gridDim.y overflow for large row counts (#45255)
Signed-off-by: Jason Li <li.jason.cs@gmail.com>
|
2026-06-19 23:27:45 -04:00 |
|
 Jee Jee LiandGitHub
|
22cc891108
|
[Kernel] Add PDL support for DeepGEMM kernel (#46006)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-18 20:49:01 +08:00 |
|
 Micah WilliamsonandGitHub
|
e945169207
|
Revert "[Kernel] Add PDL support for DeepGEMM kernel" (#45999)
|
2026-06-17 22:59:48 -07:00 |
|
 Jee Jee LiandGitHub
|
4403af8fb5
|
[Kernel] Add PDL support for DeepGEMM kernel (#42996)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-17 20:37:17 -07:00 |
|
 Charlie FuandGitHub
|
71df063c49
|
Enable perf_token_group_quant/_C_stable_libtorch for ROCm (#42758)
Signed-off-by: charlifu <charlifu@amd.com>
|
2026-06-02 23:23:28 -07:00 |
|
  
|
07aeaf9d4d
|
[6/n] Migrate activation kernels, gptq, gguf, non cutlass w8a8 to libtorch stable ABI (continued) (#42663)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-20 00:18:12 -07:00 |
|
 
|
dd6b3a5ef5
|
[Perf] Use 2D-grid to eliminate divmod in W8W8 group quant (#42153)
Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-12 10:01:30 -04:00 |
|
 
|
a3c83ff2fd
|
Faster per-token fp8 group quant packed kernel for blackwell (#41326)
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-04-30 18:09:55 -07:00 |
|
 Wei ZhaoandGitHub
|
59b2f7b640
|
[Perf] Fuse Zero Initializer for FP8 DeepGemm Block Quant Kernel (#39547)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-04-11 07:16:51 -07:00 |
|
 mikaylagawareckiandGitHub
|
bf4cc9ed2d
|
[2/n] Migrate per_token_group_quant to torch stable ABI (#36058)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
|
2026-03-25 10:15:13 -07:00 |
|