 Chris LeonardandGitHub
|
b9a7cd464c
|
[12/n] final _C library kernel migration (#45415)
|
2026-06-19 06:57:26 -07:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
5ed15f42b9
|
Fix the E8M0 scale computation in the MXFP4 (W4A4) MOE CUTLASS kernel (#43557)
Signed-off-by: Xin He <xin3.he@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-15 06:04:54 -07:00 |
|
   
|
78e7293bb1
|
[Build] Fix CUDA arch build coverage gaps (#45277)
Signed-off-by: Shengqi Chen <harry-chen@outlook.com>
Co-authored-by: Xin Li <xinli-sw@users.noreply.github.com>
Co-authored-by: ShawRong <ShawRong@users.noreply.github.com>
Co-authored-by: Change72 <Change72@users.noreply.github.com>
|
2026-06-13 22:09:20 -07:00 |
|
 Chris LeonardandGitHub
|
56aff0dd15
|
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to torch stale ABI. (#44334)
|
2026-06-04 20:14:43 -07:00 |
|
 Chris LeonardandGitHub
|
4d93bc35c9
|
Migrate header files to torch stable abi (#44013)
|
2026-06-02 08:09:52 -07:00 |
|
  
|
22a58640b4
|
[9/n] Migrate attention and cache kernels to torch stable ABI (continued) (#43717)
Signed-off-by: Chris Leonard <chleonar@redhat.com>
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-05-29 04:44:45 +00:00 |
|
  
|
5bb8d2767a
|
[Kernel] Batch invariant NVFP4 linear using cutlass (#39912)
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: Yongye Zhu <zyy1102000@gmail.com>
|
2026-05-23 09:41:12 -04:00 |
|
 Wentao YeandGitHub
|
37ece593c1
|
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement (#42774)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-05-18 16:38:12 -07:00 |
|
 Jakub ZakrzewskiandGitHub
|
6fbec8ed47
|
[Bugfix][Kernel] nvfp4 cutlass MoE: fix nvfp4 experts quant out-of-bounds read for expert counts not divisible by 4 or 16 (#40351)
Signed-off-by: Jakub Zakrzewski <jzakrzewski@nvidia.com>
|
2026-04-21 19:06:09 +00:00 |
|
  
|
d0359f3e04
|
[Bugfix] Guard mxfp4_experts_quant bindings on ENABLE_NVFP4_SM100 (#40191)
Signed-off-by: ultranationalism <www913363043@gmail.com>
Signed-off-by: mgoin <mike.goin12@gmail.com>
Co-authored-by: mgoin <mike.goin12@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-18 13:58:46 -07:00 |
|
 Michael GoinandGitHub
|
a8bffaa133
|
[Kernel] Add MXFP4 W4A4 CUTLASS MoE kernel for SM100 (#37463)
Signed-off-by: mgoin <mgoin64@gmail.com>
|
2026-04-17 16:42:32 -07:00 |
|
 Wentao YeandGitHub
|
c9a9db0e02
|
[Compile] Fix nvfp4 compile warning (#38573)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-04-01 18:28:57 +00:00 |
|
 mikaylagawareckiandGitHub
|
7c080dd3c5
|
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI (#37503)
Signed-off-by: Mikayla Gawarecki <mikaylagawarecki@gmail.com>
|
2026-03-31 10:21:13 -07:00 |
|