This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
afd7b1dce94fed484351fafd5bf5ea6601ac621e
vllm
/
csrc
/
libtorch_stable
/
quantization
T
History
Wentao Ye
and
GitHub
37ece593c1
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement (
#42774
)
...
Signed-off-by: yewentao256 <
zhyanwentao@126.com
>
2026-05-18 16:38:12 -07:00
..
awq
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (
#42339
)
2026-05-13 07:24:39 +00:00
cutlass_w4a8
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI (
#37503
)
2026-03-31 10:21:13 -07:00
fp4
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement (
#42774
)
2026-05-18 16:38:12 -07:00
gptq_allspark
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (
#42339
)
2026-05-13 07:24:39 +00:00
hadamard
/hadacore
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (
#42339
)
2026-05-13 07:24:39 +00:00
w8a8
[Bugfix] Fix SM121 (DGX Spark) exclusion from Marlin/CUTLASS FP8 paths (
#35568
)
2026-05-15 10:59:00 -07:00
vectorization_utils.cuh
[2/n] Migrate per_token_group_quant to torch stable ABI (
#36058
)
2026-03-25 10:15:13 -07:00
vectorization.cuh
[2/n] Migrate per_token_group_quant to torch stable ABI (
#36058
)
2026-03-25 10:15:13 -07:00