Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
e20a7be4d06a1d6e063cfa00fc93ec1a3980f7f7
vllm/csrc/libtorch_stable/quantization
T
History
yewentao256 e20a7be4d0 optimize per token group quant
Signed-off-by: yewentao256 <zhyanwentao@126.com>
2026-05-19 20:13:19 +00:00
..
awq
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (#42339)
2026-05-13 07:24:39 +00:00
cutlass_w4a8
[4/n] Migrate FP4/W4A8 CUTLASS kernels to torch stable ABI (#37503)
2026-03-31 10:21:13 -07:00
fp4
[Perf] Padded nvfp4 quant kernel to remove additional copy, 2.4%~5.7% e2e performance improvement (#42774)
2026-05-18 16:38:12 -07:00
gptq_allspark
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (#42339)
2026-05-13 07:24:39 +00:00
hadamard/hadacore
[5/n] Migrate CUTLASS MLA, hadamard, awq, allspark and DSV3 fused a gemm to torch stable ABI (continued) (#42339)
2026-05-13 07:24:39 +00:00
w8a8
optimize per token group quant
2026-05-19 20:13:19 +00:00
vectorization_utils.cuh
[2/n] Migrate per_token_group_quant to torch stable ABI (#36058)
2026-03-25 10:15:13 -07:00
vectorization.cuh
[2/n] Migrate per_token_group_quant to torch stable ABI (#36058)
2026-03-25 10:15:13 -07:00
Powered by Gitea Version: 1.27.1 Page: 234ms Template: 4ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API