This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
2eb9f4ff262bb39859baebf8d2109abcdadee860
vllm
/
csrc
/
attention
T
History
Michael Goin
and
GitHub
978aed5300
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
..
attention_dtypes.h
Enable scaled FP8 (e4m3fn) KV cache on ROCm (AMD GPU) (
#3290
)
2024-04-03 14:15:55 -07:00
attention_generic.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00
attention_kernels.cu
[Kernel][Attention] Separate
Attention.kv_scale
into
k_scale
and
v_scale
(
#6081
)
2024-07-16 15:31:32 -07:00
attention_utils.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00
dtype_bfloat16.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00
dtype_float16.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00
dtype_float32.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00
dtype_fp8.cuh
[CI/Build] Enforce style for C++ and CUDA code with
clang-format
(
#4722
)
2024-05-22 07:18:41 +00:00