This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
2aa9831dd380ccfcf219068ee015e00a018fa4ca
vllm
/
csrc
/
quantization
/
gptq
T
History
Antoni Baum
and
GitHub
a10d3056da
[Core] Set
linear_weights
directly on the layer (
#3977
)
2024-04-11 16:35:51 -04:00
..
compat.cuh
Add GPTQ support (
#916
)
2023-12-15 03:04:22 -08:00
matrix_view.cuh
Add Support for 2/3/8-bit GPTQ Quantization Models (
#2330
)
2024-02-28 21:52:23 -08:00
q_gemm.cu
[Core] Set
linear_weights
directly on the layer (
#3977
)
2024-04-11 16:35:51 -04:00
qdq_2.cuh
Add Support for 2/3/8-bit GPTQ Quantization Models (
#2330
)
2024-02-28 21:52:23 -08:00
qdq_3.cuh
Add Support for 2/3/8-bit GPTQ Quantization Models (
#2330
)
2024-02-28 21:52:23 -08:00
qdq_4.cuh
Add Support for 2/3/8-bit GPTQ Quantization Models (
#2330
)
2024-02-28 21:52:23 -08:00
qdq_8.cuh
Add Support for 2/3/8-bit GPTQ Quantization Models (
#2330
)
2024-02-28 21:52:23 -08:00
qdq_util.cuh
Add GPTQ support (
#916
)
2023-12-15 03:04:22 -08:00