This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
eb33ff34dd6501f23f5be92a5c842d34ab05d7fd
vllm
/
tests
/
models
/
quantization
T
History
3 people
fxmarty-amd
GitHub
Andreas Karatzas
03c6d01c30
[OCP MX ] Add back emulation to available OCP MX backends list (
#46629
)
...
Signed-off-by: Felix Marty <
Felix.Marty@amd.com
> Co-authored-by: Andreas Karatzas <
akaratza@amd.com
>
2026-06-28 12:43:19 -05:00
..
__init__.py
…
test_awq.py
[BugFix] Fix Gemma4 'layers.0.moe.experts.0.down_proj_packed' KeyError issue (
#40708
)
2026-05-09 17:20:44 +00:00
test_bitsandbytes.py
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads (
#45308
)
2026-06-11 20:21:46 -07:00
test_fp8_per_channel.py
[Quantization] add online fp8 ptpc (
#44132
)
2026-06-08 22:42:11 +08:00
test_fp8.py
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants (
#42685
)
2026-05-15 15:34:59 -04:00
test_gpt_oss.py
[OCP MX ] Add back emulation to available OCP MX backends list (
#46629
)
2026-06-28 12:43:19 -05:00
test_gptq_marlin.py
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py (
#38288
)
2026-05-15 08:25:52 +08:00
test_modelopt.py
…
test_mxfp4.py
…
test_mxfp8.py
[Kernel] Add MXFP8 to Marlin GEMM/MoE and refactor Mxfp8LinearOp (
#34664
)
2026-04-01 09:41:42 -07:00
test_nvfp4.py
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend (
#42235
)
2026-06-22 16:17:18 -04:00
test_per_token_kv_cache.py
[Feature] Triton INT4 per-token-head KV cache quantization (
#40835
)
2026-06-24 10:21:25 +00:00