Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
eb33ff34dd6501f23f5be92a5c842d34ab05d7fd
vllm/tests/models/quantization
T
History
fxmarty-amdGitHubAndreas Karatzas
03c6d01c30 [OCP MX ] Add back emulation to available OCP MX backends list (#46629)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-06-28 12:43:19 -05:00
..
__init__.py
…
test_awq.py
[BugFix] Fix Gemma4 'layers.0.moe.experts.0.down_proj_packed' KeyError issue (#40708)
2026-05-09 17:20:44 +00:00
test_bitsandbytes.py
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads (#45308)
2026-06-11 20:21:46 -07:00
test_fp8_per_channel.py
[Quantization] add online fp8 ptpc (#44132)
2026-06-08 22:42:11 +08:00
test_fp8.py
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants (#42685)
2026-05-15 15:34:59 -04:00
test_gpt_oss.py
[OCP MX ] Add back emulation to available OCP MX backends list (#46629)
2026-06-28 12:43:19 -05:00
test_gptq_marlin.py
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py (#38288)
2026-05-15 08:25:52 +08:00
test_modelopt.py
…
test_mxfp4.py
…
test_mxfp8.py
[Kernel] Add MXFP8 to Marlin GEMM/MoE and refactor Mxfp8LinearOp (#34664)
2026-04-01 09:41:42 -07:00
test_nvfp4.py
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend (#42235)
2026-06-22 16:17:18 -04:00
test_per_token_kv_cache.py
[Feature] Triton INT4 per-token-head KV cache quantization (#40835)
2026-06-24 10:21:25 +00:00
Powered by Gitea Version: 1.27.1 Page: 1409ms Template: 3ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API