Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
592ae6805cb87c9a44c29bd9c6eef9d04d91e39b
vllm/tests/models/quantization
T
History
fxmarty-amdGitHubKyle Sayers
d622e27d2b [NVFP4] NVFP4 MOE emulation fallback for H100/MI300/MI350, standardize TritonExperts usage for OCP MX emulation (#35737)
Signed-off-by: Felix Marty <Felix.Marty@amd.com>
Signed-off-by: fxmarty-amd <felmarty@amd.com>
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com>
2026-04-22 08:58:54 -07:00
..
__init__.py
…
test_awq.py
[Renderer] Define render_cmpl and render_chat (#34039)
2026-02-07 05:24:40 -08:00
test_bitsandbytes.py
[bnb] Skip moe + bnb test (#36896)
2026-03-12 18:03:25 +00:00
test_fp8.py
[1/N][Attention] Restructure attention: move files (#31916)
2026-01-09 13:10:24 -08:00
test_gguf.py
[Bugfix][Quantization] Support BF16 tensors on GGUF (#29948)
2025-12-03 10:33:46 +00:00
test_gpt_oss.py
replace cuda_device_count_stateless() to current_platform.device_count() (#37841)
2026-03-31 22:32:54 +08:00
test_gptq_marlin.py
…
test_modelopt.py
…
test_mxfp4.py
…
test_mxfp8.py
[Kernel] Add MXFP8 to Marlin GEMM/MoE and refactor Mxfp8LinearOp (#34664)
2026-04-01 09:41:42 -07:00
test_nvfp4.py
[NVFP4] NVFP4 MOE emulation fallback for H100/MI300/MI350, standardize TritonExperts usage for OCP MX emulation (#35737)
2026-04-22 08:58:54 -07:00
test_per_token_kv_cache.py
[Feature] KV cache per-token-head INT8/FP8 quantization (#38378)
2026-04-02 08:13:26 -04:00
Powered by Gitea Version: 1.27.1 Page: 4764ms Template: 4ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API