This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
091bc1026ea4e5120893ae9a730fdb1bf5e873e2
vllm
/
tests
/
models
/
quantization
T
History
Mohammad Miadh Angkad
and
GitHub
d1a38c2762
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend (
#42235
)
...
Signed-off-by: Mohammad Miadh Angkad <
176301910+mmangkad@users.noreply.github.com
>
2026-06-22 16:17:18 -04:00
..
__init__.py
[CI/Build] Reorganize models tests (
#17459
)
2025-04-30 23:03:08 -07:00
test_awq.py
[BugFix] Fix Gemma4 'layers.0.moe.experts.0.down_proj_packed' KeyError issue (
#40708
)
2026-05-09 17:20:44 +00:00
test_bitsandbytes.py
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads (
#45308
)
2026-06-11 20:21:46 -07:00
test_fp8_per_channel.py
[Quantization] add online fp8 ptpc (
#44132
)
2026-06-08 22:42:11 +08:00
test_fp8.py
[FlashAttn] Fix supports_kv_cache_dtype() accepting unhandled fp8 kv-cache dtype variants (
#42685
)
2026-05-15 15:34:59 -04:00
test_gpt_oss.py
[BugFix][Platform] Fix import vllm.platforms.rocm error on non-CUDA test_gpt_oss.py (
#43571
)
2026-05-29 23:16:49 -07:00
test_gptq_marlin.py
[Quant] Consolidate GPTQ: rename gptq_marlin.py to auto_gptq.py (
#38288
)
2026-05-15 08:25:52 +08:00
test_modelopt.py
Convert formatting to use
ruff
instead of
yapf
+
isort
(
#26247
)
2025-10-05 07:06:22 -07:00
test_mxfp4.py
Convert formatting to use
ruff
instead of
yapf
+
isort
(
#26247
)
2025-10-05 07:06:22 -07:00
test_mxfp8.py
[Kernel] Add MXFP8 to Marlin GEMM/MoE and refactor Mxfp8LinearOp (
#34664
)
2026-04-01 09:41:42 -07:00
test_nvfp4.py
[Kernel][Performance] Add FlashInfer cutedsl NVFP4 GEMM backend (
#42235
)
2026-06-22 16:17:18 -04:00
test_per_token_kv_cache.py
[Feature] KV cache per-token-head INT8/FP8 quantization (
#38378
)
2026-04-02 08:13:26 -04:00