This website requires JavaScript.
Explore
Help
Sign In
Karylab-cklius
/
vllm
Watch
1
Star
0
Fork
1
Code
Issues
Pull Requests
1
Actions
Packages
Projects
Releases
Wiki
Activity
Files
f53fa26e05c476a43f6db048a9e3b43bcb2b72fb
vllm
/
tests
/
kernels
T
History
shunting314
and
GitHub
8b141ed8c3
full cudagraph for flex-attn (
#36298
)
...
Signed-off-by: shunting314 <
shunting@meta.com
>
2026-04-02 21:15:01 -07:00
..
attention
[Kernel] Fuse FP8 output quantization into merge_attn_states (
#36518
)
2026-04-03 01:47:04 +00:00
core
Feature/silu block quant fusion v1 (
#32996
)
2026-04-01 18:50:43 +00:00
helion
…
ir
…
mamba
…
moe
[CPU] Support gelu act in cpu_fused_moe (
#38770
)
2026-04-02 14:14:32 +08:00
quantization
…
__init__.py
…
allclose_default.py
…
quant_utils.py
…
test_apply_repetition_penalties.py
…
test_awq_int4_to_int8.py
…
test_cache_kernels.py
…
test_concat_mla_q.py
…
test_cp_gather_fp8.py
…
test_fla_layernorm_guard.py
…
test_flex_attention.py
full cudagraph for flex-attn (
#36298
)
2026-04-02 21:15:01 -07:00
test_fused_gdn_post_conv.py
[Perf] fuse kernels in gdn (
#37813
)
2026-04-02 11:52:18 +00:00
test_fused_quant_activation.py
…
test_fused_recurrent_packed_decode.py
…
test_fused_sigmoid_gating_delta_rule.py
…
test_onednn.py
…
test_shuffle_rows.py
…
test_top_k_per_row.py
…
utils.py
…