This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
cutlass_fa3_mla_sparse
vllm
/
cmake
/
external_projects
T
Add File
New File
Upload File
Apply Patch
Copy Permalink
Download directory as ZIP
Download directory as TAR.GZ
Delete Directory
History
Alexander Matveev
a60418e6fb
Sparse MLA on Hopper: Use SGLang's kernel for the sparse mla low latency runs
...
Signed-off-by: Alexander Matveev <
amatveev@redhat.com
>
2026-04-17 16:51:32 +00:00
..
cutlass_fa3.cmake
Sparse MLA on Hopper: Use SGLang's kernel for the sparse mla low latency runs
2026-04-17 16:51:32 +00:00
deepgemm.cmake
[UX] Integrate DeepGEMM into vLLM wheel via CMake (
#37980
)
2026-04-08 18:56:32 -07:00
flashmla.cmake
[BugFix] Fix Python 3.13 FlashMLA import error (
#34548
)
2026-02-15 20:09:18 -08:00
qutlass.cmake
[NVIDIA] Fix DGX Spark logic (
#38126
)
2026-03-27 15:26:07 -07:00
triton_kernels.cmake
[Release 2.10] Update to Torch 2.10 - final release (
#30525
)
2026-02-08 13:51:09 -08:00
vllm_flash_attn.cmake
[Attention] relax the head dim 512 and paged kv for sm90+FA4 (
#38835
)
2026-04-08 18:23:18 +00:00