This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
2
Packages
Activity
Files
kernel-block-size-alignment-ssm
vllm
/
tests
/
compile
/
fusions_e2e
T
Add File
New File
Upload File
Apply Patch
Copy Permalink
Download directory as ZIP
Download directory as TAR.GZ
Delete Directory
History
elvischenv
and
GitHub
296839a1b0
[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE (
#30647
)
...
Signed-off-by: elvischenv <
219235043+elvischenv@users.noreply.github.com
>
2026-03-18 15:01:26 +00:00
..
__init__.py
[CI][torch.compile] Reduce e2e fusion test time (
#33293
)
2026-02-04 19:09:03 -05:00
common.py
[ROCm] [CI] Add new fusion test cases that are relevant to vLLM IR Ops (
#34307
)
2026-03-03 06:24:21 -08:00
conftest.py
[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE (
#30647
)
2026-03-18 15:01:26 +00:00
models.py
[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE (
#30647
)
2026-03-18 15:01:26 +00:00
test_tp1_quant.py
[torch.compile] Add support for non-contiguous fused RMSNorm + group quant (
#36551
)
2026-03-11 10:56:55 -07:00
test_tp2_ar_rms.py
[Perf] Eliminate padding and slicing op for GPT-OSS with Flashinfer MXFP4 MXFP8 MoE (
#30647
)
2026-03-18 15:01:26 +00:00
test_tp2_async_tp.py
[ROCm] [CI] Add new fusion test cases that are relevant to vLLM IR Ops (
#34307
)
2026-03-03 06:24:21 -08:00