This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
d79855eaacfd284689fe4ff6eefc280e1bb2473c
vllm
/
vllm
/
platforms
T
History
liuzhenwei
and
GitHub
736f1a5907
[XPU] Route mm_prefix models to Triton attention backend (
#47688
)
...
Signed-off-by: zhenwei-intel <
zhenwei.liu@intel.com
>
2026-07-06 16:52:44 +08:00
..
__init__.py
[Core] Ensure memory is pinned prior to async h2d copy (
#45424
)
2026-06-20 20:02:24 -07:00
cpu.py
[CPU][Perf]Added tanh AOR for faster gelu activations. (
#44639
)
2026-06-30 23:24:40 -07:00
cuda.py
New stable abi cleanup (
#46656
)
2026-07-03 14:02:26 +08:00
interface.py
Support nvfp4 kv with kv-cache-dtype-skip-layers sliding_window (
#42890
)
2026-07-04 02:29:13 +00:00
rocm.py
[ROCm] Remove erroneous inclusion of gptq_marlin as supported quant scheme on ROCm (
#46655
)
2026-06-24 21:27:19 -05:00
tpu.py
[Refactor][TPU] Remove torch_xla path and use tpu-inference (
#30808
)
2026-01-07 16:07:16 +08:00
xpu.py
[XPU] Route mm_prefix models to Triton attention backend (
#47688
)
2026-07-06 16:52:44 +08:00
zen_cpu.py
[ZenCPU] AMD Zen CPU Backend with supported dtypes via zentorch weekly (
#39967
)
2026-04-18 06:22:37 +00:00