Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 2 Packages Activity
Files
67eb6083e38d1a65ae41cd00a573e6e95859751a
vllm/vllm/model_executor
T
History
Harry MellorandGitHub 6ee081d1d0 Add new tp plan styles to the Transformers modelling backend (#40467)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-04-21 08:51:30 -07:00
..
kernels
Properly enable wvSplitK fp8 path for RDNA (#37712)
2026-04-20 10:09:24 -05:00
layers
[MoE] Triton MoE Perf regression - restore low latency path (#39016)
2026-04-21 02:37:11 -04:00
model_loader
[Quantization] - Layerwise reloading of Attention/KV quantized models (#38995)
2026-04-15 18:03:32 -07:00
models
Add new tp plan styles to the Transformers modelling backend (#40467)
2026-04-21 08:51:30 -07:00
offloader
[Bugfix] Respect VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY in prefetch offloader (#37699)
2026-04-14 20:43:29 -07:00
warmup
mxfp8 online quant move to new frontend (#40152)
2026-04-20 06:26:12 -07:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
[Mypy] Fix mypy for vllm/model_executor (except vllm/model_executor/layers) (#37904)
2026-03-24 17:14:01 +00:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (#33262)
2026-01-29 16:52:11 +08:00
Powered by Gitea Version: 1.27.1 Page: 314ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API