Logo
Explore Help
Sign In
Karylab-cklius/vllm
Watch 1
Star 0
Fork 1
Code Issues Pull Requests 1 Actions Packages Projects Releases Wiki Activity
Files
87c5064d0bf539b7c84ab537a679cf1f657ecffa
vllm/vllm/model_executor
T
History
Woosuk Kwon 87c5064d0b revert bf16 fusion
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-04-17 03:50:43 +00:00
..
kernels
[ROCm][FEAT] Integrate aiter gemm w8a8 ptpc (#33773)
2026-04-16 09:55:29 +08:00
layers
[Quantization] Consolidate experts_int8 with fp8 online quantization (#38463)
2026-04-16 13:12:20 -07:00
model_loader
Merge branch 'main' into woosuk/ds-exp-2
2026-04-17 03:49:34 +00:00
models
revert bf16 fusion
2026-04-17 03:50:43 +00:00
offloader
[Bugfix] Respect VLLM_WEIGHT_OFFLOADING_DISABLE_PIN_MEMORY in prefetch offloader (#37699)
2026-04-14 20:43:29 -07:00
specialized_models
Add fusion back
2026-04-17 02:35:38 +00:00
warmup
[MoE] Move DEEP_GEMM into experts/ subdirectory (#39005)
2026-04-08 19:23:08 +00:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
[Mypy] Fix mypy for vllm/model_executor (except vllm/model_executor/layers) (#37904)
2026-03-24 17:14:01 +00:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (#33262)
2026-01-29 16:52:11 +08:00
Powered by Gitea Version: 1.27.1 Page: 281ms Template: 3ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API