Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
29982d48b3640bb527bffd37fb02e06deb934849
vllm/vllm/model_executor
T
History
Lucas Wilkinsonandkhluu 0ee3b7fc3d [Bugfix][MLA] Add logits size budget to sparse indexer prefill chunking (#36178)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
(cherry picked from commit eb47454987)
2026-04-01 01:02:58 -07:00
..
kernels
[CPU] Support CT W4A16 on CPU MP kernel (#38219)
2026-03-27 14:15:28 +08:00
layers
[Bugfix][MLA] Add logits size budget to sparse indexer prefill chunking (#36178)
2026-04-01 01:02:58 -07:00
model_loader
[QeRL] Fix online quantized reloading (#38442)
2026-03-29 14:56:41 -06:00
models
[Perf] Remove redundant device copies for CPU-only pooling token IDs, 48.9% E2E throughput improvement (#38139)
2026-03-29 18:12:50 +00:00
offloader
Bugfix for offloading+prefetch for GLM-4.7-FP8 (#37178)
2026-03-17 21:22:09 +08:00
warmup
[Bugfix] Fix AttributeError when serving MXFP8 models with DeepGEMM installed (#37358)
2026-03-19 17:58:33 +00:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
[Mypy] Fix mypy for vllm/model_executor (except vllm/model_executor/layers) (#37904)
2026-03-24 17:14:01 +00:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (#33262)
2026-01-29 16:52:11 +08:00
Powered by Gitea Version: 1.27.1 Page: 210ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API