Logo
Explore Help
Sign In
Karylab-cklius/vllm
Watch 1
Star 0
Fork 1
Code Issues Pull Requests 1 Actions Packages Projects Releases Wiki Activity
Files
b8fb56d970a4fa2a8d668d7e09ee5c7fbb7b680b
vllm/vllm/model_executor
T
History
Aritra Roy GosthipatyGitHubHarry Mellor
33178f9006 Fix Qwen3-VL M-RoPE on the Transformers modeling backend (grids + compile) (#49292)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-07-21 18:28:39 +00:00
..
kernels
[XPU] FP8 o_proj with fp8_bmm and load-time scale transpose (#48334)
2026-07-20 16:32:03 +08:00
layers
[3/N][KV-Cache Layout Refactor] Standardize Mamba cache; drop get_transfer_cache_regions (#44456)
2026-07-21 09:16:15 +00:00
model_loader
[Bugfix] Re-sync parameter tp_rank after process_weights_after_loading (fix replicated / disable_tp weight reload) (#48025)
2026-07-18 08:40:06 +00:00
models
Fix Qwen3-VL M-RoPE on the Transformers modeling backend (grids + compile) (#49292)
2026-07-21 18:28:39 +00:00
offloader
Revert "[Platform] Replace torch.cuda.Event with torch.Event (#47140)" (#47668)
2026-07-06 14:18:33 +01:00
warmup
[Bugfix] Handle MLA fallback during FA4 JIT warmup (#49306)
2026-07-21 13:58:05 +00:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
Speed up docs build (#44635)
2026-06-05 14:51:44 +00:00
utils.py
Remove more unnecessary load_weights methods (#47058)
2026-06-30 15:22:16 +01:00
Powered by Gitea Version: 1.27.1 Page: 476ms Template: 4ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API