Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 2 Packages Activity
Files
b0aa5e815f059c75da92ba5b71db1b1ece8a7e0b
vllm/vllm/model_executor
T
History
Nick HillandGitHub 443e68cfa6 [Bugfix] Fix pooled Whisper encoder sliding-window kernel size (#47437)
Signed-off-by: Nick Hill <nickhill123@gmail.com>
2026-07-02 10:19:11 -07:00
..
kernels
[Attention][DSA] support dcp for FLASHINFER_MLA_SPARSE (#46076)
2026-07-01 00:32:20 -04:00
layers
support GLM-5.2 gate use FP32 (#47410)
2026-07-02 22:45:39 +08:00
model_loader
Remove more unnecessary load_weights methods (#47058)
2026-06-30 15:22:16 +01:00
models
[Bugfix] Fix pooled Whisper encoder sliding-window kernel size (#47437)
2026-07-02 10:19:11 -07:00
offloader
[Platform] Replace torch.cuda.Event with torch.Event (#47140)
2026-06-30 08:39:59 -07:00
warmup
[Model Runner V2][Perf] Warm up GLM-5.2 DSA indexer prefill metadata kernel (#47285)
2026-07-02 16:31:31 +00:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
Speed up docs build (#44635)
2026-06-05 14:51:44 +00:00
utils.py
Remove more unnecessary load_weights methods (#47058)
2026-06-30 15:22:16 +01:00
Powered by Gitea Version: 1.27.1 Page: 357ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API