Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
5719a4e4e601fb91274294d25370b7aad656d629
vllm/vllm/model_executor
T
History
Kata CoderandGitHub 5719a4e4e6 [Frontend] Support multimodal inputs for late-interaction scoring (ColQwen3) + NewModel: nvidia/nemotron-colembed (#34574)
Signed-off-by: craftsangjae <craftsangjae@gmail.com>
2026-02-20 20:01:40 -08:00
..
layers
[ROCm][Bugfix]: Only save unpadded sizes for shared_experts in MoERunner to fix rmsnorm pad fusion (#34636)
2026-02-20 19:56:16 -08:00
model_loader
[Quantization] - Added uses_meta_device_weights to quant config (#34645)
2026-02-17 23:43:44 -08:00
models
[Frontend] Support multimodal inputs for late-interaction scoring (ColQwen3) + NewModel: nvidia/nemotron-colembed (#34574)
2026-02-20 20:01:40 -08:00
warmup
[Kernel] Add KernelConfig flag to enable/disable FlashInfer autotune (#34006)
2026-02-07 05:24:44 -08:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
[torch.compile] Compile CustomOp.forward_native for SiluAndMul and QuantFP8 to avoid raw torch ops inside opaque custom ops (#32806)
2026-01-22 19:52:26 -08:00
parameter.py
[QeRL] Layerwise Reloading (#32133)
2026-01-30 08:50:05 -07:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (#33262)
2026-01-29 16:52:11 +08:00
Powered by Gitea Version: 1.27.1 Page: 144ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API