This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
5719a4e4e601fb91274294d25370b7aad656d629
vllm
/
vllm
/
model_executor
T
History
Kata Coder
and
GitHub
5719a4e4e6
[Frontend] Support multimodal inputs for late-interaction scoring (ColQwen3) + NewModel: nvidia/nemotron-colembed (
#34574
)
...
Signed-off-by: craftsangjae <
craftsangjae@gmail.com
>
2026-02-20 20:01:40 -08:00
..
layers
[ROCm][Bugfix]: Only save unpadded sizes for shared_experts in MoERunner to fix rmsnorm pad fusion (
#34636
)
2026-02-20 19:56:16 -08:00
model_loader
[Quantization] - Added uses_meta_device_weights to quant config (
#34645
)
2026-02-17 23:43:44 -08:00
models
[Frontend] Support multimodal inputs for late-interaction scoring (ColQwen3) + NewModel: nvidia/nemotron-colembed (
#34574
)
2026-02-20 20:01:40 -08:00
warmup
[Kernel] Add KernelConfig flag to enable/disable FlashInfer autotune (
#34006
)
2026-02-07 05:24:44 -08:00
__init__.py
[Platform] Deprecate seed_everything (
#31659
)
2026-01-04 18:34:04 -08:00
custom_op.py
[torch.compile] Compile
CustomOp.forward_native
for
SiluAndMul
and
QuantFP8
to avoid raw torch ops inside opaque custom ops (
#32806
)
2026-01-22 19:52:26 -08:00
parameter.py
[QeRL] Layerwise Reloading (
#32133
)
2026-01-30 08:50:05 -07:00
utils.py
[BugFix] Fix EPLB fail for MoeFP4 model with Marlin backend (
#33262
)
2026-01-29 16:52:11 +08:00