Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
efa2e424f659dd6bbc07982bc155f265e32532cd
vllm/vllm/model_executor
T
History
xiao-llmGitHubAndreas Karatzas
02bf9c7907 Fix Quark mxfp4 quantized model loading issue under mtp (#46757)
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
2026-07-16 14:58:47 -04:00
..
kernels
[ROCm][BugFix] Triton W4A16 handling for GPTQ/AutoGPTQ qzeros layout (#47770)
2026-07-15 11:55:47 -05:00
layers
Fix Quark mxfp4 quantized model loading issue under mtp (#46757)
2026-07-16 14:58:47 -04:00
model_loader
[CI][AMD] Configure MI300 tests for native execution without DinD (#48387)
2026-07-15 02:14:16 +00:00
models
[Model] Add Inkling model support [1/N] (#48799)
2026-07-15 23:40:07 -07:00
offloader
Revert "[Platform] Replace torch.cuda.Event with torch.Event (#47140)" (#47668)
2026-07-06 14:18:33 +01:00
warmup
[Refactor] Move fla to third party (#48500)
2026-07-16 19:22:36 +01:00
__init__.py
[Platform] Deprecate seed_everything (#31659)
2026-01-04 18:34:04 -08:00
custom_op.py
Add ability to replace oot ops when using lora (#37181)
2026-03-16 18:04:15 -07:00
parameter.py
Speed up docs build (#44635)
2026-06-05 14:51:44 +00:00
utils.py
Remove more unnecessary load_weights methods (#47058)
2026-06-30 15:22:16 +01:00
Powered by Gitea Version: 1.27.1 Page: 413ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API