Logo
Explore Help
Sign In
karylab_agents/vllm
Watch 1
Star 0
Fork 0
forked from Karylab-cklius/vllm
Code Pull Requests 2 Actions 1 Packages Activity
Files
63f1fde277d063fbd36ccf43cb709fafca754ed5
vllm/vllm/model_executor
T
History
Sky LeeandGitHub 343041c4c4 [model] Reduce medusa weight (#10454)
Signed-off-by: skylee-01 <497627264@qq.com>
2024-11-20 06:05:55 +00:00
..
guided_decoding
[Frontend] Bad words sampling parameter (#9717)
2024-10-26 16:29:38 +00:00
layers
[Model][Quantization] HQQ support through Marlin kernel expansion (#9766)
2024-11-19 13:31:12 -08:00
model_loader
[Doc] fix link for page that was renamed (#10455)
2024-11-19 09:48:30 -08:00
models
[model] Reduce medusa weight (#10454)
2024-11-20 06:05:55 +00:00
__init__.py
[Performance] Optimize e2e overheads: Reduce python allocations (#7162)
2024-08-08 21:34:28 -07:00
custom_op.py
[3/N][torch.compile] consolidate custom op logging (#10399)
2024-11-18 15:14:59 -08:00
parameter.py
[Kernel] (2/N) Machete - Integrate into CompressedTensorsWNA16 and GPTQMarlin (#7701)
2024-09-23 13:46:26 -04:00
pooling_metadata.py
[Model][Misc] Add e5-mistral-7b-instruct and Embedding API (#3734)
2024-05-11 11:30:37 -07:00
sampling_metadata.py
[Hardware][Intel-Gaudi] Add Intel Gaudi (HPU) inference backend (#6143)
2024-11-06 01:09:10 -08:00
utils.py
[Hardware] using current_platform.seed_everything (#9785)
2024-10-29 14:47:44 +00:00
Powered by Gitea Version: 1.27.1 Page: 120ms Template: 2ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API