 Ting SUNandGitHub
|
91b5647300
|
[Bugfix][Model] Allow Run:ai memory_limit sentinel values (#47337)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-07-05 00:08:34 +00:00 |
|
 Harry MellorandGitHub
|
1ab9522935
|
Remove more unnecessary load_weights methods (#47058)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-30 15:22:16 +01:00 |
|
  
|
8cf7c4d8ad
|
[Attention Backend] add HPC-Ops Attention backend (#46020)
Signed-off-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-30 18:17:43 +08:00 |
|
 Harry MellorandGitHub
|
5051698e41
|
Remove unnecessary load_weights methods (#44589)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-29 01:52:23 -07:00 |
|
 Yan MaandGitHub
|
acce57d8dd
|
Deprecate old FP8 online MoE quantization class (#44514)
Signed-off-by: Yan Ma <yan.ma@intel.com>
|
2026-06-23 11:53:38 -07:00 |
|
 Ting SUNandGitHub
|
9d4b87f4f0
|
[Bugfix][Model] Validate DefaultModelLoader / LoadConfig and fail with clear errors (#45196)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-06-17 21:46:33 +00:00 |
|
  
|
9d808e2309
|
[Core] Use fastsafetensors ParallelLoader for weight loading (#40183)
Signed-off-by: Git Bisector <gitbisector@gmail.com>
Signed-off-by: gitbisector <gitbisector@gmail.com>
Signed-off-by: git bisector <gitbisector@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-06-15 22:32:05 -07:00 |
|
 Ting SUNandGitHub
|
3d6ce816f0
|
[Bugfix][Model] Validate runai_streamer model_loader_extra_config (#45291)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-06-14 19:23:30 -07:00 |
|
 Isotr0pyandGitHub
|
6635279d8a
|
[Migration] Migrate GGUF quantization support to plugin (#39612)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-12 12:02:21 -07:00 |
|
 Ting SUNandGitHub
|
c1076839c9
|
[Bugfix][Model] Pass revision by name in Run:ai and bitsandbytes index downloads (#45308)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
2026-06-11 20:21:46 -07:00 |
|
 
|
5edf7ff489
|
[Core] Release cached device memory under pressure on UMA GPUs during weight loading (#45179)
Signed-off-by: mgoin <mgoin64@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-11 17:49:50 +03:00 |
|
 Harry MellorandGitHub
|
18d87a87dc
|
Deprecate Transformers v4 support (#45161)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-11 11:04:01 +08:00 |
|
 
|
fa8c868a3c
|
[Bugfix] Fix Llama4 weight loading (#45047)
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-06-10 13:40:45 -04:00 |
|
  
|
c9e5bf8135
|
[Bugfix] Fix layerwise reload dropping params after a composed weight loader (#44814)
Signed-off-by: hallerite <git@hallerite.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-06-10 06:42:05 -07:00 |
|
 bnellnmandGitHub
|
dc68bd8c41
|
[MoE Refactor] FusedMoE/MoERunner inversion refactor (#41184)
Signed-off-by: Bill Nell <bnell@redhat.com>
|
2026-06-08 10:42:58 -04:00 |
|
 Harry MellorandGitHub
|
ef0df7dbd6
|
[CI] Bump mypy version 1.19.1 -> 1.20.2 (#44647)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:56:27 +00:00 |
|
 Harry MellorandGitHub
|
a80af24356
|
Speed up docs build (#44635)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-05 14:51:44 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
62215e72c6
|
Remove KV cache scale boilerplate from model weight loading methods (#43167)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-05 05:19:04 -07:00 |
|
  
|
f91fb2fcf3
|
[Bugfix] Convert Gemma4-MM ViT linear layers to vllm native impl (#43798)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: ZiTian Zhao <zitian.zhao@tencentmusic.com>
Co-authored-by: B-201 <Joy25810@foxmail.com>
|
2026-06-01 21:41:16 -07:00 |
|
 
|
11dfa3169d
|
Add vLLM library info to Hugging Face Hub requests (#43857)
Signed-off-by: Wauplin <lucainp@gmail.com>
Signed-off-by: Lucain Pouget <lucain@huggingface.co>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-05-29 14:04:58 +00:00 |
|
 ![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
9aa131f944
|
Add Cosmos3 Reasoner model (#43356)
Signed-off-by: Maciej Bala <mbala@nvidia.com>
Signed-off-by: MaciejBalaNV <mbala@nvidia.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Isotr0py <2037008807@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-28 09:43:55 -07:00 |
|
 Benjamin BartelsandGitHub
|
05eec7120e
|
Fix RunAI streamer tensor buffer reuse during weight loading (#43464)
Signed-off-by: bbartels <benjamin@bartels.dev>
|
2026-05-27 19:16:52 -07:00 |
|
  
|
17b69828a0
|
[Core] Add native ModelExpress load format (#43105)
Signed-off-by: Zheng Luo <zheluo@nvidia.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com>
|
2026-05-21 16:05:01 -04:00 |
|
  
|
1ccdf87507
|
[Bugfix] Fix layerwise reload alias-buffer corruption (#42481)
Signed-off-by: rasdani <73563550+rasdani@users.noreply.github.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
Co-authored-by: Roger Wang <hey@rogerw.io>
|
2026-05-15 15:20:53 -07:00 |
|
 
|
d26a28ab03
|
fix: propagate revision/code_revision pins to all artifact boundaries (#42616)
Signed-off-by: jperezde <jperezde@redhat.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-05-15 02:31:54 -07:00 |
|
 Michael GoinandGitHub
|
8efd508204
|
[Quantization] Rework quantization_config to use QuantKey and allow for activation override (#41566)
|
2026-05-13 16:58:32 -04:00 |
|
 Mohammad Miadh AngkadandGitHub
|
21943d4c25
|
[Performance] Make safetensors checkpoint prefetch settings configurable (#41499)
Signed-off-by: Mohammad Miadh Angkad <MAngkad.BSDSBA2027@aim.edu>
|
2026-05-10 15:55:15 +00:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
1d694e78c9
|
[Examples][last/6] Resettle examples. (#41084)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-07 19:42:12 -07:00 |
|
 Alex BrooksandGitHub
|
db9a84e0cd
|
[Bugfix] Fix FP8 Bias Loading (#41424)
Signed-off-by: Alex Brooks <albrooks@redhat.com>
|
2026-05-03 20:30:04 +00:00 |
|
 
|
a085b5257d
|
[Docs] [QeRL] Layerwise Reloading Documentation (#40317)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-04-28 21:06:38 -07:00 |
|
 
|
99255f3cb5
|
[UX] Allow enable/disable model weights loading tracking by config (#41086)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Copilot <copilot@github.com>
|
2026-04-28 21:04:49 -07:00 |
|
 wang.yuqiandGitHub
|
a8208e6a81
|
[Examples] Resettle features examples. (#40995)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-04-28 00:33:41 -07:00 |
|
 Andreas KaratzasandGitHub
|
5e2c37facd
|
[ROCm][CI] Add missing quantization methods and fix online quant test failures (#39801)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-04-27 15:08:57 -05:00 |
|
 
|
5d5c776444
|
[Perf] FP8 FlashInfer Attn for ViT (#38065)
Signed-off-by: Zhanda Zhu <zhandazhu@gmail.com>
Co-authored-by: Yubo Gao <ybgao-nvidia@users.noreply.github.com>
|
2026-04-27 13:44:15 +08:00 |
|
 
|
d0009ddb0b
|
[Model] Support Hy3 preview (#40681)
Signed-off-by: stevenkuang <stevenkuang@tencent.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-04-23 22:08:26 +08:00 |
|
![gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
8f87eb4622
|
[Refactor] Clean up log once scope="local" (#40540)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Signed-off-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-22 16:42:43 -04:00 |
|
 Asaf GardinandGitHub
|
19fa90ed0d
|
[Quantization] - Layerwise reloading of Attention/KV quantized models (#38995)
Signed-off-by: Josephasafg <ajgard7@gmail.com>
|
2026-04-15 18:03:32 -07:00 |
|
   
|
03f8d3a548
|
Update to transformers v5 (#30566)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Signed-off-by: khluu <khluu000@gmail.com>
Signed-off-by: Kevin H. Luu <khluu000@gmail.com>
Signed-off-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
Co-authored-by: khluu <khluu000@gmail.com>
Co-authored-by: Cyrus Leung <cyrus.tl.leung@gmail.com>
Co-authored-by: jiang1.li <jiang1.li@intel.com>
|
2026-04-15 16:29:15 -07:00 |
|
 Tihomir ElekandGitHub
|
8d825b87d6
|
[Bug] Fix TypeError when hf_config.architectures is None during model loading (#38849)
Signed-off-by: Tihomir Elek <tiho.elek@gmail.com>
|
2026-04-13 12:13:21 +01:00 |
|
 Artem PerevedentsevandGitHub
|
bb6047db13
|
[Model][Perf] Enable checkpoints prefetching for Lustre FS by default (#39422)
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com>
|
2026-04-10 00:47:59 +00:00 |
|
 Vasiliy KuznetsovandGitHub
|
7b1a7423be
|
[Frontend] new online quantization frontend (#38138)
Signed-off-by: Vasiliy Kuznetsov <vasiliy@meta.com>
|
2026-04-03 11:58:39 -04:00 |
|
 Aaron HaoandGitHub
|
4729b90838
|
[Bug] Add e_score_correction_bias to SKIP_TENSORS (#38746)
Signed-off-by: ahao-anyscale <ahao@anyscale.com>
|
2026-04-02 21:15:05 -07:00 |
|
 Asaf GardinandGitHub
|
3dc01ef352
|
[Quantization] Consolidate dummy format logic into DummyModelLoader (#38637)
Signed-off-by: Josephasafg <ajgard7@gmail.com>
|
2026-03-31 22:20:45 +00:00 |
|
 Kyle SayersandGitHub
|
5869f69c5f
|
[Online Quant] [QeRL] Minor code cleanup (#38574)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-03-31 14:56:43 +00:00 |
|
 Asaf GardinandGitHub
|
1fc69f59bb
|
[Bug fix][Quantization] Fix dummy weight loading (#38478)
Signed-off-by: Josephasafg <ajgard7@gmail.com>
|
2026-03-30 16:38:02 -04:00 |
|
 PikaPikachuandGitHub
|
63babd17f1
|
[Model][Quantization] Add GGUF support for MiniMax-M2.1 (#36965)
Signed-off-by: kangletian <Letian.Kang@amd.com>
|
2026-03-30 14:24:06 +08:00 |
|
 Kyle SayersandGitHub
|
d28d86e8a3
|
[QeRL] Fix online quantized reloading (#38442)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-03-29 14:56:41 -06:00 |
|
 
|
aa4eb0db78
|
[CI]revert initialize_model context manager (#38426)
Signed-off-by: Kunshang Ji <jikunshang95@gmail.com>
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-03-28 16:56:50 +00:00 |
|
 Kyle SayersandGitHub
|
648edcf729
|
[QeRL] Compose online quantization with quantized reloading (#38032)
Signed-off-by: Kyle Sayers <kylesayrs@gmail.com>
|
2026-03-27 13:22:33 -07:00 |
|
 Andreas KaratzasandGitHub
|
f2d16207c7
|
[ROCm][CI] Fix flaky GPTQ compile correctness test (#38161)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-03-26 19:57:00 +08:00 |
|