 Micah WilliamsonandGitHub
|
6f612fbedf
|
[ROCm][CI] Patch conftest to resolve occasional OOMs (#45722)
Signed-off-by: Micah Williamson <micah.williamson@amd.com>
|
2026-06-16 10:00:15 -05:00 |
|
 Sting LinandGitHub
|
506ec6d656
|
Upgrade tpu-inference to v0.22.1 (#45793)
|
2026-06-16 07:54:57 -07:00 |
|
 
|
a52205bccf
|
[Model] Add HrmTextForCausalLM (Hierarchical Reasoning Model — Text) (#43098)
Signed-off-by: Wuyifei <wuyifei@me.com>
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
|
2026-06-16 22:41:41 +08:00 |
|
 
|
3d34f8cbdc
|
[ROCm][Cleanup] Remove stale AITER FA hybrid KV-cache TODO (#44178)
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-06-16 07:28:06 -07:00 |
|
 Carl YandGitHub
|
eb04c769d3
|
feat: MLA prefill enable FA4 fp8 output (#43050)
Signed-off-by: Carl You <4531192+carlyou@users.noreply.github.com>
|
2026-06-16 07:10:59 -07:00 |
|
 
|
ce3ef17bec
|
[Kernel][Helion][1/N] Add Helion kernel for rms_norm_per_block_quant (#36895)
Signed-off-by: Sean Chen <seachen@redhat.com>
Co-authored-by: Yanan Cao <gmagogsfm@gmail.com>
|
2026-06-16 22:09:52 +08:00 |
|
 
|
bf5149b516
|
[Bugfix] Fix FlashMLA sparse accuracy with topk_length and zero-init padding (#36616)
Signed-off-by: AjAnubolu <anuboluajay@gmail.com>
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Matthew Bonanni <mbonanni@redhat.com>
|
2026-06-16 07:09:00 -07:00 |
|
 Tahsin TunanandGitHub
|
cca3365b73
|
[Rust Frontend] Add CORS support (#45753)
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com>
|
2026-06-16 13:47:11 +00:00 |
|
 
|
040df8f2ea
|
[CI] Fix attention benchmark smoke test (#45728)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-16 13:43:37 +00:00 |
|
 
|
ced32bb474
|
[Perf] Add VLLM_TRITON_FORCE_FIRST_CONFIG to skip Triton autotuning (#42425)
Signed-off-by: Francesco Fusco <ffu@zurich.ibm.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-16 15:16:45 +02:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
c5e5c33fcd
|
[Bugfix][MoE] Restore routed output unpadding before shared expert add (#45707)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-16 16:06:28 +03:00 |
|
 Mike GandGitHub
|
a8c86eeb16
|
[Quant] Support modelopt_mixed on Ampere (SM80/SM86) (#45306)
Signed-off-by: Mike G <180722391+mikekg@users.noreply.github.com>
|
2026-06-16 08:43:44 -04:00 |
|
 Andreas KaratzasandGitHub
|
7e179e4bc0
|
[ROCm][CI] Gate incompatible HF references on Transformers v5 (#41532)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
|
2026-06-16 20:34:11 +08:00 |
|
   
|
405c7cf283
|
[ZenCPU] Add zencpu Platform Runtime Logging and Docs (#42726)
Signed-off-by: Lalithnarayan C <Lalithnarayan.C@amd.com>
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2026-06-16 08:23:12 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
3f53e2138f
|
[Refactor] Remove Fp8OnlineLinearMethod as scheduled (#45463)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-16 04:35:58 -07:00 |
|
 Hank HanandGitHub
|
d53f4593ce
|
[KV Connector][Mooncake] Pipeline-parallel support for PD-disaggregated serving with Mooncake connector (#44528)
Signed-off-by: hanhan.hank <hanhan.hank@bytedance.com>
Signed-off-by: Hank Han <hanhan7630@outlook.com>
|
2026-06-16 04:35:38 -07:00 |
|
  
|
ad32608e24
|
[MM][Perf][CG] Support dual-path ViT full CUDA graph for DeepSeek-OCR (#43586)
Signed-off-by: shen-shanshan <467638484@qq.com>
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
Co-authored-by: Roger Wang <hey@rogerw.io>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-16 04:35:20 -07:00 |
|
 Thien TranandGitHub
|
b2cfae777d
|
Add Triton recompile detection (#45631)
Signed-off-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2026-06-16 18:25:28 +08:00 |
|
 wangxiyuanandGitHub
|
3f1ff1ff14
|
[Misc]Clean up useless test (#45792)
Signed-off-by: wangxiyuan <wangxiyuan1007@gmail.com>
|
2026-06-16 09:53:08 +00:00 |
|
  
|
c69c73418a
|
[XPU][CI] add intel xpu cases for nightly CI (#44372)
Signed-off-by: wenjun.liu <wenjun.liu@intel.com>
Signed-off-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: zengxian <xiangdong.zeng@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2026-06-16 16:35:08 +08:00 |
|
 Thomas ParnellandGitHub
|
ebf3a6d705
|
[Bugfix] Fix trtllm fused allreduce+rms_norm for transformers backend (#45307)
Signed-off-by: Thomas Parnell <tpa@zurich.ibm.com>
|
2026-06-16 08:34:27 +00:00 |
|
 wang.yuqiandGitHub
|
c4fd9794e9
|
[Frontend] Remove AsyncMicrobatchTokenizer. (#45759)
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
|
2026-06-16 08:02:11 +00:00 |
|
  
|
7ad894c86a
|
[Bugfix] Prevent cuMemcpyBatchAsync segfault with MTP and KV offloading (#44784)
Signed-off-by: joshua <joshua.abraham@multicorewareinc.com>
Co-authored-by: joshua <joshua.abraham@multicorewareinc.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-06-16 07:58:39 +00:00 |
|
 Li, JiangandGitHub
|
a7fdfeef72
|
[CPU] Support Gemma Diffusion (#45690)
Signed-off-by: jiang1.li <jiang1.li@intel.com>
|
2026-06-16 14:39:56 +08:00 |
|
 Jimmy LeeandGitHub
|
8bf374955f
|
[Bug Fix] Allow pinned memory for WSL2 (#41496)
Signed-off-by: Jimmy Lee <hirejimmylee@gmail.com>
|
2026-06-16 05:56:26 +00:00 |
|
 Cyrus LeungandGitHub
|
9096659edb
|
[Cleanup] Remove dead env (#45777)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2026-06-15 22:56:23 -07:00 |
|
 Taneem IbrahimandGitHub
|
81d8f4ebac
|
[Misc] Added validation for Cohere /v2/embed input field exclusivity (#45640)
Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>
|
2026-06-16 05:42:43 +00:00 |
|
 
|
a9a8a32dcd
|
Register parsed config classes before tokenizer init (#40299)
Signed-off-by: Bortlesboat <bortstheboat@gmail.com>
Co-authored-by: OpenAI Codex <codex@openai.com>
|
2026-06-16 05:33:08 +00:00 |
|
  
|
9d808e2309
|
[Core] Use fastsafetensors ParallelLoader for weight loading (#40183)
Signed-off-by: Git Bisector <gitbisector@gmail.com>
Signed-off-by: gitbisector <gitbisector@gmail.com>
Signed-off-by: git bisector <gitbisector@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
|
2026-06-15 22:32:05 -07:00 |
|
 Ben BrowningandGitHub
|
f3858d5422
|
[Frontend] [Parser] Migrate Nemotron V3 to streaming parser engine (#45755)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-16 05:31:21 +00:00 |
|
 Bugen ZhaoandGitHub
|
259ff891be
|
[Rust Frontend] Require ModelConfig.vocab_size to be present (#45696)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-16 05:30:25 +00:00 |
|
 
|
6607a80dab
|
[Bugfix][Gemma4] Fix offline parser truncation, adjust_request token leak, and chat template sync (#45553)
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com>
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com>
|
2026-06-16 04:31:53 +00:00 |
|
 liuzhenweiandGitHub
|
b8bd773fe4
|
[XPU] Fix Triton attn fp8/bf16 check failing (#45758)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
|
2026-06-16 12:31:20 +08:00 |
|
 Ruinan MaandGitHub
|
2addbb9cc9
|
[BugFix] Support async scheduling with prompt embeds for multimodal models (#45673)
Signed-off-by: Ruinan Ma <r7ma3088@gmail.com>
|
2026-06-16 04:12:54 +00:00 |
|
 Isotr0pyandGitHub
|
e3cfea2e1b
|
[Multimodal] Add Qwen3-VL video loader (#44412)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-16 03:45:34 +00:00 |
|
 Bugen ZhaoandGitHub
|
f99260d2aa
|
[Rust Frontend] Lower out-of-vocab validation to text layer (#45685)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-16 03:37:58 +00:00 |
|
 Bugen ZhaoandGitHub
|
3f65e21e32
|
[Rust Frontend] Support max_logprobs validation (#45674)
Signed-off-by: Bugen Zhao <i@bugenzhao.com>
|
2026-06-16 10:57:56 +08:00 |
|
 xx-thomasandGitHub
|
b00e76ff72
|
[Misc][Model] add io processor for query/document embeddings from ColBERT (jinaai/jina-colbert-v2) (#45210)
Signed-off-by: thomas <thomas.varghese@columbia.edu>
|
2026-06-16 01:32:32 +00:00 |
|
 Woosuk KwonandGitHub
|
f4359a70f9
|
[DSV4][Minor] Fix supported KV cache dtypes (#44892)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-16 00:14:51 +00:00 |
|
 Itay AlroyandGitHub
|
3afe659b6b
|
[EP] Enable DBO with NIXL EP (#45275)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-06-15 23:37:22 +00:00 |
|
 Itay AlroyandGitHub
|
16e91176cf
|
[EP] Query NIXL EP top-k index dtype (#45298)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-06-15 22:50:18 +00:00 |
|
 Itay AlroyandGitHub
|
ab8b0fe338
|
nixl_ep: Skip post-receive quantization for NVFP4 (#45606)
Signed-off-by: Itay Alroy <ialroy@nvidia.com>
|
2026-06-15 22:42:05 +00:00 |
|
  
|
d467a2a7f2
|
[Bugfix] Defer block freeing until in-flight steps finish under async scheduling + PD KV consumer (#45357)
Signed-off-by: llx-08 <2596671364@qq.com>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com>
|
2026-06-15 21:36:09 +00:00 |
|
 
|
76a373eff4
|
[Frontend] Replace legacy Gemma4 parsers with engine-based implementation (#45588)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-06-15 21:34:07 +00:00 |
|
 Zang PeiyuandGitHub
|
25ee659db0
|
Fix parallel_tool_calls: null treated as false instead of default true (#44955)
Signed-off-by: factnn <166481866+factnn@users.noreply.github.com>
|
2026-06-15 21:14:10 +00:00 |
|
  
|
eacff17c8d
|
[Model Runner V2][Bugfix] Fix MRV2 LoRA warmup (#35536)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-15 13:17:23 -07:00 |
|
 Flora FengandGitHub
|
cd9078fe59
|
[Frontend] Skip structural tags for auto tool_choice without strict mode (#45600)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-06-15 19:55:31 +00:00 |
|
 Wentao YeandGitHub
|
e18fe932ca
|
[Perf] Optimize DSv4 prefill chunk planning, 4.0% E2E Throughput Improvement (#45061)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-15 19:50:21 +00:00 |
|
 
|
51ec5cf08f
|
[Bugfix] Chat Completions Harmony Refactor Clean up (#45464)
Signed-off-by: Yifan Zong <yzong@redhat.com>
Co-authored-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-15 14:45:19 -04:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
7e612a0f06
|
[KV Offloading] Implement reset_cache for TieringOffloadingManager (#44541)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-15 18:42:53 +00:00 |
|