Woosuk Kwon
|
a06a16ff0a
|
minor
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-19 03:31:41 +00:00 |
|
Woosuk Kwon
|
a1d80989d9
|
Plumb is_padding
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-19 03:23:23 +00:00 |
|
 Samuel ShenandGitHub
|
c9135db27c
|
[Docs] Update stale LMCache examples (#45762)
Signed-off-by: Samuel Shen <slshen@tensormesh.ai>
|
2026-06-19 03:21:36 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
2a6c6b9429
|
[DeepSeek-V4] Support TEP=16 for the block-FP8 shared expert (#46001)
Signed-off-by: Jeff Ma <jeffjma@umich.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-18 20:10:12 -07:00 |
|
 Jared WenandGitHub
|
ab66606993
|
[bugfix]Indexer init skip and MTP TopK share for iteration (#45895)
Signed-off-by: JaredforReal <w13431838023@gmail.com>
|
2026-06-19 09:57:51 +08:00 |
|
  
|
9ea3a4015b
|
[Bugfix] Fix corrupt outputs in MoE FP8 LoRA responses and MoE base model responses when LoRAs are loaded (#42120)
Signed-off-by: Nicholas Edelman <nedelman@nvidia.com>
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: Jee Jee Li <jeejeelee@inferact.ai>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
|
2026-06-18 18:26:09 -07:00 |
|
 Flora FengandGitHub
|
560fb8b867
|
[Cohere] Remove dead prepare_structured_tag override in Cohere parser (#46099)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
|
2026-06-19 01:02:11 +00:00 |
|
 Wentao YeandGitHub
|
675cd5d228
|
[Model Runner V2] Fix MRv2 memory leak test (#46095)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-19 00:36:40 +00:00 |
|
 
|
7f616c327d
|
[Bugfix] [Parser] Fix empty tool block silently dropping subsequent content (#46091)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
Co-authored-by: Flora Feng <4florafeng@gmail.com>
|
2026-06-18 23:17:18 +00:00 |
|
 Ivy XuandGitHub
|
c3c6d723fd
|
[Perf] Remove unused loggers in reasoning/ (#45988)
Signed-off-by: Ivy <fakeshadow1337@gmail.com>
|
2026-06-18 22:24:29 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)  
|
41dcf49ca5
|
[Bugfix][KV Connector] Disable Mooncake TP put-striding when DCP > 1 (#45371)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Co-authored-by: Jingyi Yang <girasoleyang@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-18 15:13:44 -07:00 |
|
 
|
35e4dd4a69
|
[KV Connector][Mooncake] Async lookup to reduce scheduler overhead (#45659)
Signed-off-by: Yifan Qiao <yifanqiao@inferact.ai>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
|
2026-06-18 21:44:02 +00:00 |
|
  
|
4ce2d01453
|
fix(anthropic): auto-detect template support for mid-conversation system messages (#46025)
Signed-off-by: felix0080 <felix0080@users.noreply.github.com>
Signed-off-by: Ben Browning <bbrownin@redhat.com>
Co-authored-by: felix0080 <felix0080@users.noreply.github.com>
Co-authored-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-18 16:19:11 -04:00 |
|
 Woosuk KwonandGitHub
|
16908e132e
|
[MRV2] Make FP32 Gumbel sampling more accurate (#45996)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
|
2026-06-18 19:42:09 +00:00 |
|
 Wentao YeandGitHub
|
225936a1dd
|
[CI Bug] Revert #42379 to fix CI Multi-Modal Models (Extended Generation 1) (#46070)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-18 12:37:39 -07:00 |
|
 
|
f6ba720963
|
(security) Upgrade Starlette to >= 1.0.1 to fix CVE-2026-48710 (#45675)
Signed-off-by: jperezde <jperezde@redhat.com>
Co-authored-by: Isotr0py <mozf@mail2.sysu.edu.cn>
|
2026-06-18 12:35:13 -07:00 |
|
 Wentao YeandGitHub
|
b53b1c7ffe
|
[Model Runner V2] Migration to support quantized model by default [5/N] (#44446)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
|
2026-06-18 12:20:44 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
79ca54d221
|
[Bugfix][Quantization] Don't reject fp8_e5m2 KV cache for non-fp8 quantized checkpoints (#45040)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-18 14:18:25 -04:00 |
|
 Ben BrowningandGitHub
|
09f3cd5c10
|
[Bugfix] [Parser] Fix Qwen3 latent bug in partial params dropping values containing < (#46047)
Signed-off-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-18 18:04:06 +00:00 |
|
 
|
ea6078fe6a
|
[KV Connector][Offloading] Disable parallel-agnostic fs-tier cache on V2 model runner (#46044)
Signed-off-by: Itay Etelis <etelis2019@gmail.com>
Co-authored-by: Itay Etelis <etelis2019@gmail.com>
|
2026-06-18 20:43:35 +03:00 |
|
 Palaiologos1453andGitHub
|
a0df04e477
|
[Tests] Add Qwen3 streaming parser delta boundary cases (#45708)
Signed-off-by: test test <2260891073@qq.com>
|
2026-06-18 17:37:39 +00:00 |
|
 stefankoncarevicandGitHub
|
e2352c2974
|
[ROCm][Spec Decode] Fix probabilistic draft probs test attention backend (#45706)
Signed-off-by: Stefan Koncarevic <stefan.koncarevic@amd.com>
|
2026-06-18 11:59:37 -05:00 |
|
 qli88andGitHub
|
25faa1f4cc
|
[CI]Enable mxfp4 lora test for ROCm platform (#43802)
Signed-off-by: Qiang Li <qiang.li2@amd.com>
|
2026-06-18 16:59:09 +00:00 |
|
 HumphreyandGitHub
|
4583630b56
|
[Bugfix][Kernel] Check output alignment in vectorize_with_alignment (fixes misaligned-address crash for non-multiple-of-8 head sizes) (#45466)
Signed-off-by: HumphreySun98 <humphreysun98@gmail.com>
|
2026-06-18 16:58:22 +00:00 |
|
 Divakar VermaandGitHub
|
21da47dabe
|
[ROCm][CI] move lora%N test to mi300 and gate (#45970)
Signed-off-by: Divakar Verma <divakar.verma@amd.com>
|
2026-06-19 00:50:32 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
6c379b9e54
|
[Frontend] Add Streaming Parser Engine and new GLM4.7/GLM5.1/GLM5.2 Parser (#45915)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-19 00:42:10 +08:00 |
|
 Rohan PotdarandGitHub
|
5099474633
|
[Bugfix][ROCm] Fix rocm_aiter_per_tensor_quant custom op aliasing (#45747)
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com>
|
2026-06-18 11:30:21 -05:00 |
|
 Yuwen ZhouandGitHub
|
058cc0a8b6
|
[Bugfix] Restore is_sym guard for zp in GPTQ/CT MoE to fix symmetric quant regression (#45656)
Signed-off-by: yuwenzho <yuwen.zhou@intel.com>
|
2026-06-18 16:20:29 +00:00 |
|
Woosuk Kwon
|
40a19bed77
|
Merge branch 'main' into woosuk/triton-fix
|
2026-06-18 16:06:08 +00:00 |
|
 
|
837db7605e
|
[Bugfix][Tool Parser] Handle non-finite numbers in coerce_to_schema_type (#43984)
Signed-off-by: ashishpatel26 <shriganesh.patel@gmail.com>
Co-authored-by: Ben Browning <bbrownin@redhat.com>
|
2026-06-18 16:00:20 +00:00 |
|
 Mark McLoughlinandGitHub
|
bf2a393034
|
Temporarily remove @markmc from CODEOWNERS (#46053)
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
|
2026-06-18 14:15:43 +00:00 |
|
 
|
d682968aa9
|
[Model] Remove BambaForCausalLM (#45990)
Signed-off-by: Xianbao QIAN <xianbao.qian@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-06-18 06:51:00 -07:00 |
|
 
|
021cdf72bc
|
Fix _riscv_supports_rvv_vlen128() to detect RVV on hardware without zvl flags (#43179)
Signed-off-by: liuyudong <liuyudong@iscas.ac.cn>
Co-authored-by: YuanSheng <yuansheng@isrc.iscas.ac.cn>
|
2026-06-18 21:22:35 +08:00 |
|
 AsharandGitHub
|
4cb5e746b6
|
[Rust Frontend]: Add /get_world_size route with static parallel size (#44801)
|
2026-06-18 13:10:20 +00:00 |
|
 Jee Jee LiandGitHub
|
22cc891108
|
[Kernel] Add PDL support for DeepGEMM kernel (#46006)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-18 20:49:01 +08:00 |
|
    
|
afdcbd5d39
|
[ROCm][DSv4] Functional fixes for DeepSeek V4 on MI300X/MI325X (#45681)
Signed-off-by: ganyi <ygan@amd.com>
Signed-off-by: Markus Hartikainen <markus.hartikainen@amd.com>
Signed-off-by: Tuukka Sarvi <tuukka.sarvi@amd.com>
Co-authored-by: ganyi <ygan@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Markus Hartikainen <markus.hartikainen@amd.com>
Co-authored-by: Jin Tao <jintao12@amd.com>
|
2026-06-18 12:21:14 +00:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png)   
|
8d4f54966c
|
fix(quantization): Fix AWQ dequantize on Intel XPU and refactor AutoAWQ config (#42727)
Signed-off-by: Alex <alex.tech.lab@outlook.com>
Signed-off-by: AlexHuang <jihuihuang@tencent.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-18 20:12:28 +08:00 |
|
 
|
351c72d6e5
|
[CPU] Skip Triton kernel monkey-patches when Triton-CPU is available (#44991)
Signed-off-by: jmamou <jonathan.mamou@intel.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
|
2026-06-18 18:59:30 +08:00 |
|
 Tahsin TunanandGitHub
|
7299e6509e
|
[Rust Frontend] Return model metadata fields in /v1/models (#45950)
Signed-off-by: Tahsin Tunan <tahsintunan@gmail.com>
|
2026-06-18 10:29:21 +00:00 |
|
 littlecircle0730andGitHub
|
08985351f3
|
Fix Stale Encoder Cache After Weight Update (#45093)
Signed-off-by: littlecircle0730 <littlecircle0730@gmail.com>
|
2026-06-18 09:32:10 +00:00 |
|
 Wei ZhaoandGitHub
|
5fd3b276f8
|
[Mooncake] Skip KV lookup for non-reachable SWA blocks (#45444)
Signed-off-by: wzhao18 <wzhao18.sz@gmail.com>
|
2026-06-18 02:23:20 -07:00 |
|
 
|
1e9f04da14
|
fix(anthropic): preserve inline system message position for prefix caching (#44602)
Signed-off-by: felix0080 <felix0080@users.noreply.github.com>
Co-authored-by: felix0080 <felix0080@users.noreply.github.com>
|
2026-06-18 15:58:11 +08:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
702214146c
|
[Bugfix][Frontend] Fix Anthropic count_tokens decorator order driving server load negative (#44725)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-17 23:56:46 -07:00 |
|
![mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>](/assets/img/avatar_default.png) 
|
a331589394
|
[XPU] Update nixl to v0.10.1 in Dockerfile (#40287)
Signed-off-by: zhenwei-intel <zhenwei.liu@intel.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
|
2026-06-18 14:01:26 +08:00 |
|
 Micah WilliamsonandGitHub
|
e945169207
|
Revert "[Kernel] Add PDL support for DeepGEMM kernel" (#45999)
|
2026-06-17 22:59:48 -07:00 |
|
 
|
554352a311
|
[Test][KV Connector] Add request_finished fence population tests for offloading scheduler (#45679)
Signed-off-by: Alex <alex.tech.lab@outlook.com>
Signed-off-by: AlexHuang <jihuihuang@future.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-06-18 08:13:52 +03:00 |
|
 
|
421c1ec448
|
[KV Offloading] Remove dummy worker-side stats from OffloadingConnector (#45905)
Signed-off-by: Alex <alex.tech.lab@outlook.com>
Signed-off-by: AlexHuang <jihuihuang@alexai.com>
Co-authored-by: Or Ozeri <oro@il.ibm.com>
|
2026-06-18 08:13:28 +03:00 |
|
 
|
b4c80ec0fd
|
[Refactor] Remove dead cutlass mxfp8 code (#44681)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
Co-authored-by: Shengqi Chen <harry-chen@outlook.com>
|
2026-06-17 21:18:25 -07:00 |
|
 Ronen SchafferandGitHub
|
f428718ffe
|
[Fix][KV offload] Defer on_request_finished until in-flight transfers drain (#45823)
Signed-off-by: Ronen Schaffer <ronen.schaffer@ibm.com>
|
2026-06-18 07:05:46 +03:00 |
|
 Jee Jee LiandGitHub
|
4403af8fb5
|
[Kernel] Add PDL support for DeepGEMM kernel (#42996)
Signed-off-by: Jee Jee Li <jeejeelee@inferact.ai>
|
2026-06-17 20:37:17 -07:00 |
|