Yongye Zhu
97cd2c41ad
fix precommit
...
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 03:14:49 +00:00
Yongye Zhu and Claude Opus 4.7
c001535038
[Attention][TokenSpeed MLA] Force prefill V tensor contiguous before kernel call
...
`v` arrives at both `run_prefill_new_tokens` and `run_prefill_context_chunk`
as the second half of `kv_nope.split([qk_nope_head_dim, v_head_dim], dim=-1)`
in mla_attention.py — a non-contiguous view along the last dim. The kernel
internally does `v.reshape(1, total_kv, h_k, 1, d_v)` which silently copies
when the input is non-contiguous; pull that copy out so the layout is
predictable at the kernel boundary.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:53:46 +00:00
Yongye Zhu and Claude Opus 4.7
de6bc297df
[Attention][TokenSpeed MLA] Surface install hint when package missing
...
Previously a user explicitly selecting TOKENSPEED_MLA without `tokenspeed_mla`
installed got either a generic "required dependencies not available" message
(prefill backend) or a raw ModuleNotFoundError deep inside forward_mqa at the
first request (decode backend). Now both backends fail at startup with the
exact install command: `uv pip install tokenspeed-mla`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:49:09 +00:00
Yongye Zhu and Claude Opus 4.7
0012818287
[Attention][TokenSpeed MLA] Fix trtllm LSE parity test: log2 → natural log
...
trtllm_ragged_attention_deepseek returns LSE in log2; tokenspeed and
merge_attn_states use natural log. Multiply the trtllm reference by ln 2
before comparison.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:49:09 +00:00
Yongye Zhu and Claude Opus 4.7
73cd7e25ae
[Attention][TokenSpeed MLA] Warm up BF16 prefill compile, drop seq_lens computation
...
Pre-JIT both BF16 and FP8 prefill kernels at backend init since the dtype
isn't visible from `__init__` — depends on `use_prefill_query_quantization`.
Move the per-forward `seq_lens` computation into `prepare_metadata` and
document the cuda-graph padding interaction with `query_start_loc`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:49:09 +00:00
Yongye Zhu and Claude Opus 4.7
964c6eb485
[Attention][TokenSpeed MLA] Fix decode FP8 numerics: pass output_scale and assert FP8 Q
...
The decode kernel needs both bmm scales to recover correct outputs from an
FP8 KV cache: bmm1 (softmax_scale = scale * q_scale * k_scale) and bmm2
(output_scale = k_scale, since V is stored as V_real / k_scale). We were
only passing bmm1, which left bmm2 = 1.0 and produced silently wrong output.
Also assert query dtype is float8_e4m3fn on entry to forward_mqa.
supports_quant_query_input=True (inherited from MLACommonImpl) tells the
upstream pipeline to FP8-quantize Q via _decode_concat_quant_fp8_op; the
kernel is shape-specialized for FP8 Q + FP8 KV, so any other dtype here
means the upstream quant path didn't run and the kernel will produce
garbage. Failing loud beats failing silent.
Verified: gsm8k matches reference with TOKENSPEED_MLA decode +
FLASH_ATTN prefill on Kimi-K2.5-NVFP4 / TP=4 / B200.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:49:09 +00:00
Yongye Zhu and Claude Opus 4.7
d0e6514bf8
[Attention] Add TOKENSPEED_MLA backend for DeepSeek R1 prefill + decode on Blackwell
...
Wires the tokenspeed_mla CuTe DSL kernels into vLLM as a new MLA backend,
covering both prefill (tokenspeed_mla_prefill) and decode
(tokenspeed_mla_decode). Targets Blackwell (SM100) with FP8 KV cache and
DeepSeek R1 MLA dimensions; users opt in via -ac
'{"backend":"TOKENSPEED_MLA","mla_prefill_backend":"TOKENSPEED_MLA"}'.
Includes numeric parity tests against the trtllm reference kernels.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com >
Signed-off-by: Yongye Zhu <zyy1102000@gmail.com >
2026-05-06 02:49:09 +00:00
Chauncey and GitHub
c7aa186d67
[Frontend] Supports resubmitting output items with missing fields in Responses API ( #41355 )
...
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com >
2026-05-05 22:21:33 -04:00
f653761252
[CI] Route part of B200 jobs to b200-k8s ( #41453 )
...
Signed-off-by: khluu <khluu000@gmail.com >
Co-authored-by: OpenAI Codex <noreply@openai.com >
2026-05-05 19:00:30 -07:00
Andreas Karatzas and GitHub
4a8ae26e53
[ROCm][CI] Use vLLM generation defaults for DeepSeek prefetch-offload eval ( #41575 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-06 01:08:12 +00:00
Kevin H. Luu and GitHub
1333864408
[CI] Automate Docker Hub release image publishing ( #40415 )
...
Signed-off-by: khluu <khluu000@gmail.com >
2026-05-06 00:15:23 +00:00
Matthew Bonanni and GitHub
01b9b5af67
[Attention] Minor refactor: layer takes ownership of the MLA prefill backend ( #41744 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-05 23:22:41 +00:00
8c57b6e7bc
Bump model-hosting-container-standards to >= 0.1.14 ( #39755 )
...
Signed-off-by: EC2 Default User <ec2-user@ip-172-31-20-13.us-west-2.compute.internal >
Co-authored-by: EC2 Default User <ec2-user@ip-172-31-20-13.us-west-2.compute.internal >
Co-authored-by: Robert Shaw <114415538+robertgshaw2-redhat@users.noreply.github.com >
2026-05-05 19:09:57 -04:00
Lanze Liu and GitHub
79246b5ea6
[Spec Decode] Fix max_model_len logging in speculative config for draft model ( #41571 )
...
Signed-off-by: Lanze Liu <lanzetech@gmail.com >
2026-05-05 21:56:06 +00:00
48954de237
Fix DeepGEMM ep_scatter output address overflow ( #39213 )
...
Signed-off-by: S1ro1 <matej.sirovatka@gmail.com >
Co-authored-by: Tyler Michael Smith <tyler@neuralmagic.com >
2026-05-05 18:56:56 +00:00
Julien Denize and GitHub
c6235ed180
[BUGFIX] Support streamed_args_for_tool in MistralToolParser ( #41730 )
...
Signed-off-by: juliendenize <julien.denize@mistral.ai >
2026-05-05 17:48:53 +00:00
628c436301
[New Model][ROCm] Add AMD support for DeepSeek V4 ( #40871 )
...
Signed-off-by: ganyi <ygan@amd.com >
Signed-off-by: whx-sjtu <xiaowang990929@gmail.com >
Signed-off-by: tjtanaa <tunjian.tan@embeddedllm.com >
Signed-off-by: tjtanaavllm <tunjian.tan@amd.com >
Co-authored-by: ganyi <ygan@amd.com >
Co-authored-by: tjtanaa <tunjian.tan@embeddedllm.com >
Co-authored-by: tjtanaavllm <tunjian.tan@amd.com >
2026-05-05 08:55:37 -07:00
Canlin Guo and GitHub
2228fe6868
[Attention] Move FA3→FA4 upgrade into get_flash_attn_version() ( #40815 )
...
Signed-off-by: gcanlin <canlinguosdu@gmail.com >
2026-05-05 15:43:03 +00:00
Harry Mellor and GitHub
84bd8a3c1e
Remove unnecessary runtime asserts from linear layers ( #41729 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-05 14:42:56 +00:00
Lidang Jiang and GitHub
b786ec8e74
[Bugfix] Suggest upgrading Transformers for tokenizer class errors ( #38099 )
...
Signed-off-by: Lidang-Jiang <lidangjiang@gmail.com >
2026-05-05 14:10:45 +00:00
20dcd984f9
[Bugfix] Fix RuntimeError: Already borrowed by adding thread-safe Hugging Face fast-tokenizer wrappers ( #41181 )
...
Signed-off-by: Yifan Zong <yzong@redhat.com >
Co-authored-by: wang.yuqi <yuqi.wang@daocloud.io >
2026-05-05 14:04:01 +00:00
Martin Hickey and GitHub
6fca518157
[BugFix][MyPy]: Module has no attribute "sched_getaffinity" [attr-defined] ( #41465 )
...
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com >
2026-05-05 13:20:37 +00:00
98661fe012
[Bugfix][KVConnector] Support DCP/PCP in OffloadingConnector ( #41549 )
...
Signed-off-by: Itay Etelis <itay.etelis@ibm.com >
Co-authored-by: Itay Etelis <itay.etelis@ibm.com >
2026-05-05 14:54:29 +03:00
Harry Mellor and GitHub
b0765bee17
Fix DeepSeek-OCR for Transformers v4 ( #41460 )
...
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com >
2026-05-05 11:11:21 +00:00
0a201b60cf
[Model] support Qianfan-OCR model ( #40136 )
...
Signed-off-by: bairongz <baiyuu.cs@gmail.com >
Signed-off-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: zhuangbairong <zhuangbairong@baidu.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-05 10:51:25 +00:00
8b9ea2f881
[Feature] Add Triton kernel JIT compilation monitor for inference ( #40137 )
...
Signed-off-by: Artem Perevedentsev <aperevedents@nvidia.com >
Co-authored-by: Jiangyun Zhu <riverclouds.zhu@qq.com >
2026-05-05 14:08:57 +04:00
Kunshang Ji and GitHub
2ceea42958
[XPU] use xpu topk topp sample kernel ( #39285 )
...
Signed-off-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-05 18:05:17 +08:00
bee126165f
[P/D][Mooncake] Add KVConnectorStats for transfer observability ( #40414 )
...
Signed-off-by: Zhewen Li <zhewenli@inferact.ai >
Co-authored-by: Zhewen Li <zhewenli@inferact.ai >
2026-05-05 02:17:38 -07:00
BitToby and GitHub
27cc676be3
[Model] Use AutoWeightsLoader for Plamo2 ( #41699 )
...
Signed-off-by: bittoby <218712309+bittoby@users.noreply.github.com >
2026-05-05 08:56:24 +00:00
4845aee6b7
[Benchmark] Add --trust-remote-code flag to multi-turn benchmark ( #41661 )
...
Signed-off-by: Dao Le <daole@inferact.ai >
Signed-off-by: Dao Le <Dao007forever@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-05 01:00:37 -07:00
BitToby and GitHub
0c620d2e08
[Model] Use AutoWeightsLoader for CohereMoe ( #41690 )
...
Signed-off-by: bittoby <218712309+bittoby@users.noreply.github.com >
2026-05-05 04:44:15 +00:00
6bb924bbf3
[Model] Fix Gemma4 MoE activation mismatch ( #41574 )
...
Signed-off-by: Luciano Martins <lucianommartins@users.noreply.github.com >
Co-authored-by: Luciano Martins <lucianommartins@users.noreply.github.com >
2026-05-05 04:34:11 +00:00
czhu-cohere and GitHub
eaec7be446
[BugFix] Preserve max_seq_len in ubatch metadata during CUDA graph capture ( #40961 )
...
Signed-off-by: root <conway.zhu@cohere.com >
Signed-off-by: <conway.zhu@cohere.com >
2026-05-05 04:29:34 +00:00
Jeffrey Wang and GitHub
f04fd1677b
[Ray] Enable RayExecutorV2 by default ( #41421 )
...
Signed-off-by: Jeffrey Wang <jeffreywang@anyscale.com >
2026-05-05 04:27:34 +00:00
420b0a5c95
[Hardware][Power]Add Power VSX Attention Backend and fix l2 Cache Crash ( #40451 )
...
Signed-off-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Signed-off-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Signed-off-by: Akash kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Akash Kaothalkar <akashkaothalkar@akashs-mbp.bl1-in.ibm.com >
Co-authored-by: Akash Kaothalkar <akash.kaothalkar@ibm.com >
Co-authored-by: Li, Jiang <jiang1.li@intel.com >
2026-05-04 20:51:09 -07:00
Bowen Bao and GitHub
1e9500410a
[ROCm][Quantization][2/N] Refactor quark_moe w4a8 w/ oracle ( #39136 )
...
Signed-off-by: Bowen Bao <bowenbao@amd.com >
2026-05-04 19:50:38 -07:00
Nick Hill and GitHub
416f9cdede
[Perf][2/n] Eliminate GPU<->CPU syncs in pooling code ( #41433 )
...
Signed-off-by: Nick Hill <nickhill123@gmail.com >
2026-05-05 02:43:25 +00:00
685bf811d6
[XPU] enable is_act_and_mul for xpu ( #37481 )
...
Signed-off-by: Chendi Xue <chendi.xue@intel.com >
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com >
2026-05-05 01:07:39 +00:00
Giancarlo Delfin and GitHub
e1e4646b06
[Model Runner V2] Rebuild attn metadata between draft decode steps ( #41162 )
...
Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai >
2026-05-05 00:44:55 +00:00
4f2af1a7c0
[Feature] TurboQuant: support hybrid models and uniform quantization ( #39931 )
...
Signed-off-by: JartX <sagformas@epdcenter.es >
Signed-off-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Jim Smith <jhsmith0@me.com >
Co-authored-by: Sandermage <sandermage@users.noreply.github.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-05-04 20:14:01 -04:00
Wentao Ye and GitHub
577b9623e6
[Bug] Fix status update address for non-MOE model within external dp mode ( #40839 )
...
Signed-off-by: yewentao256 <zhyanwentao@126.com >
2026-05-04 16:37:16 -07:00
Andreas Karatzas and GitHub
1cb0838721
[ROCm][CI] Fix MLA prefill scale for DeepSeek GSM8K ( #41569 )
...
Signed-off-by: Andreas Karatzas <akaratza@amd.com >
2026-05-04 16:32:55 -07:00
Matthew Bonanni and GitHub
be5983b874
[Docs] Add non-causal support to attention backend docs ( #41643 )
...
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com >
2026-05-04 20:35:15 +00:00
fxmarty-amd and GitHub
9c07342fdc
[NVFP4][fix] Fix layer.weight -> w13 typo in NVFP4 MOE emulation kernel preparation ( #41630 )
...
Signed-off-by: Felix Marty <Felix.Marty@amd.com >
2026-05-04 20:13:37 +00:00
844df54269
feat: update xgrammar==0.2.0 to use structural tags for strict tool calling + reasoning for more models ( #40894 )
...
Signed-off-by: Yuchuan <yuchuan.7streams@gmail.com >
Signed-off-by: Michael Goin <mgoin64@gmail.com >
Signed-off-by: mgoin <mgoin64@gmail.com >
Signed-off-by: Ubospica <ubospica@gmail.com >
Signed-off-by: sfeng33 <4florafeng@gmail.com >
Co-authored-by: Michael Goin <mgoin64@gmail.com >
Co-authored-by: Ubospica <ubospica@gmail.com >
Co-authored-by: sfeng33 <4florafeng@gmail.com >
2026-05-04 12:45:24 -07:00
422dd02598
[bugfix] Fix prompt logprobs on request eviction during chunked prefill ( #41411 )
...
Signed-off-by: Joachim Studnia <joachim@mistral.ai >
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-05-04 11:46:00 -07:00
8c780943b4
Fix Nano Nemotron text-only weight loading ( #41205 )
...
Signed-off-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Signed-off-by: Baekpica <35071468+Baekpica@users.noreply.github.com >
Signed-off-by: sunghoon.baek <seanbb93@gmail.com >
Co-authored-by: sunghoon.baek <sunghoon.baek@connectfy.cloud >
Co-authored-by: OpenAI Codex <codex@openai.com >
Co-authored-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com >
2026-05-04 21:43:07 +03:00
e724b0ea8d
[ROCm] ROCm7.2.2 + profiler fix + AITER 0.1.12.post2 ( #41386 )
...
Signed-off-by: Rohan138 <rohanpotdar138@gmail.com >
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com >
Co-authored-by: Rohan138 <rohanpotdar138@gmail.com >
2026-05-04 13:07:19 -05:00
712ad0286c
[Bugfix] KimiK2ReasoningParser: guard against buffered end-token in streaming ( #41068 )
...
Signed-off-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Keyi Li <likey6688@gmail.com >
Co-authored-by: Claude <noreply@anthropic.com >
Co-authored-by: Flora Feng <4florafeng@gmail.com >
2026-05-04 17:42:05 +00:00
Ekagra Ranjan and GitHub
321fa2d6d1
Limit gpu utils and lower max BS on test_transcription_api_correctness.py ( #41649 )
...
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com >
2026-05-04 10:30:02 -07:00