forked from Karylab-cklius/vllm
Support meituan-longcat/LongCat-2.0-FP8: the LongCat-Flash-Lite backbone plus a DeepSeek-V3.2-style sparse attention indexer, reusing vLLM's existing DSA infrastructure. - Resolve the checkpoint's null model_type via a registered config class (arch LongcatCausalLM); alias the oe_* n-gram config fields. - Wire the indexer into the dual-attention layer: only the first attention of each layer computes top-k indices (cli_factor), the second reuses them via the shared buffer. Adds a generic skip_topk override to DeepseekV2MLAAttention; the schedule lives in the model. - Streaming-aware indexing: force the first index_init_tokens and last index_local_tokens into the top-k set (sparse_attn_indexer). - MTP speculative decoding (one module iterated up to 3 draft steps, plain token embedding, FP8 indexer weights). - Fix a latent LongCat weight-loading bug: the post-load mla_scale_q_lora/kv_lora fold breaks under incremental load_weights calls; fold at weight-load time instead. - Fix two n-gram embedding bugs (also affect LongCat-Flash-Lite): rebuild the left-context from the accepted-token history so rejected draft tokens don't pollute it, and hash EOS-current positions with full look-back per the HF reference (unit test included). - Add --reasoning-parser longcat (<longcat_think> tags) and docs rows. - Add longcat model types to the DeepGEMM Blackwell blocklist (non-ue8m0 scales; measured GSM8K-neutral on vLLM but matches SGLang's guard). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: mgoin <mgoin64@gmail.com>