Support meituan-longcat/LongCat-2.0-FP8: the LongCat-Flash-Lite backbone
plus a DeepSeek-V3.2-style sparse attention indexer, reusing vLLM's
existing DSA infrastructure.
- Resolve the checkpoint's null model_type via a registered config class
(arch LongcatCausalLM); alias the oe_* n-gram config fields.
- Wire the indexer into the dual-attention layer: only the first
attention of each layer computes top-k indices (cli_factor), the
second reuses them via the shared buffer. Adds a generic skip_topk
override to DeepseekV2MLAAttention; the schedule lives in the model.
- Streaming-aware indexing: force the first index_init_tokens and last
index_local_tokens into the top-k set (sparse_attn_indexer).
- MTP speculative decoding (one module iterated up to 3 draft steps,
plain token embedding, FP8 indexer weights).
- Fix a latent LongCat weight-loading bug: the post-load
mla_scale_q_lora/kv_lora fold breaks under incremental load_weights
calls; fold at weight-load time instead.
- Fix two n-gram embedding bugs (also affect LongCat-Flash-Lite):
rebuild the left-context from the accepted-token history so rejected
draft tokens don't pollute it, and hash EOS-current positions with
full look-back per the HF reference (unit test included).
- Add --reasoning-parser longcat (<longcat_think> tags) and docs rows.
- Add longcat model types to the DeepGEMM Blackwell blocklist (non-ue8m0
scales; measured GSM8K-neutral on vLLM but matches SGLang's guard).
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: mgoin <mgoin64@gmail.com>