forked from Karylab-cklius/vllm
- Responses API: Qwen3-1.7B too small for reasoning + tool calling, bump to Qwen3-4B (~8 GiB, fits on 17GB MIG slice) - MM cache stats: llava-onevision produces more multimodal tokens per image than llava-1.5, increase max_model_len from 4096 to 8192 Signed-off-by: Kevin H. Luu <khluu000@gmail.com> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>