This website requires JavaScript.
Explore
Help
Sign In
Karylab-cklius
/
vllm
Watch
1
Star
0
Fork
1
Code
Issues
Pull Requests
1
Actions
Packages
Projects
Releases
Wiki
Activity
Files
c798593f0d88cec583c599ea7ea40a2cc26c312b
vllm
/
examples
/
offline_inference
T
History
Shanshan Shen
and
GitHub
fe57be7809
[MM][CG] Support
--enable-vit-cuda-graph
option for VLM examples (
#40580
)
...
Signed-off-by: shen-shanshan <
467638484@qq.com
>
2026-04-22 22:46:14 -07:00
..
disaggregated-prefill-v1
…
kv_load_failure_recovery
…
logits_processor
…
openai_batch
…
qwen2_5_omni
…
qwen3_omni
…
async_llm_streaming.py
…
audio_language.py
[model] support FireRedLID (
#39290
)
2026-04-10 08:43:58 +00:00
automatic_prefix_caching.py
…
batch_llm_inference.py
…
chat_with_tools.py
…
context_extension.py
…
data_parallel.py
…
disaggregated_prefill.py
…
encoder_decoder_multimodal.py
[model] support FireRedLID (
#39290
)
2026-04-10 08:43:58 +00:00
extract_hidden_states.py
Fix shape comment in extract_hidden_states example (
#38723
)
2026-04-01 07:29:33 -07:00
llm_engine_example.py
…
llm_engine_reset_kv.py
…
load_sharded_state.py
…
lora_with_quantization_inference.py
…
mistral-small.py
…
mlpspeculator.py
…
multilora_inference.py
…
pause_resume.py
…
prefix_caching_flexkv.py
…
prefix_caching.py
…
prompt_embed_inference.py
…
qwen_1m.py
…
reproducibility.py
…
routed_experts_e2e.py
…
run_one_batch.py
…
save_sharded_state.py
…
simple_profiling.py
…
skip_loading_weights_in_engine_init.py
…
spec_decode.py
…
structured_outputs.py
…
torchrun_dp_example.py
…
torchrun_example.py
…
vision_language_multi_image.py
Add Granite 4.1 Vision as built-in multimodal model (
#40282
)
2026-04-21 05:43:39 -07:00
vision_language.py
[MM][CG] Support
--enable-vit-cuda-graph
option for VLM examples (
#40580
)
2026-04-22 22:46:14 -07:00