This website requires JavaScript.
Explore
Help
Sign In
karylab_agents
/
vllm
Watch
1
Star
0
Fork
0
forked from
Karylab-cklius/vllm
Code
Pull Requests
2
Actions
1
Packages
Activity
Files
e1bf04b6c27a070859264290ffdccbf333f27fa6
vllm
/
examples
/
online_serving
T
History
Patrick von Platen
and
GitHub
3f7662d650
[Voxtral Realtime] Change name (
#33716
)
...
Signed-off-by: Patrick von Platen <
patrick.v.platen@gmail.com
>
2026-02-03 13:03:28 -08:00
..
chart-helm
…
dashboards
[Metrics] Complete removal of deprecated vllm:time_per_output_token_seconds metric (
#32661
)
2026-01-20 12:28:41 +00:00
disaggregated_encoder
[Doc] Update docs for MM model development with context usage (
#32691
)
2026-01-20 10:37:35 -08:00
disaggregated_serving
[P/D] rework mooncake connector and introduce its bootstrap server (
#31034
)
2026-02-03 08:08:25 -08:00
disaggregated_serving_p2p_nccl_xpyd
…
elastic_ep
…
opentelemetry
…
prometheus_grafana
…
structured_outputs
…
api_client.py
…
disaggregated_prefill.sh
…
gradio_openai_chatbot_webserver.py
…
gradio_webserver.py
…
kv_events_subscriber.py
…
multi_instance_data_parallel.py
…
multi-node-serving.sh
…
openai_chat_completion_client_for_multimodal.py
…
openai_chat_completion_client_with_tools_required.py
…
openai_chat_completion_client_with_tools_xlam_streaming.py
…
openai_chat_completion_client_with_tools_xlam.py
…
openai_chat_completion_client_with_tools.py
…
openai_chat_completion_client.py
…
openai_chat_completion_tool_calls_with_reasoning.py
…
openai_chat_completion_with_reasoning_streaming.py
…
openai_chat_completion_with_reasoning.py
…
openai_completion_client.py
…
openai_realtime_client.py
[Voxtral Realtime] Change name (
#33716
)
2026-02-03 13:03:28 -08:00
openai_realtime_microphone_client.py
[Voxtral Realtime] Change name (
#33716
)
2026-02-03 13:03:28 -08:00
openai_responses_client_with_mcp_tools.py
…
openai_responses_client_with_tools.py
…
openai_responses_client.py
…
openai_transcription_client.py
…
openai_translation_client.py
[Doc] Remove hardcoded Whisper in example openai translation client (
#32027
)
2026-01-09 14:44:52 +00:00
prompt_embed_inference_with_openai_client.py
[Frontend] Use new Renderer for Completions and Tokenize API (
#32863
)
2026-01-31 04:51:15 -08:00
ray_serve_deepseek.py
…
retrieval_augmented_generation_with_langchain.py
…
retrieval_augmented_generation_with_llamaindex.py
…
run_cluster.sh
…
sagemaker-entrypoint.sh
[Misc] Add In-Container restart capability through supervisord for sagemaker entrypoint (
#28502
)
2026-01-13 13:06:10 -08:00
streamlit_openai_chatbot_webserver.py
…
token_generation_client.py
Explicitly set
return_dict
for
apply_chat_template
(
#33372
)
2026-01-30 07:27:04 +00:00
utils.py
…