Compare commits

...
Author SHA1 Message Date
Tyler Michael Smith 9df42d2000 [Build] Update pre-commit pip-compile hook to use cu128 torch backend
Matches the test.txt lockfile change to CUDA 12.8.

Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
2026-02-18 16:41:53 -05:00
Tyler Michael SmithandClaude Opus 4.6 8d151aa148 [Build] Recompile test.txt lockfile with cu128 torch backend
The lockfile was compiled with --torch-backend cu129, pinning
torch==2.10.0+cu129. This breaks Docker builds that use CUDA_VERSION=12.8
because the cu128 PyTorch index does not carry +cu129 wheels.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
2026-02-18 14:38:23 -05:00
Tyler Michael SmithandClaude Opus 4.6 11a13b0fd3 [Docker] Add BUILDER_CUDA_VERSION to decouple build and runtime CUDA versions
Allow compiling csrc/ and extensions (DeepGEMM, EP kernels) with a
different CUDA toolkit than the one shipped in the final runtime image.
BUILDER_CUDA_VERSION controls the devel base image used for compilation,
while CUDA_VERSION selects the runtime base image and PyTorch wheel index.

Override the CUDA_VERSION env var inherited from the nvidia base image in
the build stages so PyTorch index URLs resolve to the runtime version.
Update the CUDA 13.0 release pipeline entries to pass the new arg.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>
2026-02-18 14:38:23 -05:00
Amr MahdiandGitHub df3f537a66 [CI] Remove unused precompiled wheel args from image build (#34767)
Signed-off-by: Amr Mahdi <amrmahdi@meta.com>
2026-02-17 18:58:18 -08:00
Matthew BonanniandGitHub 7743152957 [Attention] Refactor check_and_update_config (#33600)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
2026-02-17 17:06:54 -08:00
Wentao YeandGitHub ab33d2a629 [Feature] Decode Context Parallel support for GPU model runner v2 (#34179)
Signed-off-by: yewentao256 <zhyanwentao@126.com>
2026-02-17 16:27:15 -08:00
Woosuk KwonandGitHub be3af2d29e [Model Runner V2] Further simplification for PP (#34724)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-17 15:18:18 -08:00
c656ba3b4d [Kernel] Triton-based Top-k and Top-p sampler kernels (#33538)
Signed-off-by: js_park <cakeng@naver.com>
Signed-off-by: Jongseok Park <37990712+cakeng@users.noreply.github.com>
Signed-off-by: Sunga Kim <sunga.kim@berkeley.edu>
Signed-off-by: Nick Hill <nickhill123@gmail.com>
Co-authored-by: Sunga Kim <sunga.kim@berkeley.edu>
Co-authored-by: Nick Hill <nickhill123@gmail.com>
2026-02-17 23:14:30 +00:00
dc5fa77a4e [Bugfix][MTP][Sparse MLA] Allow sparse MLA with MTP to run with FULL cudagraphs (#34457)
Signed-off-by: Matthew Bonanni <mbonanni@redhat.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
2026-02-17 14:01:27 -05:00
Flora FengandGitHub 1e4a084c8e [CI] Fix flaky test_parsable_context (#34717)
Signed-off-by: sfeng33 <4florafeng@gmail.com>
2026-02-17 18:42:52 +00:00
Richard ZouandGitHub 7967e854da [BugFix] Fix sp tests (#34716)
Signed-off-by: Richard Zou <zou3519@gmail.com>
2026-02-17 17:07:56 +00:00
6bd6d0c3c1 Fixed whisper CPU test that does not spawn properly. (#34324)
Signed-off-by: Anna Mayne <anna.mayne@arm.com>
Co-authored-by: Cyrus Leung <tlleungac@connect.ust.hk>
2026-02-17 06:46:23 -08:00
Nicolò LucchesiandGitHub 8e962fef5f [CI][Nixl] Add CrossLayer KV layout tests (#34615)
Signed-off-by: NickLucche <nlucches@redhat.com>
2026-02-17 21:35:40 +08:00
Cyrus LeungandGitHub 574fe75245 [Renderer] Move InputPreprocessor into Renderer (2/2) (#34560)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
2026-02-17 05:29:01 -08:00
junuxyzandGitHub c61a98f529 [CI][BugFix] ShellCheck cleanup to remove baseline and preserve runtime behavior (#34514)
Signed-off-by: junuxyz <216036880+junuxyz@users.noreply.github.com>
2026-02-17 12:22:56 +00:00
Harry MellorandGitHub 28bffe9466 Fix docs build warning (#34686)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-02-17 02:31:40 -08:00
ad65177a19 [Bugfix] Fix 'remove_instance_endpoint' method logic in disagg_proxy_demo (#32922)
Signed-off-by: ChenqianCao <39755070+ChenqianCao@users.noreply.github.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-02-17 10:06:53 +00:00
d44a5b6c47 Remove dead bitsandbytes CxB code from 8-bit inference path (#34633)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-17 01:49:14 -08:00
Jiangyun ZhuandGitHub 1d65283e95 Revert "[Models] Fuse Qwen3.5 GDN's qkvz_proj and ba_proj" (#34683) 2026-02-17 01:29:27 -08:00
c464b57374 [Ray] Propagate third-party env vars to Ray workers via prefix matching (#34383)
Signed-off-by: Kourosh Hakhamaneshi <kourosh@anyscale.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 01:08:42 -08:00
Amr MahdiandGitHub c5c38e152a [CI] Fix bake config artifact path for AMI rebuild pipeline (#34656)
Signed-off-by: Amr Mahdi <amrmahdi@meta.com>
2026-02-17 06:39:44 +00:00
Woosuk KwonandGitHub d00df624f3 [Model Runner V2] Minor refactoring for penalties (#34662)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-16 21:43:00 -08:00
Woosuk KwonandGitHub 9752da9d9c [Model Runner V2] Minor simplification for BadWordsState (#34669)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-16 21:27:24 -08:00
Woosuk KwonandGitHub 04925b2202 [Model Runner V2] Minor cleanup for PP (#34666)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-16 19:15:31 -08:00
Woosuk KwonandGitHub d74278fb67 [Model Runner V2] Fix unintended CPU-GPU sync in make_dummy (#34667)
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-16 19:00:29 -08:00
b68fd899d1 [Bugfix] Fix fused MoE int32 overflow in stride*offset without perf regression (#34507)
Signed-off-by: haosdent <haosdent@gmail.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
2026-02-16 17:58:49 -08:00
Aneesh PutturandGitHub 0b5f9b7204 [CI] Enable mypy import following for vllm/v1/kv_offload (#34639)
Signed-off-by: Aneesh Puttur <aneeshputtur@gmail.com>
2026-02-17 09:58:15 +08:00
zhanqiuhuandGitHub 9a8853f781 [Core] Pipeline Parallel support for Model Runner V2 (#33960)
Signed-off-by: Zhanqiu Hu <zh338@cornell.edu>
2026-02-16 17:48:16 -08:00
387a1898d9 [Model Runner V2] support bad_words sampling param (#33433)
Signed-off-by: zhuhaoran <zhuhaoran.zhr@alibaba-inc.com>
Signed-off-by: Woosuk Kwon <woosuk@inferact.ai>
Co-authored-by: Woosuk Kwon <woosuk@inferact.ai>
2026-02-16 16:36:06 -08:00
roikoren755andGitHub 3b30e61507 [NemotronH] Do not force router to run in fp32 (#34582)
Signed-off-by: Roi Koren <roik@nvidia.com>
2026-02-16 10:15:32 -08:00
Alexei-V-Ivanov-AMDandGitHub 824f9e8f3c Targeting the MI355 agent pool with all existing tests (#34629)
Signed-off-by: Alexei V. Ivanov <alexei.ivanov@amd.com>
2026-02-16 17:02:27 +00:00
Nicolò LucchesiandGitHub 6cc403e67d [Bugfix][CI] Fix flaky entrypoints/openai/test_response_api_with_harmony.py::test_function_calling[openai/gpt-oss-20b] (#34624)
Signed-off-by: NickLucche <nlucches@redhat.com>
2026-02-16 16:11:07 +00:00
Almog TavorandGitHub 72d5951d02 [Bugfix] Treat generation_config max_tokens as default not ceiling (#34063)
Signed-off-by: almogtavor <almogtavor@gmail.com>
2026-02-16 07:58:24 -08:00
a3205beffb [CI] Enable mypy coverage for individual excluded files (#34292)
Signed-off-by: Lucas Kabela <lucaskabela@meta.com>
Co-authored-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-02-16 07:34:29 -08:00
Christian PintoandGitHub 6930becd45 (bugfix): Fixed encode in LLM entrypoint for IOProcessr plugin prompts (#34618)
Signed-off-by: Christian Pinto <christian.pinto@ibm.com>
2026-02-16 07:33:55 -08:00
Andreas KaratzasandGitHub 03a8770a6d [ROCm][CI] Fix plugins test group; updating terratorch and dependencies (#34589)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
2026-02-16 07:33:42 -08:00
Yiqi XueandGitHub bc56a1d56e [Bugfix] Fix ARC touch KeyError for non-ready T1 blocks in kv offload (#34576)
Signed-off-by: Yiqi Xue <xuey666@gmail.com>
2026-02-16 07:33:19 -08:00
daniserebandGitHub ec7d9e6745 Fix call to moe_mk in modelopt MoE modules (required for LoRA) (#34575)
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
2026-02-16 07:33:09 -08:00
Isotr0pyandGitHub 3bb4e4311c [Models] Fuse Qwen3.5 GDN's qkvz_proj and ba_proj (#34492)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
2026-02-16 07:32:51 -08:00
Amr MahdiandGitHub 08f8c198ae [CI] Disable precompiled wheel path in CI image builds (#34606)
Signed-off-by: Amr Mahdi <amrmahdi@meta.com>
2026-02-16 15:14:43 +00:00
Harry MellorandGitHub a21cedf4ff Bump lm-eval version for Transformers v5 compatibility (#33994)
Signed-off-by: Harry Mellor <19981378+hmellor@users.noreply.github.com>
2026-02-16 05:24:35 -08:00
emricksini-handGitHub 3ef74cde5d [CI][Tracing] Fix race condition by adding server readiness check (#34364)
Attempt to resolve #34284: "Metrics Tracing (2GPU)" fails with a
segmentation fault.

Signed-off-by: emricksini-h <emrick.birivoutin@hcompany.ai>
2026-02-16 12:57:39 +00:00
cd81cdb399 [Scheduler][ASR] Fix CrossAttn blocks per-request for Variable length encoder inputs (#31058)
Signed-off-by: Ekagra Ranjan <3116519+ekagra-ranjan@users.noreply.github.com>
Co-authored-by: Nicolò Lucchesi <nlucches@redhat.com>
2026-02-16 11:08:44 +00:00
Andreas KaratzasandGitHub 1e828573b4 [CI][Metrics] Stabilize tests with polling and subprocess guards (#34566)
test_abort_metrics_reset is flaky due to hardware-dependent
fixed sleeps: replace fixed sleeps with polling.

test_metrics_exist_run_batch passes even when the engine crashes
on startup (false positive): add subprocess lifecycle guards.

Signed-off-by: Andreas Karatzas <akaratza@amd.com>
2026-02-16 10:52:02 +00:00
a5ccc85c8c [Bugfix] Fix Dynamo unexpected keyword argument (#34320)
Signed-off-by: Samu Tamminen <stammine@amd.com>
Co-authored-by: Wentao Ye <44945378+yewentao256@users.noreply.github.com>
2026-02-16 01:32:30 -08:00
Roger WangandGitHub b5475d0534 Revert "[Misc] fix qwen3.5 config" (#34610) 2026-02-16 01:06:05 -08:00
JJJYmmmandGitHub 9521002f0a [Misc] fix qwen3.5 config (#34604) 2026-02-16 00:25:38 -08:00
Cyrus LeungandGitHub ec17bdd894 [Renderer] Move InputPreprocessor into Renderer (1.5/2) (#34598)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
2026-02-15 23:46:33 -08:00
Amr MahdiandGitHub bb59c90248 [CI] Write bake config to temp directory instead of repo root (#34569)
Signed-off-by: Amr Mahdi <amrmahdi@meta.com>
2026-02-15 22:15:47 -08:00
bnellnmandGitHub 5bff999d12 [Bugfix] Add method to swap quant_method on FusedMoE to fix LoRA issues (#34453)
Signed-off-by: Bill Nell <bnell@redhat.com>
2026-02-15 20:10:50 -08:00
bb85929aa6 [BugFix] Fix Python 3.13 FlashMLA import error (#34548)
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
2026-02-15 20:09:18 -08:00
Parth BansalandGitHub 5653021094 [Doc] Add Mistral-7b-v0.3 model to the batch invariance validated model (#34584)
Signed-off-by: Parth Bansal <parthbansal127@gmail.com>
2026-02-16 12:09:00 +08:00
Andreas KaratzasandGitHub 974d829b05 [CI][Frontend] Return 422 instead of 500 for invalid Anthropic tool_choice (#34590)
Signed-off-by: Andreas Karatzas <akaratza@amd.com>
2026-02-15 20:06:48 -08:00
Isotr0pyandGitHub 91ac5d9bfd [CI/Build] Enable tests for recent day-0 new models (#34585)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
2026-02-15 18:17:04 -08:00
Luka GovedičGitHubgemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
23d825aba1 [torch.compile] Disable ar-rms fusion for ds3-fp4 & DP, fix CI test (#34392)
Signed-off-by: Luka Govedič <lgovedic@redhat.com>
Signed-off-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-15 06:33:57 -08:00
Maryam TahhanGitHubLi, Jiang <jiang1.li@intel.com>
f07a128413 [CPU][ARM] Add ARM BF16 cross-compilation support and improve documen… (#33079)
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-authored-by: Li, Jiang <jiang1.li@intel.com>
2026-02-15 06:33:08 -08:00
Isotr0pyandGitHub 71cd89264f [MM Encoder] Add Triton ViT attention backend (#32183)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
2026-02-15 06:32:47 -08:00
Isotr0pyandGitHub 19fab44152 [Doc] Update Encoder-Decoder models support doc with Florence-2 (#34581)
Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
2026-02-15 04:18:57 -08:00
201 changed files with 7554 additions and 2323 deletions
+11 -12
View File
@@ -8,7 +8,7 @@ clean_docker_tag() {
} }
print_usage_and_exit() { print_usage_and_exit() {
echo "Usage: $0 <registry> <repo> <commit> <branch> <vllm_use_precompiled> <vllm_merge_base_commit> <cache_from> <cache_to>" echo "Usage: $0 <registry> <repo> <commit> <branch> <image_tag> [<image_tag_latest>]"
exit 1 exit 1
} }
@@ -142,11 +142,16 @@ resolve_parent_commit() {
print_bake_config() { print_bake_config() {
echo "--- :page_facing_up: Resolved bake configuration" echo "--- :page_facing_up: Resolved bake configuration"
BAKE_CONFIG_FILE="bake-config-build-${BUILDKITE_BUILD_NUMBER:-local}.json" # Write to a temp directory to avoid polluting the repo root (which is the
# Docker build context). Files left in the repo root get COPY'd into the
# image and can cause duplicate artifact uploads from downstream steps.
local bake_tmp
bake_tmp="$(mktemp -d)"
BAKE_CONFIG_FILE="${bake_tmp}/bake-config-build-${BUILDKITE_BUILD_NUMBER:-local}.json"
docker buildx bake -f "${VLLM_BAKE_FILE_PATH}" -f "${CI_HCL_PATH}" --print "${TARGET}" | tee "${BAKE_CONFIG_FILE}" || true docker buildx bake -f "${VLLM_BAKE_FILE_PATH}" -f "${CI_HCL_PATH}" --print "${TARGET}" | tee "${BAKE_CONFIG_FILE}" || true
echo "Saved bake config to ${BAKE_CONFIG_FILE}" echo "Saved bake config to ${BAKE_CONFIG_FILE}"
echo "--- :arrow_down: Uploading bake config to Buildkite" echo "--- :arrow_down: Uploading bake config to Buildkite"
buildkite-agent artifact upload "${BAKE_CONFIG_FILE}" (cd "$(dirname "${BAKE_CONFIG_FILE}")" && buildkite-agent artifact upload "$(basename "${BAKE_CONFIG_FILE}")")
} }
################################# #################################
@@ -154,7 +159,7 @@ print_bake_config() {
################################# #################################
print_instance_info print_instance_info
if [[ $# -lt 7 ]]; then if [[ $# -lt 5 ]]; then
print_usage_and_exit print_usage_and_exit
fi fi
@@ -163,10 +168,8 @@ REGISTRY=$1
REPO=$2 REPO=$2
BUILDKITE_COMMIT=$3 BUILDKITE_COMMIT=$3
BRANCH=$4 BRANCH=$4
VLLM_USE_PRECOMPILED=$5 IMAGE_TAG=$5
VLLM_MERGE_BASE_COMMIT=$6 IMAGE_TAG_LATEST=${6:-} # only used for main branch, optional
IMAGE_TAG=$7
IMAGE_TAG_LATEST=${8:-} # only used for main branch, optional
# build config # build config
TARGET="test-ci" TARGET="test-ci"
@@ -193,8 +196,6 @@ export CACHE_FROM
export CACHE_FROM_BASE_BRANCH export CACHE_FROM_BASE_BRANCH
export CACHE_FROM_MAIN export CACHE_FROM_MAIN
export CACHE_TO export CACHE_TO
export VLLM_USE_PRECOMPILED
export VLLM_MERGE_BASE_COMMIT
# print args # print args
echo "--- :mag: Arguments" echo "--- :mag: Arguments"
@@ -202,8 +203,6 @@ echo "REGISTRY: ${REGISTRY}"
echo "REPO: ${REPO}" echo "REPO: ${REPO}"
echo "BUILDKITE_COMMIT: ${BUILDKITE_COMMIT}" echo "BUILDKITE_COMMIT: ${BUILDKITE_COMMIT}"
echo "BRANCH: ${BRANCH}" echo "BRANCH: ${BRANCH}"
echo "VLLM_USE_PRECOMPILED: ${VLLM_USE_PRECOMPILED}"
echo "VLLM_MERGE_BASE_COMMIT: ${VLLM_MERGE_BASE_COMMIT}"
echo "IMAGE_TAG: ${IMAGE_TAG}" echo "IMAGE_TAG: ${IMAGE_TAG}"
echo "IMAGE_TAG_LATEST: ${IMAGE_TAG_LATEST}" echo "IMAGE_TAG_LATEST: ${IMAGE_TAG_LATEST}"
+1 -2
View File
@@ -5,8 +5,7 @@ steps:
depends_on: [] depends_on: []
timeout_in_minutes: 600 timeout_in_minutes: 600
commands: commands:
- if [[ "$BUILDKITE_BRANCH" != "main" ]]; then .buildkite/image_build/image_build.sh $REGISTRY $REPO $BUILDKITE_COMMIT $BRANCH $VLLM_USE_PRECOMPILED $VLLM_MERGE_BASE_COMMIT $IMAGE_TAG; fi - if [[ "$BUILDKITE_BRANCH" == "main" ]]; then .buildkite/image_build/image_build.sh $REGISTRY $REPO $BUILDKITE_COMMIT $BRANCH $IMAGE_TAG $IMAGE_TAG_LATEST; else .buildkite/image_build/image_build.sh $REGISTRY $REPO $BUILDKITE_COMMIT $BRANCH $IMAGE_TAG; fi
- if [[ "$BUILDKITE_BRANCH" == "main" ]]; then .buildkite/image_build/image_build.sh $REGISTRY $REPO $BUILDKITE_COMMIT $BRANCH $VLLM_USE_PRECOMPILED $VLLM_MERGE_BASE_COMMIT $IMAGE_TAG $IMAGE_TAG_LATEST; fi
retry: retry:
automatic: automatic:
- exit_status: -1 # Agent was lost - exit_status: -1 # Agent was lost
+5 -5
View File
@@ -11,10 +11,10 @@ REPO=$2
BUILDKITE_COMMIT=$3 BUILDKITE_COMMIT=$3
# authenticate with AWS ECR # authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin $REGISTRY aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY"
# skip build if image already exists # skip build if image already exists
if [[ -z $(docker manifest inspect $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu) ]]; then if [[ -z $(docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu) ]]; then
echo "Image not found, proceeding with build..." echo "Image not found, proceeding with build..."
else else
echo "Image found" echo "Image found"
@@ -24,13 +24,13 @@ fi
# build # build
docker build --file docker/Dockerfile.cpu \ docker build --file docker/Dockerfile.cpu \
--build-arg max_jobs=16 \ --build-arg max_jobs=16 \
--build-arg buildkite_commit=$BUILDKITE_COMMIT \ --build-arg buildkite_commit="$BUILDKITE_COMMIT" \
--build-arg VLLM_CPU_AVX512BF16=true \ --build-arg VLLM_CPU_AVX512BF16=true \
--build-arg VLLM_CPU_AVX512VNNI=true \ --build-arg VLLM_CPU_AVX512VNNI=true \
--build-arg VLLM_CPU_AMXBF16=true \ --build-arg VLLM_CPU_AMXBF16=true \
--tag $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu \ --tag "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu \
--target vllm-test \ --target vllm-test \
--progress plain . --progress plain .
# push # push
docker push $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu docker push "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu
@@ -11,10 +11,10 @@ REPO=$2
BUILDKITE_COMMIT=$3 BUILDKITE_COMMIT=$3
# authenticate with AWS ECR # authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin $REGISTRY aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY"
# skip build if image already exists # skip build if image already exists
if [[ -z $(docker manifest inspect $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu) ]]; then if [[ -z $(docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu) ]]; then
echo "Image not found, proceeding with build..." echo "Image not found, proceeding with build..."
else else
echo "Image found" echo "Image found"
@@ -24,10 +24,10 @@ fi
# build # build
docker build --file docker/Dockerfile.cpu \ docker build --file docker/Dockerfile.cpu \
--build-arg max_jobs=16 \ --build-arg max_jobs=16 \
--build-arg buildkite_commit=$BUILDKITE_COMMIT \ --build-arg buildkite_commit="$BUILDKITE_COMMIT" \
--tag $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu \ --tag "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu \
--target vllm-test \ --target vllm-test \
--progress plain . --progress plain .
# push # push
docker push $REGISTRY/$REPO:$BUILDKITE_COMMIT-cpu docker push "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-cpu
+5 -5
View File
@@ -11,10 +11,10 @@ REPO=$2
BUILDKITE_COMMIT=$3 BUILDKITE_COMMIT=$3
# authenticate with AWS ECR # authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin $REGISTRY aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY"
# skip build if image already exists # skip build if image already exists
if [[ -z $(docker manifest inspect $REGISTRY/$REPO:$BUILDKITE_COMMIT-hpu) ]]; then if [[ -z $(docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-hpu) ]]; then
echo "Image not found, proceeding with build..." echo "Image not found, proceeding with build..."
else else
echo "Image found" echo "Image found"
@@ -25,10 +25,10 @@ fi
docker build \ docker build \
--file tests/pytorch_ci_hud_benchmark/Dockerfile.hpu \ --file tests/pytorch_ci_hud_benchmark/Dockerfile.hpu \
--build-arg max_jobs=16 \ --build-arg max_jobs=16 \
--build-arg buildkite_commit=$BUILDKITE_COMMIT \ --build-arg buildkite_commit="$BUILDKITE_COMMIT" \
--tag $REGISTRY/$REPO:$BUILDKITE_COMMIT-hpu \ --tag "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-hpu \
--progress plain \ --progress plain \
https://github.com/vllm-project/vllm-gaudi.git https://github.com/vllm-project/vllm-gaudi.git
# push # push
docker push $REGISTRY/$REPO:$BUILDKITE_COMMIT-hpu docker push "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-hpu
@@ -2,7 +2,7 @@
# We can use this script to compute baseline accuracy on chartqa for vllm. # We can use this script to compute baseline accuracy on chartqa for vllm.
# #
# Make sure you have lm-eval-harness installed: # Make sure you have lm-eval-harness installed:
# pip install "lm-eval[api]>=0.4.9.2" # pip install "lm-eval[api]>=0.4.11"
usage() { usage() {
echo`` echo``
@@ -41,4 +41,4 @@ lm_eval --model vllm-vlm \
--tasks chartqa \ --tasks chartqa \
--batch_size auto \ --batch_size auto \
--apply_chat_template \ --apply_chat_template \
--limit $LIMIT --limit "$LIMIT"
@@ -2,7 +2,7 @@
# We can use this script to compute baseline accuracy on GSM for transformers. # We can use this script to compute baseline accuracy on GSM for transformers.
# #
# Make sure you have lm-eval-harness installed: # Make sure you have lm-eval-harness installed:
# pip install "lm-eval[api]>=0.4.9.2" # pip install "lm-eval[api]>=0.4.11"
usage() { usage() {
echo`` echo``
@@ -3,7 +3,7 @@
# We use this for fp8, which HF does not support. # We use this for fp8, which HF does not support.
# #
# Make sure you have lm-eval-harness installed: # Make sure you have lm-eval-harness installed:
# pip install "lm-eval[api]>=0.4.9.2" # pip install "lm-eval[api]>=0.4.11"
usage() { usage() {
echo`` echo``
@@ -3,7 +3,7 @@
# We use this for fp8, which HF does not support. # We use this for fp8, which HF does not support.
# #
# Make sure you have lm-eval-harness installed: # Make sure you have lm-eval-harness installed:
# pip install "lm-eval[api]>=0.4.9.2" # pip install "lm-eval[api]>=0.4.11"
usage() { usage() {
echo`` echo``
@@ -20,14 +20,11 @@ usage() {
echo echo
} }
while getopts "m:b:l:f:t:" OPT; do while getopts "m:l:f:t:" OPT; do
case ${OPT} in case ${OPT} in
m ) m )
MODEL="$OPTARG" MODEL="$OPTARG"
;; ;;
b )
BATCH_SIZE="$OPTARG"
;;
l ) l )
LIMIT="$OPTARG" LIMIT="$OPTARG"
;; ;;
@@ -15,11 +15,11 @@ DTYPE_FILTER="${DTYPE_FILTER:-}"
check_gpus() { check_gpus() {
if command -v nvidia-smi; then if command -v nvidia-smi; then
# check the number of GPUs and GPU type. # check the number of GPUs and GPU type.
declare -g gpu_count=$(nvidia-smi --list-gpus | wc -l) declare -g gpu_count=$(nvidia-smi --list-gpus | grep -c . || true)
elif command -v amd-smi; then elif command -v amd-smi; then
declare -g gpu_count=$(amd-smi list | grep 'GPU' | wc -l) declare -g gpu_count=$(amd-smi list | grep -c 'GPU' || true)
elif command -v hl-smi; then elif command -v hl-smi; then
declare -g gpu_count=$(hl-smi --list | grep -i "Module ID" | wc -l) declare -g gpu_count=$(hl-smi --list | grep -ci "Module ID" || true)
fi fi
if [[ $gpu_count -gt 0 ]]; then if [[ $gpu_count -gt 0 ]]; then
@@ -47,7 +47,7 @@ check_cpus() {
declare -g numa_count=$(lscpu | grep "NUMA node(s):" | awk '{print $3}') declare -g numa_count=$(lscpu | grep "NUMA node(s):" | awk '{print $3}')
if [[ $numa_count -gt 0 ]]; then if [[ $numa_count -gt 0 ]]; then
echo "NUMA found." echo "NUMA found."
echo $numa_count echo "$numa_count"
else else
echo "Need at least 1 NUMA to run benchmarking." echo "Need at least 1 NUMA to run benchmarking."
exit 1 exit 1
@@ -434,7 +434,7 @@ run_serving_tests() {
# iterate over different max_concurrency # iterate over different max_concurrency
for max_concurrency in $max_concurrency_list; do for max_concurrency in $max_concurrency_list; do
new_test_name=$test_name"_qps_"$qps"_concurrency_"$max_concurrency new_test_name="${test_name}_qps_${qps}_concurrency_${max_concurrency}"
echo " new test name $new_test_name" echo " new test name $new_test_name"
# pass the tensor parallel size, the compilation mode, and the optimization # pass the tensor parallel size, the compilation mode, and the optimization
# level to the client so that they can be used on the benchmark dashboard # level to the client so that they can be used on the benchmark dashboard
@@ -471,7 +471,7 @@ run_serving_tests() {
# clean up # clean up
if [[ "${DRY_RUN:-0}" != "1" ]]; then if [[ "${DRY_RUN:-0}" != "1" ]]; then
kill -9 $server_pid kill -9 "$server_pid"
kill_gpu_processes kill_gpu_processes
fi fi
done done
+4 -4
View File
@@ -31,7 +31,7 @@ steps:
commands: commands:
# #NOTE: torch_cuda_arch_list is derived from upstream PyTorch build files here: # #NOTE: torch_cuda_arch_list is derived from upstream PyTorch build files here:
# https://github.com/pytorch/pytorch/blob/main/.ci/aarch64_linux/aarch64_ci_build.sh#L7 # https://github.com/pytorch/pytorch/blob/main/.ci/aarch64_linux/aarch64_ci_build.sh#L7
- "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg torch_cuda_arch_list='8.7 8.9 9.0 10.0+PTX 12.0' --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag vllm-ci:build-image --target build --progress plain -f docker/Dockerfile ." - "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg BUILDER_CUDA_VERSION=13.0.1 --build-arg torch_cuda_arch_list='8.7 8.9 9.0 10.0+PTX 12.0' --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag vllm-ci:build-image --target build --progress plain -f docker/Dockerfile ."
- "mkdir artifacts" - "mkdir artifacts"
- "docker run --rm -v $(pwd)/artifacts:/artifacts_host vllm-ci:build-image bash -c 'cp -r dist /artifacts_host && chmod -R a+rw /artifacts_host'" - "docker run --rm -v $(pwd)/artifacts:/artifacts_host vllm-ci:build-image bash -c 'cp -r dist /artifacts_host && chmod -R a+rw /artifacts_host'"
- "bash .buildkite/scripts/upload-nightly-wheels.sh manylinux_2_35" - "bash .buildkite/scripts/upload-nightly-wheels.sh manylinux_2_35"
@@ -70,7 +70,7 @@ steps:
agents: agents:
queue: cpu_queue_postmerge queue: cpu_queue_postmerge
commands: commands:
- "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag vllm-ci:build-image --target build --progress plain -f docker/Dockerfile ." - "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg BUILDER_CUDA_VERSION=13.0.1 --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag vllm-ci:build-image --target build --progress plain -f docker/Dockerfile ."
- "mkdir artifacts" - "mkdir artifacts"
- "docker run --rm -v $(pwd)/artifacts:/artifacts_host vllm-ci:build-image bash -c 'cp -r dist /artifacts_host && chmod -R a+rw /artifacts_host'" - "docker run --rm -v $(pwd)/artifacts:/artifacts_host vllm-ci:build-image bash -c 'cp -r dist /artifacts_host && chmod -R a+rw /artifacts_host'"
- "bash .buildkite/scripts/upload-nightly-wheels.sh manylinux_2_35" - "bash .buildkite/scripts/upload-nightly-wheels.sh manylinux_2_35"
@@ -123,7 +123,7 @@ steps:
queue: cpu_queue_postmerge queue: cpu_queue_postmerge
commands: commands:
- "aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7" - "aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7"
- "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg INSTALL_KV_CONNECTORS=true --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 --target vllm-openai --progress plain -f docker/Dockerfile ." - "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg BUILDER_CUDA_VERSION=13.0.1 --build-arg INSTALL_KV_CONNECTORS=true --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 --target vllm-openai --progress plain -f docker/Dockerfile ."
- "docker push public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130" - "docker push public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130"
# re-tag to default image tag and push, just in case arm64 build fails # re-tag to default image tag and push, just in case arm64 build fails
- "docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-cu130" - "docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-cu130"
@@ -137,7 +137,7 @@ steps:
commands: commands:
- "aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7" - "aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7"
# compute capability 12.0 for RTX-50 series / RTX PRO 6000 Blackwell, 12.1 for DGX Spark # compute capability 12.0 for RTX-50 series / RTX PRO 6000 Blackwell, 12.1 for DGX Spark
- "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg torch_cuda_arch_list='8.7 8.9 9.0 10.0+PTX 12.0 12.1' --build-arg INSTALL_KV_CONNECTORS=true --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 --target vllm-openai --progress plain -f docker/Dockerfile ." - "DOCKER_BUILDKIT=1 docker build --build-arg max_jobs=16 --build-arg USE_SCCACHE=1 --build-arg GIT_REPO_CHECK=1 --build-arg CUDA_VERSION=13.0.1 --build-arg BUILDER_CUDA_VERSION=13.0.1 --build-arg torch_cuda_arch_list='8.7 8.9 9.0 10.0+PTX 12.0 12.1' --build-arg INSTALL_KV_CONNECTORS=true --build-arg BUILD_BASE_IMAGE=nvidia/cuda:13.0.1-devel-ubuntu22.04 --tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130 --target vllm-openai --progress plain -f docker/Dockerfile ."
- "docker push public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130" - "docker push public.ecr.aws/q9t5s3a7/vllm-release-repo:$BUILDKITE_COMMIT-$(uname -m)-cu130"
- block: "Build release image for x86_64 CPU" - block: "Build release image for x86_64 CPU"
+1 -1
View File
@@ -25,7 +25,7 @@ S3_REGION="${AWS_DEFAULT_REGION:-us-west-2}"
S3_URL="http://${S3_BUCKET}.s3-website-${S3_REGION}.amazonaws.com" S3_URL="http://${S3_BUCKET}.s3-website-${S3_REGION}.amazonaws.com"
# Format ROCm version for path (e.g., "7.1" -> "rocm710") # Format ROCm version for path (e.g., "7.1" -> "rocm710")
ROCM_VERSION_PATH="rocm$(echo ${ROCM_VERSION} | tr -d '.')" ROCM_VERSION_PATH="rocm$(echo "${ROCM_VERSION}" | tr -d '.')"
ROCM_PATH="rocm/${BUILDKITE_COMMIT}/${ROCM_VERSION_PATH}" ROCM_PATH="rocm/${BUILDKITE_COMMIT}/${ROCM_VERSION_PATH}"
buildkite-agent annotate --style 'success' --context 'rocm-release-workflow' << EOF buildkite-agent annotate --style 'success' --context 'rocm-release-workflow' << EOF
## ROCm Wheel and Docker Image Releases ## ROCm Wheel and Docker Image Releases
+3 -3
View File
@@ -83,7 +83,7 @@ case "${1:-}" in
exit 1 exit 1
fi fi
WHEEL_COUNT=$(ls artifacts/rocm-base-wheels/*.whl 2>/dev/null | wc -l) WHEEL_COUNT=$(find artifacts/rocm-base-wheels -maxdepth 1 -name '*.whl' 2>/dev/null | wc -l)
if [[ "$WHEEL_COUNT" -eq 0 ]]; then if [[ "$WHEEL_COUNT" -eq 0 ]]; then
echo "ERROR: No wheels found in artifacts/rocm-base-wheels/" >&2 echo "ERROR: No wheels found in artifacts/rocm-base-wheels/" >&2
exit 1 exit 1
@@ -110,9 +110,9 @@ case "${1:-}" in
echo "" echo ""
echo "Downloaded wheels:" echo "Downloaded wheels:"
ls -lh artifacts/rocm-base-wheels/ find artifacts/rocm-base-wheels -maxdepth 1 -name '*.whl' -exec ls -lh {} \;
WHEEL_COUNT=$(ls artifacts/rocm-base-wheels/*.whl 2>/dev/null | wc -l) WHEEL_COUNT=$(find artifacts/rocm-base-wheels -maxdepth 1 -name '*.whl' 2>/dev/null | wc -l)
echo "" echo ""
echo "Total: $WHEEL_COUNT wheels" echo "Total: $WHEEL_COUNT wheels"
echo "========================================" echo "========================================"
@@ -134,7 +134,7 @@ log_info "Fetching merged PRs from milestone '${MILESTONE}'..."
# Store PR data in a temp file # Store PR data in a temp file
PR_DATA=$(mktemp) PR_DATA=$(mktemp)
trap "rm -f $PR_DATA" EXIT trap 'rm -f "$PR_DATA"' EXIT
if ! gh pr list --state merged --search "milestone:${MILESTONE}" \ if ! gh pr list --state merged --search "milestone:${MILESTONE}" \
--limit 1000 \ --limit 1000 \
@@ -27,7 +27,7 @@ function cpu_tests() {
podman exec -it "$container_id" bash -c " podman exec -it "$container_id" bash -c "
export TORCH_COMPILE_DISABLE=1 export TORCH_COMPILE_DISABLE=1
set -xve set -xve
python3 examples/offline_inference/basic/generate.py --model facebook/opt-125m" >> $HOME/test_basic.log python3 examples/offline_inference/basic/generate.py --model facebook/opt-125m" >> "$HOME"/test_basic.log
# Run basic model test # Run basic model test
podman exec -it "$container_id" bash -c " podman exec -it "$container_id" bash -c "
@@ -43,7 +43,7 @@ function cpu_tests() {
pytest -v -s tests/models/language/generation/test_common.py::test_models[False-False-5-32-google/gemma-1.1-2b-it] pytest -v -s tests/models/language/generation/test_common.py::test_models[False-False-5-32-google/gemma-1.1-2b-it]
pytest -v -s tests/models/language/pooling/test_classification.py::test_models[float-jason9693/Qwen2.5-1.5B-apeach] pytest -v -s tests/models/language/pooling/test_classification.py::test_models[float-jason9693/Qwen2.5-1.5B-apeach]
# TODO: Below test case tests/models/language/pooling/test_embedding.py::test_models[True-ssmits/Qwen2-7B-Instruct-embed-base] fails on ppc64le. Disabling it for time being. # TODO: Below test case tests/models/language/pooling/test_embedding.py::test_models[True-ssmits/Qwen2-7B-Instruct-embed-base] fails on ppc64le. Disabling it for time being.
# pytest -v -s tests/models/language/pooling/test_embedding.py -m cpu_model" >> $HOME/test_rest.log # pytest -v -s tests/models/language/pooling/test_embedding.py -m cpu_model" >> "$HOME"/test_rest.log
} }
# All of CPU tests are expected to be finished less than 40 mins. # All of CPU tests are expected to be finished less than 40 mins.
@@ -16,5 +16,5 @@ echo "--- :docker: Building Docker image"
docker build --progress plain --tag "$IMAGE_NAME" --target vllm-test -f docker/Dockerfile.cpu . docker build --progress plain --tag "$IMAGE_NAME" --target vllm-test -f docker/Dockerfile.cpu .
# Run the image, setting --shm-size=4g for tensor parallel. # Run the image, setting --shm-size=4g for tensor parallel.
docker run --rm --cpuset-cpus=$CORE_RANGE --cpuset-mems=$NUMA_NODE -v ~/.cache/huggingface:/root/.cache/huggingface --privileged=true -e HF_TOKEN -e VLLM_CPU_KVCACHE_SPACE=16 -e VLLM_CPU_CI_ENV=1 -e VLLM_CPU_SIM_MULTI_NUMA=1 --shm-size=4g $IMAGE_NAME \ docker run --rm --cpuset-cpus="$CORE_RANGE" --cpuset-mems="$NUMA_NODE" -v ~/.cache/huggingface:/root/.cache/huggingface --privileged=true -e HF_TOKEN -e VLLM_CPU_KVCACHE_SPACE=16 -e VLLM_CPU_CI_ENV=1 -e VLLM_CPU_SIM_MULTI_NUMA=1 --shm-size=4g "$IMAGE_NAME" \
timeout $TIMEOUT_VAL bash -c "set -euox pipefail; echo \"--- Print packages\"; pip list; echo \"--- Running tests\"; ${TEST_COMMAND}" timeout "$TIMEOUT_VAL" bash -c "set -euox pipefail; echo \"--- Print packages\"; pip list; echo \"--- Running tests\"; ${TEST_COMMAND}"
@@ -7,7 +7,7 @@ set -exuo pipefail
# Try building the docker image # Try building the docker image
image_name="hpu/upstream-vllm-ci:${BUILDKITE_COMMIT}" image_name="hpu/upstream-vllm-ci:${BUILDKITE_COMMIT}"
container_name="hpu-upstream-vllm-ci-${BUILDKITE_COMMIT}-container" container_name="hpu-upstream-vllm-ci-${BUILDKITE_COMMIT}-container"
cat <<EOF | docker build -t ${image_name} -f - . cat <<EOF | docker build -t "${image_name}" -f - .
FROM gaudi-base-image:latest FROM gaudi-base-image:latest
COPY ./ /workspace/vllm COPY ./ /workspace/vllm
@@ -39,12 +39,12 @@ EOF
# functions, while other platforms only need one remove_docker_container # functions, while other platforms only need one remove_docker_container
# function. # function.
EXITCODE=1 EXITCODE=1
remove_docker_containers() { docker rm -f ${container_name} || true; } remove_docker_containers() { docker rm -f "${container_name}" || true; }
trap 'remove_docker_containers; exit $EXITCODE;' EXIT trap 'remove_docker_containers; exit $EXITCODE;' EXIT
remove_docker_containers remove_docker_containers
echo "Running HPU plugin v1 test" echo "Running HPU plugin v1 test"
docker run --rm --runtime=habana --name=${container_name} --network=host \ docker run --rm --runtime=habana --name="${container_name}" --network=host \
-e HABANA_VISIBLE_DEVICES=all \ -e HABANA_VISIBLE_DEVICES=all \
-e VLLM_SKIP_WARMUP=true \ -e VLLM_SKIP_WARMUP=true \
-e PT_HPU_ENABLE_LAZY_COLLECTIVES=true \ -e PT_HPU_ENABLE_LAZY_COLLECTIVES=true \
+15 -20
View File
@@ -41,6 +41,7 @@ get_config() {
echo "Error: file '${TEST_RUN_CONFIG_FILE}' does not exist in the warehouse" >&2 echo "Error: file '${TEST_RUN_CONFIG_FILE}' does not exist in the warehouse" >&2
exit 1 exit 1
fi fi
# shellcheck source=/dev/null
source "${TEST_RUN_CONFIG_FILE}" source "${TEST_RUN_CONFIG_FILE}"
echo "Base docker image name that get from configuration: ${BASE_IMAGE_NAME}" echo "Base docker image name that get from configuration: ${BASE_IMAGE_NAME}"
return 0 return 0
@@ -48,9 +49,8 @@ get_config() {
# get test running configuration. # get test running configuration.
fetch_vllm_test_cfg fetch_vllm_test_cfg
get_config
# Check if the function call was successful. If not, exit the script. # Check if the function call was successful. If not, exit the script.
if [ $? -ne 0 ]; then if ! get_config; then
exit 1 exit 1
fi fi
@@ -62,14 +62,14 @@ agent_idx=$(echo "${BUILDKITE_AGENT_NAME}" | awk -F'-' '{print $(NF-1)}')
echo "agent_idx: ${agent_idx}" echo "agent_idx: ${agent_idx}"
builder_name="cachebuilder${agent_idx}" builder_name="cachebuilder${agent_idx}"
builder_cache_dir="/mnt/docker-cache${agent_idx}" builder_cache_dir="/mnt/docker-cache${agent_idx}"
mkdir -p ${builder_cache_dir} mkdir -p "${builder_cache_dir}"
# Try building the docker image # Try building the docker image
cat <<EOF | DOCKER_BUILDKIT=1 docker build \ cat <<EOF | DOCKER_BUILDKIT=1 docker build \
--add-host cache-service-vllm.nginx-pypi-cache.svc.cluster.local:${PYPI_CACHE_HOST} \ --add-host cache-service-vllm.nginx-pypi-cache.svc.cluster.local:"${PYPI_CACHE_HOST}" \
--builder ${builder_name} --cache-from type=local,src=${builder_cache_dir} \ --builder "${builder_name}" --cache-from type=local,src="${builder_cache_dir}" \
--cache-to type=local,dest=${builder_cache_dir},mode=max \ --cache-to type=local,dest="${builder_cache_dir}",mode=max \
--progress=plain --load -t ${image_name} -f - . --progress=plain --load -t "${image_name}" -f - .
FROM ${BASE_IMAGE_NAME} FROM ${BASE_IMAGE_NAME}
# Define environments # Define environments
@@ -116,7 +116,7 @@ RUN --mount=type=cache,target=/root/.cache/pip \
export PIP_EXTRA_INDEX_URL=https://mirrors.huaweicloud.com/ascend/repos/pypi && \ export PIP_EXTRA_INDEX_URL=https://mirrors.huaweicloud.com/ascend/repos/pypi && \
source /usr/local/Ascend/ascend-toolkit/set_env.sh && \ source /usr/local/Ascend/ascend-toolkit/set_env.sh && \
source /usr/local/Ascend/nnal/atb/set_env.sh && \ source /usr/local/Ascend/nnal/atb/set_env.sh && \
export LD_LIBRARY_PATH=\$LD_LIBRARY_PATH:/usr/local/Ascend/ascend-toolkit/latest/`uname -i`-linux/devlib && \ export LD_LIBRARY_PATH=\$LD_LIBRARY_PATH:/usr/local/Ascend/ascend-toolkit/latest/$(uname -i)-linux/devlib && \
python3 -m pip install -v -e /workspace/vllm-ascend/ --extra-index https://download.pytorch.org/whl/cpu/ python3 -m pip install -v -e /workspace/vllm-ascend/ --extra-index https://download.pytorch.org/whl/cpu/
ENV VLLM_WORKER_MULTIPROC_METHOD=spawn ENV VLLM_WORKER_MULTIPROC_METHOD=spawn
@@ -139,7 +139,7 @@ trap remove_docker_container EXIT
# Generate corresponding --device args based on BUILDKITE_AGENT_NAME # Generate corresponding --device args based on BUILDKITE_AGENT_NAME
# Ascend NPU BUILDKITE_AGENT_NAME format is {hostname}-{agent_idx}-{npu_card_num}cards, and agent_idx starts from 1. # Ascend NPU BUILDKITE_AGENT_NAME format is {hostname}-{agent_idx}-{npu_card_num}cards, and agent_idx starts from 1.
# e.g. atlas-a2-001-1-2cards means this is the 1-th agent on atlas-a2-001 host, and it has 2 NPU cards. # e.g. atlas-a2-001-1-2cards means this is the 1-th agent on atlas-a2-001 host, and it has 2 NPU cards.
# returns --device /dev/davinci0 --device /dev/davinci1 # returns one argument per line: --device, /dev/davinciX, ...
parse_and_gen_devices() { parse_and_gen_devices() {
local input="$1" local input="$1"
local index cards_num local index cards_num
@@ -151,29 +151,24 @@ parse_and_gen_devices() {
return 1 return 1
fi fi
local devices=""
local i=0 local i=0
while (( i < cards_num )); do while (( i < cards_num )); do
local dev_idx=$(((index - 1)*cards_num + i )) local dev_idx=$(((index - 1)*cards_num + i ))
devices="$devices --device /dev/davinci${dev_idx}" printf '%s\n' "--device"
printf '%s\n' "/dev/davinci${dev_idx}"
((i++)) ((i++))
done done
# trim leading space
devices="${devices#"${devices%%[![:space:]]*}"}"
# Output devices: assigned to the caller variable
printf '%s' "$devices"
} }
devices=$(parse_and_gen_devices "${BUILDKITE_AGENT_NAME}") || exit 1 mapfile -t device_args < <(parse_and_gen_devices "${BUILDKITE_AGENT_NAME}") || exit 1
# Run the image and execute the Out-Of-Tree (OOT) platform interface test case on Ascend NPU hardware. # Run the image and execute the Out-Of-Tree (OOT) platform interface test case on Ascend NPU hardware.
# This test checks whether the OOT platform interface is functioning properly in conjunction with # This test checks whether the OOT platform interface is functioning properly in conjunction with
# the hardware plugin vllm-ascend. # the hardware plugin vllm-ascend.
model_cache_dir=/mnt/modelscope${agent_idx} model_cache_dir=/mnt/modelscope${agent_idx}
mkdir -p ${model_cache_dir} mkdir -p "${model_cache_dir}"
docker run \ docker run \
${devices} \ "${device_args[@]}" \
--device /dev/davinci_manager \ --device /dev/davinci_manager \
--device /dev/devmm_svm \ --device /dev/devmm_svm \
--device /dev/hisi_hdc \ --device /dev/hisi_hdc \
@@ -182,7 +177,7 @@ docker run \
-v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \ -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
-v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \ -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
-v /etc/ascend_install.info:/etc/ascend_install.info \ -v /etc/ascend_install.info:/etc/ascend_install.info \
-v ${model_cache_dir}:/root/.cache/modelscope \ -v "${model_cache_dir}":/root/.cache/modelscope \
--entrypoint="" \ --entrypoint="" \
--name "${container_name}" \ --name "${container_name}" \
"${image_name}" \ "${image_name}" \
@@ -61,7 +61,7 @@ echo "Results will be stored in: $RESULTS_DIR"
echo "--- Installing Python dependencies ---" echo "--- Installing Python dependencies ---"
python3 -m pip install --progress-bar off git+https://github.com/thuml/depyf.git \ python3 -m pip install --progress-bar off git+https://github.com/thuml/depyf.git \
&& python3 -m pip install --progress-bar off pytest pytest-asyncio tpu-info \ && python3 -m pip install --progress-bar off pytest pytest-asyncio tpu-info \
&& python3 -m pip install --progress-bar off "lm-eval[api]>=0.4.9.2" \ && python3 -m pip install --progress-bar off "lm-eval[api]>=0.4.11" \
&& python3 -m pip install --progress-bar off hf-transfer tblib==3.1.0 && python3 -m pip install --progress-bar off hf-transfer tblib==3.1.0
echo "--- Python dependencies installed ---" echo "--- Python dependencies installed ---"
@@ -61,7 +61,7 @@ echo "Results will be stored in: $RESULTS_DIR"
echo "--- Installing Python dependencies ---" echo "--- Installing Python dependencies ---"
python3 -m pip install --progress-bar off git+https://github.com/thuml/depyf.git \ python3 -m pip install --progress-bar off git+https://github.com/thuml/depyf.git \
&& python3 -m pip install --progress-bar off pytest pytest-asyncio tpu-info \ && python3 -m pip install --progress-bar off pytest pytest-asyncio tpu-info \
&& python3 -m pip install --progress-bar off "lm-eval[api]>=0.4.9.2" \ && python3 -m pip install --progress-bar off "lm-eval[api]>=0.4.11" \
&& python3 -m pip install --progress-bar off hf-transfer tblib==3.1.0 && python3 -m pip install --progress-bar off hf-transfer tblib==3.1.0
echo "--- Python dependencies installed ---" echo "--- Python dependencies installed ---"
@@ -8,7 +8,7 @@ image_name="xpu/vllm-ci:${BUILDKITE_COMMIT}"
container_name="xpu_${BUILDKITE_COMMIT}_$(tr -dc A-Za-z0-9 < /dev/urandom | head -c 10; echo)" container_name="xpu_${BUILDKITE_COMMIT}_$(tr -dc A-Za-z0-9 < /dev/urandom | head -c 10; echo)"
# Try building the docker image # Try building the docker image
docker build -t ${image_name} -f docker/Dockerfile.xpu . docker build -t "${image_name}" -f docker/Dockerfile.xpu .
# Setup cleanup # Setup cleanup
remove_docker_container() { remove_docker_container() {
+10 -10
View File
@@ -21,16 +21,16 @@ echo "Pushing original tag $ORIG_TAG_NAME$ORIG_TAG_SUFFIX to new nightly tag nam
# pull original arch-dependent images from AWS ECR Public # pull original arch-dependent images from AWS ECR Public
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7 aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/q9t5s3a7
docker pull public.ecr.aws/q9t5s3a7/vllm-release-repo:$ORIG_TAG_NAME-x86_64$ORIG_TAG_SUFFIX docker pull public.ecr.aws/q9t5s3a7/vllm-release-repo:"$ORIG_TAG_NAME"-x86_64"$ORIG_TAG_SUFFIX"
docker pull public.ecr.aws/q9t5s3a7/vllm-release-repo:$ORIG_TAG_NAME-aarch64$ORIG_TAG_SUFFIX docker pull public.ecr.aws/q9t5s3a7/vllm-release-repo:"$ORIG_TAG_NAME"-aarch64"$ORIG_TAG_SUFFIX"
# tag arch-dependent images # tag arch-dependent images
docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$ORIG_TAG_NAME-x86_64$ORIG_TAG_SUFFIX vllm/vllm-openai:$TAG_NAME-x86_64 docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:"$ORIG_TAG_NAME"-x86_64"$ORIG_TAG_SUFFIX" vllm/vllm-openai:"$TAG_NAME"-x86_64
docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:$ORIG_TAG_NAME-aarch64$ORIG_TAG_SUFFIX vllm/vllm-openai:$TAG_NAME-aarch64 docker tag public.ecr.aws/q9t5s3a7/vllm-release-repo:"$ORIG_TAG_NAME"-aarch64"$ORIG_TAG_SUFFIX" vllm/vllm-openai:"$TAG_NAME"-aarch64
# push arch-dependent images to DockerHub # push arch-dependent images to DockerHub
docker push vllm/vllm-openai:$TAG_NAME-x86_64 docker push vllm/vllm-openai:"$TAG_NAME"-x86_64
docker push vllm/vllm-openai:$TAG_NAME-aarch64 docker push vllm/vllm-openai:"$TAG_NAME"-aarch64
# push arch-independent manifest to DockerHub # push arch-independent manifest to DockerHub
docker manifest create vllm/vllm-openai:$TAG_NAME vllm/vllm-openai:$TAG_NAME-x86_64 vllm/vllm-openai:$TAG_NAME-aarch64 --amend docker manifest create vllm/vllm-openai:"$TAG_NAME" vllm/vllm-openai:"$TAG_NAME"-x86_64 vllm/vllm-openai:"$TAG_NAME"-aarch64 --amend
docker manifest create vllm/vllm-openai:$TAG_NAME-$BUILDKITE_COMMIT vllm/vllm-openai:$TAG_NAME-x86_64 vllm/vllm-openai:$TAG_NAME-aarch64 --amend docker manifest create vllm/vllm-openai:"$TAG_NAME"-"$BUILDKITE_COMMIT" vllm/vllm-openai:"$TAG_NAME"-x86_64 vllm/vllm-openai:"$TAG_NAME"-aarch64 --amend
docker manifest push vllm/vllm-openai:$TAG_NAME docker manifest push vllm/vllm-openai:"$TAG_NAME"
docker manifest push vllm/vllm-openai:$TAG_NAME-$BUILDKITE_COMMIT docker manifest push vllm/vllm-openai:"$TAG_NAME"-"$BUILDKITE_COMMIT"
+1 -1
View File
@@ -67,7 +67,7 @@ start_nodes() {
# 3. map the huggingface cache directory to the container # 3. map the huggingface cache directory to the container
# 3. assign ip addresses to the containers (head node: 192.168.10.10, worker nodes: # 3. assign ip addresses to the containers (head node: 192.168.10.10, worker nodes:
# starting from 192.168.10.11) # starting from 192.168.10.11)
docker run -d $GPU_DEVICES --shm-size=10.24gb -e HF_TOKEN \ docker run -d "$GPU_DEVICES" --shm-size=10.24gb -e HF_TOKEN \
-v ~/.cache/huggingface:/root/.cache/huggingface --name "node$node" \ -v ~/.cache/huggingface:/root/.cache/huggingface --name "node$node" \
--network docker-net --ip 192.168.10.$((10 + $node)) --rm "$DOCKER_IMAGE" \ --network docker-net --ip 192.168.10.$((10 + $node)) --rm "$DOCKER_IMAGE" \
/bin/bash -c "tail -f /dev/null" /bin/bash -c "tail -f /dev/null"
+1 -1
View File
@@ -29,7 +29,7 @@ fi
if ! command -v uv &> /dev/null; then if ! command -v uv &> /dev/null; then
echo "Installing UV package manager..." echo "Installing UV package manager..."
curl -LsSf https://astral.sh/uv/install.sh | sh curl -LsSf https://astral.sh/uv/install.sh | sh
source $HOME/.local/bin/env source "$HOME"/.local/bin/env
fi fi
# Clone Prime-RL repository at specific branch for reproducible tests # Clone Prime-RL repository at specific branch for reproducible tests
@@ -51,14 +51,14 @@ for BACK in "${BACKENDS[@]}"; do
--enable-eplb \ --enable-eplb \
--trust-remote-code \ --trust-remote-code \
--max-model-len 2048 \ --max-model-len 2048 \
--all2all-backend $BACK \ --all2all-backend "$BACK" \
--port $PORT & --port "$PORT" &
SERVER_PID=$! SERVER_PID=$!
wait_for_server $PORT wait_for_server "$PORT"
TAG=$(echo "$MODEL" | tr '/: \\n' '_____') TAG=$(echo "$MODEL" | tr '/: \\n' '_____')
OUT="${OUT_DIR}/${TAG}_${BACK}.json" OUT="${OUT_DIR}/${TAG}_${BACK}.json"
python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port $PORT --num-questions ${NUM_Q} --save-results ${OUT} python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port "$PORT" --num-questions "${NUM_Q}" --save-results "${OUT}"
python3 - <<PY python3 - <<PY
import json; acc=json.load(open('${OUT}'))['accuracy'] import json; acc=json.load(open('${OUT}'))['accuracy']
print(f"${MODEL} ${BACK}: accuracy {acc:.3f}") print(f"${MODEL} ${BACK}: accuracy {acc:.3f}")
@@ -47,20 +47,20 @@ for BACK in "${BACKENDS[@]}"; do
vllm serve "$MODEL" \ vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--enable-eplb \ --enable-eplb \
--all2all-backend $BACK \ --all2all-backend "$BACK" \
--eplb-config '{"window_size":10, "step_interval":100, "num_redundant_experts":0, "log_balancedness":true}' \ --eplb-config '{"window_size":10, "step_interval":100, "num_redundant_experts":0, "log_balancedness":true}' \
--tensor-parallel-size ${TENSOR_PARALLEL_SIZE} \ --tensor-parallel-size "${TENSOR_PARALLEL_SIZE}" \
--data-parallel-size ${DATA_PARALLEL_SIZE} \ --data-parallel-size "${DATA_PARALLEL_SIZE}" \
--enable-expert-parallel \ --enable-expert-parallel \
--trust-remote-code \ --trust-remote-code \
--max-model-len 2048 \ --max-model-len 2048 \
--port $PORT & --port "$PORT" &
SERVER_PID=$! SERVER_PID=$!
wait_for_server $PORT wait_for_server "$PORT"
TAG=$(echo "$MODEL" | tr '/: \\n' '_____') TAG=$(echo "$MODEL" | tr '/: \\n' '_____')
OUT="${OUT_DIR}/${TAG}_${BACK}.json" OUT="${OUT_DIR}/${TAG}_${BACK}.json"
python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port $PORT --num-questions ${NUM_Q} --save-results ${OUT} python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port "$PORT" --num-questions "${NUM_Q}" --save-results "${OUT}"
python3 - <<PY python3 - <<PY
import json; acc=json.load(open('${OUT}'))['accuracy'] import json; acc=json.load(open('${OUT}'))['accuracy']
print(f"${MODEL} ${BACK}: accuracy {acc:.3f}") print(f"${MODEL} ${BACK}: accuracy {acc:.3f}")
@@ -51,20 +51,20 @@ for BACK in "${BACKENDS[@]}"; do
--tensor-parallel-size 4 \ --tensor-parallel-size 4 \
--enable-expert-parallel \ --enable-expert-parallel \
--enable-eplb \ --enable-eplb \
--all2all-backend $BACK \ --all2all-backend "$BACK" \
--eplb-config '{"window_size":200,"step_interval":600,"use_async":true}' \ --eplb-config '{"window_size":200,"step_interval":600,"use_async":true}' \
--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":1}' \ --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":1}' \
--trust-remote-code \ --trust-remote-code \
--max-model-len 2048 \ --max-model-len 2048 \
--gpu-memory-utilization 0.9 \ --gpu-memory-utilization 0.9 \
"${PLATFORM_ARGS[@]}" \ "${PLATFORM_ARGS[@]}" \
--port $PORT & --port "$PORT" &
SERVER_PID=$! SERVER_PID=$!
wait_for_server $PORT wait_for_server "$PORT"
TAG=$(echo "$MODEL" | tr '/: \\n' '_____') TAG=$(echo "$MODEL" | tr '/: \\n' '_____')
OUT="${OUT_DIR}/${TAG}_${BACK}.json" OUT="${OUT_DIR}/${TAG}_${BACK}.json"
python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port $PORT --num-questions ${NUM_Q} --save-results ${OUT} python3 tests/evals/gsm8k/gsm8k_eval.py --host http://127.0.0.1 --port "$PORT" --num-questions "${NUM_Q}" --save-results "${OUT}"
python3 - <<PY python3 - <<PY
import json; acc=json.load(open('${OUT}'))['accuracy'] import json; acc=json.load(open('${OUT}'))['accuracy']
print(f"${MODEL} ${BACK}: accuracy {acc:.3f}") print(f"${MODEL} ${BACK}: accuracy {acc:.3f}")
+8 -7
View File
@@ -9,10 +9,11 @@ ENV_FILE=$1
# For testing on local vm, use `set -a` to export all variables # For testing on local vm, use `set -a` to export all variables
source /etc/environment source /etc/environment
source $ENV_FILE # shellcheck source=/dev/null
source "$ENV_FILE"
remove_docker_container() { remove_docker_container() {
docker rm -f $CONTAINER_NAME || true; docker rm -f "$CONTAINER_NAME" || true;
} }
trap remove_docker_container EXIT trap remove_docker_container EXIT
@@ -41,13 +42,13 @@ echo
echo "starting docker...$CONTAINER_NAME" echo "starting docker...$CONTAINER_NAME"
echo echo
docker run \ docker run \
-v $DOWNLOAD_DIR:$DOWNLOAD_DIR \ -v "$DOWNLOAD_DIR":"$DOWNLOAD_DIR" \
--env-file $ENV_FILE \ --env-file "$ENV_FILE" \
-e HF_TOKEN="$HF_TOKEN" \ -e HF_TOKEN="$HF_TOKEN" \
-e TARGET_COMMIT=$BUILDKITE_COMMIT \ -e TARGET_COMMIT="$BUILDKITE_COMMIT" \
-e MODEL=$MODEL \ -e MODEL="$MODEL" \
-e WORKSPACE=/workspace \ -e WORKSPACE=/workspace \
--name $CONTAINER_NAME \ --name "$CONTAINER_NAME" \
-d \ -d \
--privileged \ --privileged \
--network host \ --network host \
+10 -10
View File
@@ -42,21 +42,21 @@ echo "lanching vllm..."
echo "logging to $VLLM_LOG" echo "logging to $VLLM_LOG"
echo echo
vllm serve $MODEL \ vllm serve "$MODEL" \
--seed 42 \ --seed 42 \
--max-num-seqs $MAX_NUM_SEQS \ --max-num-seqs "$MAX_NUM_SEQS" \
--max-num-batched-tokens $MAX_NUM_BATCHED_TOKENS \ --max-num-batched-tokens "$MAX_NUM_BATCHED_TOKENS" \
--tensor-parallel-size $TENSOR_PARALLEL_SIZE \ --tensor-parallel-size "$TENSOR_PARALLEL_SIZE" \
--no-enable-prefix-caching \ --no-enable-prefix-caching \
--download_dir $DOWNLOAD_DIR \ --download_dir "$DOWNLOAD_DIR" \
--max-model-len $MAX_MODEL_LEN > "$VLLM_LOG" 2>&1 & --max-model-len "$MAX_MODEL_LEN" > "$VLLM_LOG" 2>&1 &
echo "wait for 20 minutes.." echo "wait for 20 minutes.."
echo echo
# sleep 1200 # sleep 1200
# wait for 10 minutes... # wait for 10 minutes...
for i in {1..120}; do for _ in {1..120}; do
# TODO: detect other type of errors. # TODO: detect other type of errors.
if grep -Fq "raise RuntimeError" "$VLLM_LOG"; then if grep -Fq "raise RuntimeError" "$VLLM_LOG"; then
echo "Detected RuntimeError, exiting." echo "Detected RuntimeError, exiting."
@@ -78,11 +78,11 @@ echo "logging to $BM_LOG"
echo echo
vllm bench serve \ vllm bench serve \
--backend vllm \ --backend vllm \
--model $MODEL \ --model "$MODEL" \
--dataset-name sonnet \ --dataset-name sonnet \
--dataset-path benchmarks/sonnet_4x.txt \ --dataset-path benchmarks/sonnet_4x.txt \
--sonnet-input-len $INPUT_LEN \ --sonnet-input-len "$INPUT_LEN" \
--sonnet-output-len $OUTPUT_LEN \ --sonnet-output-len "$OUTPUT_LEN" \
--ignore-eos > "$BM_LOG" --ignore-eos > "$BM_LOG"
echo "completed..." echo "completed..."
+6 -7
View File
@@ -76,16 +76,15 @@ mkdir -p "$INDICES_OUTPUT_DIR"
# this indices have relative paths that could work as long as it is next to the wheel directory in s3 # this indices have relative paths that could work as long as it is next to the wheel directory in s3
# i.e., the wheels are always in s3://vllm-wheels/<commit>/ # i.e., the wheels are always in s3://vllm-wheels/<commit>/
# and indices can be placed in /<commit>/, or /nightly/, or /<version>/ # and indices can be placed in /<commit>/, or /nightly/, or /<version>/
if [[ ! -z "$DEFAULT_VARIANT_ALIAS" ]]; then alias_args=()
alias_arg="--alias-to-default $DEFAULT_VARIANT_ALIAS" if [[ -n "$DEFAULT_VARIANT_ALIAS" ]]; then
else alias_args=(--alias-to-default "$DEFAULT_VARIANT_ALIAS")
alias_arg=""
fi fi
# HACK: we do not need regex module here, but it is required by pre-commit hook # HACK: we do not need regex module here, but it is required by pre-commit hook
# To avoid any external dependency, we simply replace it back to the stdlib re module # To avoid any external dependency, we simply replace it back to the stdlib re module
sed -i 's/import regex as re/import re/g' .buildkite/scripts/generate-nightly-index.py sed -i 's/import regex as re/import re/g' .buildkite/scripts/generate-nightly-index.py
$PYTHON .buildkite/scripts/generate-nightly-index.py --version "$SUBPATH" --current-objects "$obj_json" --output-dir "$INDICES_OUTPUT_DIR" --comment "commit $BUILDKITE_COMMIT" $alias_arg $PYTHON .buildkite/scripts/generate-nightly-index.py --version "$SUBPATH" --current-objects "$obj_json" --output-dir "$INDICES_OUTPUT_DIR" --comment "commit $BUILDKITE_COMMIT" "${alias_args[@]}"
# copy indices to /<commit>/ unconditionally # copy indices to /<commit>/ unconditionally
echo "Uploading indices to $S3_COMMIT_PREFIX" echo "Uploading indices to $S3_COMMIT_PREFIX"
@@ -100,9 +99,9 @@ fi
# re-generate and copy to /<pure_version>/ only if it does not have "dev" in the version # re-generate and copy to /<pure_version>/ only if it does not have "dev" in the version
if [[ "$version" != *"dev"* ]]; then if [[ "$version" != *"dev"* ]]; then
echo "Re-generating indices for /$pure_version/" echo "Re-generating indices for /$pure_version/"
rm -rf "$INDICES_OUTPUT_DIR/*" rm -rf "${INDICES_OUTPUT_DIR:?}/*"
mkdir -p "$INDICES_OUTPUT_DIR" mkdir -p "$INDICES_OUTPUT_DIR"
# wheel-dir is overridden to be the commit directory, so that the indices point to the correct wheel path # wheel-dir is overridden to be the commit directory, so that the indices point to the correct wheel path
$PYTHON .buildkite/scripts/generate-nightly-index.py --version "$pure_version" --wheel-dir "$SUBPATH" --current-objects "$obj_json" --output-dir "$INDICES_OUTPUT_DIR" --comment "version $pure_version" $alias_arg $PYTHON .buildkite/scripts/generate-nightly-index.py --version "$pure_version" --wheel-dir "$SUBPATH" --current-objects "$obj_json" --output-dir "$INDICES_OUTPUT_DIR" --comment "version $pure_version" "${alias_args[@]}"
aws s3 cp --recursive "$INDICES_OUTPUT_DIR/" "s3://$BUCKET/$pure_version/" aws s3 cp --recursive "$INDICES_OUTPUT_DIR/" "s3://$BUCKET/$pure_version/"
fi fi
@@ -7,7 +7,7 @@ SUBPATH=$BUILDKITE_COMMIT
S3_COMMIT_PREFIX="s3://$BUCKET/$SUBPATH/" S3_COMMIT_PREFIX="s3://$BUCKET/$SUBPATH/"
RELEASE_VERSION=$(buildkite-agent meta-data get release-version) RELEASE_VERSION=$(buildkite-agent meta-data get release-version)
GIT_VERSION=$(git describe --exact-match --tags $BUILDKITE_COMMIT 2>/dev/null) GIT_VERSION=$(git describe --exact-match --tags "$BUILDKITE_COMMIT" 2>/dev/null)
echo "Release version from Buildkite: $RELEASE_VERSION" echo "Release version from Buildkite: $RELEASE_VERSION"
@@ -55,7 +55,7 @@ mkdir -p $DIST_DIR
aws s3 cp --recursive --exclude "*" --include "vllm-${PURE_VERSION}*.whl" --exclude "*dev*" --exclude "*rc[0-9]*" "$S3_COMMIT_PREFIX" $DIST_DIR aws s3 cp --recursive --exclude "*" --include "vllm-${PURE_VERSION}*.whl" --exclude "*dev*" --exclude "*rc[0-9]*" "$S3_COMMIT_PREFIX" $DIST_DIR
echo "Wheels copied to local directory" echo "Wheels copied to local directory"
# generate source tarball # generate source tarball
git archive --format=tar.gz --output="$DIST_DIR/vllm-${PURE_VERSION}.tar.gz" $BUILDKITE_COMMIT git archive --format=tar.gz --output="$DIST_DIR/vllm-${PURE_VERSION}.tar.gz" "$BUILDKITE_COMMIT"
ls -la $DIST_DIR ls -la $DIST_DIR
# upload wheels to PyPI (only default variant, i.e. files without '+' in the name) # upload wheels to PyPI (only default variant, i.e. files without '+' in the name)
@@ -65,6 +65,6 @@ if [[ -z "$PYPI_WHEEL_FILES" ]]; then
exit 1 exit 1
fi fi
python3 -m twine check $PYPI_WHEEL_FILES python3 -m twine check "$PYPI_WHEEL_FILES"
python3 -m twine upload --non-interactive --verbose $PYPI_WHEEL_FILES python3 -m twine upload --non-interactive --verbose "$PYPI_WHEEL_FILES"
echo "Wheels uploaded to PyPI" echo "Wheels uploaded to PyPI"
+2 -2
View File
@@ -55,7 +55,7 @@ mkdir -p all-rocm-wheels
cp artifacts/rocm-base-wheels/*.whl all-rocm-wheels/ 2>/dev/null || true cp artifacts/rocm-base-wheels/*.whl all-rocm-wheels/ 2>/dev/null || true
cp artifacts/rocm-vllm-wheel/*.whl all-rocm-wheels/ 2>/dev/null || true cp artifacts/rocm-vllm-wheel/*.whl all-rocm-wheels/ 2>/dev/null || true
WHEEL_COUNT=$(ls all-rocm-wheels/*.whl 2>/dev/null | wc -l) WHEEL_COUNT=$(find all-rocm-wheels -maxdepth 1 -name '*.whl' 2>/dev/null | wc -l)
echo "Total wheels to upload: $WHEEL_COUNT" echo "Total wheels to upload: $WHEEL_COUNT"
if [ "$WHEEL_COUNT" -eq 0 ]; then if [ "$WHEEL_COUNT" -eq 0 ]; then
@@ -115,7 +115,7 @@ if [[ "$BUILDKITE_BRANCH" == "main" && "$BUILDKITE_PULL_REQUEST" == "false" ]] |
fi fi
# Extract version from vLLM wheel and update version-specific index # Extract version from vLLM wheel and update version-specific index
VLLM_WHEEL=$(ls all-rocm-wheels/vllm*.whl 2>/dev/null | head -1) VLLM_WHEEL=$(find all-rocm-wheels -maxdepth 1 -name 'vllm*.whl' 2>/dev/null | head -1)
if [ -n "$VLLM_WHEEL" ]; then if [ -n "$VLLM_WHEEL" ]; then
VERSION=$(unzip -p "$VLLM_WHEEL" '**/METADATA' | grep '^Version: ' | cut -d' ' -f2) VERSION=$(unzip -p "$VLLM_WHEEL" '**/METADATA' | grep '^Version: ' | cut -d' ' -f2)
echo "Version in wheel: $VERSION" echo "Version in wheel: $VERSION"
File diff suppressed because it is too large Load Diff
+11
View File
@@ -197,6 +197,17 @@ steps:
- uv pip install --system -r /vllm-workspace/requirements/kv_connectors.txt - uv pip install --system -r /vllm-workspace/requirements/kv_connectors.txt
- DP_EP=1 bash v1/kv_connector/nixl_integration/config_sweep_accuracy_test.sh - DP_EP=1 bash v1/kv_connector/nixl_integration/config_sweep_accuracy_test.sh
- label: CrossLayer KV layout Distributed NixlConnector PD accuracy tests (4 GPUs)
timeout_in_minutes: 30
working_dir: "/vllm-workspace/tests"
num_devices: 4
source_file_dependencies:
- vllm/distributed/kv_transfer/kv_connector/v1/nixl_connector.py
- tests/v1/kv_connector/nixl_integration/
commands:
- uv pip install --system -r /vllm-workspace/requirements/kv_connectors.txt
- CROSS_LAYERS_BLOCKS=True bash v1/kv_connector/nixl_integration/config_sweep_accuracy_test.sh
- label: Pipeline + Context Parallelism (4 GPUs)) - label: Pipeline + Context Parallelism (4 GPUs))
timeout_in_minutes: 60 timeout_in_minutes: 60
working_dir: "/vllm-workspace/tests" working_dir: "/vllm-workspace/tests"
+2
View File
@@ -123,6 +123,7 @@ steps:
- tests/test_inputs.py - tests/test_inputs.py
- tests/test_outputs.py - tests/test_outputs.py
- tests/test_pooling_params.py - tests/test_pooling_params.py
- tests/test_ray_env.py
- tests/multimodal - tests/multimodal
- tests/renderers - tests/renderers
- tests/standalone_tests/lazy_imports.py - tests/standalone_tests/lazy_imports.py
@@ -136,6 +137,7 @@ steps:
- pytest -v -s test_inputs.py - pytest -v -s test_inputs.py
- pytest -v -s test_outputs.py - pytest -v -s test_outputs.py
- pytest -v -s test_pooling_params.py - pytest -v -s test_pooling_params.py
- pytest -v -s test_ray_env.py
- pytest -v -s -m 'cpu_test' multimodal - pytest -v -s -m 'cpu_test' multimodal
- pytest -v -s renderers - pytest -v -s renderers
- pytest -v -s tokenizers_ - pytest -v -s tokenizers_
+1 -1
View File
@@ -38,7 +38,7 @@ repos:
rev: 0.9.1 rev: 0.9.1
hooks: hooks:
- id: pip-compile - id: pip-compile
args: [requirements/test.in, -o, requirements/test.txt, --index-strategy, unsafe-best-match, --torch-backend, cu129, --python-platform, x86_64-manylinux_2_28, --python-version, "3.12"] args: [requirements/test.in, -o, requirements/test.txt, --index-strategy, unsafe-best-match, --torch-backend, cu128, --python-platform, x86_64-manylinux_2_28, --python-version, "3.12"]
files: ^requirements/test\.(in|txt)$ files: ^requirements/test\.(in|txt)$
- repo: local - repo: local
hooks: hooks:
+22 -22
View File
@@ -46,10 +46,10 @@ echo "VLLM_LOGGING_LEVEL=$VLLM_LOGGING_LEVEL"
echo "RESULT_FILE=$RESULT" echo "RESULT_FILE=$RESULT"
echo "====================== AUTO TUNEPARAMETERS ====================" echo "====================== AUTO TUNEPARAMETERS ===================="
rm -rf $LOG_FOLDER rm -rf "$LOG_FOLDER"
rm -rf $PROFILE_PATH rm -rf "$PROFILE_PATH"
mkdir -p $LOG_FOLDER mkdir -p "$LOG_FOLDER"
mkdir -p $PROFILE_PATH mkdir -p "$PROFILE_PATH"
cd "$BASE/vllm" cd "$BASE/vllm"
@@ -114,7 +114,7 @@ start_server() {
# wait for 10 minutes... # wait for 10 minutes...
server_started=0 server_started=0
for i in {1..60}; do for _ in {1..60}; do
# This line checks whether the server is still alive or not, # This line checks whether the server is still alive or not,
# since that we should always have permission to send signal to the server process. # since that we should always have permission to send signal to the server process.
kill -0 $server_pid 2> /dev/null || break kill -0 $server_pid 2> /dev/null || break
@@ -145,12 +145,12 @@ run_benchmark() {
local vllm_log="$LOG_FOLDER/vllm_log_${max_num_seqs}_${max_num_batched_tokens}.txt" local vllm_log="$LOG_FOLDER/vllm_log_${max_num_seqs}_${max_num_batched_tokens}.txt"
echo "vllm_log: $vllm_log" echo "vllm_log: $vllm_log"
echo echo
rm -f $vllm_log rm -f "$vllm_log"
pkill -if "vllm serve" || true pkill -if "vllm serve" || true
echo "starting server..." echo "starting server..."
# Call start_server without a profile_dir to avoid profiling overhead # Call start_server without a profile_dir to avoid profiling overhead
start_server $gpu_memory_utilization $max_num_seqs $max_num_batched_tokens $vllm_log "" start_server "$gpu_memory_utilization" "$max_num_seqs" "$max_num_batched_tokens" "$vllm_log" ""
result=$? result=$?
if [[ "$result" -eq 1 ]]; then if [[ "$result" -eq 1 ]]; then
echo "server failed to start. gpu_memory_utilization:$gpu_memory_utilization, max_num_seqs:$max_num_seqs, max_num_batched_tokens: $max_num_batched_tokens" echo "server failed to start. gpu_memory_utilization:$gpu_memory_utilization, max_num_seqs:$max_num_seqs, max_num_batched_tokens: $max_num_batched_tokens"
@@ -168,15 +168,15 @@ run_benchmark() {
# --profile flag is removed from this call # --profile flag is removed from this call
vllm bench serve \ vllm bench serve \
--backend vllm \ --backend vllm \
--model $MODEL \ --model "$MODEL" \
--dataset-name random \ --dataset-name random \
--random-input-len $adjusted_input_len \ --random-input-len $adjusted_input_len \
--random-output-len $OUTPUT_LEN \ --random-output-len "$OUTPUT_LEN" \
--ignore-eos \ --ignore-eos \
--disable-tqdm \ --disable-tqdm \
--request-rate inf \ --request-rate inf \
--percentile-metrics ttft,tpot,itl,e2el \ --percentile-metrics ttft,tpot,itl,e2el \
--goodput e2el:$MAX_LATENCY_ALLOWED_MS \ --goodput e2el:"$MAX_LATENCY_ALLOWED_MS" \
--num-prompts 1000 \ --num-prompts 1000 \
--random-prefix-len $prefix_len \ --random-prefix-len $prefix_len \
--host "$HOSTNAME" \ --host "$HOSTNAME" \
@@ -195,20 +195,20 @@ run_benchmark() {
request_rate=$((${throughput%.*} + 1)) request_rate=$((${throughput%.*} + 1))
while ((request_rate > 0)); do while ((request_rate > 0)); do
# clear prefix cache # clear prefix cache
curl -X POST http://${HOSTNAME}:8004/reset_prefix_cache curl -X POST http://"${HOSTNAME}":8004/reset_prefix_cache
sleep 5 sleep 5
bm_log="$LOG_FOLDER/bm_log_${max_num_seqs}_${max_num_batched_tokens}_requestrate_${request_rate}.txt" bm_log="$LOG_FOLDER/bm_log_${max_num_seqs}_${max_num_batched_tokens}_requestrate_${request_rate}.txt"
vllm bench serve \ vllm bench serve \
--backend vllm \ --backend vllm \
--model $MODEL \ --model "$MODEL" \
--dataset-name random \ --dataset-name random \
--random-input-len $adjusted_input_len \ --random-input-len $adjusted_input_len \
--random-output-len $OUTPUT_LEN \ --random-output-len "$OUTPUT_LEN" \
--ignore-eos \ --ignore-eos \
--disable-tqdm \ --disable-tqdm \
--request-rate $request_rate \ --request-rate $request_rate \
--percentile-metrics ttft,tpot,itl,e2el \ --percentile-metrics ttft,tpot,itl,e2el \
--goodput e2el:$MAX_LATENCY_ALLOWED_MS \ --goodput e2el:"$MAX_LATENCY_ALLOWED_MS" \
--num-prompts 100 \ --num-prompts 100 \
--random-prefix-len $prefix_len \ --random-prefix-len $prefix_len \
--host "$HOSTNAME" \ --host "$HOSTNAME" \
@@ -255,7 +255,7 @@ gpu_memory_utilization=0.98
find_gpu_memory_utilization=0 find_gpu_memory_utilization=0
while (( $(echo "$gpu_memory_utilization >= 0.9" | bc -l) )); do while (( $(echo "$gpu_memory_utilization >= 0.9" | bc -l) )); do
# Pass empty string for profile_dir argument # Pass empty string for profile_dir argument
start_server $gpu_memory_utilization "${num_seqs_list[-1]}" "${num_batched_tokens_list[-1]}" "$LOG_FOLDER/vllm_log_gpu_memory_utilization_$gpu_memory_utilization.log" "" start_server "$gpu_memory_utilization" "${num_seqs_list[-1]}" "${num_batched_tokens_list[-1]}" "$LOG_FOLDER/vllm_log_gpu_memory_utilization_$gpu_memory_utilization.log" ""
result=$? result=$?
if [[ "$result" -eq 0 ]]; then if [[ "$result" -eq 0 ]]; then
find_gpu_memory_utilization=1 find_gpu_memory_utilization=1
@@ -274,7 +274,7 @@ fi
for num_seqs in "${num_seqs_list[@]}"; do for num_seqs in "${num_seqs_list[@]}"; do
for num_batched_tokens in "${num_batched_tokens_list[@]}"; do for num_batched_tokens in "${num_batched_tokens_list[@]}"; do
run_benchmark $num_seqs $num_batched_tokens $gpu_memory_utilization run_benchmark "$num_seqs" "$num_batched_tokens" "$gpu_memory_utilization"
done done
done done
echo "finish permutations" echo "finish permutations"
@@ -285,7 +285,7 @@ echo "finish permutations"
if (( $(echo "$best_throughput > 0" | bc -l) )); then if (( $(echo "$best_throughput > 0" | bc -l) )); then
echo echo
echo "Benchmark tuning finished. Now running profiling on the best configuration found..." echo "Benchmark tuning finished. Now running profiling on the best configuration found..."
echo "Best config: max_num_seqs: $best_max_num_seqs, max_num_batched_tokens: $best_num_batched_tokens, throughput: $best_throughput" echo "Best config: max_num_seqs: $best_max_num_seqs, max_num_batched_tokens: $best_num_batched_tokens, throughput: $best_throughput, goodput: $best_goodput"
echo echo
vllm_log="$LOG_FOLDER/vllm_log_BEST_PROFILE.txt" vllm_log="$LOG_FOLDER/vllm_log_BEST_PROFILE.txt"
@@ -293,7 +293,7 @@ if (( $(echo "$best_throughput > 0" | bc -l) )); then
# Start server with the best params and profiling ENABLED # Start server with the best params and profiling ENABLED
echo "Starting server for profiling..." echo "Starting server for profiling..."
start_server $gpu_memory_utilization $best_max_num_seqs $best_num_batched_tokens "$vllm_log" "$PROFILE_PATH" start_server "$gpu_memory_utilization" "$best_max_num_seqs" "$best_num_batched_tokens" "$vllm_log" "$PROFILE_PATH"
# Run benchmark with the best params and the --profile flag # Run benchmark with the best params and the --profile flag
echo "Running benchmark with profiling..." echo "Running benchmark with profiling..."
@@ -301,15 +301,15 @@ if (( $(echo "$best_throughput > 0" | bc -l) )); then
adjusted_input_len=$(( INPUT_LEN - prefix_len )) adjusted_input_len=$(( INPUT_LEN - prefix_len ))
vllm bench serve \ vllm bench serve \
--backend vllm \ --backend vllm \
--model $MODEL \ --model "$MODEL" \
--dataset-name random \ --dataset-name random \
--random-input-len $adjusted_input_len \ --random-input-len $adjusted_input_len \
--random-output-len $OUTPUT_LEN \ --random-output-len "$OUTPUT_LEN" \
--ignore-eos \ --ignore-eos \
--disable-tqdm \ --disable-tqdm \
--request-rate $best_request_rate \ --request-rate "$best_request_rate" \
--percentile-metrics ttft,tpot,itl,e2el \ --percentile-metrics ttft,tpot,itl,e2el \
--goodput e2el:$MAX_LATENCY_ALLOWED_MS \ --goodput e2el:"$MAX_LATENCY_ALLOWED_MS" \
--num-prompts 100 \ --num-prompts 100 \
--random-prefix-len $prefix_len \ --random-prefix-len $prefix_len \
--host "$HOSTNAME" \ --host "$HOSTNAME" \
+1 -1
View File
@@ -64,7 +64,7 @@ for i in $(seq 0 $(($num_runs - 1))); do
else else
STATUS="FAILURE" STATUS="FAILURE"
((FAILURE_COUNT++)) ((FAILURE_COUNT++))
FAILED_RUNS+=("Run #$((i+1)): $(echo $run_object | jq -c .)") FAILED_RUNS+=("Run #$((i+1)): $(echo "$run_object" | jq -c .)")
fi fi
RUN_OUTPUT=$(<"$RUN_OUTPUT_FILE") RUN_OUTPUT=$(<"$RUN_OUTPUT_FILE")
+471
View File
@@ -0,0 +1,471 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""
Benchmark comparing Triton vs PyTorch sort-based top-k/top-p implementations.
Compares:
- apply_top_k_top_p_triton (Triton binary search)
- apply_top_k_top_p (PyTorch sort-based)
Scenarios:
- top_k only (whole batch, partial batch)
- top_p only (whole batch, partial batch)
- mix of top_k and top_p
"""
import argparse
import gc
from dataclasses import dataclass
import torch
from vllm.v1.sample.ops.topk_topp_sampler import apply_top_k_top_p_pytorch
from vllm.v1.sample.ops.topk_topp_triton import (
apply_top_k_top_p_triton,
reset_buffer_cache,
)
@dataclass
class BenchmarkConfig:
"""Configuration for a benchmark run."""
name: str
batch_size: int
vocab_size: int
# k and p can be tensors or None
k_values: torch.Tensor | None # [batch_size] or None
p_values: torch.Tensor | None # [batch_size] or None
description: str
ops_pct: float = 0.0 # Percentage of ops relative to batch size
def calculate_ops_pct(
k_values: torch.Tensor | None,
p_values: torch.Tensor | None,
vocab_size: int,
batch_size: int,
) -> float:
"""
Calculate the percentage of active top-k and top-p operations.
Returns percentage where 100% = batch_size ops.
E.g., if all rows have both top-k and top-p active, returns 200%.
"""
active_ops = 0
if k_values is not None:
# Count rows where k < vocab_size (active top-k filtering)
active_ops += (k_values < vocab_size).sum().item()
if p_values is not None:
# Count rows where p < 1.0 (active top-p filtering)
active_ops += (p_values < 1.0).sum().item()
return (active_ops / batch_size) * 100 if batch_size > 0 else 0.0
def create_logits(
batch_size: int, vocab_size: int, device: str = "cuda"
) -> torch.Tensor:
"""Create random logits mimicking a realistic LLM distribution.
Uses a Zipf-like probability distribution (rank^-1.1) converted to logits
via log, then randomly permuted per row. This produces a peaked distribution
where a small number of tokens capture most probability mass, similar to
real model outputs.
"""
# Create Zipf-like probabilities: p(rank) ~ rank^(-alpha)
ranks = torch.arange(1, vocab_size + 1, dtype=torch.float32, device=device)
probs = ranks.pow(-1.1)
probs = probs / probs.sum()
# Convert to logits (log-probabilities, unnormalized is fine)
base_logits = probs.log()
# Broadcast to batch and randomly permute each row
logits = base_logits.unsqueeze(0).expand(batch_size, -1).clone()
for i in range(batch_size):
logits[i] = logits[i, torch.randperm(vocab_size, device=device)]
return logits
def measure_memory() -> tuple[int, int]:
"""Return (allocated, reserved) memory in bytes."""
torch.cuda.synchronize()
return torch.cuda.memory_allocated(), torch.cuda.max_memory_allocated()
def reset_memory_stats():
"""Reset peak memory statistics."""
reset_buffer_cache()
torch.cuda.reset_peak_memory_stats()
torch.cuda.empty_cache()
gc.collect()
def benchmark_function(
func,
logits: torch.Tensor,
k: torch.Tensor | None,
p: torch.Tensor | None,
warmup_iters: int = 5,
benchmark_iters: int = 20,
) -> tuple[float, int]:
"""
Benchmark a function and return (avg_time_ms, peak_memory_bytes).
Returns average time in milliseconds and peak memory usage.
"""
# Warmup
for _ in range(warmup_iters):
logits_copy = logits.clone()
func(logits_copy, k, p)
torch.cuda.synchronize()
# Reset memory stats before benchmark
reset_memory_stats()
# Benchmark
start_events = [
torch.cuda.Event(enable_timing=True) for _ in range(benchmark_iters)
]
end_events = [torch.cuda.Event(enable_timing=True) for _ in range(benchmark_iters)]
for i in range(benchmark_iters):
logits_copy = logits.clone()
start_events[i].record()
func(logits_copy, k, p)
end_events[i].record()
torch.cuda.synchronize()
# Calculate timing
times = [
start_events[i].elapsed_time(end_events[i]) for i in range(benchmark_iters)
]
avg_time = sum(times) / len(times)
# Get peak memory
_, peak_memory = measure_memory()
return avg_time, peak_memory
def create_benchmark_configs(
batch_sizes: list[int],
vocab_sizes: list[int],
device: str = "cuda",
) -> list[BenchmarkConfig]:
"""Create all benchmark configurations."""
configs = []
for vocab_size in vocab_sizes:
for batch_size in batch_sizes:
# 1. Top-k only - whole batch (all rows have k < vocab_size)
k_all = torch.full((batch_size,), 50, dtype=torch.int32, device=device)
configs.append(
BenchmarkConfig(
name=f"topk_whole_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=k_all,
p_values=None,
description=f"Top-k only (whole batch, k=50), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(k_all, None, vocab_size, batch_size),
)
)
# 2. Top-k only - partial batch (half have k=50, half have k=vocab_size)
k_partial = torch.full((batch_size,), 50, dtype=torch.int32, device=device)
k_partial[batch_size // 2 :] = vocab_size # No filtering for second half
configs.append(
BenchmarkConfig(
name=f"topk_partial_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=k_partial,
p_values=None,
description=f"Top-k only (partial batch, 50% k=50, 50% k=vocab), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(k_partial, None, vocab_size, batch_size),
)
)
# 3. Top-p only - whole batch (all rows have p < 1.0)
p_all = torch.full((batch_size,), 0.9, dtype=torch.float32, device=device)
configs.append(
BenchmarkConfig(
name=f"topp_whole_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=None,
p_values=p_all,
description=f"Top-p only (whole batch, p=0.9), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(None, p_all, vocab_size, batch_size),
)
)
# 4. Top-p only - partial batch (half have p=0.9, half have p=1.0)
p_partial = torch.full(
(batch_size,), 0.9, dtype=torch.float32, device=device
)
p_partial[batch_size // 2 :] = 1.0 # No filtering for second half
configs.append(
BenchmarkConfig(
name=f"topp_partial_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=None,
p_values=p_partial,
description=f"Top-p only (partial batch, 50% p=0.9, 50% p=1.0), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(None, p_partial, vocab_size, batch_size),
)
)
# 5. Mix of top-k and top-p (both applied to whole batch)
k_mix = torch.full((batch_size,), 100, dtype=torch.int32, device=device)
p_mix = torch.full((batch_size,), 0.9, dtype=torch.float32, device=device)
configs.append(
BenchmarkConfig(
name=f"topk_topp_whole_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=k_mix,
p_values=p_mix,
description=f"Top-k + Top-p (whole batch, k=100, p=0.9), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(k_mix, p_mix, vocab_size, batch_size),
)
)
# 6. Mix with partial application (some rows k only, some p only, some both)
k_mixed = torch.full(
(batch_size,), vocab_size, dtype=torch.int32, device=device
)
p_mixed = torch.full((batch_size,), 1.0, dtype=torch.float32, device=device)
# First third: k only
third = batch_size // 3
k_mixed[:third] = 50
# Second third: p only
p_mixed[third : 2 * third] = 0.5
# Last third: both k and p
k_mixed[2 * third :] = 100
p_mixed[2 * third :] = 0.9
configs.append(
BenchmarkConfig(
name=f"mixed_partial_b{batch_size}_v{vocab_size // 1000}k",
batch_size=batch_size,
vocab_size=vocab_size,
k_values=k_mixed,
p_values=p_mixed,
description=f"Mixed partial (1/3 k=50, 1/3 p=0.9, 1/3 both), "
f"batch={batch_size}, vocab={vocab_size}",
ops_pct=calculate_ops_pct(k_mixed, p_mixed, vocab_size, batch_size),
)
)
return configs
def format_memory(bytes_val: int) -> str:
"""Format memory in human-readable form."""
if bytes_val >= 1024**3:
return f"{bytes_val / (1024**3):.2f} GB"
elif bytes_val >= 1024**2:
return f"{bytes_val / (1024**2):.2f} MB"
elif bytes_val >= 1024:
return f"{bytes_val / 1024:.2f} KB"
return f"{bytes_val} B"
def run_benchmark(
configs: list[BenchmarkConfig],
warmup_iters: int = 5,
benchmark_iters: int = 20,
verbose: bool = True,
):
"""Run all benchmarks and print results."""
results = []
print("=" * 100)
print("Top-k/Top-p Benchmark: Triton vs PyTorch Sort-based")
print("=" * 100)
print()
for config in configs:
if verbose:
print(f"Running: {config.description}")
# Create fresh logits for this config
logits = create_logits(config.batch_size, config.vocab_size)
# Benchmark Triton
reset_memory_stats()
triton_time, triton_mem = benchmark_function(
apply_top_k_top_p_triton,
logits,
config.k_values,
config.p_values,
warmup_iters,
benchmark_iters,
)
# Benchmark PyTorch
reset_memory_stats()
pytorch_time, pytorch_mem = benchmark_function(
apply_top_k_top_p_pytorch,
logits,
config.k_values,
config.p_values,
warmup_iters,
benchmark_iters,
)
speedup = pytorch_time / triton_time if triton_time > 0 else float("inf")
mem_ratio = pytorch_mem / triton_mem if triton_mem > 0 else float("inf")
result = {
"config": config,
"triton_time_ms": triton_time,
"pytorch_time_ms": pytorch_time,
"triton_mem": triton_mem,
"pytorch_mem": pytorch_mem,
"speedup": speedup,
"mem_ratio": mem_ratio,
}
results.append(result)
if verbose:
print(f" Triton: {triton_time:.3f} ms, {format_memory(triton_mem)}")
print(f" PyTorch: {pytorch_time:.3f} ms, {format_memory(pytorch_mem)}")
print(f" Speedup: {speedup:.2f}x, Memory ratio: {mem_ratio:.2f}x")
print()
# Clean up
del logits
reset_memory_stats()
return results
def print_summary_table(results: list[dict]):
"""Print a summary table of results."""
print()
print("=" * 130)
print("SUMMARY TABLE")
print("=" * 130)
print()
# Header
header = (
f"{'Scenario':<40} {'Batch':>6} {'Vocab':>7} {'Ops%':>6} "
f"{'Triton (ms)':>12} {'PyTorch (ms)':>13} {'Speedup':>8} "
f"{'Tri Mem':>10} {'Pyt Mem':>10}"
)
print(header)
print("-" * 130)
# Group by scenario type
current_vocab = None
for result in results:
config = result["config"]
# Add separator between vocab sizes
if current_vocab != config.vocab_size:
if current_vocab is not None:
print("-" * 130)
current_vocab = config.vocab_size
scenario = config.name.split("_b")[0] # Extract scenario name
print(
f"{scenario:<40} {config.batch_size:>6} {config.vocab_size:>7} "
f"{config.ops_pct:>5.0f}% "
f"{result['triton_time_ms']:>12.3f} {result['pytorch_time_ms']:>13.3f} "
f"{result['speedup']:>7.2f}x "
f"{format_memory(result['triton_mem']):>10} "
f"{format_memory(result['pytorch_mem']):>10}"
)
print("=" * 130)
def main():
parser = argparse.ArgumentParser(
description="Benchmark Triton vs PyTorch sort-based top-k/top-p implementations"
)
parser.add_argument(
"--batch-sizes",
type=int,
nargs="+",
default=[1, 4, 16, 64, 128, 512, 1024, 2048],
help="Batch sizes to test (default: 1 4 16 64)",
)
parser.add_argument(
"--vocab-sizes",
type=int,
nargs="+",
default=[32768, 131072], # 32k, 128k
help="Vocabulary sizes to test (default: 32768 131072)",
)
parser.add_argument(
"--warmup-iters",
type=int,
default=5,
help="Number of warmup iterations (default: 5)",
)
parser.add_argument(
"--benchmark-iters",
type=int,
default=20,
help="Number of benchmark iterations (default: 20)",
)
parser.add_argument(
"--quiet",
action="store_true",
help="Only print summary table",
)
args = parser.parse_args()
# Print configuration
print(f"Batch sizes: {args.batch_sizes}")
print(f"Vocab sizes: {args.vocab_sizes}")
print(f"Warmup iterations: {args.warmup_iters}")
print(f"Benchmark iterations: {args.benchmark_iters}")
print()
# Check CUDA
if not torch.cuda.is_available():
print("ERROR: CUDA is not available. This benchmark requires a GPU.")
return
device_name = torch.cuda.get_device_name(0)
print(f"GPU: {device_name}")
print()
# Create configs
configs = create_benchmark_configs(
args.batch_sizes,
args.vocab_sizes,
)
# Run benchmarks
results = run_benchmark(
configs,
warmup_iters=args.warmup_iters,
benchmark_iters=args.benchmark_iters,
verbose=not args.quiet,
)
# Print summary
print_summary_table(results)
if __name__ == "__main__":
main()
+16 -14
View File
@@ -71,7 +71,7 @@ while [[ $# -gt 0 ]]; do
usage usage
;; ;;
*) *)
echo "Unknown argument: $1\n" printf "Unknown argument: %s\n" "$1"
usage usage
;; ;;
esac esac
@@ -84,15 +84,17 @@ mkdir -p "$OUTPUT_DIR"
QPS_VALUES=(25 20 15 10 5 1) QPS_VALUES=(25 20 15 10 5 1)
# Common parameters # Common parameters
COMMON_PARAMS="--backend $BACKEND \ COMMON_PARAMS=(
--model $MODEL \ --backend "$BACKEND"
--dataset $DATASET \ --model "$MODEL"
--structured-output-ratio $STRUCTURED_OUTPUT_RATIO \ --dataset "$DATASET"
--save-results \ --structured-output-ratio "$STRUCTURED_OUTPUT_RATIO"
--result-dir $OUTPUT_DIR \ --save-results
--output-len $MAX_NEW_TOKENS \ --result-dir "$OUTPUT_DIR"
--port $PORT \ --output-len "$MAX_NEW_TOKENS"
--tokenizer-mode $TOKENIZER_MODE" --port "$PORT"
--tokenizer-mode "$TOKENIZER_MODE"
)
echo "Starting structured output benchmark with model: $MODEL" echo "Starting structured output benchmark with model: $MODEL"
echo "Backend: $BACKEND" echo "Backend: $BACKEND"
@@ -109,17 +111,17 @@ for qps in "${QPS_VALUES[@]}"; do
GIT_BRANCH=$(git rev-parse --abbrev-ref HEAD 2>/dev/null || echo "unknown") GIT_BRANCH=$(git rev-parse --abbrev-ref HEAD 2>/dev/null || echo "unknown")
# Construct filename for this run # Construct filename for this run
FILENAME="${BACKEND}_${qps}qps_$(basename $MODEL)_${DATASET}_${GIT_HASH}.json" FILENAME="${BACKEND}_${qps}qps_$(basename "$MODEL")_${DATASET}_${GIT_HASH}_${GIT_BRANCH}.json"
NUM_PROMPTS=$(echo "$TOTAL_SECONDS * $qps" | bc) NUM_PROMPTS=$(echo "$TOTAL_SECONDS * $qps" | bc)
NUM_PROMPTS=${NUM_PROMPTS%.*} # Remove fractional part NUM_PROMPTS=${NUM_PROMPTS%.*} # Remove fractional part
echo "Running benchmark with $NUM_PROMPTS prompts" echo "Running benchmark with $NUM_PROMPTS prompts"
# Run the benchmark # Run the benchmark
python "$SCRIPT_DIR/benchmark_serving_structured_output.py" $COMMON_PARAMS \ python "$SCRIPT_DIR/benchmark_serving_structured_output.py" "${COMMON_PARAMS[@]}" \
--request-rate $qps \ --request-rate "$qps" \
--result-filename "$FILENAME" \ --result-filename "$FILENAME" \
--num-prompts $NUM_PROMPTS --num-prompts "$NUM_PROMPTS"
echo "Completed benchmark with QPS: $qps" echo "Completed benchmark with QPS: $qps"
echo "----------------------------------------" echo "----------------------------------------"
+5
View File
@@ -18,6 +18,7 @@ set(ENABLE_AVX512 $ENV{VLLM_CPU_AVX512})
set(ENABLE_AVX512BF16 $ENV{VLLM_CPU_AVX512BF16}) set(ENABLE_AVX512BF16 $ENV{VLLM_CPU_AVX512BF16})
set(ENABLE_AVX512VNNI $ENV{VLLM_CPU_AVX512VNNI}) set(ENABLE_AVX512VNNI $ENV{VLLM_CPU_AVX512VNNI})
set(ENABLE_AMXBF16 $ENV{VLLM_CPU_AMXBF16}) set(ENABLE_AMXBF16 $ENV{VLLM_CPU_AMXBF16})
set(ENABLE_ARM_BF16 $ENV{VLLM_CPU_ARM_BF16})
include_directories("${CMAKE_SOURCE_DIR}/csrc") include_directories("${CMAKE_SOURCE_DIR}/csrc")
@@ -115,6 +116,10 @@ else()
set(AVX512_FOUND ON) set(AVX512_FOUND ON)
message(STATUS "AVX512 support enabled via VLLM_CPU_AVX512 environment variable") message(STATUS "AVX512 support enabled via VLLM_CPU_AVX512 environment variable")
endif() endif()
if (ENABLE_ARM_BF16)
set(ARM_BF16_FOUND ON)
message(STATUS "ARM BF16 support enabled via VLLM_CPU_ARM_BF16 environment variable")
endif()
endif() endif()
if (AVX512_FOUND AND NOT AVX512_DISABLED) if (AVX512_FOUND AND NOT AVX512_DISABLED)
+1 -1
View File
@@ -19,7 +19,7 @@ else()
FetchContent_Declare( FetchContent_Declare(
flashmla flashmla
GIT_REPOSITORY https://github.com/vllm-project/FlashMLA GIT_REPOSITORY https://github.com/vllm-project/FlashMLA
GIT_TAG c2afa9cb93e674d5a9120a170a6da57b89267208 GIT_TAG 692917b1cda61b93ac9ee2d846ec54e75afe87b1
GIT_PROGRESS TRUE GIT_PROGRESS TRUE
CONFIGURE_COMMAND "" CONFIGURE_COMMAND ""
BUILD_COMMAND "" BUILD_COMMAND ""
+15 -5
View File
@@ -22,7 +22,12 @@
# docker buildx bake -f docker/docker-bake.hcl -f docker/versions.json # docker buildx bake -f docker/docker-bake.hcl -f docker/versions.json
# ============================================================================= # =============================================================================
ARG CUDA_VERSION=12.9.1 ARG CUDA_VERSION=12.8.1
# BUILDER_CUDA_VERSION controls the CUDA toolkit used to compile csrc/ and
# extensions (DeepGEMM, EP kernels). It can differ from CUDA_VERSION, which
# is the CUDA version shipped in the final runtime image and used to select
# the matching PyTorch wheel.
ARG BUILDER_CUDA_VERSION=12.9.1
ARG PYTHON_VERSION=3.12 ARG PYTHON_VERSION=3.12
# By parameterizing the base images, we allow third-party to use their own # By parameterizing the base images, we allow third-party to use their own
@@ -36,7 +41,7 @@ ARG PYTHON_VERSION=3.12
# compatibility with other Linux OSes. The main reason for this is that the # compatibility with other Linux OSes. The main reason for this is that the
# glibc version is baked into the distro, and binaries built with one glibc # glibc version is baked into the distro, and binaries built with one glibc
# version are not backwards compatible with OSes that use an earlier version. # version are not backwards compatible with OSes that use an earlier version.
ARG BUILD_BASE_IMAGE=nvidia/cuda:${CUDA_VERSION}-devel-ubuntu20.04 ARG BUILD_BASE_IMAGE=nvidia/cuda:${BUILDER_CUDA_VERSION}-devel-ubuntu20.04
# Using cuda base image with minimal dependencies necessary for JIT compilation (FlashInfer, DeepGEMM, EP kernels) # Using cuda base image with minimal dependencies necessary for JIT compilation (FlashInfer, DeepGEMM, EP kernels)
ARG FINAL_BASE_IMAGE=nvidia/cuda:${CUDA_VERSION}-base-ubuntu22.04 ARG FINAL_BASE_IMAGE=nvidia/cuda:${CUDA_VERSION}-base-ubuntu22.04
@@ -92,8 +97,13 @@ ARG INSTALL_KV_CONNECTORS=false
FROM ${BUILD_BASE_IMAGE} AS base FROM ${BUILD_BASE_IMAGE} AS base
ARG CUDA_VERSION ARG CUDA_VERSION
ARG BUILDER_CUDA_VERSION
ARG PYTHON_VERSION ARG PYTHON_VERSION
# Override the CUDA_VERSION env var inherited from the nvidia base image
# (which equals BUILDER_CUDA_VERSION) so that $CUDA_VERSION in RUN commands
# resolves to the runtime CUDA version used for PyTorch wheel selection.
ENV CUDA_VERSION=${CUDA_VERSION}
ENV DEBIAN_FRONTEND=noninteractive ENV DEBIAN_FRONTEND=noninteractive
# Install system dependencies including build tools # Install system dependencies including build tools
@@ -133,7 +143,7 @@ ENV UV_LINK_MODE=copy
RUN gcc --version RUN gcc --version
# Ensure CUDA compatibility library is loaded # Ensure CUDA compatibility library is loaded
RUN echo "/usr/local/cuda-$(echo "$CUDA_VERSION" | cut -d. -f1,2)/compat/" > /etc/ld.so.conf.d/cuda-compat.conf && ldconfig RUN echo "/usr/local/cuda-$(echo "$BUILDER_CUDA_VERSION" | cut -d. -f1,2)/compat/" > /etc/ld.so.conf.d/cuda-compat.conf && ldconfig
# ============================================================ # ============================================================
# SLOW-CHANGING DEPENDENCIES BELOW # SLOW-CHANGING DEPENDENCIES BELOW
@@ -309,7 +319,7 @@ RUN --mount=type=cache,target=/root/.cache/ccache \
# Build DeepGEMM, pplx-kernels, DeepEP - runs in PARALLEL with csrc-build # Build DeepGEMM, pplx-kernels, DeepEP - runs in PARALLEL with csrc-build
# This stage is independent and doesn't affect csrc cache # This stage is independent and doesn't affect csrc cache
FROM base AS extensions-build FROM base AS extensions-build
ARG CUDA_VERSION ARG BUILDER_CUDA_VERSION
# This timeout (in seconds) is necessary when installing some dependencies via uv since it's likely to time out # This timeout (in seconds) is necessary when installing some dependencies via uv since it's likely to time out
ENV UV_HTTP_TIMEOUT=500 ENV UV_HTTP_TIMEOUT=500
@@ -325,7 +335,7 @@ COPY tools/install_deepgemm.sh /tmp/install_deepgemm.sh
RUN --mount=type=cache,target=/root/.cache/uv \ RUN --mount=type=cache,target=/root/.cache/uv \
mkdir -p /tmp/deepgemm/dist && \ mkdir -p /tmp/deepgemm/dist && \
VLLM_DOCKER_BUILD_CONTEXT=1 TORCH_CUDA_ARCH_LIST="9.0a 10.0a" /tmp/install_deepgemm.sh \ VLLM_DOCKER_BUILD_CONTEXT=1 TORCH_CUDA_ARCH_LIST="9.0a 10.0a" /tmp/install_deepgemm.sh \
--cuda-version "${CUDA_VERSION}" \ --cuda-version "${BUILDER_CUDA_VERSION}" \
${DEEPGEMM_GIT_REF:+--ref "$DEEPGEMM_GIT_REF"} \ ${DEEPGEMM_GIT_REF:+--ref "$DEEPGEMM_GIT_REF"} \
--wheel-dir /tmp/deepgemm/dist || \ --wheel-dir /tmp/deepgemm/dist || \
echo "DeepGEMM build skipped (CUDA version requirement not met)" echo "DeepGEMM build skipped (CUDA version requirement not met)"
+16
View File
@@ -20,6 +20,7 @@
# VLLM_CPU_AVX512BF16=false (default)|true (for cross-compilation) # VLLM_CPU_AVX512BF16=false (default)|true (for cross-compilation)
# VLLM_CPU_AVX512VNNI=false (default)|true (for cross-compilation) # VLLM_CPU_AVX512VNNI=false (default)|true (for cross-compilation)
# VLLM_CPU_AMXBF16=false (default)|true (for cross-compilation) # VLLM_CPU_AMXBF16=false (default)|true (for cross-compilation)
# VLLM_CPU_ARM_BF16=false (default)|true (for cross-compilation)
# #
######################### COMMON BASE IMAGE ######################### ######################### COMMON BASE IMAGE #########################
@@ -108,9 +109,22 @@ ENV VLLM_CPU_AVX512VNNI=${VLLM_CPU_AVX512VNNI}
# Support for building with AMXBF16 ISA: docker build --build-arg VLLM_CPU_AMXBF16="true" ... # Support for building with AMXBF16 ISA: docker build --build-arg VLLM_CPU_AMXBF16="true" ...
ARG VLLM_CPU_AMXBF16=1 ARG VLLM_CPU_AMXBF16=1
ENV VLLM_CPU_AMXBF16=${VLLM_CPU_AMXBF16} ENV VLLM_CPU_AMXBF16=${VLLM_CPU_AMXBF16}
# Support for cross-compilation with ARM BF16 ISA: docker build --build-arg VLLM_CPU_ARM_BF16="true" ...
ARG VLLM_CPU_ARM_BF16=0
ENV VLLM_CPU_ARM_BF16=${VLLM_CPU_ARM_BF16}
WORKDIR /vllm-workspace WORKDIR /vllm-workspace
# Validate build arguments - prevent mixing incompatible ISA flags
RUN if [ "$TARGETARCH" = "arm64" ] && { [ "$VLLM_CPU_AVX2" != "0" ] || [ "$VLLM_CPU_AVX512" != "0" ] || [ "$VLLM_CPU_AVX512BF16" != "0" ] || [ "$VLLM_CPU_AVX512VNNI" != "0" ]; }; then \
echo "ERROR: Cannot use x86-specific ISA flags (AVX2, AVX512, etc.) when building for ARM64 (--platform=linux/arm64)"; \
exit 1; \
fi && \
if [ "$TARGETARCH" = "amd64" ] && [ "$VLLM_CPU_ARM_BF16" != "0" ]; then \
echo "ERROR: Cannot use ARM-specific ISA flags (ARM_BF16) when building for x86_64 (--platform=linux/amd64)"; \
exit 1; \
fi
# Copy build requirements # Copy build requirements
COPY requirements/cpu-build.txt requirements/build.txt COPY requirements/cpu-build.txt requirements/build.txt
@@ -224,6 +238,7 @@ ARG VLLM_CPU_AVX512
ARG VLLM_CPU_AVX512BF16 ARG VLLM_CPU_AVX512BF16
ARG VLLM_CPU_AVX512VNNI ARG VLLM_CPU_AVX512VNNI
ARG VLLM_CPU_AMXBF16 ARG VLLM_CPU_AMXBF16
ARG VLLM_CPU_ARM_BF16
ARG PYTHON_VERSION ARG PYTHON_VERSION
LABEL ai.vllm.build.target-arch="${TARGETARCH}" LABEL ai.vllm.build.target-arch="${TARGETARCH}"
@@ -233,6 +248,7 @@ LABEL ai.vllm.build.cpu-avx512="${VLLM_CPU_AVX512:-false}"
LABEL ai.vllm.build.cpu-avx512bf16="${VLLM_CPU_AVX512BF16:-false}" LABEL ai.vllm.build.cpu-avx512bf16="${VLLM_CPU_AVX512BF16:-false}"
LABEL ai.vllm.build.cpu-avx512vnni="${VLLM_CPU_AVX512VNNI:-false}" LABEL ai.vllm.build.cpu-avx512vnni="${VLLM_CPU_AVX512VNNI:-false}"
LABEL ai.vllm.build.cpu-amxbf16="${VLLM_CPU_AMXBF16:-false}" LABEL ai.vllm.build.cpu-amxbf16="${VLLM_CPU_AMXBF16:-false}"
LABEL ai.vllm.build.cpu-arm-bf16="${VLLM_CPU_ARM_BF16:-false}"
LABEL ai.vllm.build.python-version="${PYTHON_VERSION:-3.12}" LABEL ai.vllm.build.python-version="${PYTHON_VERSION:-3.12}"
ENTRYPOINT ["vllm", "serve"] ENTRYPOINT ["vllm", "serve"]
+4 -1
View File
@@ -2,6 +2,9 @@
"_comment": "Auto-generated from Dockerfile ARGs. Do not edit manually. Run: python tools/generate_versions_json.py", "_comment": "Auto-generated from Dockerfile ARGs. Do not edit manually. Run: python tools/generate_versions_json.py",
"variable": { "variable": {
"CUDA_VERSION": { "CUDA_VERSION": {
"default": "12.8.1"
},
"BUILDER_CUDA_VERSION": {
"default": "12.9.1" "default": "12.9.1"
}, },
"PYTHON_VERSION": { "PYTHON_VERSION": {
@@ -11,7 +14,7 @@
"default": "nvidia/cuda:12.9.1-devel-ubuntu20.04" "default": "nvidia/cuda:12.9.1-devel-ubuntu20.04"
}, },
"FINAL_BASE_IMAGE": { "FINAL_BASE_IMAGE": {
"default": "nvidia/cuda:12.9.1-base-ubuntu22.04" "default": "nvidia/cuda:12.8.1-base-ubuntu22.04"
}, },
"GET_PIP_URL": { "GET_PIP_URL": {
"default": "https://bootstrap.pypa.io/get-pip.py" "default": "https://bootstrap.pypa.io/get-pip.py"
+1
View File
@@ -182,6 +182,7 @@ The following table lists backends that support full CUDA Graphs at the time of
| FlashInfer | `UNIFORM_SINGLE_TOKEN_DECODE` | Will be set to `UNIFORM_BATCH` when using TRTLLM attention on Blackwell | | FlashInfer | `UNIFORM_SINGLE_TOKEN_DECODE` | Will be set to `UNIFORM_BATCH` when using TRTLLM attention on Blackwell |
| FlashMLA | `UNIFORM_BATCH` | | | FlashMLA | `UNIFORM_BATCH` | |
| FlashInferMLA | `UNIFORM_BATCH` | | | FlashInferMLA | `UNIFORM_BATCH` | |
| FlashInferMLASparse | `UNIFORM_BATCH` | |
| AITER MLA | `UNIFORM_SINGLE_TOKEN_DECODE` | | | AITER MLA | `UNIFORM_SINGLE_TOKEN_DECODE` | |
| CUTLASS MLA | `UNIFORM_SINGLE_TOKEN_DECODE` | | | CUTLASS MLA | `UNIFORM_SINGLE_TOKEN_DECODE` | |
| Mamba attention| `UNIFORM_SINGLE_TOKEN_DECODE` | | | Mamba attention| `UNIFORM_SINGLE_TOKEN_DECODE` | |
+1
View File
@@ -109,6 +109,7 @@ Batch invariance has been tested and verified on the following models:
- **Qwen2.5**: `Qwen/Qwen2.5-0.5B-Instruct`, `Qwen/Qwen2.5-1.5B-Instruct`, `Qwen/Qwen2.5-3B-Instruct`, `Qwen/Qwen2.5-7B-Instruct`, `Qwen/Qwen2.5-14B-Instruct`, `Qwen/Qwen2.5-32B-Instruct` - **Qwen2.5**: `Qwen/Qwen2.5-0.5B-Instruct`, `Qwen/Qwen2.5-1.5B-Instruct`, `Qwen/Qwen2.5-3B-Instruct`, `Qwen/Qwen2.5-7B-Instruct`, `Qwen/Qwen2.5-14B-Instruct`, `Qwen/Qwen2.5-32B-Instruct`
- **Llama 3**: `meta-llama/Llama-3.1-8B-Instruct`, `meta-llama/Llama-3.2-1B-Instruct` - **Llama 3**: `meta-llama/Llama-3.1-8B-Instruct`, `meta-llama/Llama-3.2-1B-Instruct`
- **GPT-OSS**: `openai/gpt-oss-20b`, `openai/gpt-oss-120b` - **GPT-OSS**: `openai/gpt-oss-20b`, `openai/gpt-oss-120b`
- **Mistral**: `mistralai/Mistral-7B-v0.3`
Other models may also work, but these have been explicitly validated. If you encounter issues with a specific model, please report them on the [GitHub issue tracker](https://github.com/vllm-project/vllm/issues/new/choose). Other models may also work, but these have been explicitly validated. If you encounter issues with a specific model, please report them on the [GitHub issue tracker](https://github.com/vllm-project/vllm/issues/new/choose).
+1 -1
View File
@@ -84,7 +84,7 @@ Since simple RTN does not require data for weight quantization and the activatio
Install `vllm` and `lm-evaluation-harness` for evaluation: Install `vllm` and `lm-evaluation-harness` for evaluation:
```bash ```bash
pip install vllm "lm-eval[api]>=0.4.9.2" pip install vllm "lm-eval[api]>=0.4.11"
``` ```
Load and run the model in `vllm`: Load and run the model in `vllm`:
+1 -1
View File
@@ -18,7 +18,7 @@ pip install llmcompressor
Additionally, install `vllm` and `lm-evaluation-harness` for evaluation: Additionally, install `vllm` and `lm-evaluation-harness` for evaluation:
```bash ```bash
pip install vllm "lm-eval[api]>=0.4.9.2" pip install vllm "lm-eval[api]>=0.4.11"
``` ```
## Quantization Process ## Quantization Process
+1 -1
View File
@@ -23,7 +23,7 @@ pip install llmcompressor
Additionally, install `vllm` and `lm-evaluation-harness` for evaluation: Additionally, install `vllm` and `lm-evaluation-harness` for evaluation:
```bash ```bash
pip install vllm "lm-eval[api]>=0.4.9.2" pip install vllm "lm-eval[api]>=0.4.11"
``` ```
## Quantization Process ## Quantization Process
+1 -1
View File
@@ -20,7 +20,7 @@ for more installation details.
Additionally, install `vllm` and `lm-evaluation-harness` for evaluation: Additionally, install `vllm` and `lm-evaluation-harness` for evaluation:
```bash ```bash
pip install vllm "lm-eval[api]>=0.4.9.2" pip install vllm "lm-eval[api]>=0.4.11"
``` ```
## Quantization Process ## Quantization Process
@@ -172,25 +172,78 @@ docker pull public.ecr.aws/q9t5s3a7/vllm-ci-postmerge-repo:${VLLM_COMMIT}-arm64-
# --8<-- [end:pre-built-images] # --8<-- [end:pre-built-images]
# --8<-- [start:build-image-from-source] # --8<-- [start:build-image-from-source]
## Building for your target ARM CPU
```bash ```bash
docker build -f docker/Dockerfile.cpu \ docker build -f docker/Dockerfile.cpu \
--tag vllm-cpu-env . --platform=linux/arm64 \
--build-arg VLLM_CPU_ARM_BF16=<false (default)|true> \
--tag vllm-cpu-env \
--target vllm-openai .
```
# Launching OpenAI server !!! note "Auto-detection by default"
By default, ARM CPU instruction sets (BF16, NEON, etc.) are automatically detected from the build system's CPU flags. The `VLLM_CPU_ARM_BF16` build argument is used for cross-compilation:
- `VLLM_CPU_ARM_BF16=true` - Force-enable ARM BF16 support (build with BF16 regardless of build system capabilities)
- `VLLM_CPU_ARM_BF16=false` - Rely on auto-detection (default)
### Examples
**Auto-detection build (native ARM)**
```bash
# Building on ARM64 system - platform auto-detected
docker build -f docker/Dockerfile.cpu \
--tag vllm-cpu-arm64 \
--target vllm-openai .
```
**Cross-compile for ARM with BF16 support**
```bash
# Building on ARM64 for newer ARM CPUs with BF16
docker build -f docker/Dockerfile.cpu \
--build-arg VLLM_CPU_ARM_BF16=true \
--tag vllm-cpu-arm64-bf16 \
--target vllm-openai .
```
**Cross-compile from x86_64 to ARM64 with BF16**
```bash
# Requires Docker buildx with ARM emulation (QEMU)
docker buildx build -f docker/Dockerfile.cpu \
--platform=linux/arm64 \
--build-arg VLLM_CPU_ARM_BF16=true \
--build-arg max_jobs=4 \
--tag vllm-cpu-arm64-bf16 \
--target vllm-openai \
--load .
```
!!! note "ARM BF16 requirements"
ARM BF16 support requires ARMv8.6-A or later (FEAT_BF16). Supported on AWS Graviton3/4, AmpereOne, and other recent ARM processors.
## Launching the OpenAI server
```bash
docker run --rm \ docker run --rm \
--privileged=true \ --security-opt seccomp=unconfined \
--cap-add SYS_NICE \
--shm-size=4g \ --shm-size=4g \
-p 8000:8000 \ -p 8000:8000 \
-e VLLM_CPU_KVCACHE_SPACE=<KV cache space> \ -e VLLM_CPU_KVCACHE_SPACE=<KV cache space> \
-e VLLM_CPU_OMP_THREADS_BIND=<CPU cores for inference> \ -e VLLM_CPU_OMP_THREADS_BIND=<CPU cores for inference> \
vllm-cpu-env \ vllm-cpu-arm64 \
--model=meta-llama/Llama-3.2-1B-Instruct \ meta-llama/Llama-3.2-1B-Instruct \
--dtype=bfloat16 \ --dtype=bfloat16 \
other vLLM OpenAI server arguments other vLLM OpenAI server arguments
``` ```
!!! tip !!! tip "Alternative to --privileged"
An alternative of `--privileged=true` is `--cap-add SYS_NICE --security-opt seccomp=unconfined`. Instead of `--privileged=true`, use `--cap-add SYS_NICE --security-opt seccomp=unconfined` for better security.
# --8<-- [end:build-image-from-source] # --8<-- [end:build-image-from-source]
# --8<-- [start:extra-information] # --8<-- [start:extra-information]
+1
View File
@@ -181,6 +181,7 @@ Some model architectures are supported via vLLM plugins. These plugins extend vL
| Architecture | Models | Plugin Repository | | Architecture | Models | Plugin Repository |
|--------------|--------|-------------------| |--------------|--------|-------------------|
| `BartForConditionalGeneration` | BART | [bart-plugin](https://github.com/vllm-project/bart-plugin) | | `BartForConditionalGeneration` | BART | [bart-plugin](https://github.com/vllm-project/bart-plugin) |
| `Florence2ForConditionalGeneration` | Florence-2 | [bart-plugin](https://github.com/vllm-project/bart-plugin) |
For other model architectures not natively supported, in particular for Encoder-Decoder models, we recommend following a similar pattern by implementing support through the plugin system. For other model architectures not natively supported, in particular for Encoder-Decoder models, we recommend following a similar pattern by implementing support through the plugin system.
+1
View File
@@ -137,6 +137,7 @@ Please note that prefix caching is not yet supported for any of the above models
Whisper is supported natively. Other encoder-decoder models are supported via the plugin system: Whisper is supported natively. Other encoder-decoder models are supported via the plugin system:
- **BART**: `BartForConditionalGeneration` is supported via the official [bart-plugin](https://github.com/vllm-project/bart-plugin). - **BART**: `BartForConditionalGeneration` is supported via the official [bart-plugin](https://github.com/vllm-project/bart-plugin).
- **Florence-2**: `Florence2ForConditionalGeneration` is supported via the official [bart-plugin](https://github.com/vllm-project/bart-plugin).
For other encoder-decoder models (e.g., `MllamaForConditionalGeneration`), we recommend For other encoder-decoder models (e.g., `MllamaForConditionalGeneration`), we recommend
following a similar pattern by implementing support through the [plugin system](../design/plugin_system.md). following a similar pattern by implementing support through the [plugin system](../design/plugin_system.md).
@@ -8,7 +8,7 @@ declare -a PIDS=()
############################################################################### ###############################################################################
MODEL="${MODEL:-Qwen/Qwen2.5-VL-3B-Instruct}" MODEL="${MODEL:-Qwen/Qwen2.5-VL-3B-Instruct}"
LOG_PATH="${LOG_PATH:-./logs}" LOG_PATH="${LOG_PATH:-./logs}"
mkdir -p $LOG_PATH mkdir -p "$LOG_PATH"
ENCODE_PORT="${ENCODE_PORT:-19534}" ENCODE_PORT="${ENCODE_PORT:-19534}"
PREFILL_PORT="${PREFILL_PORT:-19535}" PREFILL_PORT="${PREFILL_PORT:-19535}"
@@ -84,10 +84,10 @@ trap cleanup TERM
# clear previous cache # clear previous cache
echo "remove previous ec cache folder" echo "remove previous ec cache folder"
rm -rf $EC_SHARED_STORAGE_PATH rm -rf "$EC_SHARED_STORAGE_PATH"
echo "make ec cache folder" echo "make ec cache folder"
mkdir -p $EC_SHARED_STORAGE_PATH mkdir -p "$EC_SHARED_STORAGE_PATH"
############################################################################### ###############################################################################
# Encoder worker # Encoder worker
@@ -100,7 +100,7 @@ CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \
--no-enable-prefix-caching \ --no-enable-prefix-caching \
--max-num-batched-tokens 114688 \ --max-num-batched-tokens 114688 \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_producer", "ec_role": "ec_producer",
@@ -124,7 +124,7 @@ vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_consumer", "ec_role": "ec_consumer",
@@ -152,7 +152,7 @@ vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--kv-transfer-config '{ --kv-transfer-config '{
"kv_connector": "NixlConnector", "kv_connector": "NixlConnector",
"kv_role": "kv_consumer" "kv_role": "kv_consumer"
@@ -162,9 +162,9 @@ vllm serve "$MODEL" \
PIDS+=($!) PIDS+=($!)
# Wait for workers # Wait for workers
wait_for_server $ENCODE_PORT wait_for_server "$ENCODE_PORT"
wait_for_server $PREFILL_PORT wait_for_server "$PREFILL_PORT"
wait_for_server $DECODE_PORT wait_for_server "$DECODE_PORT"
############################################################################### ###############################################################################
# Proxy # Proxy
@@ -179,7 +179,7 @@ python disagg_epd_proxy.py \
PIDS+=($!) PIDS+=($!)
wait_for_server $PROXY_PORT wait_for_server "$PROXY_PORT"
echo "All services are up!" echo "All services are up!"
############################################################################### ###############################################################################
@@ -187,14 +187,14 @@ echo "All services are up!"
############################################################################### ###############################################################################
echo "Running benchmark (stream)..." echo "Running benchmark (stream)..."
vllm bench serve \ vllm bench serve \
--model $MODEL \ --model "$MODEL" \
--backend openai-chat \ --backend openai-chat \
--endpoint /v1/chat/completions \ --endpoint /v1/chat/completions \
--dataset-name hf \ --dataset-name hf \
--dataset-path lmarena-ai/VisionArena-Chat \ --dataset-path lmarena-ai/VisionArena-Chat \
--seed 0 \ --seed 0 \
--num-prompts $NUM_PROMPTS \ --num-prompts "$NUM_PROMPTS" \
--port $PROXY_PORT --port "$PROXY_PORT"
PIDS+=($!) PIDS+=($!)
@@ -202,10 +202,10 @@ PIDS+=($!)
# Single request with local image # Single request with local image
############################################################################### ###############################################################################
echo "Running single request with local image (non-stream)..." echo "Running single request with local image (non-stream)..."
curl http://127.0.0.1:${PROXY_PORT}/v1/chat/completions \ curl http://127.0.0.1:"${PROXY_PORT}"/v1/chat/completions \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{ -d '{
"model": "'${MODEL}'", "model": "'"${MODEL}"'",
"messages": [ "messages": [
{"role": "system", "content": "You are a helpful assistant."}, {"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [ {"role": "user", "content": [
@@ -8,7 +8,7 @@ declare -a PIDS=()
############################################################################### ###############################################################################
MODEL="${MODEL:-Qwen/Qwen2.5-VL-3B-Instruct}" MODEL="${MODEL:-Qwen/Qwen2.5-VL-3B-Instruct}"
LOG_PATH="${LOG_PATH:-./logs}" LOG_PATH="${LOG_PATH:-./logs}"
mkdir -p $LOG_PATH mkdir -p "$LOG_PATH"
ENCODE_PORT="${ENCODE_PORT:-19534}" ENCODE_PORT="${ENCODE_PORT:-19534}"
PREFILL_DECODE_PORT="${PREFILL_DECODE_PORT:-19535}" PREFILL_DECODE_PORT="${PREFILL_DECODE_PORT:-19535}"
@@ -78,10 +78,10 @@ trap cleanup TERM
# clear previous cache # clear previous cache
echo "remove previous ec cache folder" echo "remove previous ec cache folder"
rm -rf $EC_SHARED_STORAGE_PATH rm -rf "$EC_SHARED_STORAGE_PATH"
echo "make ec cache folder" echo "make ec cache folder"
mkdir -p $EC_SHARED_STORAGE_PATH mkdir -p "$EC_SHARED_STORAGE_PATH"
############################################################################### ###############################################################################
# Encoder worker # Encoder worker
@@ -94,7 +94,7 @@ CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \
--no-enable-prefix-caching \ --no-enable-prefix-caching \
--max-num-batched-tokens 114688 \ --max-num-batched-tokens 114688 \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_producer", "ec_role": "ec_producer",
@@ -115,7 +115,7 @@ CUDA_VISIBLE_DEVICES="$GPU_PD" vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_consumer", "ec_role": "ec_consumer",
@@ -128,8 +128,8 @@ CUDA_VISIBLE_DEVICES="$GPU_PD" vllm serve "$MODEL" \
PIDS+=($!) PIDS+=($!)
# Wait for workers # Wait for workers
wait_for_server $ENCODE_PORT wait_for_server "$ENCODE_PORT"
wait_for_server $PREFILL_DECODE_PORT wait_for_server "$PREFILL_DECODE_PORT"
############################################################################### ###############################################################################
# Proxy # Proxy
@@ -144,7 +144,7 @@ python disagg_epd_proxy.py \
PIDS+=($!) PIDS+=($!)
wait_for_server $PROXY_PORT wait_for_server "$PROXY_PORT"
echo "All services are up!" echo "All services are up!"
############################################################################### ###############################################################################
@@ -152,14 +152,14 @@ echo "All services are up!"
############################################################################### ###############################################################################
echo "Running benchmark (stream)..." echo "Running benchmark (stream)..."
vllm bench serve \ vllm bench serve \
--model $MODEL \ --model "$MODEL" \
--backend openai-chat \ --backend openai-chat \
--endpoint /v1/chat/completions \ --endpoint /v1/chat/completions \
--dataset-name hf \ --dataset-name hf \
--dataset-path lmarena-ai/VisionArena-Chat \ --dataset-path lmarena-ai/VisionArena-Chat \
--seed 0 \ --seed 0 \
--num-prompts $NUM_PROMPTS \ --num-prompts "$NUM_PROMPTS" \
--port $PROXY_PORT --port "$PROXY_PORT"
PIDS+=($!) PIDS+=($!)
@@ -167,10 +167,10 @@ PIDS+=($!)
# Single request with local image # Single request with local image
############################################################################### ###############################################################################
echo "Running single request with local image (non-stream)..." echo "Running single request with local image (non-stream)..."
curl http://127.0.0.1:${PROXY_PORT}/v1/chat/completions \ curl http://127.0.0.1:"${PROXY_PORT}"/v1/chat/completions \
-H "Content-Type: application/json" \ -H "Content-Type: application/json" \
-d '{ -d '{
"model": "'${MODEL}'", "model": "'"${MODEL}"'",
"messages": [ "messages": [
{"role": "system", "content": "You are a helpful assistant."}, {"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": [ {"role": "user", "content": [
@@ -54,7 +54,7 @@ wait_for_server() {
# You can also adjust --kv-ip and --kv-port for distributed inference. # You can also adjust --kv-ip and --kv-port for distributed inference.
# prefilling instance, which is the KV producer # prefilling instance, which is the KV producer
CUDA_VISIBLE_DEVICES=0 vllm serve $MODEL_NAME \ CUDA_VISIBLE_DEVICES=0 vllm serve "$MODEL_NAME" \
--host 0.0.0.0 \ --host 0.0.0.0 \
--port 8100 \ --port 8100 \
--max-model-len 100 \ --max-model-len 100 \
@@ -64,7 +64,7 @@ CUDA_VISIBLE_DEVICES=0 vllm serve $MODEL_NAME \
'{"kv_connector":"P2pNcclConnector","kv_role":"kv_producer","kv_rank":0,"kv_parallel_size":2,"kv_buffer_size":"1e9","kv_port":"14579","kv_connector_extra_config":{"proxy_ip":"'"$VLLM_HOST_IP"'","proxy_port":"30001","http_ip":"'"$VLLM_HOST_IP"'","http_port":"8100","send_type":"PUT_ASYNC"}}' & '{"kv_connector":"P2pNcclConnector","kv_role":"kv_producer","kv_rank":0,"kv_parallel_size":2,"kv_buffer_size":"1e9","kv_port":"14579","kv_connector_extra_config":{"proxy_ip":"'"$VLLM_HOST_IP"'","proxy_port":"30001","http_ip":"'"$VLLM_HOST_IP"'","http_port":"8100","send_type":"PUT_ASYNC"}}' &
# decoding instance, which is the KV consumer # decoding instance, which is the KV consumer
CUDA_VISIBLE_DEVICES=1 vllm serve $MODEL_NAME \ CUDA_VISIBLE_DEVICES=1 vllm serve "$MODEL_NAME" \
--host 0.0.0.0 \ --host 0.0.0.0 \
--port 8200 \ --port 8200 \
--max-model-len 100 \ --max-model-len 100 \
@@ -328,9 +328,9 @@ class Proxy:
if instance_type == "decode" and instance in self.decode_instances: if instance_type == "decode" and instance in self.decode_instances:
self.decode_instances.remove(instance) self.decode_instances.remove(instance)
self.decode_cycler = itertools.cycle(self.decode_instances) self.decode_cycler = itertools.cycle(self.decode_instances)
if instance_type == "prefill" and instance in self.decode_instances: if instance_type == "prefill" and instance in self.prefill_instances:
self.prefill_instances.remove(instance) self.prefill_instances.remove(instance)
self.prefill_cycler = itertools.cycle(self.decode_instances) self.prefill_cycler = itertools.cycle(self.prefill_instances)
class RoundRobinSchedulingPolicy(SchedulingPolicy): class RoundRobinSchedulingPolicy(SchedulingPolicy):
@@ -34,7 +34,7 @@ wait_for_server() {
done" && return 0 || return 1 done" && return 0 || return 1
} }
vllm serve $MODEL_NAME \ vllm serve "$MODEL_NAME" \
--port 8100 \ --port 8100 \
--max-model-len 100 \ --max-model-len 100 \
--enforce-eager \ --enforce-eager \
@@ -143,7 +143,7 @@ main() {
IFS=',' read -ra BOOTSTRAP_PORT_ARRAY <<< "$BOOTSTRAP_PORTS" IFS=',' read -ra BOOTSTRAP_PORT_ARRAY <<< "$BOOTSTRAP_PORTS"
IFS=',' read -ra DECODE_PORT_ARRAY <<< "$DECODE_PORTS" IFS=',' read -ra DECODE_PORT_ARRAY <<< "$DECODE_PORTS"
proxy_param="" proxy_args=()
# ============================================================================= # =============================================================================
# Launch Prefill Servers (X Producers) # Launch Prefill Servers (X Producers)
@@ -156,12 +156,12 @@ main() {
local bootstrap_port=${BOOTSTRAP_PORT_ARRAY[$i]} local bootstrap_port=${BOOTSTRAP_PORT_ARRAY[$i]}
echo " Prefill server $((i+1)): GPU $gpu_id, Port $port, Bootstrap Port $bootstrap_port" echo " Prefill server $((i+1)): GPU $gpu_id, Port $port, Bootstrap Port $bootstrap_port"
VLLM_MOONCAKE_BOOTSTRAP_PORT=$bootstrap_port CUDA_VISIBLE_DEVICES=$gpu_id vllm serve $MODEL \ VLLM_MOONCAKE_BOOTSTRAP_PORT=$bootstrap_port CUDA_VISIBLE_DEVICES=$gpu_id vllm serve "$MODEL" \
--port $port \ --port "$port" \
--kv-transfer-config \ --kv-transfer-config \
"{\"kv_connector\":\"MooncakeConnector\",\"kv_role\":\"kv_producer\"}" > prefill$((i+1)).log 2>&1 & "{\"kv_connector\":\"MooncakeConnector\",\"kv_role\":\"kv_producer\"}" > prefill$((i+1)).log 2>&1 &
PIDS+=($!) PIDS+=($!)
proxy_param="${proxy_param} --prefill http://0.0.0.0:${port} $bootstrap_port" proxy_args+=(--prefill "http://0.0.0.0:${port}" "$bootstrap_port")
done done
# ============================================================================= # =============================================================================
@@ -174,12 +174,12 @@ main() {
local port=${DECODE_PORT_ARRAY[$i]} local port=${DECODE_PORT_ARRAY[$i]}
echo " Decode server $((i+1)): GPU $gpu_id, Port $port" echo " Decode server $((i+1)): GPU $gpu_id, Port $port"
CUDA_VISIBLE_DEVICES=$gpu_id vllm serve $MODEL \ CUDA_VISIBLE_DEVICES=$gpu_id vllm serve "$MODEL" \
--port $port \ --port "$port" \
--kv-transfer-config \ --kv-transfer-config \
"{\"kv_connector\":\"MooncakeConnector\",\"kv_role\":\"kv_consumer\"}" > decode$((i+1)).log 2>&1 & "{\"kv_connector\":\"MooncakeConnector\",\"kv_role\":\"kv_consumer\"}" > decode$((i+1)).log 2>&1 &
PIDS+=($!) PIDS+=($!)
proxy_param="${proxy_param} --decode http://0.0.0.0:${port}" proxy_args+=(--decode "http://0.0.0.0:${port}")
done done
# ============================================================================= # =============================================================================
@@ -187,7 +187,7 @@ main() {
# ============================================================================= # =============================================================================
echo "" echo ""
echo "Starting proxy server on port $PROXY_PORT..." echo "Starting proxy server on port $PROXY_PORT..."
python3 mooncake_connector_proxy.py $proxy_param --port $PROXY_PORT > proxy.log 2>&1 & python3 mooncake_connector_proxy.py "${proxy_args[@]}" --port "$PROXY_PORT" > proxy.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# ============================================================================= # =============================================================================
@@ -196,9 +196,10 @@ main() {
echo "" echo ""
echo "Waiting for all servers to start..." echo "Waiting for all servers to start..."
for port in "${PREFILL_PORT_ARRAY[@]}" "${DECODE_PORT_ARRAY[@]}"; do for port in "${PREFILL_PORT_ARRAY[@]}" "${DECODE_PORT_ARRAY[@]}"; do
if ! wait_for_server $port; then if ! wait_for_server "$port"; then
echo "Failed to start server on port $port" echo "Failed to start server on port $port"
cleanup cleanup
# shellcheck disable=SC2317
exit 1 exit 1
fi fi
done done
@@ -209,8 +210,8 @@ main() {
# ============================================================================= # =============================================================================
# Run Benchmark # Run Benchmark
# ============================================================================= # =============================================================================
vllm bench serve --port $PROXY_PORT --seed $(date +%s) \ vllm bench serve --port "$PROXY_PORT" --seed "$(date +%s)" \
--backend vllm --model $MODEL \ --backend vllm --model "$MODEL" \
--dataset-name random --random-input-len 7500 --random-output-len 200 \ --dataset-name random --random-input-len 7500 --random-output-len 200 \
--num-prompts 200 --burstiness 100 --request-rate 2 | tee benchmark.log --num-prompts 200 --burstiness 100 --request-rate 2 | tee benchmark.log
@@ -166,10 +166,10 @@ main() {
local kv_port=$((21001 + i)) local kv_port=$((21001 + i))
echo " Prefill server $((i+1)): GPU $gpu_id, Port $port, KV Port $kv_port" echo " Prefill server $((i+1)): GPU $gpu_id, Port $port, KV Port $kv_port"
CUDA_VISIBLE_DEVICES=$gpu_id vllm serve $MODEL \ CUDA_VISIBLE_DEVICES=$gpu_id vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--host 0.0.0.0 \ --host 0.0.0.0 \
--port $port \ --port "$port" \
--tensor-parallel-size 1 \ --tensor-parallel-size 1 \
--seed 1024 \ --seed 1024 \
--dtype float16 \ --dtype float16 \
@@ -194,10 +194,10 @@ main() {
local kv_port=$((22001 + i)) local kv_port=$((22001 + i))
echo " Decode server $((i+1)): GPU $gpu_id, Port $port, KV Port $kv_port" echo " Decode server $((i+1)): GPU $gpu_id, Port $port, KV Port $kv_port"
CUDA_VISIBLE_DEVICES=$gpu_id vllm serve $MODEL \ CUDA_VISIBLE_DEVICES=$gpu_id vllm serve "$MODEL" \
--enforce-eager \ --enforce-eager \
--host 0.0.0.0 \ --host 0.0.0.0 \
--port $port \ --port "$port" \
--tensor-parallel-size 1 \ --tensor-parallel-size 1 \
--seed 1024 \ --seed 1024 \
--dtype float16 \ --dtype float16 \
@@ -217,9 +217,10 @@ main() {
echo "" echo ""
echo "Waiting for all servers to start..." echo "Waiting for all servers to start..."
for port in "${PREFILL_PORT_ARRAY[@]}" "${DECODE_PORT_ARRAY[@]}"; do for port in "${PREFILL_PORT_ARRAY[@]}" "${DECODE_PORT_ARRAY[@]}"; do
if ! wait_for_server $port; then if ! wait_for_server "$port"; then
echo "Failed to start server on port $port" echo "Failed to start server on port $port"
cleanup cleanup
# shellcheck disable=SC2317
exit 1 exit 1
fi fi
done done
@@ -231,8 +232,8 @@ main() {
# Run Benchmark # Run Benchmark
# ============================================================================= # =============================================================================
cd ../../../benchmarks/ cd ../../../benchmarks/
vllm bench serve --port 10001 --seed $(date +%s) \ vllm bench serve --port 10001 --seed "$(date +%s)" \
--model $MODEL \ --model "$MODEL" \
--dataset-name random --random-input-len 7500 --random-output-len 200 \ --dataset-name random --random-input-len 7500 --random-output-len 200 \
--num-prompts 200 --burstiness 100 --request-rate 2 | tee benchmark.log --num-prompts 200 --burstiness 100 --request-rate 2 | tee benchmark.log
+5 -5
View File
@@ -50,8 +50,8 @@ while [[ $# -gt 0 ]]; do
done done
vllm bench serve \ vllm bench serve \
--model $MODEL_NAME \ --model "$MODEL_NAME" \
--host $HOST \ --host "$HOST" \
--port $PORT \ --port "$PORT" \
--num-prompts $NUM_PROMPTS \ --num-prompts "$NUM_PROMPTS" \
--request-rate $REQUEST_RATE --request-rate "$REQUEST_RATE"
@@ -57,15 +57,15 @@ echo "Starting vLLM server for $MODEL_NAME with data parallel size: $DATA_PARALL
export RAY_DEDUP_LOGS=0 export RAY_DEDUP_LOGS=0
export VLLM_USE_DEEP_GEMM=1 export VLLM_USE_DEEP_GEMM=1
vllm serve $MODEL_NAME \ vllm serve "$MODEL_NAME" \
--data-parallel-size $DATA_PARALLEL_SIZE \ --data-parallel-size "$DATA_PARALLEL_SIZE" \
--data-parallel-size-local $DATA_PARALLEL_SIZE \ --data-parallel-size-local "$DATA_PARALLEL_SIZE" \
--data-parallel-backend ray \ --data-parallel-backend ray \
--enforce-eager \ --enforce-eager \
--enable-expert-parallel \ --enable-expert-parallel \
--enable-eplb \ --enable-eplb \
--all2all-backend pplx \ --all2all-backend pplx \
--num-redundant-experts $REDUNDANT_EXPERTS \ --num-redundant-experts "$REDUNDANT_EXPERTS" \
--trust-remote-code \ --trust-remote-code \
--host $HOST \ --host "$HOST" \
--port $PORT --port "$PORT"
@@ -57,8 +57,7 @@ case "$subcommand" in
# Retry until the worker node connects to the head node or the timeout expires. # Retry until the worker node connects to the head node or the timeout expires.
for (( i=0; i < $ray_init_timeout; i+=5 )); do for (( i=0; i < $ray_init_timeout; i+=5 )); do
ray start --address=$ray_address:$ray_port --block "${start_params[@]}" if ray start --address="$ray_address":"$ray_port" --block "${start_params[@]}"; then
if [ $? -eq 0 ]; then
echo "Worker: Ray runtime started with head address $ray_address:$ray_port" echo "Worker: Ray runtime started with head address $ray_address:$ray_port"
exit 0 exit 0
fi fi
@@ -95,12 +94,12 @@ case "$subcommand" in
fi fi
# Start the Ray head node. # Start the Ray head node.
ray start --head --port=$ray_port "${start_params[@]}" ray start --head --port="$ray_port" "${start_params[@]}"
# Poll Ray until every worker node is active. # Poll Ray until every worker node is active.
for (( i=0; i < $ray_init_timeout; i+=5 )); do for (( i=0; i < $ray_init_timeout; i+=5 )); do
active_nodes=`python3 -c 'import ray; ray.init(); print(sum(node["Alive"] for node in ray.nodes()))'` active_nodes=$(python3 -c 'import ray; ray.init(); print(sum(node["Alive"] for node in ray.nodes()))')
if [ $active_nodes -eq $ray_cluster_size ]; then if [ "$active_nodes" -eq "$ray_cluster_size" ]; then
echo "All ray workers are active and the ray cluster is initialized successfully." echo "All ray workers are active and the ray cluster is initialized successfully."
exit 0 exit 0
fi fi
@@ -22,11 +22,10 @@ check_hf_token() {
check_num_gpus() { check_num_gpus() {
# can you check if the number of GPUs are >=2 via nvidia-smi/rocm-smi? # can you check if the number of GPUs are >=2 via nvidia-smi/rocm-smi?
which rocm-smi > /dev/null 2>&1 if ! which rocm-smi > /dev/null 2>&1; then
if [ $? -ne 0 ]; then
num_gpus=$(nvidia-smi --query-gpu=name --format=csv,noheader | wc -l) num_gpus=$(nvidia-smi --query-gpu=name --format=csv,noheader | wc -l)
else else
num_gpus=$(rocm-smi --showid | grep Instinct | wc -l) num_gpus=$(rocm-smi --showid | grep -c Instinct)
fi fi
if [ "$num_gpus" -lt 2 ]; then if [ "$num_gpus" -lt 2 ]; then
@@ -39,8 +38,7 @@ check_num_gpus() {
ensure_python_library_installed() { ensure_python_library_installed() {
echo "Checking if $1 is installed..." echo "Checking if $1 is installed..."
python3 -c "import $1" > /dev/null 2>&1 if ! python3 -c "import $1" > /dev/null 2>&1; then
if [ $? -ne 0 ]; then
if [ "$1" == "nixl" ]; then if [ "$1" == "nixl" ]; then
echo "$1 is not installed. Please refer to https://github.com/ai-dynamo/nixl for installation." echo "$1 is not installed. Please refer to https://github.com/ai-dynamo/nixl for installation."
else else
@@ -102,12 +100,12 @@ main() {
bash disagg_vllm_launcher.sh prefiller \ bash disagg_vllm_launcher.sh prefiller \
> >(tee prefiller.log) 2>&1 & > >(tee prefiller.log) 2>&1 &
prefiller_pid=$! prefiller_pid=$!
PIDS+=($prefiller_pid) PIDS+=("$prefiller_pid")
bash disagg_vllm_launcher.sh decoder \ bash disagg_vllm_launcher.sh decoder \
> >(tee decoder.log) 2>&1 & > >(tee decoder.log) 2>&1 &
decoder_pid=$! decoder_pid=$!
PIDS+=($decoder_pid) PIDS+=("$decoder_pid")
python3 disagg_proxy_server.py \ python3 disagg_proxy_server.py \
--host localhost \ --host localhost \
@@ -118,7 +116,7 @@ main() {
--decoder-port 8200 \ --decoder-port 8200 \
> >(tee proxy.log) 2>&1 & > >(tee proxy.log) 2>&1 &
proxy_pid=$! proxy_pid=$!
PIDS+=($proxy_pid) PIDS+=("$proxy_pid")
wait_for_server 8100 wait_for_server 8100
wait_for_server 8200 wait_for_server 8200
@@ -128,7 +126,7 @@ main() {
# begin benchmark # begin benchmark
cd ../../../../benchmarks/ cd ../../../../benchmarks/
vllm bench serve --port 9000 --seed $(date +%s) \ vllm bench serve --port 9000 --seed "$(date +%s)" \
--model meta-llama/Llama-3.1-8B-Instruct \ --model meta-llama/Llama-3.1-8B-Instruct \
--dataset-name random --random-input-len 7500 --random-output-len 200 \ --dataset-name random --random-input-len 7500 --random-output-len 200 \
--num-prompts 200 --burstiness 100 --request-rate 3.6 | tee benchmark.log --num-prompts 200 --burstiness 100 --request-rate 3.6 | tee benchmark.log
@@ -34,7 +34,7 @@ if [[ $1 == "prefiller" ]]; then
VLLM_ENABLE_V1_MULTIPROCESSING=1 \ VLLM_ENABLE_V1_MULTIPROCESSING=1 \
VLLM_WORKER_MULTIPROC_METHOD=spawn \ VLLM_WORKER_MULTIPROC_METHOD=spawn \
CUDA_VISIBLE_DEVICES=0 \ CUDA_VISIBLE_DEVICES=0 \
vllm serve $MODEL \ vllm serve "$MODEL" \
--port 8100 \ --port 8100 \
--enforce-eager \ --enforce-eager \
--kv-transfer-config \ --kv-transfer-config \
@@ -51,7 +51,7 @@ elif [[ $1 == "decoder" ]]; then
VLLM_ENABLE_V1_MULTIPROCESSING=1 \ VLLM_ENABLE_V1_MULTIPROCESSING=1 \
VLLM_WORKER_MULTIPROC_METHOD=spawn \ VLLM_WORKER_MULTIPROC_METHOD=spawn \
CUDA_VISIBLE_DEVICES=1 \ CUDA_VISIBLE_DEVICES=1 \
vllm serve $MODEL \ vllm serve "$MODEL" \
--port 8200 \ --port 8200 \
--enforce-eager \ --enforce-eager \
--kv-transfer-config \ --kv-transfer-config \
@@ -103,7 +103,7 @@ vllm serve "$MODEL_NAME" \
--tensor-parallel-size "$GPU_COUNT" \ --tensor-parallel-size "$GPU_COUNT" \
--enforce-eager \ --enforce-eager \
--pooler-config "$POOLER_CONFIG" \ --pooler-config "$POOLER_CONFIG" \
--served-model-name ${MODEL_CODE} \ --served-model-name "${MODEL_CODE}" \
--api-key "$API_KEY" \ --api-key "$API_KEY" \
--trust-remote-code \ --trust-remote-code \
--port "$PORT" \ --port "$PORT" \
@@ -20,13 +20,15 @@ def main():
torch.set_default_dtype(torch.float16) torch.set_default_dtype(torch.float16)
image_url = "https://huggingface.co/christian-pinto/Prithvi-EO-2.0-300M-TL-VLLM/resolve/main/valencia_example_2024-10-26.tiff" # noqa: E501 image_url = "https://huggingface.co/christian-pinto/Prithvi-EO-2.0-300M-TL-VLLM/resolve/main/valencia_example_2024-10-26.tiff" # noqa: E501
img_prompt = dict( img_data = dict(
data=image_url, data=image_url,
data_format="url", data_format="url",
image_format="tiff", image_format="tiff",
out_data_format="b64_json", out_data_format="b64_json",
) )
prompt = dict(data=img_data)
llm = LLM( llm = LLM(
model="ibm-nasa-geospatial/Prithvi-EO-2.0-300M-TL-Sen1Floods11", model="ibm-nasa-geospatial/Prithvi-EO-2.0-300M-TL-Sen1Floods11",
skip_tokenizer_init=True, skip_tokenizer_init=True,
@@ -41,7 +43,7 @@ def main():
enable_mm_embeds=True, enable_mm_embeds=True,
) )
pooler_output = llm.encode(img_prompt, pooling_task="plugin") pooler_output = llm.encode(prompt, pooling_task="plugin")
output = pooler_output[0].outputs output = pooler_output[0].outputs
print(output) print(output)
+1 -1
View File
@@ -27,7 +27,7 @@ mistral_common[image,audio] >= 1.9.1 # required for voxtral test
num2words # required for smolvlm test num2words # required for smolvlm test
opencv-python-headless >= 4.13.0 # required for video test opencv-python-headless >= 4.13.0 # required for video test
datamodel_code_generator # required for minicpm3 test datamodel_code_generator # required for minicpm3 test
lm-eval[api]>=0.4.9.2 # required for model evaluation test lm-eval[api]>=0.4.11 # required for model evaluation test
mteb>=1.38.11, <2 # required for mteb test mteb>=1.38.11, <2 # required for mteb test
transformers==4.57.5 transformers==4.57.5
tokenizers==0.22.0 tokenizers==0.22.0
+7 -3
View File
@@ -58,7 +58,7 @@ schemathesis==3.39.15
# OpenAI schema test # OpenAI schema test
# Evaluation and benchmarking # Evaluation and benchmarking
lm-eval[api]==0.4.9.2 lm-eval[api]==0.4.11
jiwer==4.0.0 jiwer==4.0.0
# Required for multiprocessed tests that use spawn method, Datasets and Evaluate Test # Required for multiprocessed tests that use spawn method, Datasets and Evaluate Test
@@ -67,8 +67,6 @@ multiprocess==0.70.16
# Required for v1/metrics/test_engine_logger_apis.py # Required for v1/metrics/test_engine_logger_apis.py
ray[cgraph,default]>=2.48.0 ray[cgraph,default]>=2.48.0
# Plugins test
terratorch @ git+https://github.com/IBM/terratorch.git@07184fcf91a1324f831ff521dd238d97fe350e3e
torchgeo==0.7.0 torchgeo==0.7.0
# via terratorch # via terratorch
# MTEB Benchmark Test # MTEB Benchmark Test
@@ -98,3 +96,9 @@ transformers==4.57.3
huggingface-hub==0.36.2 huggingface-hub==0.36.2
# Pin Mistral Common # Pin Mistral Common
mistral-common[image,audio]==1.9.1 mistral-common[image,audio]==1.9.1
# Required for Prithvi tests
terratorch==1.2.2
# Required for Prithvi tests
segmentation-models-pytorch==0.5.0
# Required for Prithvi tests
imagehash==4.3.2
+1 -1
View File
@@ -35,7 +35,7 @@ num2words # required for smolvlm test
open_clip_torch==2.32.0 # Required for nemotron_vl test, Nemotron Parse in test_common.py open_clip_torch==2.32.0 # Required for nemotron_vl test, Nemotron Parse in test_common.py
opencv-python-headless >= 4.13.0 # required for video test opencv-python-headless >= 4.13.0 # required for video test
datamodel_code_generator # required for minicpm3 test datamodel_code_generator # required for minicpm3 test
lm-eval[api]>=0.4.9.2 # required for model evaluation test lm-eval[api]>=0.4.11 # required for model evaluation test
mteb[bm25s]>=2, <3 # required for mteb test mteb[bm25s]>=2, <3 # required for mteb test
transformers==4.57.5 transformers==4.57.5
tokenizers==0.22.0 tokenizers==0.22.0
+21 -33
View File
@@ -1,13 +1,11 @@
# This file was autogenerated by uv via the following command: # This file was autogenerated by uv via the following command:
# uv pip compile requirements/test.in -o requirements/test.txt --index-strategy unsafe-best-match --torch-backend cu129 --python-platform x86_64-manylinux_2_28 --python-version 3.12 # uv pip compile requirements/test.in -o requirements/test.txt --index-strategy unsafe-best-match --torch-backend cu128 --python-platform x86_64-manylinux_2_28 --python-version 3.12
absl-py==2.1.0 absl-py==2.1.0
# via # via
# rouge-score # rouge-score
# tensorboard # tensorboard
accelerate==1.0.1 accelerate==1.0.1
# via # via peft
# lm-eval
# peft
aenum==3.1.16 aenum==3.1.16
# via lightly # via lightly
affine==2.4.0 affine==2.4.0
@@ -138,7 +136,6 @@ colorama==0.4.6
# perceptron # perceptron
# sacrebleu # sacrebleu
# schemathesis # schemathesis
# tqdm-multiprocess
colorful==0.5.6 colorful==0.5.6
# via ray # via ray
colorlog==6.10.1 colorlog==6.10.1
@@ -383,6 +380,7 @@ jinja2==3.1.6
# via # via
# datamodel-code-generator # datamodel-code-generator
# genai-perf # genai-perf
# lm-eval
# torch # torch
jiwer==3.0.5 jiwer==3.0.5
# via -r requirements/test.in # via -r requirements/test.in
@@ -448,7 +446,7 @@ lightning-utilities==0.14.3
# torchmetrics # torchmetrics
llvmlite==0.44.0 llvmlite==0.44.0
# via numba # via numba
lm-eval==0.4.9.2 lm-eval==0.4.11
# via -r requirements/test.in # via -r requirements/test.in
lxml==5.3.0 lxml==5.3.0
# via # via
@@ -513,8 +511,6 @@ numba==0.61.2
# via # via
# -r requirements/test.in # -r requirements/test.in
# librosa # librosa
numexpr==2.10.1
# via lm-eval
numpy==2.2.6 numpy==2.2.6
# via # via
# -r requirements/test.in # -r requirements/test.in
@@ -540,11 +536,11 @@ numpy==2.2.6
# librosa # librosa
# lightly # lightly
# lightly-utils # lightly-utils
# lm-eval
# matplotlib # matplotlib
# mistral-common # mistral-common
# mteb # mteb
# numba # numba
# numexpr
# opencv-python-headless # opencv-python-headless
# optuna # optuna
# pandas # pandas
@@ -578,28 +574,28 @@ numpy==2.2.6
# tritonclient # tritonclient
# vocos # vocos
# xarray # xarray
nvidia-cublas-cu12==12.9.1.4 nvidia-cublas-cu12==12.8.4.1
# via # via
# nvidia-cudnn-cu12 # nvidia-cudnn-cu12
# nvidia-cusolver-cu12 # nvidia-cusolver-cu12
# torch # torch
nvidia-cuda-cupti-cu12==12.9.79 nvidia-cuda-cupti-cu12==12.8.90
# via torch # via torch
nvidia-cuda-nvrtc-cu12==12.9.86 nvidia-cuda-nvrtc-cu12==12.8.93
# via torch # via torch
nvidia-cuda-runtime-cu12==12.9.79 nvidia-cuda-runtime-cu12==12.8.90
# via torch # via torch
nvidia-cudnn-cu12==9.10.2.21 nvidia-cudnn-cu12==9.10.2.21
# via torch # via torch
nvidia-cufft-cu12==11.4.1.4 nvidia-cufft-cu12==11.3.3.83
# via torch # via torch
nvidia-cufile-cu12==1.14.1.1 nvidia-cufile-cu12==1.13.1.3
# via torch # via torch
nvidia-curand-cu12==10.3.10.19 nvidia-curand-cu12==10.3.9.90
# via torch # via torch
nvidia-cusolver-cu12==11.7.5.82 nvidia-cusolver-cu12==11.7.3.90
# via torch # via torch
nvidia-cusparse-cu12==12.5.10.65 nvidia-cusparse-cu12==12.5.8.93
# via # via
# nvidia-cusolver-cu12 # nvidia-cusolver-cu12
# torch # torch
@@ -607,7 +603,7 @@ nvidia-cusparselt-cu12==0.7.1
# via torch # via torch
nvidia-nccl-cu12==2.27.5 nvidia-nccl-cu12==2.27.5
# via torch # via torch
nvidia-nvjitlink-cu12==12.9.86 nvidia-nvjitlink-cu12==12.8.93
# via # via
# nvidia-cufft-cu12 # nvidia-cufft-cu12
# nvidia-cusolver-cu12 # nvidia-cusolver-cu12
@@ -615,7 +611,7 @@ nvidia-nvjitlink-cu12==12.9.86
# torch # torch
nvidia-nvshmem-cu12==3.4.5 nvidia-nvshmem-cu12==3.4.5
# via torch # via torch
nvidia-nvtx-cu12==12.9.79 nvidia-nvtx-cu12==12.8.90
# via torch # via torch
omegaconf==2.3.0 omegaconf==2.3.0
# via # via
@@ -707,9 +703,7 @@ pathvalidate==3.2.1
patsy==1.0.1 patsy==1.0.1
# via statsmodels # via statsmodels
peft==0.16.0 peft==0.16.0
# via # via -r requirements/test.in
# -r requirements/test.in
# lm-eval
perceptron==0.1.4 perceptron==0.1.4
# via -r requirements/test.in # via -r requirements/test.in
perf-analyzer==0.1.0 perf-analyzer==0.1.0
@@ -792,8 +786,6 @@ pyasn1==0.6.1
# rsa # rsa
pyasn1-modules==0.4.2 pyasn1-modules==0.4.2
# via google-auth # via google-auth
pybind11==2.13.6
# via lm-eval
pycocotools==2.0.8 pycocotools==2.0.8
# via terratorch # via terratorch
pycountry==24.6.1 pycountry==24.6.1
@@ -1162,7 +1154,7 @@ tomli==2.2.1
# via schemathesis # via schemathesis
tomli-w==1.2.0 tomli-w==1.2.0
# via schemathesis # via schemathesis
torch==2.10.0+cu129 torch==2.10.0+cu128
# via # via
# -r requirements/test.in # -r requirements/test.in
# accelerate # accelerate
@@ -1171,7 +1163,6 @@ torch==2.10.0+cu129
# kornia # kornia
# lightly # lightly
# lightning # lightning
# lm-eval
# mteb # mteb
# open-clip-torch # open-clip-torch
# peft # peft
@@ -1188,7 +1179,7 @@ torch==2.10.0+cu129
# torchvision # torchvision
# vector-quantize-pytorch # vector-quantize-pytorch
# vocos # vocos
torchaudio==2.10.0+cu129 torchaudio==2.10.0+cu128
# via # via
# -r requirements/test.in # -r requirements/test.in
# encodec # encodec
@@ -1201,7 +1192,7 @@ torchmetrics==1.7.4
# pytorch-lightning # pytorch-lightning
# terratorch # terratorch
# torchgeo # torchgeo
torchvision==0.25.0+cu129 torchvision==0.25.0+cu128
# via # via
# -r requirements/test.in # -r requirements/test.in
# lightly # lightly
@@ -1229,15 +1220,11 @@ tqdm==4.67.3
# sentence-transformers # sentence-transformers
# tacoreader # tacoreader
# terratorch # terratorch
# tqdm-multiprocess
# transformers # transformers
tqdm-multiprocess==0.0.11
# via lm-eval
transformers==4.57.5 transformers==4.57.5
# via # via
# -r requirements/test.in # -r requirements/test.in
# genai-perf # genai-perf
# lm-eval
# peft # peft
# sentence-transformers # sentence-transformers
# transformers-stream-generator # transformers-stream-generator
@@ -1272,6 +1259,7 @@ typing-extensions==4.15.0
# librosa # librosa
# lightning # lightning
# lightning-utilities # lightning-utilities
# lm-eval
# mistral-common # mistral-common
# mteb # mteb
# opentelemetry-api # opentelemetry-api
@@ -229,7 +229,7 @@ def _compare_sp(
if chunked_prefill: if chunked_prefill:
common_args.append("--enable-chunked-prefill") common_args.append("--enable-chunked-prefill")
if eager_mode: if eager_mode:
common_args.append("--enforce-eager") common_args.append("-cc.cudagraph_mode=none")
if runner != "auto": if runner != "auto":
common_args.extend(["--runner", runner]) common_args.extend(["--runner", runner])
if trust_remote_code: if trust_remote_code:
@@ -145,6 +145,7 @@ async def test_request_cancellation(server: RemoteOpenAIServer):
model=MODEL_NAME, model=MODEL_NAME,
max_tokens=10000, max_tokens=10000,
extra_body={"min_tokens": 10000}, extra_body={"min_tokens": 10000},
temperature=0.0,
) )
) )
tasks.append(task) tasks.append(task)
@@ -163,7 +164,7 @@ async def test_request_cancellation(server: RemoteOpenAIServer):
# be able to respond to this one within the timeout # be able to respond to this one within the timeout
client = server.get_async_client(timeout=5) client = server.get_async_client(timeout=5)
response = await client.chat.completions.create( response = await client.chat.completions.create(
messages=chat_input, model=MODEL_NAME, max_tokens=10 messages=chat_input, model=MODEL_NAME, max_tokens=10, temperature=0.0
) )
assert len(response.choices) == 1 assert len(response.choices) == 1
@@ -17,6 +17,7 @@ from transformers import AutoTokenizer
from tests.conftest import LocalAssetServer from tests.conftest import LocalAssetServer
from tests.utils import RemoteOpenAIServer from tests.utils import RemoteOpenAIServer
from vllm import version from vllm import version
from vllm.utils.network_utils import get_open_port
MODELS = { MODELS = {
"text": "TinyLlama/TinyLlama-1.1B-Chat-v1.0", "text": "TinyLlama/TinyLlama-1.1B-Chat-v1.0",
@@ -315,14 +316,26 @@ async def test_abort_metrics_reset(
client.completions.create( client.completions.create(
model=model_name, model=model_name,
prompt=prompt_ids, prompt=prompt_ids,
max_tokens=100, # Long generation to give time to abort max_tokens=500, # Long generation to give time to abort
temperature=0.0, temperature=0.0,
) )
) )
tasks.append(task) tasks.append(task)
# Wait a bit for requests to start processing # Poll until we see running requests rather than using a fixed sleep,
await asyncio.sleep(0.5) # since generation speed varies across hardware.
try:
await _poll_until(
lambda: _get_running_metrics_from_api(server)[0] > 0,
timeout=10.0,
interval=0.1,
description="running_requests > 0",
)
except TimeoutError:
for task in tasks:
task.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
pytest.fail("Requests never appeared as running in metrics")
# Check that we have running requests # Check that we have running requests
running_requests, waiting_requests, kv_cache_usage = _get_running_metrics_from_api( running_requests, waiting_requests, kv_cache_usage = _get_running_metrics_from_api(
@@ -336,13 +349,15 @@ async def test_abort_metrics_reset(
# Cancel all tasks to abort the requests # Cancel all tasks to abort the requests
for task in tasks: for task in tasks:
task.cancel() task.cancel()
await asyncio.gather(*tasks, return_exceptions=True)
# Wait for cancellations to be processed # Poll until metrics reset rather than using a fixed sleep
await asyncio.sleep(1.0) await _poll_until(
lambda: _get_running_metrics_from_api(server) == (0, 0, 0),
# Check that metrics have reset to zero timeout=10.0,
response = requests.get(server.url_for("metrics")) interval=0.2,
assert response.status_code == HTTPStatus.OK description="gauge metrics back to zero",
)
# Verify running and waiting requests counts and KV cache usage are zero # Verify running and waiting requests counts and KV cache usage are zero
running_requests_after, waiting_requests_after, kv_cache_usage_after = ( running_requests_after, waiting_requests_after, kv_cache_usage_after = (
@@ -360,6 +375,18 @@ async def test_abort_metrics_reset(
) )
async def _poll_until(
predicate, *, timeout: float, interval: float = 0.5, description: str = "condition"
):
"""Poll until predicate() returns True, or raise TimeoutError."""
start = time.time()
while time.time() - start < timeout:
if predicate():
return
await asyncio.sleep(interval)
raise TimeoutError(f"Timed out after {timeout}s waiting for: {description}")
def _get_running_metrics_from_api(server: RemoteOpenAIServer): def _get_running_metrics_from_api(server: RemoteOpenAIServer):
"""Return (running_count, waiting_count, kv_cache_usage)""" """Return (running_count, waiting_count, kv_cache_usage)"""
@@ -399,7 +426,7 @@ def test_metrics_exist_run_batch():
input_batch = """{"custom_id": "request-0", "method": "POST", "url": "/v1/embeddings", "body": {"model": "intfloat/multilingual-e5-small", "input": "You are a helpful assistant."}}""" # noqa: E501 input_batch = """{"custom_id": "request-0", "method": "POST", "url": "/v1/embeddings", "body": {"model": "intfloat/multilingual-e5-small", "input": "You are a helpful assistant."}}""" # noqa: E501
base_url = "0.0.0.0" base_url = "0.0.0.0"
port = "8001" port = str(get_open_port())
server_url = f"http://{base_url}:{port}" server_url = f"http://{base_url}:{port}"
with ( with (
@@ -427,17 +454,32 @@ def test_metrics_exist_run_batch():
], ],
) )
def is_server_up(url): try:
def is_server_up(url):
try:
response = requests.get(url)
return response.status_code == 200
except requests.ConnectionError:
return False
start = time.time()
timeout = 120
while not is_server_up(server_url):
if proc.poll() is not None:
pytest.fail(
f"Batch process exited early with returncode={proc.returncode}"
)
if time.time() - start > timeout:
pytest.fail("Batch server did not start within timeout")
time.sleep(1)
response = requests.get(server_url + "/metrics")
assert response.status_code == HTTPStatus.OK
finally:
proc.terminate()
try: try:
response = requests.get(url) proc.wait(timeout=15)
return response.status_code == 200 except subprocess.TimeoutExpired:
except requests.ConnectionError: proc.kill()
return False proc.wait(timeout=5)
while not is_server_up(server_url):
time.sleep(1)
response = requests.get(server_url + "/metrics")
assert response.status_code == HTTPStatus.OK
proc.wait()
+6 -9
View File
@@ -195,18 +195,15 @@ def test_chat_batch_failure_cleanup(llm_for_failure_test):
valid_msg = [{"role": "user", "content": "Hello"}] valid_msg = [{"role": "user", "content": "Hello"}]
long_text = "This is a very long text to test the error " * 50 long_text = "This is a very long text to test the error " * 50
invalid_msg = [{"role": "user", "content": long_text}] invalid_msg = [{"role": "user", "content": long_text}]
batch_1 = [
valid_msg, batch_1 = [valid_msg, valid_msg, invalid_msg]
valid_msg, batch_2 = [valid_msg, valid_msg]
invalid_msg,
]
batch_2 = [
valid_msg,
valid_msg,
]
sampling_params = SamplingParams(temperature=0, max_tokens=10) sampling_params = SamplingParams(temperature=0, max_tokens=10)
with pytest.raises(ValueError, match="context length is only"): with pytest.raises(ValueError, match="context length is only"):
llm.chat(batch_1, sampling_params=sampling_params) llm.chat(batch_1, sampling_params=sampling_params)
assert llm.llm_engine.get_num_unfinished_requests() == 0
outputs_2 = llm.chat(batch_2, sampling_params=sampling_params) outputs_2 = llm.chat(batch_2, sampling_params=sampling_params)
assert len(outputs_2) == len(batch_2) assert len(outputs_2) == len(batch_2)
assert llm.llm_engine.get_num_unfinished_requests() == 0 assert llm.llm_engine.get_num_unfinished_requests() == 0
@@ -6,7 +6,6 @@ import time
import pytest import pytest
import pytest_asyncio import pytest_asyncio
import requests
from openai import BadRequestError, NotFoundError, OpenAI from openai import BadRequestError, NotFoundError, OpenAI
from openai_harmony import ( from openai_harmony import (
Message, Message,
@@ -513,11 +512,9 @@ async def test_code_interpreter(client: OpenAI, model_name: str):
def get_weather(latitude, longitude): def get_weather(latitude, longitude):
response = requests.get( # Return a static temperature value to avoid flaky SSL/network errors
f"https://api.open-meteo.com/v1/forecast?latitude={latitude}&longitude={longitude}&current=temperature_2m,wind_speed_10m&hourly=temperature_2m,relative_humidity_2m,wind_speed_10m" # noqa # from calling the external api.open-meteo.com API in CI.
) return 15.0
data = response.json()
return data["current"]["temperature_2m"]
def get_place_to_travel(): def get_place_to_travel():
@@ -172,19 +172,26 @@ async def test_mcp_tool_call(client: OpenAI, model_name: str):
assert response is not None assert response is not None
assert response.status == "completed" assert response.status == "completed"
assert response.output[0].type == "reasoning"
assert response.output[1].type == "mcp_call"
assert type(response.output[1].arguments) is str
assert type(response.output[1].output) is str
assert response.output[2].type == "reasoning"
# make sure the correct math is in the final output
assert response.output[3].type == "message"
assert any(s in response.output[3].content[0].text for s in ("56088", "56,088"))
# test raw input_messages / output_messages # The model may produce multiple reasoning/mcp_call rounds before the
# final message, so validate structurally rather than by exact index.
output_types = [o.type for o in response.output]
assert "reasoning" in output_types
mcp_calls = [o for o in response.output if o.type == "mcp_call"]
assert len(mcp_calls) >= 1
assert type(mcp_calls[0].arguments) is str
assert type(mcp_calls[0].output) is str
# The final output should be a message containing the correct answer
assert response.output[-1].type == "message"
assert any(s in response.output[-1].content[0].text for s in ("56088", "56,088"))
# Test raw input_messages / output_messages
assert len(response.input_messages) == 1 assert len(response.input_messages) == 1
assert len(response.output_messages) == 3 assert len(response.output_messages) >= 3
assert any(s in response.output_messages[2]["message"] for s in ("56088", "56,088")) assert any(
s in response.output_messages[-1]["message"] for s in ("56088", "56,088")
)
@pytest.mark.asyncio @pytest.mark.asyncio
@@ -151,6 +151,7 @@ def test_openapi_stateless(case: schemathesis.Case):
# requires a longer timeout # requires a longer timeout
("POST", "/v1/chat/completions"): LONG_TIMEOUT_SECONDS, ("POST", "/v1/chat/completions"): LONG_TIMEOUT_SECONDS,
("POST", "/v1/completions"): LONG_TIMEOUT_SECONDS, ("POST", "/v1/completions"): LONG_TIMEOUT_SECONDS,
("POST", "/v1/messages"): LONG_TIMEOUT_SECONDS,
}.get(key, DEFAULT_TIMEOUT_SECONDS) }.get(key, DEFAULT_TIMEOUT_SECONDS)
# No need to verify SSL certificate for localhost # No need to verify SSL certificate for localhost
+50 -10
View File
@@ -526,6 +526,7 @@ class MockModelConfig:
allowed_media_domains: list[str] | None = None allowed_media_domains: list[str] | None = None
encoder_config = None encoder_config = None
generation_config: str = "auto" generation_config: str = "auto"
override_generation_config: dict[str, Any] = field(default_factory=dict)
media_io_kwargs: dict[str, dict[str, Any]] = field(default_factory=dict) media_io_kwargs: dict[str, dict[str, Any]] = field(default_factory=dict)
skip_tokenizer_init: bool = False skip_tokenizer_init: bool = False
is_encoder_decoder: bool = False is_encoder_decoder: bool = False
@@ -651,12 +652,10 @@ async def test_serving_chat_should_set_correct_max_tokens():
assert mock_engine.generate.call_args.args[1].max_tokens == 10 assert mock_engine.generate.call_args.args[1].max_tokens == 10
# Setting server's max_tokens in the generation_config.json # Model author's generation_config.json sets max_tokens (auto, no override)
# lower than context_window - prompt_tokens # — should act as fallback only, not ceiling
mock_model_config = MockModelConfig() mock_model_config = MockModelConfig()
mock_model_config.diff_sampling_param = { mock_model_config.diff_sampling_param = {"max_tokens": 10}
"max_tokens": 10 # Setting server-side max_tokens limit
}
# Reinitialize the engine with new settings # Reinitialize the engine with new settings
mock_engine = MagicMock(spec=AsyncLLM) mock_engine = MagicMock(spec=AsyncLLM)
@@ -680,7 +679,50 @@ async def test_serving_chat_should_set_correct_max_tokens():
assert mock_engine.generate.call_args.args[1].max_tokens == 10 assert mock_engine.generate.call_args.args[1].max_tokens == 10
# Test Case 2: Request's max_tokens set higher than server accepts # Test Case 2: Request's max_tokens set higher than generation_config
# default so request-provided max_tokens takes precedence
req.max_tokens = 15
with suppress(Exception):
await serving_chat.create_chat_completion(req)
assert mock_engine.generate.call_args.args[1].max_tokens == 15
# Test Case 3: Request's max_tokens set lower than server accepts
req.max_tokens = 5
with suppress(Exception):
await serving_chat.create_chat_completion(req)
assert mock_engine.generate.call_args.args[1].max_tokens == 5
# User explicitly sets max_tokens via --override-generation-config
# — should act as a ceiling
mock_model_config = MockModelConfig()
mock_model_config.diff_sampling_param = {"max_tokens": 10}
mock_model_config.override_generation_config = {"max_new_tokens": 10}
mock_engine = MagicMock(spec=AsyncLLM)
mock_engine.errored = False
mock_engine.model_config = mock_model_config
mock_engine.input_processor = MagicMock()
mock_engine.io_processor = MagicMock()
mock_engine.renderer = _build_renderer(mock_engine.model_config)
serving_chat = _build_serving_chat(mock_engine)
# Test Case 3.1: No max_tokens — uses override as default
req = ChatCompletionRequest(
model=MODEL_NAME,
messages=[{"role": "user", "content": "what is 1+1?"}],
)
with suppress(Exception):
await serving_chat.create_chat_completion(req)
assert mock_engine.generate.call_args.args[1].max_tokens == 10
# Test Case 3.2: Request max_tokens higher — capped by user ceiling from override
req.max_tokens = 15 req.max_tokens = 15
with suppress(Exception): with suppress(Exception):
@@ -688,7 +730,7 @@ async def test_serving_chat_should_set_correct_max_tokens():
assert mock_engine.generate.call_args.args[1].max_tokens == 10 assert mock_engine.generate.call_args.args[1].max_tokens == 10
# Test Case 3: Request's max_tokens set lower than server accepts # Test Case 3.3: Request max_tokens lower — respected
req.max_tokens = 5 req.max_tokens = 5
with suppress(Exception): with suppress(Exception):
@@ -699,9 +741,7 @@ async def test_serving_chat_should_set_correct_max_tokens():
# Setting server's max_tokens in the generation_config.json # Setting server's max_tokens in the generation_config.json
# higher than context_window - prompt_tokens # higher than context_window - prompt_tokens
mock_model_config = MockModelConfig() mock_model_config = MockModelConfig()
mock_model_config.diff_sampling_param = { mock_model_config.diff_sampling_param = {"max_tokens": 200}
"max_tokens": 200 # Setting server-side max_tokens limit
}
# Reinitialize the engine with new settings # Reinitialize the engine with new settings
mock_engine = MagicMock(spec=AsyncLLM) mock_engine = MagicMock(spec=AsyncLLM)
+73 -1
View File
@@ -1,6 +1,7 @@
# SPDX-License-Identifier: Apache-2.0 # SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project # SPDX-FileCopyrightText: Copyright contributors to the vLLM project
from vllm.entrypoints.utils import sanitize_message
from vllm.entrypoints.utils import get_max_tokens, sanitize_message
def test_sanitize_message(): def test_sanitize_message():
@@ -8,3 +9,74 @@ def test_sanitize_message():
sanitize_message("<_io.BytesIO object at 0x7a95e299e750>") sanitize_message("<_io.BytesIO object at 0x7a95e299e750>")
== "<_io.BytesIO object>" == "<_io.BytesIO object>"
) )
class TestGetMaxTokens:
"""Tests for get_max_tokens() to ensure generation_config's max_tokens
acts as a default when from model author, and as a ceiling when
explicitly set by the user."""
def test_default_sampling_params_used_when_no_request_max_tokens(self):
"""When user doesn't specify max_tokens, generation_config default
should apply."""
result = get_max_tokens(
max_model_len=24000,
max_tokens=None,
input_length=100,
default_sampling_params={"max_tokens": 2048},
)
assert result == 2048
def test_request_max_tokens_not_capped_by_default_sampling_params(self):
"""When user specifies max_tokens in request, model author's
generation_config max_tokens must NOT cap it (fixes #34005)."""
result = get_max_tokens(
max_model_len=24000,
max_tokens=5000,
input_length=100,
default_sampling_params={"max_tokens": 2048},
)
assert result == 5000
def test_override_max_tokens_caps_request(self):
"""When user explicitly sets max_tokens, it acts as a ceiling."""
result = get_max_tokens(
max_model_len=24000,
max_tokens=5000,
input_length=100,
default_sampling_params={"max_tokens": 2048},
override_max_tokens=2048,
)
assert result == 2048
def test_override_max_tokens_used_as_default(self):
"""When no request max_tokens, override still applies as default."""
result = get_max_tokens(
max_model_len=24000,
max_tokens=None,
input_length=100,
default_sampling_params={"max_tokens": 2048},
override_max_tokens=2048,
)
assert result == 2048
def test_max_model_len_still_caps_output(self):
"""max_model_len - input_length is always the hard ceiling."""
result = get_max_tokens(
max_model_len=3000,
max_tokens=5000,
input_length=100,
default_sampling_params={"max_tokens": 2048},
)
assert result == 2900 # 3000 - 100
def test_request_max_tokens_smaller_than_default(self):
"""When user explicitly requests fewer tokens than gen_config default,
that should be respected."""
result = get_max_tokens(
max_model_len=24000,
max_tokens=512,
input_length=100,
default_sampling_params={"max_tokens": 2048},
)
assert result == 512
+16 -2
View File
@@ -17,7 +17,7 @@ from vllm.platforms import current_platform
from vllm.platforms.cpu import CpuPlatform from vllm.platforms.cpu import CpuPlatform
from vllm.platforms.cuda import CudaPlatform from vllm.platforms.cuda import CudaPlatform
from vllm.platforms.rocm import RocmPlatform from vllm.platforms.rocm import RocmPlatform
from vllm.utils.torch_utils import set_random_seed from vllm.utils.torch_utils import set_default_torch_dtype, set_random_seed
from vllm.v1.attention.backends.registry import AttentionBackendEnum from vllm.v1.attention.backends.registry import AttentionBackendEnum
from vllm.v1.attention.selector import _cached_get_attn_backend from vllm.v1.attention.selector import _cached_get_attn_backend
@@ -71,6 +71,15 @@ def test_mha_attn_platform(default_vllm_config, device: str):
attn = MMEncoderAttention(16, 72, scale=1) attn = MMEncoderAttention(16, 72, scale=1)
assert attn.attn_backend == AttentionBackendEnum.FLASH_ATTN assert attn.attn_backend == AttentionBackendEnum.FLASH_ATTN
# Test CUDA with head_size=72 (not divisible by 32)
# - should use vLLM's FlashAttention
with (
patch("vllm.model_executor.models.vision.current_platform", CudaPlatform()),
set_default_torch_dtype(torch.float32),
):
attn = MMEncoderAttention(16, 72, scale=1)
assert attn.attn_backend == AttentionBackendEnum.TRITON_ATTN
def ref_attention( def ref_attention(
query: torch.Tensor, query: torch.Tensor,
@@ -153,7 +162,12 @@ def test_mha_attn_forward(
v, v,
scale=scale, scale=scale,
).reshape(batch_size, seq_len, num_heads * head_size) ).reshape(batch_size, seq_len, num_heads * head_size)
torch.testing.assert_close(output, ref_output) tol_kwargs = (
dict(rtol=1e-3, atol=1e-3)
if attn.attn_backend == AttentionBackendEnum.TRITON_ATTN
else {}
)
torch.testing.assert_close(output, ref_output, **tol_kwargs)
@pytest.mark.parametrize("var_seq_len", VAR_SEQ_LENS) @pytest.mark.parametrize("var_seq_len", VAR_SEQ_LENS)
+49
View File
@@ -396,6 +396,55 @@ def test_fused_moe(
) )
def test_fused_moe_int64_overflow(monkeypatch, workspace_init):
"""Regression test for int32 overflow in stride*offset products.
When chunking is disabled and M is large, stride_cm * offs_token can
exceed int32 max. Verifies the offs_token int64 cast (fix for #34413)
prevents overflow and produces correct results.
Reproduces the scenario from PR #34279.
"""
# ~12 GB GPU memory needed for intermediate caches
free_mem = torch.cuda.mem_get_info()[0]
if free_mem < 12 * 1024**3:
pytest.skip("Insufficient GPU memory for overflow test")
set_random_seed(7)
m, n, k, e, topk = 100000, 2048, 1024, 8, 6
dtype = torch.bfloat16
# Disable chunking to expose the overflow-prone code path
monkeypatch.setenv("VLLM_FUSED_MOE_CHUNK_SIZE", "10000000")
a = torch.randn((m, k), device="cuda", dtype=dtype) / 10
w1 = torch.randn((e, 2 * n, k), device="cuda", dtype=dtype) / 10
w2 = torch.randn((e, k, n), device="cuda", dtype=dtype) / 10
score = torch.randn((m, e), device="cuda", dtype=dtype)
# Verify the test exercises the overflow condition:
# C has shape (M, topk, N) where N = w1.size(1) = 2*n
# stride_cm = C.stride(1) = N, max offs_token = M * topk
# Product must exceed int32 max for this test to be meaningful
N = w1.size(1)
assert N * m * topk > 2**31 - 1, "Test params don't trigger int32 overflow"
fused_moe_fn = functools.partial(fused_moe, renormalize=False)
with set_current_vllm_config(vllm_config):
run_moe_test(
torch_moe,
fused_moe_fn,
a=a,
w1=w1,
w2=w2,
score=score,
topk=topk,
global_num_experts=e,
)
@pytest.mark.parametrize("m,n,k", FUSED_MOE_MNK_FACTORS_SMALL_M) @pytest.mark.parametrize("m,n,k", FUSED_MOE_MNK_FACTORS_SMALL_M)
@pytest.mark.parametrize("e", NUM_EXPERTS_LARGE) @pytest.mark.parametrize("e", NUM_EXPERTS_LARGE)
@pytest.mark.parametrize("topk", TOP_KS_SMALL) @pytest.mark.parametrize("topk", TOP_KS_SMALL)
@@ -117,7 +117,6 @@ def check_model_available(model: str) -> None:
@pytest.mark.parametrize("dtype", ["half", "float"]) @pytest.mark.parametrize("dtype", ["half", "float"])
@pytest.mark.parametrize("num_logprobs", [5]) @pytest.mark.parametrize("num_logprobs", [5])
@pytest.mark.parametrize("enforce_eager", [True, False]) @pytest.mark.parametrize("enforce_eager", [True, False])
@create_new_process_for_each_test("spawn")
def test_models( def test_models(
hf_runner, hf_runner,
vllm_runner, vllm_runner,
@@ -102,13 +102,13 @@ def glmasr_patch_mm_data(mm_data: MultiModalDataDict) -> MultiModalDataDict:
# incorrect token ids. So we need use `add_special_tokens=False` here # incorrect token ids. So we need use `add_special_tokens=False` here
# to leave bos_token to be added by the processor. # to leave bos_token to be added by the processor.
_ADD_SPECIAL_TOKENS_OVERRIDES = { _ADD_SPECIAL_TOKENS_OVERRIDES = {
"lfm2_vl": False,
"nemotron_parse": False, "nemotron_parse": False,
"ovis": False, "ovis": False,
"ovis2_5": False, "ovis2_5": False,
"paligemma": False, "paligemma": False,
"ultravox": False, "ultravox": False,
"whisper": False, "whisper": False,
"lfm2_vl": False,
} }
_IGNORE_MM_KEYS = { _IGNORE_MM_KEYS = {
@@ -450,6 +450,8 @@ def test_processing_correctness(
num_batches: int, num_batches: int,
simplify_rate: float, simplify_rate: float,
): ):
if model_id == "allendou/Fun-ASR-Nano-2512-vllm":
pytest.skip("Cached audio `input_features` not matched. Fix later.")
if model_id == "google/gemma-3n-E2B-it": if model_id == "google/gemma-3n-E2B-it":
pytest.skip("Fix later") pytest.skip("Fix later")
if model_id == "OpenGVLab/InternVL2-2B": if model_id == "OpenGVLab/InternVL2-2B":
@@ -468,9 +470,6 @@ def test_processing_correctness(
"correctness test as is. Let's revisit adapting this " "correctness test as is. Let's revisit adapting this "
"test once more realtime models exist." "test once more realtime models exist."
) )
if model_id == "internlm/Intern-S1-Pro":
# FIXME(Isotr0py): Fix later.
pytest.skip("Tokenization issue. Fix later")
_test_processing_correctness( _test_processing_correctness(
model_id, model_id,
@@ -490,8 +489,9 @@ def _assert_inputs_equal(
if ignore_mm_keys is None: if ignore_mm_keys is None:
ignore_mm_keys = set() ignore_mm_keys = set()
a_rest = {k: v for k, v in a.items() if k != "mm_kwargs"} ignore_prompt_keys = ("prompt", "mm_kwargs")
b_rest = {k: v for k, v in b.items() if k != "mm_kwargs"} a_rest = {k: v for k, v in a.items() if k not in ignore_prompt_keys}
b_rest = {k: v for k, v in b.items() if k not in ignore_prompt_keys}
assert a_rest == b_rest, msg assert a_rest == b_rest, msg
@@ -160,9 +160,6 @@ def test_model_tensor_schema(model_id: str):
pytest.skip( pytest.skip(
"Kimi-K2.5's offline inference has issues about vision chunks. Fix later." "Kimi-K2.5's offline inference has issues about vision chunks. Fix later."
) )
if model_id == "internlm/Intern-S1-Pro":
# FIXME(Isotr0py): Fix later.
pytest.skip("Intern-S1-Pro has issue to pass the test.")
model_info = HF_EXAMPLE_MODELS.find_hf_info(model_id) model_info = HF_EXAMPLE_MODELS.find_hf_info(model_id)
model_info.check_available_online(on_fail="skip") model_info.check_available_online(on_fail="skip")
-2
View File
@@ -730,7 +730,6 @@ _MULTIMODAL_EXAMPLE_MODELS = {
), ),
"FunASRForConditionalGeneration": _HfExamplesInfo( "FunASRForConditionalGeneration": _HfExamplesInfo(
"allendou/Fun-ASR-Nano-2512-vllm", "allendou/Fun-ASR-Nano-2512-vllm",
is_available_online=False,
), ),
"FunAudioChatForConditionalGeneration": _HfExamplesInfo( "FunAudioChatForConditionalGeneration": _HfExamplesInfo(
"funaudiochat", is_available_online=False "funaudiochat", is_available_online=False
@@ -755,7 +754,6 @@ _MULTIMODAL_EXAMPLE_MODELS = {
"Glm4vMoeForConditionalGeneration": _HfExamplesInfo("zai-org/GLM-4.5V"), "Glm4vMoeForConditionalGeneration": _HfExamplesInfo("zai-org/GLM-4.5V"),
"GlmOcrForConditionalGeneration": _HfExamplesInfo( "GlmOcrForConditionalGeneration": _HfExamplesInfo(
"zai-org/GLM-OCR", "zai-org/GLM-OCR",
is_available_online=False,
min_transformers_version="5.1.0", min_transformers_version="5.1.0",
), ),
"H2OVLChatModel": _HfExamplesInfo( "H2OVLChatModel": _HfExamplesInfo(
@@ -120,13 +120,15 @@ async def test_prithvi_mae_plugin_online(
def test_prithvi_mae_plugin_offline( def test_prithvi_mae_plugin_offline(
vllm_runner, model_name: str, image_url: str | dict, plugin: str, expected_hash: str vllm_runner, model_name: str, image_url: str | dict, plugin: str, expected_hash: str
): ):
img_prompt = dict( img_data = dict(
data=image_url, data=image_url,
data_format="url", data_format="url",
image_format="tiff", image_format="tiff",
out_data_format="b64_json", out_data_format="b64_json",
) )
prompt = dict(data=img_data)
with vllm_runner( with vllm_runner(
model_name, model_name,
runner="pooling", runner="pooling",
@@ -139,7 +141,7 @@ def test_prithvi_mae_plugin_offline(
io_processor_plugin=plugin, io_processor_plugin=plugin,
default_torch_num_threads=1, default_torch_num_threads=1,
) as llm_runner: ) as llm_runner:
pooler_output = llm_runner.get_llm().encode(img_prompt, pooling_task="plugin") pooler_output = llm_runner.get_llm().encode(prompt, pooling_task="plugin")
output = pooler_output[0].outputs output = pooler_output[0].outputs
# verify the output is formatted as expected for this plugin # verify the output is formatted as expected for this plugin
+2 -2
View File
@@ -93,14 +93,14 @@ def _build_renderer(
def _preprocess_prompt( def _preprocess_prompt(
mdoel_config: ModelConfig, model_config: ModelConfig,
prompt_or_prompts: SingletonPrompt | bytes | Sequence[SingletonPrompt | bytes], prompt_or_prompts: SingletonPrompt | bytes | Sequence[SingletonPrompt | bytes],
): ):
return [ return [
( (
prompt prompt
if isinstance(prompt, bytes) if isinstance(prompt, bytes)
else parse_model_prompt(mdoel_config, prompt) else parse_model_prompt(model_config, prompt)
) )
for prompt in prompt_to_seq(prompt_or_prompts) for prompt in prompt_to_seq(prompt_or_prompts)
] ]
@@ -0,0 +1,165 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
import pytest
from vllm.assets.image import ImageAsset
from vllm.assets.video import VideoAsset
from vllm.config import CacheConfig, ModelConfig, VllmConfig
from vllm.renderers.hf import HfRenderer
from vllm.tokenizers.registry import tokenizer_args_from_config
cherry_pil_image = ImageAsset("cherry_blossom").pil_image
stop_pil_image = ImageAsset("stop_sign").pil_image
baby_reading_np_ndarrays = VideoAsset("baby_reading").np_ndarrays
def _build_renderer(
*, mm_cache_gb: float = 4.0, enable_prefix_caching: bool = True
) -> HfRenderer:
model_config = ModelConfig(
model="Qwen/Qwen2.5-VL-3B-Instruct",
max_model_len=128,
mm_processor_cache_gb=mm_cache_gb,
)
vllm_config = VllmConfig(
model_config=model_config,
cache_config=CacheConfig(enable_prefix_caching=enable_prefix_caching),
)
_, tokenizer_name, _, kwargs = tokenizer_args_from_config(model_config)
return HfRenderer.from_config(
vllm_config,
tokenizer_kwargs={**kwargs, "tokenizer_name": tokenizer_name},
)
def test_multi_modal_uuids_length_mismatch_raises():
renderer = _build_renderer()
mm_data = {"image": [cherry_pil_image, stop_pil_image]}
# Mismatch: 2 items but only 1 uuid provided
mm_uuids = {"image": ["hash_cherry"]}
mm_processor = renderer.get_mm_processor()
mm_items = mm_processor.info.parse_mm_data(mm_data)
with pytest.raises(ValueError, match="must have same length as"):
renderer._process_mm_uuids(mm_data, mm_items, mm_uuids, "req-1")
def test_multi_modal_uuids_missing_modality_raises():
renderer = _build_renderer()
mm_data = {
"image": [cherry_pil_image],
"video": None,
}
# Only image uuids provided; video missing should raise
mm_uuids = {"image": ["hash_cherry"]}
mm_processor = renderer.get_mm_processor()
mm_items = mm_processor.info.parse_mm_data(mm_data)
with pytest.raises(ValueError, match="is empty but .* is missing"):
renderer._process_mm_uuids(mm_data, mm_items, mm_uuids, "req-2")
@pytest.mark.parametrize(
"mm_cache_gb, enable_prefix_caching",
[
(4.0, True), # default behavior
(4.0, False), # prefix caching disabled
(0.0, True), # processor cache disabled
],
)
def test_multi_modal_uuids_accepts_none_and_passes_through(
monkeypatch, mm_cache_gb: float, enable_prefix_caching: bool
):
renderer = _build_renderer(
mm_cache_gb=mm_cache_gb,
enable_prefix_caching=enable_prefix_caching,
)
mm_data = {
"image": [cherry_pil_image, stop_pil_image],
"video": baby_reading_np_ndarrays,
}
# Use a consistent two-image scenario across all configurations
mm_uuids = {"image": [None, "hash_stop"], "video": None}
mm_processor = renderer.get_mm_processor()
mm_items = mm_processor.info.parse_mm_data(mm_data)
processed_mm_uuids = renderer._process_mm_uuids(
mm_data, mm_items, mm_uuids, "req-3"
)
assert processed_mm_uuids == mm_uuids
@pytest.mark.parametrize(
"mm_cache_gb, enable_prefix_caching",
[
(4.0, True), # default behavior
(4.0, False), # prefix caching disabled
(0.0, True), # processor cache disabled
],
)
def test_multi_modal_uuids_accepts_empty(
monkeypatch, mm_cache_gb: float, enable_prefix_caching: bool
):
renderer = _build_renderer(
mm_cache_gb=mm_cache_gb,
enable_prefix_caching=enable_prefix_caching,
)
# While None means cached multi-modal input requiring UUIDs
# an empty list means no multi-modal input
mm_data = {"image": [], "video": []} # type: ignore[var-annotated]
mm_uuids = {"image": [], "video": None} # type: ignore[var-annotated]
mm_processor = renderer.get_mm_processor()
mm_items = mm_processor.info.parse_mm_data(mm_data)
processed_mm_uuids = renderer._process_mm_uuids(
mm_data, mm_items, mm_uuids, "req-4"
)
assert processed_mm_uuids == mm_uuids
def test_multi_modal_uuids_ignored_when_caching_disabled(monkeypatch):
# When both processor cache is 0 and prefix caching disabled, the
# processor builds overrides from request id instead of using user UUIDs.
renderer = _build_renderer(mm_cache_gb=0.0, enable_prefix_caching=False)
request_id = "req-42"
mm_data = {
"image": [cherry_pil_image, stop_pil_image],
"video": baby_reading_np_ndarrays,
}
mm_uuids = {"image": ["hash_cherry", "hash_stop"], "video": ["hash_video"]}
mm_processor = renderer.get_mm_processor()
mm_items = mm_processor.info.parse_mm_data(mm_data)
processed_mm_uuids = renderer._process_mm_uuids(
mm_data, mm_items, mm_uuids, request_id
)
# Expect request-id-based overrides are passed through
assert set(mm_uuids.keys()) == {"image", "video"}
assert len(mm_uuids["image"]) == 2
assert len(mm_uuids["video"]) == 1
assert processed_mm_uuids["image"][0].startswith(
f"{request_id}-image-"
) and processed_mm_uuids["image"][0].endswith("-0")
assert processed_mm_uuids["image"][1].startswith(
f"{request_id}-image-"
) and processed_mm_uuids["image"][1].endswith("-1")
assert processed_mm_uuids["video"][0].startswith(
f"{request_id}-video-"
) and processed_mm_uuids["video"][0].endswith("-0")
-2
View File
@@ -20,7 +20,6 @@ MM_BEAM_WIDTHS = [2]
MODELS = ["TinyLlama/TinyLlama-1.1B-Chat-v1.0"] MODELS = ["TinyLlama/TinyLlama-1.1B-Chat-v1.0"]
@pytest.mark.skip_v1 # V1 engine does not yet support beam search
@pytest.mark.parametrize("model", MODELS) @pytest.mark.parametrize("model", MODELS)
@pytest.mark.parametrize("dtype", ["half"]) @pytest.mark.parametrize("dtype", ["half"])
@pytest.mark.parametrize("max_tokens", MAX_TOKENS) @pytest.mark.parametrize("max_tokens", MAX_TOKENS)
@@ -62,7 +61,6 @@ def test_beam_search_single_input(
) )
@pytest.mark.skip_v1 # V1 engine does not yet support beam search
@pytest.mark.parametrize("model", MODELS) @pytest.mark.parametrize("model", MODELS)
@pytest.mark.parametrize("dtype", ["half"]) @pytest.mark.parametrize("dtype", ["half"])
@pytest.mark.parametrize("max_tokens", MAX_TOKENS) @pytest.mark.parametrize("max_tokens", MAX_TOKENS)
@@ -6,7 +6,7 @@ set -e
merge_base_commit=$(git merge-base HEAD origin/main) merge_base_commit=$(git merge-base HEAD origin/main)
echo "INFO: current merge base commit with main: $merge_base_commit" echo "INFO: current merge base commit with main: $merge_base_commit"
git show --oneline -s $merge_base_commit git show --oneline -s "$merge_base_commit"
# test whether the metadata.json url is valid, retry each 3 minutes up to 5 times # test whether the metadata.json url is valid, retry each 3 minutes up to 5 times
# this avoids cumbersome error messages & manual retries in case the precompiled wheel # this avoids cumbersome error messages & manual retries in case the precompiled wheel
@@ -40,7 +40,7 @@ for i in {1..5}; do
fi fi
fi fi
# failure handling & retry logic # failure handling & retry logic
if [ $i -eq 5 ]; then if [ "$i" -eq 5 ]; then
echo "ERROR: metadata is still not available after 5 attempts." echo "ERROR: metadata is still not available after 5 attempts."
echo "ERROR: Please check whether the precompiled wheel for commit $merge_base_commit is available." echo "ERROR: Please check whether the precompiled wheel for commit $merge_base_commit is available."
echo " NOTE: If $merge_base_commit is a new commit on main, maybe try again after its release pipeline finishes." echo " NOTE: If $merge_base_commit is a new commit on main, maybe try again after its release pipeline finishes."
+194
View File
@@ -0,0 +1,194 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
"""Tests for vllm.ray.ray_env — env var propagation to Ray workers."""
import os
from unittest.mock import patch
from vllm.ray.ray_env import get_env_vars_to_copy
# ---------------------------------------------------------------------------
# Default prefix matching
# ---------------------------------------------------------------------------
class TestDefaultPrefixes:
"""Built-in prefixes (VLLM_, LMCACHE_, NCCL_, UCX_, HF_, HUGGING_FACE_)
should be forwarded without any extra configuration."""
@patch.dict(os.environ, {"LMCACHE_LOCAL_CPU": "True"}, clear=False)
def test_lmcache_prefix(self):
result = get_env_vars_to_copy()
assert "LMCACHE_LOCAL_CPU" in result
@patch.dict(os.environ, {"NCCL_DEBUG": "INFO"}, clear=False)
def test_nccl_prefix(self):
result = get_env_vars_to_copy()
assert "NCCL_DEBUG" in result
@patch.dict(os.environ, {"UCX_TLS": "rc"}, clear=False)
def test_ucx_prefix(self):
result = get_env_vars_to_copy()
assert "UCX_TLS" in result
@patch.dict(os.environ, {"HF_TOKEN": "secret"}, clear=False)
def test_hf_token_via_prefix(self):
result = get_env_vars_to_copy()
assert "HF_TOKEN" in result
@patch.dict(os.environ, {"HUGGING_FACE_HUB_TOKEN": "secret"}, clear=False)
def test_hugging_face_prefix(self):
result = get_env_vars_to_copy()
assert "HUGGING_FACE_HUB_TOKEN" in result
# ---------------------------------------------------------------------------
# Default extra vars
# ---------------------------------------------------------------------------
class TestDefaultExtraVars:
"""Individual vars listed in VLLM_RAY_EXTRA_ENV_VARS_TO_COPY's default."""
def test_pythonhashseed_in_result(self):
"""PYTHONHASHSEED should always be in the result set (as a name to
copy) regardless of whether it is actually set in os.environ."""
result = get_env_vars_to_copy()
assert "PYTHONHASHSEED" in result
# ---------------------------------------------------------------------------
# User-supplied extensions
# ---------------------------------------------------------------------------
class TestUserExtensions:
"""Users can add prefixes and extra vars at deploy time."""
@patch.dict(
os.environ,
{
"VLLM_RAY_EXTRA_ENV_VAR_PREFIXES_TO_COPY": "MYLIB_",
"MYLIB_FOO": "bar",
},
clear=False,
)
def test_user_prefix(self):
"""User-supplied prefixes are additive — built-in defaults are kept."""
result = get_env_vars_to_copy()
assert "MYLIB_FOO" in result
@patch.dict(
os.environ,
{
"VLLM_RAY_EXTRA_ENV_VARS_TO_COPY": "MY_SECRET",
"MY_SECRET": "val",
},
clear=False,
)
def test_user_extra_var(self):
"""User-supplied extras are additive — PYTHONHASHSEED still included."""
result = get_env_vars_to_copy()
assert "MY_SECRET" in result
assert "PYTHONHASHSEED" in result
# ---------------------------------------------------------------------------
# Exclusion
# ---------------------------------------------------------------------------
class TestExclusion:
"""exclude_vars and RAY_NON_CARRY_OVER_ENV_VARS take precedence."""
@patch.dict(os.environ, {"CUDA_VISIBLE_DEVICES": "0,1"}, clear=False)
def test_exclude_vars(self):
result = get_env_vars_to_copy(exclude_vars={"CUDA_VISIBLE_DEVICES"})
assert "CUDA_VISIBLE_DEVICES" not in result
@patch.dict(os.environ, {"LMCACHE_LOCAL_CPU": "True"}, clear=False)
@patch(
"vllm.ray.ray_env.RAY_NON_CARRY_OVER_ENV_VARS",
{"LMCACHE_LOCAL_CPU"},
)
def test_non_carry_over_blacklist(self):
result = get_env_vars_to_copy()
assert "LMCACHE_LOCAL_CPU" not in result
# ---------------------------------------------------------------------------
# additional_vars (platform extension point)
# ---------------------------------------------------------------------------
class TestAdditionalVars:
"""The additional_vars parameter supports platform-specific vars."""
@patch.dict(os.environ, {"CUSTOM_PLATFORM_VAR": "1"}, clear=False)
def test_additional_vars_passthrough(self):
result = get_env_vars_to_copy(additional_vars={"CUSTOM_PLATFORM_VAR"})
assert "CUSTOM_PLATFORM_VAR" in result
# ---------------------------------------------------------------------------
# Edge cases
# ---------------------------------------------------------------------------
class TestEdgeCases:
"""Prefix matching should be strict (startswith, not contains)."""
@patch.dict(os.environ, {"LMCACH_TYPO": "1"}, clear=False)
def test_prefix_no_partial_match(self):
"""'LMCACH_' does not match the 'LMCACHE_' prefix."""
result = get_env_vars_to_copy()
assert "LMCACH_TYPO" not in result
@patch.dict(
os.environ,
{
"VLLM_RAY_EXTRA_ENV_VAR_PREFIXES_TO_COPY": " MYLIB_ , OTHER_ ",
},
clear=False,
)
def test_csv_whitespace_handling(self):
"""Whitespace around commas and tokens should be stripped."""
result = get_env_vars_to_copy()
# MYLIB_ and OTHER_ should be parsed as valid prefixes — no crash
assert isinstance(result, set)
@patch.dict(
os.environ,
{
"VLLM_RAY_EXTRA_ENV_VAR_PREFIXES_TO_COPY": "MYLIB_",
"LMCACHE_BACKEND": "cpu",
"NCCL_DEBUG": "INFO",
"MYLIB_FOO": "bar",
},
clear=False,
)
def test_user_prefix_additive(self):
"""Setting VLLM_RAY_EXTRA_ENV_VAR_PREFIXES_TO_COPY does NOT drop defaults."""
result = get_env_vars_to_copy()
# Built-in defaults still present
assert "LMCACHE_BACKEND" in result
assert "NCCL_DEBUG" in result
# User addition also present
assert "MYLIB_FOO" in result
@patch.dict(
os.environ,
{
"VLLM_RAY_EXTRA_ENV_VARS_TO_COPY": "MY_FLAG",
"PYTHONHASHSEED": "42",
"MY_FLAG": "1",
},
clear=False,
)
def test_user_extra_additive(self):
"""Setting VLLM_RAY_EXTRA_ENV_VARS_TO_COPY does NOT drop defaults."""
result = get_env_vars_to_copy()
# Built-in default still present
assert "PYTHONHASHSEED" in result
# User addition also present
assert "MY_FLAG" in result
+23
View File
@@ -107,6 +107,22 @@ class FakeTraceService(TraceServiceServicer):
self.evt.clear() self.evt.clear()
def _wait_for_server_ready(address: str, timeout: float = 5.0) -> bool:
"""Wait for the gRPC server to be ready to accept connections."""
import socket
import time
host, port = address.rsplit(":", 1)
deadline = time.monotonic() + timeout
while time.monotonic() < deadline:
try:
with socket.create_connection((host, int(port)), timeout=0.5):
return True
except (OSError, ConnectionRefusedError):
time.sleep(0.1)
return False
@pytest.fixture @pytest.fixture
def trace_service() -> Generator[FakeTraceService, None, None]: def trace_service() -> Generator[FakeTraceService, None, None]:
"""Fixture to set up a fake gRPC trace service.""" """Fixture to set up a fake gRPC trace service."""
@@ -116,6 +132,13 @@ def trace_service() -> Generator[FakeTraceService, None, None]:
server.add_insecure_port(FAKE_TRACE_SERVER_ADDRESS) server.add_insecure_port(FAKE_TRACE_SERVER_ADDRESS)
server.start() server.start()
# Wait for the server to be ready to accept connections
if not _wait_for_server_ready(FAKE_TRACE_SERVER_ADDRESS):
server.stop(grace=None)
raise RuntimeError(
f"Fake trace server failed to start on {FAKE_TRACE_SERVER_ADDRESS}"
)
yield service yield service
server.stop(grace=None) server.stop(grace=None)
+294
View File
@@ -3676,6 +3676,300 @@ def test_abort_request_finished_recving():
assert not scheduler.finished_recving_kv_req_ids assert not scheduler.finished_recving_kv_req_ids
# ==============================================================================
# Variable-length encoder cross-attention block allocation tests
# ==============================================================================
def _create_encoder_decoder_scheduler(
block_size: int = 16,
num_blocks: int = 10000,
max_num_batched_tokens: int = 8192,
max_num_seqs: int = 16,
) -> Scheduler:
"""Create a scheduler configured for encoder-decoder cross-attention
block allocation testing.
Constructs a scheduler with both FullAttentionSpec (self-attention) and
CrossAttentionSpec (cross-attention) KV cache groups, then patches it
to behave as an encoder-decoder model.
"""
from vllm.v1.core.encoder_cache_manager import EncoderDecoderCacheManager
from vllm.v1.kv_cache_interface import CrossAttentionSpec
model_config = ModelConfig(
model="facebook/opt-125m",
trust_remote_code=True,
dtype="float16",
seed=42,
)
scheduler_config = SchedulerConfig(
max_num_seqs=max_num_seqs,
max_num_batched_tokens=max_num_batched_tokens,
max_model_len=max_num_batched_tokens,
# is_encoder_decoder disables chunked prefill and prefix caching
is_encoder_decoder=True,
)
cache_config = CacheConfig(
block_size=block_size,
gpu_memory_utilization=0.9,
swap_space=0,
cache_dtype="auto",
enable_prefix_caching=False,
)
cache_config.num_gpu_blocks = num_blocks
vllm_config = VllmConfig(
scheduler_config=scheduler_config,
model_config=model_config,
cache_config=cache_config,
)
# KV cache config with both self-attention and cross-attention groups,
# mirroring an encoder-decoder model like Whisper.
kv_cache_config = KVCacheConfig(
num_blocks=num_blocks,
kv_cache_tensors=[],
kv_cache_groups=[
KVCacheGroupSpec(
["self_attn_layer"],
FullAttentionSpec(
block_size=block_size,
num_kv_heads=1,
head_size=1,
dtype=torch.float32,
),
),
KVCacheGroupSpec(
["cross_attn_layer"],
CrossAttentionSpec(
block_size=block_size,
num_kv_heads=1,
head_size=1,
dtype=torch.float32,
),
),
],
)
# Construct the scheduler. Since opt-125m is not truly encoder-decoder,
# the __init__ won't set up encoder-decoder internals. We patch them
# after construction.
scheduler = Scheduler(
vllm_config=vllm_config,
kv_cache_config=kv_cache_config,
block_size=block_size,
structured_output_manager=StructuredOutputManager(vllm_config),
)
# Patch to enable encoder-decoder behavior in the scheduling loop.
scheduler.is_encoder_decoder = True
scheduler.max_num_encoder_input_tokens = max_num_batched_tokens
scheduler.encoder_cache_manager = EncoderDecoderCacheManager(
cache_size=max_num_batched_tokens
)
return scheduler
def _get_num_cross_attn_blocks(scheduler: Scheduler, request_id: str) -> int:
"""Get the number of cross-attention blocks allocated for a request."""
from vllm.v1.core.single_type_kv_cache_manager import CrossAttentionManager
coordinator = scheduler.kv_cache_manager.coordinator
for manager in coordinator.single_type_managers:
if isinstance(manager, CrossAttentionManager):
blocks = manager.req_to_blocks.get(request_id, [])
return len(blocks)
raise AssertionError("No CrossAttentionManager found in coordinator")
def test_variable_length_cross_attn_block_allocation():
"""Test that cross-attention blocks are allocated per-request based on
actual encoder input length, not a fixed maximum.
Fixed max-encoder-length allocation would assign
`ceil(max_encoder_tokens / block_size)` blocks to
every request whereas with dynamic allocation, exactly
`ceil(actual_encoder_tokens / block_size)` blocks are assigned
to each request.
"""
block_size = 16
scheduler = _create_encoder_decoder_scheduler(block_size=block_size)
# Create requests with distinctly different encoder input lengths,
# simulating variable-length audio inputs to a model like Whisper.
encoder_lengths = [500, 1000, 200]
num_prompt_tokens = 100 # Decoder prompt tokens
requests = []
for i, enc_len in enumerate(encoder_lengths):
req = create_requests(
num_requests=1,
num_tokens=num_prompt_tokens,
mm_hashes_list=[[f"enc_hash_{i}"]],
mm_positions=[[PlaceholderRange(offset=0, length=enc_len)]],
req_ids=[f"req_{i}"],
)[0]
requests.append(req)
# Add and schedule all requests.
for req in requests:
scheduler.add_request(req)
output = scheduler.schedule()
# All requests should be scheduled.
assert len(output.scheduled_new_reqs) == len(requests)
# Verify cross-attention blocks per request match the actual encoder length.
from math import ceil
for req, enc_len in zip(requests, encoder_lengths):
expected_blocks = ceil(enc_len / block_size)
actual_blocks = _get_num_cross_attn_blocks(scheduler, req.request_id)
assert actual_blocks == expected_blocks, (
f"Request {req.request_id} with {enc_len} encoder tokens: "
f"expected {expected_blocks} cross-attn blocks, "
f"got {actual_blocks}"
)
# Verify that different encoder lengths produce different block counts,
# confirming variable-length (not fixed-max) allocation.
block_counts = [
_get_num_cross_attn_blocks(scheduler, req.request_id) for req in requests
]
assert len(set(block_counts)) > 1, (
"All requests have the same number of cross-attn blocks, "
"suggesting static max-based allocation instead of per-request"
)
def test_cross_attn_blocks_not_over_allocated():
"""Test that cross-attention blocks are not over-allocated compared to
what each request actually needs."""
from math import ceil
block_size = 16
max_encoder_tokens = 1500 # e.g., Whisper's max mel-spectrogram length
scheduler = _create_encoder_decoder_scheduler(block_size=block_size)
# Request with a small encoder input (much less than the max).
small_enc_len = 200
request = create_requests(
num_requests=1,
num_tokens=100,
mm_hashes_list=[["enc_small"]],
mm_positions=[[PlaceholderRange(offset=0, length=small_enc_len)]],
req_ids=["req_small"],
)[0]
scheduler.add_request(request)
output = scheduler.schedule()
assert len(output.scheduled_new_reqs) == 1
actual_blocks = _get_num_cross_attn_blocks(scheduler, request.request_id)
expected_blocks = ceil(small_enc_len / block_size)
max_blocks = ceil(max_encoder_tokens / block_size)
# Blocks should match the actual encoder length.
assert actual_blocks == expected_blocks, (
f"Expected {expected_blocks} blocks for {small_enc_len} encoder tokens, "
f"got {actual_blocks}"
)
# Blocks should be strictly less than what max-based allocation would give.
assert actual_blocks < max_blocks, (
f"Cross-attn blocks ({actual_blocks}) should be less than max "
f"({max_blocks}), indicating no over-allocation"
)
def test_cross_attn_blocks_not_under_allocated():
"""Test that cross-attention blocks are sufficient for each request's
actual encoder input length. Every encoder token must have a slot.
Tests various edge cases including exact block boundaries, off-by-one,
and the minimum/maximum encoder input sizes.
"""
from math import ceil
block_size = 16
# Test various encoder lengths including edge cases around block boundaries.
test_cases = [
1, # Minimum: single encoder token
block_size - 1, # Just under one full block
block_size, # Exactly one full block
block_size + 1, # Just over one block (needs 2 blocks)
block_size * 10, # Exact multiple of block size
block_size * 10 + 1, # One over exact multiple
1500, # Whisper's typical max
]
for enc_len in test_cases:
scheduler = _create_encoder_decoder_scheduler(block_size=block_size)
request = create_requests(
num_requests=1,
num_tokens=100,
mm_hashes_list=[[f"enc_{enc_len}"]],
mm_positions=[[PlaceholderRange(offset=0, length=enc_len)]],
req_ids=[f"req_{enc_len}"],
)[0]
scheduler.add_request(request)
output = scheduler.schedule()
assert len(output.scheduled_new_reqs) == 1
actual_blocks = _get_num_cross_attn_blocks(scheduler, request.request_id)
expected_blocks = ceil(enc_len / block_size)
# Number of blocks must be exactly ceil(enc_len / block_size).
assert actual_blocks == expected_blocks, (
f"Encoder length {enc_len}: expected {expected_blocks} blocks, "
f"got {actual_blocks}"
)
# Total available slots must be >= encoder tokens (no under-allocation).
total_slots = actual_blocks * block_size
assert total_slots >= enc_len, (
f"Encoder length {enc_len}: total slots {total_slots} < "
f"needed {enc_len} (under-allocation)"
)
def test_cross_attn_zero_blocks_without_encoder_inputs():
"""Test that requests without encoder inputs get zero cross-attention
blocks, even when the scheduler is configured for encoder-decoder."""
block_size = 16
scheduler = _create_encoder_decoder_scheduler(block_size=block_size)
# Create a text-only request (no mm_features).
request = create_requests(
num_requests=1,
num_tokens=100,
req_ids=["req_text_only"],
)[0]
# Text-only request has no encoder inputs.
assert not request.has_encoder_inputs
scheduler.add_request(request)
output = scheduler.schedule()
assert len(output.scheduled_new_reqs) == 1
# No cross-attention blocks should be allocated.
actual_blocks = _get_num_cross_attn_blocks(scheduler, request.request_id)
assert actual_blocks == 0, (
f"Text-only request should have 0 cross-attn blocks, got {actual_blocks}"
)
def test_eagle3_mm_encoder_cache_with_shift(): def test_eagle3_mm_encoder_cache_with_shift():
"""Test EAGLE3 encoder scheduling accounts for shift_computed_tokens. """Test EAGLE3 encoder scheduling accounts for shift_computed_tokens.
@@ -24,7 +24,7 @@ MODEL="${MODEL:-Qwen/Qwen2.5-VL-3B-Instruct}"
# Set 1 to use multimodal prompts; else to use text-only # Set 1 to use multimodal prompts; else to use text-only
USE_MM_PROMPTS="${USE_MM_PROMPTS:-1}" USE_MM_PROMPTS="${USE_MM_PROMPTS:-1}"
MM_FLAG="" MM_FLAG=""
if [ $USE_MM_PROMPTS = "1" ]; then if [ "$USE_MM_PROMPTS" = "1" ]; then
MM_FLAG="--use_mm_prompts" MM_FLAG="--use_mm_prompts"
fi fi
@@ -51,7 +51,7 @@ LOG_PATH="${LOG_PATH:-/tmp}"
BASELINE_FILE="${BASELINE_FILE:-/tmp/vllm_baseline.txt}" BASELINE_FILE="${BASELINE_FILE:-/tmp/vllm_baseline.txt}"
BASELINE_PD_FILE="${BASELINE_PD_FILE:-/tmp/vllm_epd_baseline.txt}" BASELINE_PD_FILE="${BASELINE_PD_FILE:-/tmp/vllm_epd_baseline.txt}"
mkdir -p $LOG_PATH mkdir -p "$LOG_PATH"
# Trap the SIGINT signal (triggered by Ctrl+C) # Trap the SIGINT signal (triggered by Ctrl+C)
trap 'kill $(jobs -pr)' SIGINT SIGTERM EXIT trap 'kill $(jobs -pr)' SIGINT SIGTERM EXIT
@@ -87,20 +87,20 @@ run_baseline() {
# Start baseline instance # Start baseline instance
echo "Starting baseline instance on GPU $GPU_SINGLE, port $PORT" echo "Starting baseline instance on GPU $GPU_SINGLE, port $PORT"
CUDA_VISIBLE_DEVICES="$GPU_SINGLE" vllm serve "$MODEL" \ CUDA_VISIBLE_DEVICES="$GPU_SINGLE" vllm serve "$MODEL" \
--port $PORT \ --port "$PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
> $LOG_PATH/baseline.log 2>&1 & > "$LOG_PATH"/baseline.log 2>&1 &
local BASELINE_PID=$! local BASELINE_PID=$!
# Wait for baseline to start # Wait for baseline to start
echo "Waiting for baseline instance to start..." echo "Waiting for baseline instance to start..."
wait_for_server $PORT wait_for_server "$PORT"
curl http://127.0.0.1:$PORT/v1/models curl http://127.0.0.1:"$PORT"/v1/models
echo "" echo ""
# Run test in baseline mode # Run test in baseline mode
@@ -139,14 +139,14 @@ run_epd_1e_1pd() {
# Start encoder instance # Start encoder instance
echo "Starting encoder instance on GPU $GPU_E, port $ENCODE_PORT" echo "Starting encoder instance on GPU $GPU_E, port $ENCODE_PORT"
CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \ CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \
--port $ENCODE_PORT \ --port "$ENCODE_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.01 \ --gpu-memory-utilization 0.01 \
--enable-request-id-headers \ --enable-request-id-headers \
--no-enable-prefix-caching \ --no-enable-prefix-caching \
--max-num-batched-tokens 114688 \ --max-num-batched-tokens 114688 \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_producer", "ec_role": "ec_producer",
@@ -154,18 +154,18 @@ run_epd_1e_1pd() {
"shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'" "shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'"
} }
}' \ }' \
> $LOG_PATH/1e1pd_encoder.log 2>&1 & > "$LOG_PATH"/1e1pd_encoder.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Start prefill+decode instance # Start prefill+decode instance
echo "Starting PD instance on GPU $GPU_PD, port $PREFILL_DECODE_PORT" echo "Starting PD instance on GPU $GPU_PD, port $PREFILL_DECODE_PORT"
CUDA_VISIBLE_DEVICES="$GPU_PD" vllm serve "$MODEL" \ CUDA_VISIBLE_DEVICES="$GPU_PD" vllm serve "$MODEL" \
--port $PREFILL_DECODE_PORT \ --port "$PREFILL_DECODE_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_consumer", "ec_role": "ec_consumer",
@@ -173,32 +173,32 @@ run_epd_1e_1pd() {
"shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'" "shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'"
} }
}' \ }' \
> $LOG_PATH/1e1pd_pd.log 2>&1 & > "$LOG_PATH"/1e1pd_pd.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for instances to start # Wait for instances to start
echo "Waiting for encoder instance..." echo "Waiting for encoder instance..."
wait_for_server $ENCODE_PORT wait_for_server "$ENCODE_PORT"
echo "Waiting for PD instance..." echo "Waiting for PD instance..."
wait_for_server $PREFILL_DECODE_PORT wait_for_server "$PREFILL_DECODE_PORT"
# Start proxy # Start proxy
echo "Starting EPD proxy on port $PROXY_PORT" echo "Starting EPD proxy on port $PROXY_PORT"
python "${GIT_ROOT}/examples/online_serving/disaggregated_encoder/disagg_epd_proxy.py" \ python "${GIT_ROOT}/examples/online_serving/disaggregated_encoder/disagg_epd_proxy.py" \
--host "0.0.0.0" \ --host "0.0.0.0" \
--port $PROXY_PORT \ --port "$PROXY_PORT" \
--encode-servers-urls "http://localhost:$ENCODE_PORT" \ --encode-servers-urls "http://localhost:$ENCODE_PORT" \
--prefill-servers-urls "disable" \ --prefill-servers-urls "disable" \
--decode-servers-urls "http://localhost:$PREFILL_DECODE_PORT" \ --decode-servers-urls "http://localhost:$PREFILL_DECODE_PORT" \
> $LOG_PATH/1e1pd_proxy.log 2>&1 & > "$LOG_PATH"/1e1pd_proxy.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for proxy # Wait for proxy
echo "Waiting for proxy..." echo "Waiting for proxy..."
wait_for_server $PROXY_PORT wait_for_server "$PROXY_PORT"
curl http://127.0.0.1:$PROXY_PORT/v1/models curl http://127.0.0.1:"$PROXY_PORT"/v1/models
curl http://127.0.0.1:$PROXY_PORT/health curl http://127.0.0.1:"$PROXY_PORT"/health
echo "" echo ""
echo "All EPD (1E+1PD) services are up!" echo "All EPD (1E+1PD) services are up!"
@@ -217,7 +217,7 @@ run_epd_1e_1pd() {
echo "✓✓ 1E+1PD Correctness Test finished" echo "✓✓ 1E+1PD Correctness Test finished"
echo "Stopping EPD (1E+1PD) instances..." echo "Stopping EPD (1E+1PD) instances..."
for pid in "${PIDS[@]}"; do for pid in "${PIDS[@]}"; do
kill $pid 2>/dev/null || true kill "$pid" 2>/dev/null || true
done done
sleep 2 sleep 2
cleanup_instances cleanup_instances
@@ -244,17 +244,17 @@ run_baseline_1p_1d() {
CUDA_VISIBLE_DEVICES="$GPU_P" \ CUDA_VISIBLE_DEVICES="$GPU_P" \
VLLM_NIXL_SIDE_CHANNEL_PORT=5559 \ VLLM_NIXL_SIDE_CHANNEL_PORT=5559 \
vllm serve "$MODEL" \ vllm serve "$MODEL" \
--port $PREFILL_PORT \ --port "$PREFILL_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--kv-transfer-config '{ --kv-transfer-config '{
"kv_connector": "NixlConnector", "kv_connector": "NixlConnector",
"kv_role": "kv_producer" "kv_role": "kv_producer"
}' \ }' \
> $LOG_PATH/1p1d_prefill.log 2>&1 & > "$LOG_PATH"/1p1d_prefill.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Start decode instance # Start decode instance
@@ -262,40 +262,40 @@ run_baseline_1p_1d() {
CUDA_VISIBLE_DEVICES="$GPU_D" \ CUDA_VISIBLE_DEVICES="$GPU_D" \
VLLM_NIXL_SIDE_CHANNEL_PORT=6000 \ VLLM_NIXL_SIDE_CHANNEL_PORT=6000 \
vllm serve "$MODEL" \ vllm serve "$MODEL" \
--port $DECODE_PORT \ --port "$DECODE_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--kv-transfer-config '{ --kv-transfer-config '{
"kv_connector": "NixlConnector", "kv_connector": "NixlConnector",
"kv_role": "kv_consumer" "kv_role": "kv_consumer"
}' \ }' \
> $LOG_PATH/1p1d_decode.log 2>&1 & > "$LOG_PATH"/1p1d_decode.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for instances to start # Wait for instances to start
echo "Waiting for prefill instance..." echo "Waiting for prefill instance..."
wait_for_server $PREFILL_PORT wait_for_server "$PREFILL_PORT"
echo "Waiting for decode instance..." echo "Waiting for decode instance..."
wait_for_server $DECODE_PORT wait_for_server "$DECODE_PORT"
# Start proxy # Start proxy
echo "Starting EPD proxy on port $PROXY_PORT" echo "Starting EPD proxy on port $PROXY_PORT"
python "${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/toy_proxy_server.py" \ python "${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/toy_proxy_server.py" \
--host "0.0.0.0" \ --host "0.0.0.0" \
--port $PROXY_PORT \ --port "$PROXY_PORT" \
--prefiller-ports $PREFILL_PORT \ --prefiller-ports "$PREFILL_PORT" \
--decoder-ports $DECODE_PORT \ --decoder-ports "$DECODE_PORT" \
> $LOG_PATH/1p1d_proxy.log 2>&1 & > "$LOG_PATH"/1p1d_proxy.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for proxy # Wait for proxy
echo "Waiting for proxy..." echo "Waiting for proxy..."
wait_for_server $PROXY_PORT wait_for_server "$PROXY_PORT"
curl http://127.0.0.1:$PROXY_PORT/healthcheck curl http://127.0.0.1:"$PROXY_PORT"/healthcheck
echo "" echo ""
echo "All PD (1P+1D) services are up!" echo "All PD (1P+1D) services are up!"
@@ -313,7 +313,7 @@ run_baseline_1p_1d() {
# Cleanup # Cleanup
echo "Stopping PD (1P+1D) instances..." echo "Stopping PD (1P+1D) instances..."
for pid in "${PIDS[@]}"; do for pid in "${PIDS[@]}"; do
kill $pid 2>/dev/null || true kill "$pid" 2>/dev/null || true
done done
sleep 2 sleep 2
cleanup_instances cleanup_instances
@@ -339,14 +339,14 @@ run_epd_1e_1p_1d() {
# Start encoder instance # Start encoder instance
echo "Starting encoder instance on GPU $GPU_E, port $ENCODE_PORT" echo "Starting encoder instance on GPU $GPU_E, port $ENCODE_PORT"
CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \ CUDA_VISIBLE_DEVICES="$GPU_E" vllm serve "$MODEL" \
--port $ENCODE_PORT \ --port "$ENCODE_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.01 \ --gpu-memory-utilization 0.01 \
--enable-request-id-headers \ --enable-request-id-headers \
--no-enable-prefix-caching \ --no-enable-prefix-caching \
--max-num-batched-tokens 114688 \ --max-num-batched-tokens 114688 \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_producer", "ec_role": "ec_producer",
@@ -354,7 +354,7 @@ run_epd_1e_1p_1d() {
"shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'" "shared_storage_path": "'"$EC_SHARED_STORAGE_PATH"'"
} }
}' \ }' \
> $LOG_PATH/1e1p1d_encoder.log 2>&1 & > "$LOG_PATH"/1e1p1d_encoder.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Start prefill instance # Start prefill instance
@@ -362,12 +362,12 @@ run_epd_1e_1p_1d() {
CUDA_VISIBLE_DEVICES="$GPU_P" \ CUDA_VISIBLE_DEVICES="$GPU_P" \
VLLM_NIXL_SIDE_CHANNEL_PORT=5559 \ VLLM_NIXL_SIDE_CHANNEL_PORT=5559 \
vllm serve "$MODEL" \ vllm serve "$MODEL" \
--port $PREFILL_PORT \ --port "$PREFILL_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--ec-transfer-config '{ --ec-transfer-config '{
"ec_connector": "ECExampleConnector", "ec_connector": "ECExampleConnector",
"ec_role": "ec_consumer", "ec_role": "ec_consumer",
@@ -379,7 +379,7 @@ run_epd_1e_1p_1d() {
"kv_connector": "NixlConnector", "kv_connector": "NixlConnector",
"kv_role": "kv_producer" "kv_role": "kv_producer"
}' \ }' \
> $LOG_PATH/1e1p1d_prefill.log 2>&1 & > "$LOG_PATH"/1e1p1d_prefill.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Start decode instance # Start decode instance
@@ -387,44 +387,44 @@ run_epd_1e_1p_1d() {
CUDA_VISIBLE_DEVICES="$GPU_D" \ CUDA_VISIBLE_DEVICES="$GPU_D" \
VLLM_NIXL_SIDE_CHANNEL_PORT=6000 \ VLLM_NIXL_SIDE_CHANNEL_PORT=6000 \
vllm serve "$MODEL" \ vllm serve "$MODEL" \
--port $DECODE_PORT \ --port "$DECODE_PORT" \
--enforce-eager \ --enforce-eager \
--gpu-memory-utilization 0.7 \ --gpu-memory-utilization 0.7 \
--enable-request-id-headers \ --enable-request-id-headers \
--max-num-seqs 128 \ --max-num-seqs 128 \
--allowed-local-media-path ${GIT_ROOT}/tests/v1/ec_connector/integration \ --allowed-local-media-path "${GIT_ROOT}"/tests/v1/ec_connector/integration \
--kv-transfer-config '{ --kv-transfer-config '{
"kv_connector": "NixlConnector", "kv_connector": "NixlConnector",
"kv_role": "kv_consumer" "kv_role": "kv_consumer"
}' \ }' \
> $LOG_PATH/1e1p1d_decode.log 2>&1 & > "$LOG_PATH"/1e1p1d_decode.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for instances to start # Wait for instances to start
echo "Waiting for encoder instance..." echo "Waiting for encoder instance..."
wait_for_server $ENCODE_PORT wait_for_server "$ENCODE_PORT"
echo "Waiting for prefill instance..." echo "Waiting for prefill instance..."
wait_for_server $PREFILL_PORT wait_for_server "$PREFILL_PORT"
echo "Waiting for decode instance..." echo "Waiting for decode instance..."
wait_for_server $DECODE_PORT wait_for_server "$DECODE_PORT"
# Start proxy # Start proxy
echo "Starting EPD proxy on port $PROXY_PORT" echo "Starting EPD proxy on port $PROXY_PORT"
python "${GIT_ROOT}/examples/online_serving/disaggregated_encoder/disagg_epd_proxy.py" \ python "${GIT_ROOT}/examples/online_serving/disaggregated_encoder/disagg_epd_proxy.py" \
--host "0.0.0.0" \ --host "0.0.0.0" \
--port $PROXY_PORT \ --port "$PROXY_PORT" \
--encode-servers-urls "http://localhost:$ENCODE_PORT" \ --encode-servers-urls "http://localhost:$ENCODE_PORT" \
--prefill-servers-urls "http://localhost:$PREFILL_PORT" \ --prefill-servers-urls "http://localhost:$PREFILL_PORT" \
--decode-servers-urls "http://localhost:$DECODE_PORT" \ --decode-servers-urls "http://localhost:$DECODE_PORT" \
> $LOG_PATH/1e1p1d_proxy.log 2>&1 & > "$LOG_PATH"/1e1p1d_proxy.log 2>&1 &
PIDS+=($!) PIDS+=($!)
# Wait for proxy # Wait for proxy
echo "Waiting for proxy..." echo "Waiting for proxy..."
wait_for_server $PROXY_PORT wait_for_server "$PROXY_PORT"
curl http://127.0.0.1:$PROXY_PORT/v1/models curl http://127.0.0.1:"$PROXY_PORT"/v1/models
curl http://127.0.0.1:$PROXY_PORT/health curl http://127.0.0.1:"$PROXY_PORT"/health
echo "" echo ""
echo "All EPD (1E+1P+1D) services are up!" echo "All EPD (1E+1P+1D) services are up!"
@@ -443,7 +443,7 @@ run_epd_1e_1p_1d() {
echo "✓✓ 1E+1P+1D Correctness Test finished" echo "✓✓ 1E+1P+1D Correctness Test finished"
echo "Stopping EPD (1E+1P+1D) instances..." echo "Stopping EPD (1E+1P+1D) instances..."
for pid in "${PIDS[@]}"; do for pid in "${PIDS[@]}"; do
kill $pid 2>/dev/null || true kill "$pid" 2>/dev/null || true
done done
sleep 2 sleep 2
cleanup_instances cleanup_instances
@@ -1,174 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
import pytest
from vllm.assets.image import ImageAsset
from vllm.assets.video import VideoAsset
from vllm.config import CacheConfig, ModelConfig, VllmConfig
from vllm.multimodal import MultiModalUUIDDict
from vllm.sampling_params import SamplingParams
from vllm.v1.engine.input_processor import InputProcessor
cherry_pil_image = ImageAsset("cherry_blossom").pil_image
stop_pil_image = ImageAsset("stop_sign").pil_image
baby_reading_np_ndarrays = VideoAsset("baby_reading").np_ndarrays
def _build_input_processor(
*, mm_cache_gb: float = 4.0, enable_prefix_caching: bool = True
) -> InputProcessor:
model_config = ModelConfig(
model="Qwen/Qwen2.5-VL-3B-Instruct",
max_model_len=128,
mm_processor_cache_gb=mm_cache_gb,
)
vllm_config = VllmConfig(
model_config=model_config,
cache_config=CacheConfig(enable_prefix_caching=enable_prefix_caching),
)
return InputProcessor(vllm_config)
def test_multi_modal_uuids_length_mismatch_raises():
input_processor = _build_input_processor()
prompt = {
"prompt": "USER: <image>\nDescribe\nASSISTANT:",
"multi_modal_data": {"image": [cherry_pil_image, stop_pil_image]},
# Mismatch: 2 items but only 1 uuid provided
"multi_modal_uuids": {"image": ["hash_cherry"]},
}
with pytest.raises(ValueError, match="must have same length as"):
input_processor.process_inputs(
request_id="req-1",
prompt=prompt, # type: ignore[arg-type]
params=SamplingParams(),
)
def test_multi_modal_uuids_missing_modality_raises():
input_processor = _build_input_processor()
prompt = {
"prompt": "USER: <image><video>\nDescribe\nASSISTANT:",
# Two modalities provided in data
"multi_modal_data": {
"image": [cherry_pil_image],
"video": None,
},
# Only image uuids provided; video missing should raise
"multi_modal_uuids": {"image": ["hash_cherry"]},
}
with pytest.raises(ValueError, match="is empty but .* is missing"):
input_processor.process_inputs(
request_id="req-2",
prompt=prompt, # type: ignore[arg-type]
params=SamplingParams(),
)
@pytest.mark.parametrize(
"mm_cache_gb, enable_prefix_caching",
[
(4.0, True), # default behavior
(4.0, False), # prefix caching disabled
(0.0, True), # processor cache disabled
],
)
def test_multi_modal_uuids_accepts_none_and_passes_through(
monkeypatch, mm_cache_gb: float, enable_prefix_caching: bool
):
input_processor = _build_input_processor(
mm_cache_gb=mm_cache_gb,
enable_prefix_caching=enable_prefix_caching,
)
# Capture the overrides passed to InputPreprocessor.preprocess
captured: dict[str, object] = {}
def fake_preprocess(
prompt, *, tokenization_kwargs=None, lora_request=None, mm_uuids=None
):
captured["mm_uuids"] = mm_uuids
# Minimal processed inputs for decoder-only flow
return {"type": "token", "prompt_token_ids": [1]}
# Monkeypatch only the bound preprocess method on this instance
monkeypatch.setattr(
input_processor.input_preprocessor, "preprocess", fake_preprocess, raising=True
)
# Use a consistent two-image scenario across all configurations
mm_uuids = {"image": [None, "hash_stop"], "video": None}
prompt = {
"prompt": "USER: <image><image>\nTwo images\nASSISTANT:",
"multi_modal_data": {
"image": [cherry_pil_image, stop_pil_image],
"video": baby_reading_np_ndarrays,
},
"multi_modal_uuids": mm_uuids,
}
input_processor.process_inputs(
request_id="req-3",
prompt=prompt, # type: ignore[arg-type]
params=SamplingParams(),
)
assert captured["mm_uuids"] == mm_uuids
def test_multi_modal_uuids_ignored_when_caching_disabled(monkeypatch):
# When both processor cache is 0 and prefix caching disabled, the
# processor builds overrides from request id instead of using user UUIDs.
input_processor = _build_input_processor(
mm_cache_gb=0.0, enable_prefix_caching=False
)
captured: dict[str, MultiModalUUIDDict] = {}
def fake_preprocess(
prompt, *, tokenization_kwargs=None, lora_request=None, mm_uuids=None
):
captured["mm_uuids"] = mm_uuids
return {"type": "token", "prompt_token_ids": [1]}
monkeypatch.setattr(
input_processor.input_preprocessor, "preprocess", fake_preprocess, raising=True
)
request_id = "req-42"
mm_uuids = {"image": ["hash_cherry", "hash_stop"], "video": ["hash_video"]}
prompt = {
"prompt": "USER: <image><image><video>\nDescribe\nASSISTANT:",
"multi_modal_data": {
"image": [cherry_pil_image, stop_pil_image],
"video": [baby_reading_np_ndarrays],
},
"multi_modal_uuids": mm_uuids,
}
input_processor.process_inputs(
request_id=request_id,
prompt=prompt, # type: ignore[arg-type]
params=SamplingParams(),
)
# Expect request-id-based overrides are passed through
assert set(mm_uuids.keys()) == {"image", "video"}
assert len(mm_uuids["image"]) == 2
assert len(mm_uuids["video"]) == 1
assert captured["mm_uuids"]["image"][0].startswith(
f"{request_id}-image-"
) and captured["mm_uuids"]["image"][0].endswith("-0")
assert captured["mm_uuids"]["image"][1].startswith(
f"{request_id}-image-"
) and captured["mm_uuids"]["image"][1].endswith("-1")
assert captured["mm_uuids"]["video"][0].startswith(
f"{request_id}-video-"
) and captured["mm_uuids"]["video"][0].endswith("-0")
@@ -32,9 +32,14 @@ run_tests() {
echo "=== Running tests (${label}) ===" echo "=== Running tests (${label}) ==="
for cfg in "${configs[@]}"; do for cfg in "${configs[@]}"; do
local -a cfg_parts extra_args_parts
read -r -a cfg_parts <<< "$cfg"
read -r -a extra_args_parts <<< "$extra_args"
echo "-> Running with ${cfg} ${extra_args:+and ${extra_args}}" echo "-> Running with ${cfg} ${extra_args:+and ${extra_args}}"
# Use 'env' to safely set variables without eval # Use 'env' to safely set variables without eval
if ! env ${cfg} bash "${SCRIPT}" ${extra_args}; then # keep argv splitting safe and SC2086-clean via arrays.
if ! env "${cfg_parts[@]}" bash "${SCRIPT}" "${extra_args_parts[@]}"; then
echo "❌ Test failed for config: ${cfg} ${extra_args:+(${extra_args})}" echo "❌ Test failed for config: ${cfg} ${extra_args:+(${extra_args})}"
exit 1 exit 1
fi fi
@@ -109,9 +109,9 @@ get_model_args() {
get_num_gpus() { get_num_gpus() {
if [[ "$SMI_BIN" == *"nvidia"* ]]; then if [[ "$SMI_BIN" == *"nvidia"* ]]; then
echo "$($SMI_BIN --query-gpu=name --format=csv,noheader | wc -l)" $SMI_BIN --query-gpu=name --format=csv,noheader | wc -l
elif [[ "$SMI_BIN" == *"rocm"* ]]; then elif [[ "$SMI_BIN" == *"rocm"* ]]; then
echo "$($SMI_BIN -l | grep GPU | wc -l)" $SMI_BIN -l | grep -c GPU
else else
# works for non-cuda platforms, # works for non-cuda platforms,
# assuming at least 1 device and # assuming at least 1 device and
@@ -182,7 +182,7 @@ run_tests_for_model() {
# Store host and port for proxy configuration # Store host and port for proxy configuration
PREFILL_HOSTS+=("localhost") PREFILL_HOSTS+=("localhost")
PREFILL_PORTS+=($PORT) PREFILL_PORTS+=("$PORT")
done done
# Start decode instances # Start decode instances
@@ -237,30 +237,30 @@ run_tests_for_model() {
# Store host and port for proxy configuration # Store host and port for proxy configuration
DECODE_HOSTS+=("localhost") DECODE_HOSTS+=("localhost")
DECODE_PORTS+=($PORT) DECODE_PORTS+=("$PORT")
done done
# Wait for all instances to start # Wait for all instances to start
for PORT in "${PREFILL_PORTS[@]}"; do for PORT in "${PREFILL_PORTS[@]}"; do
echo "Waiting for prefill instance on port $PORT to start..." echo "Waiting for prefill instance on port $PORT to start..."
wait_for_server $PORT wait_for_server "$PORT"
done done
for PORT in "${DECODE_PORTS[@]}"; do for PORT in "${DECODE_PORTS[@]}"; do
echo "Waiting for decode instance on port $PORT to start..." echo "Waiting for decode instance on port $PORT to start..."
wait_for_server $PORT wait_for_server "$PORT"
done done
# Build the command for the proxy server with all the hosts and ports # Build the command for the proxy server with all the hosts and ports
PROXY_CMD="python3 ${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/toy_proxy_server.py --port 8192" PROXY_CMD="python3 ${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/toy_proxy_server.py --port 8192"
# Add all prefill hosts and ports # Add all prefill hosts and ports
PROXY_CMD+=" --prefiller-hosts ${PREFILL_HOSTS[@]}" PROXY_CMD+=" --prefiller-hosts ${PREFILL_HOSTS[*]}"
PROXY_CMD+=" --prefiller-ports ${PREFILL_PORTS[@]}" PROXY_CMD+=" --prefiller-ports ${PREFILL_PORTS[*]}"
# Add all decode hosts and ports # Add all decode hosts and ports
PROXY_CMD+=" --decoder-hosts ${DECODE_HOSTS[@]}" PROXY_CMD+=" --decoder-hosts ${DECODE_HOSTS[*]}"
PROXY_CMD+=" --decoder-ports ${DECODE_PORTS[@]}" PROXY_CMD+=" --decoder-ports ${DECODE_PORTS[*]}"
# Start the proxy server # Start the proxy server
echo "Starting proxy server with command: $PROXY_CMD" echo "Starting proxy server with command: $PROXY_CMD"
@@ -271,7 +271,7 @@ run_tests_for_model() {
# Run lm eval for this model # Run lm eval for this model
echo "Running tests for $model_name" echo "Running tests for $model_name"
TEST_MODEL=$model_name python3 -m pytest -s -x ${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/test_accuracy.py TEST_MODEL=$model_name python3 -m pytest -s -x "${GIT_ROOT}"/tests/v1/kv_connector/nixl_integration/test_accuracy.py
# Clean up before running next model # Clean up before running next model
cleanup_instances cleanup_instances
@@ -114,10 +114,10 @@ run_tests_for_model() {
eval "$FULL_CMD &" eval "$FULL_CMD &"
# Wait for all instances to start # Wait for all instances to start
echo "Waiting for prefill instance on port $PORT to start..." echo "Waiting for prefill instance on port $PREFILL_PORT to start..."
wait_for_server $PREFILL_PORT wait_for_server "$PREFILL_PORT"
echo "Waiting for decode instance on port $PORT to start..." echo "Waiting for decode instance on port $DECODE_PORT to start..."
wait_for_server $DECODE_PORT wait_for_server "$DECODE_PORT"
# Build the command for the proxy server with all the hosts and ports # Build the command for the proxy server with all the hosts and ports
PROXY_PORT=8192 PROXY_PORT=8192
@@ -133,7 +133,7 @@ run_tests_for_model() {
# Run lm eval for this model # Run lm eval for this model
echo "Running tests for $model_name" echo "Running tests for $model_name"
PREFILL_PORT=$PREFILL_PORT DECODE_PORT=$DECODE_PORT PROXY_PORT=$PROXY_PORT python -m pytest -s -v ${GIT_ROOT}/tests/v1/kv_connector/nixl_integration/test_edge_cases.py PREFILL_PORT=$PREFILL_PORT DECODE_PORT=$DECODE_PORT PROXY_PORT=$PROXY_PORT python -m pytest -s -v "${GIT_ROOT}"/tests/v1/kv_connector/nixl_integration/test_edge_cases.py
# Clean up before running next model # Clean up before running next model
cleanup_instances cleanup_instances

Some files were not shown because too many files have changed in this diff Show More