Logo
Explore Help
Sign In
Karylab-cklius/vllm
Watch 1
Star 0
Fork 1
Code Issues Pull Requests 1 Actions Packages Projects Releases Wiki Activity
Files
0a9ef0cfce131300fea219952b6e85c8dee45578
vllm/tests
T
History
Adrian AbeytaGitHubLuka Govedič
0a9ef0cfce Move query quantization to attention layer for Flashinfer & Triton. (#26534)
Signed-off-by: adabeyta <aabeyta@redhat.com>
Signed-off-by: Adrian Abeyta <aabeyta@redhat.com>
Co-authored-by: Luka Govedič <ProExpertProg@users.noreply.github.com>
2025-10-15 19:01:38 -04:00
..
basic_correctness
…
benchmarks
…
compile
Move query quantization to attention layer for Flashinfer & Triton. (#26534)
2025-10-15 19:01:38 -04:00
config
…
cuda
…
detokenizer
…
distributed
…
engine
…
entrypoints
…
evals
…
kernels
…
kv_transfer
…
lora
…
model_executor
…
models
…
multimodal
…
plugins
…
plugins_tests
…
prompts
…
quantization
…
reasoning
…
samplers
…
speculative_decoding/speculators
…
standalone_tests
…
system_messages
…
tokenization
…
tool_use
…
tools
…
tpu
…
transformers_utils
…
utils_
…
v1
…
vllm_test_utils
…
weight_loading
…
__init__.py
…
ci_envs.py
…
conftest.py
…
test_config.py
…
test_embedded_commit.py
…
test_envs.py
…
test_inputs.py
…
test_logger.py
…
test_outputs.py
…
test_pooling_params.py
…
test_regression.py
…
test_routing_simulator.py
Convert formatting to use ruff instead of yapf + isort (#26247)
2025-10-05 07:06:22 -07:00
test_scalartype.py
…
test_seed_behavior.py
…
test_sequence.py
…
test_triton_utils.py
…
test_version.py
…
test_vllm_port.py
…
utils.py
…
Powered by Gitea Version: 1.27.1 Page: 5747ms Template: 8ms
Auto
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API