| Título | vllm-project vLLM v0.27.1 Denial of Service |
|---|
| Descrição | A denial-of-service vulnerability exists in vLLM when the OpenAI-compatible server is started with `--enable-prompt-embeds`. The issue was reproduced in vLLM 0.27.1, and the same affected input-handling logic was present on the main branch at the time of reporting.
An API client that is permitted to submit completion requests can trigger a fatal CUDA device-side assertion by sending a batch of concurrent `/v1/completions` requests containing `prompt_embeds` together with a sampling penalty, such as `repetition_penalty`, `presence_penalty`, or `frequency_penalty`. The failure occurs when requests with different prompt lengths are batched and the longest prompt in the batch is supplied through `prompt_embeds`.
The vulnerability is caused by inconsistent handling of embedding-based prompts in the request batching and sampling paths. For ordinary token-based prompts, vLLM writes the request's token IDs into the batch's `token_ids_cpu` buffer. For a `prompt_embeds` request, however, no token IDs exist, so the corresponding prompt positions in `token_ids_cpu` are not populated. vLLM separately records that these positions are embeddings rather than token IDs using an `is_token_ids` mask, and the model-input path respects this distinction. The sampling-penalty path does not.
When a repetition, presence, or frequency penalty is requested, vLLM constructs a prompt-token tensor from `token_ids_cpu` and passes its contents to a `scatter_add_` operation used to calculate token occurrence counts. The penalty path does not exclude positions representing `prompt_embeds`. As a result, values remaining in the reused `token_ids_cpu` batch rows can be interpreted as vocabulary token IDs even though they do not correspond to valid token IDs for the current request.
Under heterogeneous batching, the maximum prompt length of the batch determines how much of each row is read. If an embedding-based request establishes the batch's maximum prompt length, stale values from these unwritten positions can therefore be included in the penalty calculation. A stale value outside the valid vocabulary-index range is subsequently used as an index by the CUDA scatter/gather kernel, resulting in a `scatter gather kernel index out of bounds` device-side assertion.
The CUDA assertion is fatal to vLLM's `EngineCore`. The failure is not isolated to the request that triggers it: all in-flight requests fail, subsequent requests return HTTP 500 errors, and the affected engine does not recover automatically. Restoring service requires restarting the vLLM process. Consequently, a remote API client can cause denial of service for all users sharing the affected vLLM engine by submitting a small burst of otherwise valid embedding-based requests when prompt embeddings are enabled.
The issue does not require malformed embedding dimensions, invalid model data, chunked prefill, mixed token-and-embedding traffic, or privileged access to the host. Testing showed that either `repetition_penalty` alone or `presence_penalty`/`frequency_penalty` is sufficient to reach the vulnerable path. Token-only requests under the same batching conditions do not trigger the failure, and embedding requests without sampling penalties do not reach the vulnerable penalty-processing path. |
|---|
| Fonte | ⚠️ https://github.com/vllm-project/vllm/issues/57266 |
|---|
| Utilizador | Zyz3366 (UID 97230) |
|---|
| Submissão | 20/09/2026 00h17 (há 16 dias) |
|---|
| Moderação | 05/10/2026 22h26 (16 days later) |
|---|
| Estado | Aceite |
|---|
| Entrada VulDB | 413808 [vllm-project vLLM até 0.31.0 Completions Request mamba_mixer2.py conv_ssm_forward Divulgação de Informação] |
|---|
| Pontos | 20 |
|---|