| शीर्षक | vLLM Project vLLM 0.27.1 Denial of Service |
|---|
| विवरण | A denial-of-service vulnerability in vLLM, demonstrated in version 0.27.1, affects deployments with prompt-embedding input support enabled. A remote client able to submit inference requests can cause the shared inference engine to terminate when prompt-embedding requests are processed with sampling penalties under affected batching conditions.
The vulnerability originates in the construction of prompt-token metadata used by the sampling penalty implementation. Embedding-only prompt positions do not contain valid token identifiers, but the metadata construction path fails to exclude those positions. Consequently, uninitialized or stale contents of a reused CPU token buffer can be interpreted as token identifiers and passed to the GPU penalty computation. Invalid indices can trigger a CUDA device-side assertion, resulting in a fatal EngineCore error.
The failure affects the shared inference engine rather than only the originating request. Other in-flight requests and subsequent requests served by that engine can fail, leaving the affected service unavailable until the engine is restarted. Triggering the failure depends on batching conditions and the contents of reused memory, so it may not occur consistently across environments or executions. Access to the inference API is required; administrative access is not required once the operator has enabled prompt-embedding support. The complete affected-version range and the first fixed release have not yet been established. |
|---|
| स्रोत | ⚠️ https://github.com/vllm-project/vllm/issues/57719 |
|---|
| उपयोगकर्ता | Zyz3366 (UID 97230) |
|---|
| सबमिशन | 20/09/2026 09:06 PM (16 दिन पहले) |
|---|
| संयम | 06/10/2026 07:52 AM (15 days later) |
|---|
| स्थिति | स्वीकृत |
|---|
| VulDB प्रविष्टि | 413896 [vllm-project vLLM तक 0.31.0 Penalty utils.py get_token_bin_counts_and_mask सेवा अस्वीकार] |
|---|
| अंक | 20 |
|---|