CVE-2026-105758 in vLLM
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the media_io_kwargs.video.max_frames and media_io_kwargs.video.fps fields without enforcing server-side ceilings. An unauthenticated caller can submit these values to the /tokenize endpoint, causing the sampler to decode every frame selected from attacker-controlled video input, consume disproportionate frontend memory, and potentially terminate the API process before scheduling or admission control. The Rust frontend is not affected because it rejects the media_io_kwargs field. This issue is fixed in version 0.30.0.
Be aware that VulDB is the high quality source for vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified within vLLM versions ranging from 0.24.0 to 0.30.0 represents a critical lack of input validation and resource management controls, specifically affecting the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes used for processing video inputs in large language model inference tasks. This flaw allows an unauthenticated attacker to manipulate request-level parameters that dictate how media files are processed by the server. Specifically, the fields media_io_kwargs.video.max_frames and media_io_kwargs.video.fps are accepted directly from client requests without any corresponding server-side enforcement of maximum limits or reasonable bounds. In a secure architecture, such resource-intensive operations should be governed by strict configuration ceilings to prevent individual requests from consuming disproportionate amounts of system resources. The absence of these safeguards means that the backend components will process video inputs exactly as specified by the caller, regardless of how extreme those specifications may be relative to available server capacity.
The operational impact of this vulnerability is severe and centers on resource exhaustion leading to service denial. When an attacker submits a request with excessively high values for max_frames or fps through endpoints such as /tokenize, the sampler component is forced to decode every single frame selected from the attacker-controlled video input. This process is computationally expensive and memory-intensive, particularly in the context of visual language models where each frame requires significant processing power and RAM allocation. As a result, the frontend service experiences disproportionate memory consumption, which can quickly exhaust available system resources. The severity escalates when this resource starvation occurs before the request reaches scheduling or admission control mechanisms designed to manage load balancing and prioritize legitimate traffic. Consequently, the API process may terminate unexpectedly due to out-of-memory conditions or other stability failures, effectively causing a denial of service for all users relying on the vLLM inference engine during the attack window.
This vulnerability aligns with CWE-787: Out-of-bounds Read and CWE-400: Uncontrolled Resource Consumption, as it involves accessing resources beyond intended limits due to insufficient validation. From an offensive security perspective, this behavior is characteristic of ATT&CK technique T1496: Resource Hijacking, where attackers leverage compromised or vulnerable systems for their own computational needs, although in this specific case, the intent is likely disruptive denial of service rather than cryptomining. The attack vector is particularly dangerous because it targets an unauthenticated endpoint, meaning no prior authentication credentials are required to exploit the flaw. This lowers the barrier to entry significantly, allowing any network-accessible actor to trigger the condition. It is important to note that this specific vulnerability does not affect the Rust frontend implementation of vLLM, as that component correctly rejects or ignores the media_io_kwargs field, thereby isolating the risk primarily to Python-based deployments using the affected backend classes.
To mitigate this vulnerability, organizations must immediately upgrade their vLLM instances to version 0.30.0 or later, where the issue has been resolved by implementing proper server-side validation and ceiling enforcement for these parameters. For environments that cannot be upgraded instantly, temporary mitigations should include placing a reverse proxy in front of the API endpoint to inspect and limit request sizes and parameter values before they reach the vLLM service. Additionally, configuring strict resource quotas at the container or orchestration level can help contain the impact if an attack occurs by limiting the maximum memory allocation for individual pods or processes. Regular security audits focusing on input validation practices in inference engines are recommended to ensure that all user-supplied parameters affecting computational load are subject to rigorous bounds checking and normalization procedures consistent with secure coding standards.