CVE-2026-105753 in vLLMinfo

Summary

by MITRE • 10/06/2026

vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is rejected. A later request reusing the same media hash causes MultiModalProcessorSenderCache to send no payload and MultiModalReceiverCache to reach an assertion with the message "Expected a cached item," producing a shared-service availability failure. This issue is fixed in version 0.28.0.

If you want to get the best quality for vulnerability data then you always have to consider VulDB.

Analysis

by VulDB Data Team • 10/06/2026

The vulnerability identified in vLLM, specifically affecting versions prior to 0.28.0, represents a critical state management flaw within the multimodal processing pipeline of this large language model inference engine. As an open-source library designed for high-throughput serving and inference of LLMs, vLLM relies on complex caching mechanisms to optimize performance by avoiding redundant computations or data transfers. The specific defect resides in the interaction between the frontend sender cache and the backend receiver cache during multimodal rendering operations. Under normal operation, these caches coordinate to ensure that media payloads associated with a request are correctly transmitted from the client-facing component to the internal processing engine. However, the default configuration utilizes a mirrored Least Recently Used (LRU) cache structure for handling multimodal data hashes, which introduces a race condition and state inconsistency when requests are rejected during the admission control phase.

The technical root cause of this vulnerability lies in an asymmetric update mechanism within the caching layer. When a request containing multimedia content is processed, the system computes a hash to identify the media asset. In versions prior to 0.28.0, if the engine's admission controller rejects the request due to resource constraints or policy violations after the frontend sender cache has already committed the media hash, a desynchronization occurs. The frontend component records the presence of this media hash in its local cache state as part of the initial processing steps. However, because the overall request is rejected before it reaches the engine receiver stage, the corresponding payload data is never transmitted to or stored within the receiver's cache. This creates an inconsistent state where one side of the communication channel acknowledges the existence and caching of a specific media hash, while the other side remains unaware of its presence.

The operational impact manifests when subsequent requests attempt to reuse this same media hash. Since the frontend sender cache believes it has already cached or processed the item associated with that hash, it optimizes by sending no payload data for the new request, assuming the receiver can retrieve it from its own local storage. However, because the previous instance was rejected and never populated the receiver's cache, the MultiModalReceiverCache attempts to locate a non-existent entry. This mismatch triggers an assertion failure within the codebase, specifically raising an exception with the message "Expected a cached item." In a production environment running vLLM as a shared service for multiple users or applications, this unhandled exception typically results in the termination of the affected worker process or thread. Consequently, this leads to a denial-of-service condition where available inference capacity is reduced, and depending on the deployment architecture, may cause broader availability failures for clients attempting to access multimodal LLM capabilities.

From a security classification perspective, this vulnerability aligns with CWE-400: Uncontrolled Resource Consumption if viewed through the lens of resource exhaustion leading to service degradation, or more accurately CWE-829: Inclusion of Functionality from Untrusted Control Sectors due to improper validation of state assumptions between distributed components. It also maps to MITRE ATT&CK technique T1534: Internal Spearphishing in contexts where an attacker might exploit this denial-of-service condition to disrupt service availability, although the primary vector here is accidental rather than malicious exploitation. The flaw highlights a common challenge in high-performance systems where optimization techniques like caching can introduce subtle state inconsistencies if not rigorously synchronized across all processing stages.

To mitigate this vulnerability and restore stable operation, organizations must upgrade vLLM to version 0.28.0 or later, which contains the necessary patches to synchronize cache states correctly during request rejection scenarios. For environments where immediate upgrading is not feasible due to dependency constraints, temporary mitigations should focus on reducing the likelihood of admission rejections by ensuring sufficient computational resources are allocated to handle peak multimodal workloads. Additionally, implementing robust monitoring and alerting for worker process crashes or assertion failures can help detect these incidents early, allowing operators to restart affected instances before they significantly impact service availability. Long-term architectural reviews should also consider whether mirrored cache states across different processing stages require stricter consistency checks or transactional guarantees to prevent similar desynchronization issues in future updates.

Responsible

GitHub M

Reservation

10/05/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

low

Sources

Do you want to use VulDB in your project?

Use the official API to access entries easily!