CVE-2026-105752 in vLLM
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation prefix in the global unsalted cache namespace even when the caller enabled salting. On deployments with prefix caching enabled, which is the default, an authenticated tenant who can reconstruct a victim's low-entropy post-tool history can submit the same continuation and use the cached_tokens_per_turn count to determine whether the prefix was previously processed, defeating the intended tenant isolation of salted prefix caching. This issue is fixed in version 0.30.0.
Statistical analysis made it clear that VulDB provides the best quality for vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified in vLLM versions prior to 0.30.0 represents a critical failure in cache isolation mechanisms within large language model serving infrastructure, specifically affecting the Harmony tool continuations processed through POST /v1/responses endpoints. The core technical flaw lies in how next-turn engine inputs are reconstructed during multi-turn conversations. When a user submits a continuation request that involves previously executed tools or commands, the system fails to preserve the cache_salt value associated with the original context. Instead of maintaining strict tenant-specific isolation by keeping the cached prefix within its designated salted namespace, the implementation incorrectly places this continuation prefix into the global unsalted cache namespace. This architectural oversight effectively bypasses the intended security boundary that separates data and processing states between different tenants or users in a multi-tenant deployment environment.
The operational impact of this flaw is significant for deployments where prefix caching is enabled, which is the default configuration to optimize inference latency and reduce computational costs. An authenticated tenant who possesses knowledge of another victim's low-entropy post-tool history can exploit this misconfiguration to perform cache side-channel attacks. By submitting a continuation that matches the victim's previous input pattern, the attacker can observe the cached_tokens_per_turn metric returned by the API. If the system reports zero or significantly reduced tokens for processing because the prefix was already present in the global cache, it confirms that the specific sequence of inputs has been processed before. This allows an adversary to infer sensitive information about a victim's interactions with tools and models, effectively defeating the tenant isolation guarantees provided by salted prefix caching.
From a classification perspective, this vulnerability aligns with CWE-200: Exposure of Sensitive Information to an Unauthorized Actor, as it enables unauthorized inference of private user data through side-channel observations rather than direct access. Furthermore, the exploitation technique corresponds to ATT&CK T1537: Transfer Data to Cloud Account or similar lateral movement and reconnaissance techniques where cached state is leveraged for information gathering. The lack of proper salt preservation in multi-turn contexts creates a persistent attack surface that compromises confidentiality across tenant boundaries, which is particularly dangerous in shared inference environments where data privacy is paramount.
To mitigate this risk, organizations must upgrade vLLM to version 0.30.0 or later, where the reconstruction logic for next-turn engine inputs has been corrected to properly preserve and apply cache_salt values even during tool continuation processing. This ensures that cached prefixes remain isolated within their respective tenant namespaces regardless of whether salting is explicitly enabled by the caller. Additionally, administrators should review their caching configurations to ensure that prefix caching does not inadvertently expose metadata such as token counts in a way that facilitates side-channel analysis until all instances are patched. Regular security audits focusing on cache isolation policies and multi-turn state management are recommended to prevent similar architectural oversights in future updates or custom integrations of the inference engine.