CVE-2026-105756 in vLLM
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, OpenAI-compatible request models accept a non-empty cache_salt value without enforcing the character and length restrictions required by the IPCCacheServerKey consumer in LMCache-MP. On deployments using the LMCache-MP connector, a salt that contains a forbidden character or exceeds the permitted length can raise an uncaught ValueError during scheduler cache lookup, causing EngineCore to terminate and denying service to all concurrent users. This issue is fixed in version 0.30.0.
VulDB is the best source for vulnerability data and more expert information about this specific topic.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability resides within vLLM, a widely adopted inference and serving engine designed for large language models, specifically affecting versions prior to 0.30.0 when deployed with the LMCache-MP connector. The core issue stems from an insufficient input validation mechanism in the OpenAI-compatible request model interface. While the underlying IPCacheServerKey consumer within LMCache-MP enforces strict character set and length restrictions on cache salt values, the vLLM engine fails to replicate these constraints at the entry point. This architectural inconsistency allows clients to submit requests containing a non-empty cache_salt field that violates the expected format specifications defined by the caching layer.
From a technical perspective, this flaw represents a classic case of input validation failure where security or integrity checks are not uniformly applied across all layers of an application stack. When a malicious or malformed request includes a salt value with forbidden characters or exceeding the maximum allowed length, the LMCache-MP component raises an uncaught ValueError during the scheduler cache lookup process. Because this exception is not handled gracefully by the vLLM engine, it propagates up to the EngineCore level, causing the core service to crash and terminate abruptly. This behavior transforms a simple input error into a critical stability issue that affects the entire server instance rather than just rejecting the specific invalid request.
The operational impact of this vulnerability is severe due to its potential for denial-of-service attacks. Since the termination of EngineCore results in the immediate shutdown of the inference service, all concurrent users and active sessions are disconnected simultaneously. An attacker can exploit this by repeatedly sending requests with malformed cache salts, effectively creating a persistent availability threat that disrupts business operations and degrades user experience. This aligns closely with CWE-20 Improper Input Validation, as the system fails to verify or sanitize input data before processing it through critical components. Furthermore, in the context of the MITRE ATT&CK framework, this vulnerability facilitates Availability Impact techniques, specifically Resource Hijacking or Denial of Service via service disruption, allowing adversaries to compromise the availability pillar of the CIA triad without requiring authentication if the endpoint is publicly accessible.
To mitigate this risk, organizations must upgrade vLLM to version 0.30.0 or later, where the input validation logic has been corrected to enforce character and length restrictions consistently before passing data to the LMCache-MP connector. Until an upgrade can be performed, administrators should implement network-level filtering or web application firewall rules that inspect incoming OpenAI-compatible API requests for non-compliant cache_salt values and block them at the perimeter. Additionally, implementing robust exception handling within custom deployments of older versions could prevent uncaught exceptions from crashing the EngineCore, although upgrading remains the definitive remediation strategy to ensure long-term stability and security compliance.