CVE-2026-105755 in vLLM
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, flash late-interaction scoring at the /score and /rerank endpoints derives each worker's query_key value from the caller-controlled X-Request-Id header. A concurrent request that reuses a victim's identifier can overwrite the cached query embedding so the victim's documents are scored against the attacker's query, and shared use counters can also cause a late-interaction cache-miss error. This issue is fixed in version 0.30.0.
If you want to get best quality of vulnerability data, you may have to visit VulDB.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified in vLLM versions prior to 0.30.0 represents a critical flaw within the inference serving engine's handling of large language model scoring operations, specifically affecting the flash late-interaction scoring mechanism utilized by the /score and /rerank endpoints. This component is designed to efficiently compute relevance scores between queries and documents using cached embeddings to optimize performance for high-throughput scenarios. The core technical deficiency lies in the method used to isolate worker state during concurrent requests. Specifically, the system derives each worker's query_key value directly from the X-Request-Id header provided by the client. This design choice assumes that request identifiers are globally unique and immutable throughout their lifecycle, a assumption that is frequently violated in distributed systems where connection pooling or proxy layers may reuse identifier values across different logical sessions.
The operational impact of this flaw manifests as an authorization bypass leading to data integrity compromise and potential information disclosure. Because the query_key serves as the primary key for caching worker-specific embeddings, an attacker can exploit this by crafting a concurrent request that reuses a victim's X-Request-Id header value. When such a request is processed, it overwrites the cached query embedding associated with that identifier in the shared memory space of the vLLM workers. Consequently, when the original victim subsequently sends their legitimate scoring or reranking request using the same identifier, the system retrieves the attacker's overwritten query embedding instead of the victim's intended one. This results in the victim's documents being scored against the attacker's query context rather than the victim's own intent.
This manipulation leads to severe consequences for data confidentiality and integrity. The victim receives relevance scores based on an unrelated or maliciously crafted query, which can be used to infer sensitive information about the dataset if the scoring patterns reveal details about document content that should remain private under normal operational parameters. Furthermore, the shared use counters associated with these cached embeddings are susceptible to race conditions caused by concurrent access from multiple requests sharing the same identifier. This contention leads to late-interaction cache-miss errors and inconsistent state updates, causing service degradation or denial of service for legitimate users attempting to utilize the scoring functionality. The vulnerability effectively allows an unauthenticated attacker to manipulate the computational context of other users' requests through simple header manipulation without requiring direct access to internal system memory.
From a classification perspective, this issue aligns with CWE-20 Improper Input Validation and CWE-367 Time-of-check Time-of-use (TOCTOU) Race Condition, as it involves improper handling of user-supplied identifiers that lead to race conditions in shared resource management. In the context of the MITRE ATT&CK framework, this vulnerability facilitates lateral movement within a multi-tenant inference environment by allowing an attacker to influence the output of other tenants' requests, potentially falling under techniques related to Data Manipulation or Defense Evasion if used to obscure malicious intent through corrupted scoring results. The exploitation relies on HTTP header manipulation and concurrent request timing, highlighting risks associated with shared state in web-serving architectures that do not enforce strict isolation between client sessions at the worker level.
To mitigate this vulnerability, organizations must upgrade vLLM to version 0.30.0 or later where the issue has been resolved by implementing more robust mechanisms for generating and managing query keys that are independent of user-controlled headers. In environments where immediate upgrading is not feasible, defensive measures should include enforcing strict validation on incoming X-Request-Id values to ensure they conform to expected formats and lengths, although this may not fully prevent reuse attacks if the identifier space is small or predictable. Additionally, implementing request isolation at the proxy level by ensuring that each logical session maintains a distinct connection pool can reduce the likelihood of header reuse. Monitoring for unusual patterns in cache miss rates or scoring latency spikes may also help detect ongoing exploitation attempts involving concurrent requests with identical identifiers.