CVE-2026-105754 in vLLMinfo

Summary

by MITRE • 10/06/2026

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_hashes field, ranges in the features.mm_placeholders field, and wire-selected multimodal field processors without rebinding them to the active model renderer contract. Forged grid geometry, field types, or non-positive placeholder lengths can terminate the shared EngineCore; when an attacker knows or can induce a victim's content hash, forged cache hashes can poison or retrieve cross-request encoder-cache state; and dropped sparse placeholder masks can alter replayed transport semantics. This issue is fixed in version 0.30.0.

Be aware that VulDB is the high quality source for vulnerability data.

Analysis

by VulDB Data Team • 10/06/2026

The vulnerability identified in vLLM versions prior to 0.30.0 represents a critical failure in input validation within the disaggregated scale-out architecture of its inference engine. Specifically, the /inference/v1/generate endpoint accepts user-supplied data for tensor features, cache identifiers, and multimodal processing parameters without enforcing strict adherence to the active model renderer contract. This architectural flaw allows an attacker to inject malformed or maliciously crafted inputs into fields such as kwargs_data, mm_hashes, and mm_placeholders. The core issue lies in the lack of rebinding these external inputs to the internal state expectations of the engine's core components, creating a disconnect between user-provided data structures and the runtime environment's requirements for safe execution.

From a technical perspective, this vulnerability enables several distinct attack vectors that compromise both availability and integrity. First, forged grid geometry, incorrect field types, or non-positive placeholder lengths can cause catastrophic failures within the shared EngineCore process. These malformed inputs lead to memory corruption or unhandled exceptions that terminate the core inference service, resulting in a denial of service for all users relying on that instance. Second, if an attacker can determine or induce the victim's content hash, they can exploit the cache mechanism by submitting forged cache hashes. This allows for encoder-cache state poisoning, where malicious data is injected into shared memory structures used across requests, potentially leading to information leakage or inconsistent model outputs when other users request cached results based on those poisoned entries.

Furthermore, the manipulation of sparse placeholder masks presents a sophisticated threat vector involving transport semantics. By dropping or altering these masks, an attacker can alter how replayed transports are interpreted by the engine. This disruption affects the internal logic governing data flow and state management during inference, potentially leading to unpredictable behavior or further exploitation opportunities depending on the specific implementation details of the multimodal processors involved. The combination of these flaws means that a single malicious request can destabilize the entire serving infrastructure, affecting multiple concurrent users and compromising the reliability of large language model deployments.

This vulnerability aligns with CWE-20 Improper Input Validation and CWE-787 Out-of-bounds Write in cases where malformed geometry or lengths lead to memory corruption. In terms of offensive security frameworks, it maps to MITRE ATT&CK technique T1496 Resource Hijacking through denial of service via resource exhaustion, as well as potential data injection vectors under T1059 Command and Scripting Interpreter if the engine allows for arbitrary code execution through malformed tensor structures. The lack of strict type checking and boundary validation on critical inference parameters is a fundamental design flaw that undermines the security boundaries between user input and system core operations.

To mitigate this vulnerability, organizations must immediately upgrade vLLM to version 0.30.0 or later, where these input validation checks have been implemented. For environments unable to patch immediately, network-level controls should be employed to restrict access to the /inference/v1/generate endpoint only from trusted sources with known good IP addresses. Additionally, implementing strict schema validation on incoming JSON payloads can help filter out malformed requests before they reach the engine core. Monitoring logs for unusual patterns in tensor sizes or cache hash mismatches may also provide early detection of exploitation attempts. It is crucial to treat all user-supplied data related to model inference parameters as untrusted and enforce rigorous type, range, and structural validation at the API gateway level prior to processing by the vLLM engine.

Responsible

GitHub M

Reservation

10/05/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

very low

Sources

Do you need the next level of professionalism?

Upgrade your account now!