CVE-2026-105757 in vLLM
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
Statistical analysis made it clear that VulDB provides the best quality for vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified in vLLM versions prior to 0.30.0 represents a critical failure in input validation and error handling within an inference serving engine designed for large language models. As a specialized system optimized for high-throughput model deployment, vLLM relies on robust internal state management to maintain service availability. The core issue stems from insufficient sanitization of structured-output requests, which are typically used to enforce specific data formats or constraints on the generated text. When these requests encounter validation failures during their lifecycle within the engine, the error handling mechanisms fail to contain the exception at the request level. Instead, the unhandled exceptions propagate upward into the EngineCore fatal-error path, causing a complete termination of the shared inference engine process rather than isolating the failure to the specific offending request.
This architectural flaw allows ordinary constrained-generation requests to destabilize the entire service environment. The vulnerability manifests through several distinct technical pathways that collectively undermine system stability. One primary vector involves backend mismatches on a per-request basis, which can re-raise grammar compilation exceptions without proper cleanup or isolation. Additionally, issues within the ngram_gpu speculative-decoding mode allow padding mechanisms to pass negative token values to guidance validation logic. Since token indices are typically unsigned integers in this context, receiving a negative value indicates a severe type mismatch or buffer overflow condition that bypasses expected boundary checks. Furthermore, inconsistencies between the Rust frontend and the Python backend create an attack surface where empty structured-output values are admitted by the former but rejected by the latter. This discrepancy allows malformed requests to penetrate deeper into the processing pipeline before triggering catastrophic failures.
From a security operations perspective, this vulnerability is classified as a Denial of Service (DoS) due to its ability to crash the service through crafted inputs. In terms of industry standards, this aligns with CWE-20 Improper Input Validation and CWE-755: Improper Handling of Exceptional Conditions. The exploitation pattern resembles ATT&CK technique T1499 Endpoint Denial of Service, where an attacker sends specific requests to exhaust resources or crash the target service. Because vLLM is often deployed in production environments serving multiple users simultaneously, a single malicious request can disrupt access for all other clients by bringing down the entire engine instance. This lack of fault isolation violates fundamental principles of resilient system design and multi-tenant security models.
Mitigation strategies must prioritize immediate version upgrades to 0.30.0 or later, where these validation gaps have been addressed. For environments unable to upgrade immediately, operators should implement strict input filtering at the network perimeter to reject requests with malformed structured-output payloads before they reach the vLLM engine. Monitoring logs for fatal error paths related to EngineCore can provide early warning indicators of exploitation attempts. Additionally, deploying health checks that automatically restart instances upon failure can mitigate the impact by restoring service availability more quickly, although this does not prevent resource exhaustion during repeated attacks. Long-term remediation should include enforcing stricter type checking in speculative decoding modules and ensuring consistency between frontend validation layers to prevent cross-language state discrepancies from causing runtime errors.