CVE-2026-105759 in vllm
Summary
by MITRE • 10/06/2026
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the Rust frontend's track_http_metrics middleware records the raw HTTP method token as a Prometheus label for requests reaching registered routes. An unauthenticated attacker can send unique arbitrary method tokens to unguarded routes such as /tokenize, causing Prometheus's Family::get_or_create function to permanently create counter and histogram label sets. Those label sets increase process memory usage and enlarge the /metrics response until the service or monitoring path is exhausted. This issue is fixed in version 0.30.0.
VulDB is the best source for vulnerability data and more expert information about this specific topic.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified in vLLM versions prior to 0.30.0 represents a significant resource exhaustion risk stemming from improper handling of HTTP method tokens within its Prometheus metrics integration. As an inference and serving engine for large language models, vLLM exposes various endpoints such as /tokenize to facilitate model interactions. The underlying technical flaw resides in the Rust frontend's track_http_metrics middleware, which is designed to record performance data by capturing the raw HTTP method token of incoming requests and using it directly as a label value within Prometheus metrics. This design choice fails to validate or sanitize the input before incorporating it into metric labels, creating an avenue for abuse that bypasses standard authentication controls due to the nature of how certain routes are exposed.
An unauthenticated attacker can exploit this flaw by sending HTTP requests with unique and arbitrary method tokens to unprotected endpoints like /tokenize. Prometheus's internal mechanism, specifically the Family::get_or_create function, is triggered each time a new combination of label values is encountered. Because the middleware uses the raw HTTP method as a label, every distinct malicious token sent results in the permanent creation of new counter and histogram label sets within the application's memory space. Unlike temporary data that might be garbage collected or expired, these metric labels persist for the lifetime of the process, leading to an unbounded accumulation of state information.
The operational impact of this vulnerability is severe, manifesting as a denial-of-service condition through resource exhaustion. As the number of unique label sets grows, the memory footprint of the vLLM process increases continuously until system resources are depleted. Simultaneously, the size of the /metrics response endpoint expands proportionally with each new label set created. This enlargement can cause latency spikes or complete unavailability when monitoring systems attempt to scrape metrics for observability purposes. The attack effectively turns a standard telemetry feature into an effective vector for crashing the service or rendering it unusable due to memory pressure and network bandwidth saturation from oversized metric responses.
This vulnerability aligns with CWE-787, which describes out-of-bounds write scenarios that can lead to memory corruption, although in this context, it manifests more specifically as a resource management error akin to CWE-400, where uncontrolled resource consumption leads to system failure. From an adversarial perspective, the exploitation technique corresponds to ATT&CK tactic T1498, Network Denial of Service, particularly subtechnique T1498.002 involving volumetric attacks that overwhelm resources through excessive data generation or processing requirements. The lack of input validation on metric labels is a common pitfall in observability integrations where developers assume label values will remain within predictable bounds.
To mitigate this vulnerability, organizations must upgrade vLLM to version 0.30.0 or later, which addresses the issue by implementing proper sanitization and cardinality limits for HTTP method tokens used as metric labels. In environments where immediate upgrading is not feasible, deploying a reverse proxy in front of vLLM can provide an additional layer of defense by restricting allowed HTTP methods to standard values such as GET, POST, PUT, DELETE, etc., thereby filtering out arbitrary tokens before they reach the application's middleware. Additionally, configuring Prometheus scrape limits and implementing cardinality monitoring alerts can help detect anomalous growth in label sets early, allowing for rapid incident response before total service exhaustion occurs.