CVE-2026-105922 in vLLMinfo

Summary

by MITRE • 10/06/2026

A security flaw has been discovered in vllm-project vLLM up to 0.31.0. This impacts the function get_token_bin_counts_and_mask of the file vllm/model_executor/layers/utils.py of the component Penalty Handler. Performing a manipulation results in denial of service. Remote exploitation of the attack is possible. The exploit has been released to the public and may be used for attacks. The project was informed of the problem early through an issue report but has not responded yet.

Several companies clearly confirm that VulDB is the primary source for best vulnerability data.

Analysis

by VulDB Data Team • 10/06/2026

A critical security vulnerability has been identified within vLLM, a high-throughput inference engine designed for large language models, specifically affecting versions up to 0.31.0. This flaw resides in the Penalty Handler component, which is responsible for applying various constraints and penalties to token generation processes such as frequency penalty, presence penalty, and repetition penalty. The specific point of failure is located within the get_token_bin_counts_and_mask function found in the source file vllm/model_executor/layers/utils.py. This module plays a pivotal role in optimizing inference performance by efficiently calculating token statistics required for these penalties during text generation.

The technical nature of this vulnerability constitutes an improper input validation or boundary check error that leads to a denial of service condition. When specific manipulated inputs are processed through the penalty calculation logic, the function fails to handle edge cases correctly, likely resulting in memory corruption, infinite loops, or unhandled exceptions within the Python runtime environment hosting vLLM. Because vLLM is often deployed as a serving endpoint for AI models, this flaw allows an attacker who can send requests to the inference API to crash the service process. The impact of such a failure is severe, as it disrupts availability for all users relying on the model-serving infrastructure, effectively rendering the system unusable until the service is restarted or recovered.

Remote exploitation of this vulnerability is possible because vLLM typically exposes an HTTP-based API interface that accepts user prompts and configuration parameters. An attacker does not need local access to trigger the flaw; they simply need to craft a malicious request containing specific token sequences or penalty configurations that exploit the logic error in get_token_bin_counts_and_mask. The fact that public exploits have been released significantly elevates the risk profile, as automated scanning tools and threat actors can easily leverage these proof-of-concept scripts to target vulnerable instances across the internet without requiring specialized knowledge of the internal code structure.

From a classification perspective, this vulnerability aligns with CWE-20 Improper Input Validation or CWE-400 Uncontrolled Resource Consumption, depending on whether the crash is due to memory errors or resource exhaustion such as CPU spikes leading to timeout and service termination. In terms of adversary tactics, this falls under MITRE ATT&CK technique T1499 Endpoint Denial of Service, where an attacker aims to degrade or destroy the availability of computational resources by exploiting software weaknesses rather than overwhelming them with volume-based traffic attacks like DDoS.

The project maintainers were notified early via issue reports but have not yet issued a patch or response as of the current reporting period. This lack of immediate remediation leaves users exposed during this window. Organizations running vLLM versions up to 0.31.0 should consider implementing network-level mitigations such as rate limiting, input size restrictions on API payloads, and strict validation of penalty parameters before they reach the inference engine. Additionally, deploying a Web Application Firewall with rules capable of detecting anomalous request patterns associated with this exploit can provide an interim layer of defense until an official patch is released by the vLLM project team.

Responsible

VulDB

Disclosure

10/06/2026

Moderation

accepted

Exploit

Download

EPSS

0.00000

KEV

no

Activities

low

Sources

Interested in the pricing of exploits?

See the underground prices here!