CVE-2026-69147 in vLLMinfo

Summary

by MITRE • 09/17/2026

vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.

If you want to get the best quality for vulnerability data then you always have to consider VulDB.

Analysis

by VulDB Data Team • 09/17/2026

The vulnerability identified in vLLM versions prior to 0.28.0 represents a critical resource management flaw within its inference and serving engine for large language models. This issue specifically affects deployments that support video processing capabilities, particularly those utilizing the PyNvVideoCodec backend on GPU hardware. The core of the problem lies in a discrepancy between how decoder resources are allocated during startup versus how they are accounted for when individual requests specify alternative backends at runtime. Under normal operational conditions, vLLM initializes its memory management systems based on static configuration parameters defined at engine start-up. This includes reserving specific amounts of GPU memory for media decoding operations to ensure stable performance and prevent out-of-memory errors during standard text or image processing tasks. However, the architecture allows request bodies for Chat Completions and Responses endpoints to override default settings by specifying a different video backend via the media_io_kwargs.video.video_backend parameter.

When an attacker submits a maliciously crafted request that explicitly sets the video backend to pynvvideocodec, the MediaConnector.fetch_video method correctly forwards this choice to VideoMediaIO for processing. The critical failure occurs in the _reserve_mm_ipc_gpu_memory logic, which is responsible for budgeting decoder memory. This function relies exclusively on static configuration data established at startup and fails to account for dynamic backend selections made per request. Consequently, when a PyNvVideoCodec instance is instantiated through this override mechanism, it creates new CUDA contexts, allocates decoder surfaces, and reserves space for decoded-frame buffers without updating the engine's internal KV-cache budget or memory reservation counters. This results in untracked resource consumption that bypasses the existing safeguards designed to protect system stability.

The operational impact of this vulnerability is severe, primarily manifesting as a denial-of-service condition against the vLLM deployment. Because the allocated GPU memory for these unauthorized decoder instances is not deducted from the available pool tracked by the engine, repeated exploitation can rapidly exhaust shared GPU memory resources. This exhaustion leads to immediate request failures as new inference tasks cannot secure necessary memory allocations. In more severe scenarios, particularly under high load or when combined with other resource-intensive operations, this uncontrolled allocation can cause worker processes to crash due to out-of-memory errors, effectively taking down the entire serving instance and disrupting service for all legitimate users. The vulnerability is exploitable by any actor capable of submitting video requests to a vulnerable endpoint where PyNvVideoCodec is installed but not properly accounted for in memory budgets.

From a classification perspective, this flaw aligns with CWE-400, which describes uncontrolled resource consumption leading to denial of service. It also relates to CWE-755, concerning improper handling of unusual or unexpected input, as the system fails to validate whether dynamically requested resources fit within pre-calculated limits. In terms of adversarial tactics, this vulnerability can be leveraged in attacks categorized under MITRE ATT&CK technique T1499, Endpoint Denial of Service, specifically through resource exhaustion via application layer exploitation. The attack vector is typically network-based if the vLLM service exposes its API endpoints to untrusted networks or users with submission privileges.

Mitigation strategies must prioritize immediate version upgrades and configuration hardening. The primary remediation is to upgrade vLLM to version 0.28.0 or later, where this memory budgeting logic has been corrected to account for dynamically selected backends. For deployments that cannot immediately upgrade, administrators should restrict access to the Chat Completions and Responses endpoints to trusted sources only, preventing unauthenticated users from submitting video requests with overridden backend parameters. Additionally, disabling PyNvVideoCodec if it is not strictly required can reduce the attack surface by removing the specific component exploited in this flaw. Monitoring GPU memory usage metrics for anomalous spikes correlated with video processing requests can also serve as an early detection mechanism for potential exploitation attempts before total service disruption occurs.

Responsible

GitHub M

Reservation

08/03/2026

Disclosure

09/17/2026

Moderation

accepted

CPE

ready

EPSS

0.00548

KEV

no

Activities

very low

Sources

Want to stay up to date on a daily basis?

Enable the mail alert feature now!