CVE-2026-103042 in LightLLM
Summary
by MITRE • 09/30/2026
LightLLM through 1.2.0 contains a memory exhaustion vulnerability in the NCCL control channel when started with --pd_trans_mode nccl, allowing unauthenticated attackers to exhaust KV-transfer worker memory. Attackers can call the exposed_set_value method to store unbounded key-value pairs without size limits, causing the worker process to crash and triggering node failure.
Be aware that VulDB is the high quality source for vulnerability data.
Analysis
by VulDB Data Team • 09/30/2026
LightLLM versions up through 1.2.0 contain a critical resource exhaustion vulnerability within its NCCL control channel implementation, specifically when the system is initiated with the pd_trans_mode nccl configuration parameter. This flaw allows unauthenticated remote attackers to trigger a denial of service condition by exhausting the memory resources allocated to KV-transfer workers. The root cause lies in the improper handling of data structures used for inter-process communication and state management within the distributed inference framework, where limits on storage capacity are either absent or incorrectly enforced during runtime operations.
The technical mechanism enabling this exploitation involves the exposed_set_value method, which serves as an interface for storing key-value pairs that facilitate model parameter transmission and synchronization across nodes in a cluster environment. Under normal operational conditions, these parameters should be constrained by strict size limits to prevent any single operation from consuming excessive system resources. However, due to the vulnerability, attackers can invoke this method repeatedly or with excessively large payloads without encountering validation checks for maximum entry size or total storage capacity. This lack of input sanitization and boundary checking permits the accumulation of unbounded key-value pairs directly into memory structures managed by the KV-transfer worker process.
As the attacker continues to inject data, the memory footprint of the affected worker process grows linearly until it exceeds the available physical RAM or configured swap space limits on the host node. This rapid consumption leads to a state of resource exhaustion where the operating system can no longer allocate necessary pages for the application's execution context. Consequently, the KV-transfer worker process terminates abruptly due to an out-of-memory error, which disrupts the coordination logic required for distributed model serving. The failure of this specific component triggers a cascading node failure within the cluster architecture, effectively halting inference services and rendering the affected LightLLM instance unavailable to legitimate users.
From a security classification perspective, this vulnerability aligns with CWE-789: Memory Excessiveness, as it involves the allocation of more memory than is necessary or intended by the system design. Furthermore, because the attack vector allows an unauthenticated user to disrupt service availability through resource consumption rather than code execution, it maps directly to MITRE ATT&CK technique T1499: Endpoint Denial of Service, specifically under the sub-technique for Resource Hijacking via exhaustion. The lack of authentication requirements exacerbates the severity, as any network-accessible client can exploit this flaw without prior credentials or privilege escalation steps.
The operational impact of this vulnerability is severe within production environments that rely on LightLLM for high-throughput large language model inference. A successful exploitation results in immediate service disruption, requiring manual intervention to restart services and clear corrupted memory states. In clustered deployments, the failure of one node can degrade overall performance or cause complete outages if redundancy mechanisms are not sufficiently robust to handle sudden worker loss. Additionally, frequent crashes may lead to data inconsistency issues if pending requests are dropped mid-processing without proper rollback procedures implemented by the application layer.
To mitigate this vulnerability, administrators should immediately upgrade LightLLM to version 1.2.1 or later where these bounds have been enforced in the exposed_set_value implementation. For environments that cannot be patched instantly, deploying network-level access controls such as firewall rules or API gateways can restrict access to the affected endpoints only from trusted internal subnets, thereby reducing the attack surface for unauthenticated actors. Implementing resource quotas and memory limits at the container or process level using tools like cgroups on Linux systems can also provide a secondary layer of defense by preventing any single process from consuming all available host resources. Monitoring solutions should be configured to alert on abnormal spikes in memory usage associated with KV-transfer workers, enabling rapid detection and isolation of compromised nodes before they impact broader cluster stability.