CVE-2026-98257 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

rds: ib: use rds_conn_drop() on protocol version mismatch

rds_ib_cm_connect_complete() runs from the RDMA-CM event handler with conn->c_cm_lock held. When the peer negotiates a protocol version older than RDS_PROTOCOL_COMPAT_VERSION, the handler calls rds_conn_destroy(), which is only safe in the rmmod path: it synchronously tears the connection down and flush_work()es the shutdown work cp_down_w.

That shutdown work (rds_conn_shutdown()) needs cp_cm_lock, which is the very lock the event handler still holds, so the flush never completes: the two workers wait on each other and the RDS connection workqueues stall for good.

All other RDMA-CM failure paths (REJECTED, CONNECT_ERROR, DISCONNECTED) use rds_conn_drop(), which marks the connection RDS_CONN_ERROR and schedules the shutdown work asynchronously. Use it here as well.

If you want to get best quality of vulnerability data, you may have to visit VulDB.

Analysis

by VulDB Data Team • 10/06/2026

The vulnerability resides within the Reliable Datagram Sockets over InfiniBand implementation of the Linux kernel, specifically in the connection management logic that handles RDMA Connection Manager events. The core issue is a deadlock condition triggered during protocol version negotiation failures. When an RDS connection attempt encounters a peer with an older or incompatible protocol version, the event handler function rds_ib_cm_connect_complete executes while holding the cp_cm_lock mutex. In this specific error path, the code incorrectly invokes rds_conn_destroy to tear down the connection. This function is designed for use only during module removal and performs synchronous cleanup operations that include flushing shutdown work items via flush_work on the cp_down_w queue.

The critical technical flaw arises because the scheduled shutdown worker, specifically rds_conn_shutdown(), also requires acquisition of the same cp_cm_lock mutex to proceed with its teardown tasks. Since the event handler thread already holds this lock and is blocked waiting for the flush operation to complete, while the worker thread waits for the lock held by the event handler, a classic circular dependency deadlock occurs. This mutual blocking causes the RDS connection workqueues to stall indefinitely, effectively freezing any further processing of RDMA events or data transfers associated with that connection context until system reboot or manual intervention.

From an operational impact perspective, this vulnerability leads to local denial of service conditions for services relying on RDS over InfiniBand. The stalled workqueue prevents the kernel from properly cleaning up resources and can block subsequent network operations, potentially affecting high-performance computing clusters or distributed storage systems that depend on reliable RDMA communication. Although it does not allow remote code execution, the ability to trigger this deadlock remotely by initiating a connection with an incompatible protocol version poses a significant availability risk in production environments where such connections are established dynamically.

This issue is categorized under CWE-833, Deadlock, as it involves two or more threads waiting for each other to release resources held in a circular chain of dependencies. In the context of the MITRE ATT&CK framework, this vulnerability aligns with techniques related to resource exhaustion and denial of service, specifically within the persistence or disruption categories where system functionality is impaired through kernel-level blocking mechanisms rather than application-layer crashes.

The resolution involves replacing the synchronous rds_conn_destroy call with rds_conn_drop in the protocol version mismatch path. This function marks the connection state as RDS_CONN_ERROR and schedules the shutdown work asynchronously, allowing the event handler to release the cp_cm_lock immediately without waiting for the worker thread to complete its tasks. This change aligns the error handling logic with other RDMA-CM failure paths such as REJECTED, CONNECT_ERROR, and DISCONNECTED, which already utilize rds_conn_drop to ensure safe and non-blocking connection teardown.

To mitigate this vulnerability in systems not yet patched, administrators should monitor for stalled RDS workqueues or unresponsive InfiniBand connections that may indicate the presence of this deadlock condition. Updating the Linux kernel to a version containing the fix is the primary remediation strategy. Additionally, ensuring strict protocol version compatibility between connected endpoints can prevent triggering the specific code path responsible for the issue until patches are applied across all nodes in the cluster environment.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00184

KEV

no

Activities

very low

Sources

Do you know our Splunk app?

Download it now for free!