CVE-2026-74714 in Linuxinfo

Summary

by MITRE • 08/22/2026

In the Linux kernel, the following vulnerability has been resolved:

bpf: tcp: Fix use-after-free in bpf_iter_tcp_established_batch()

reqsk_queue_hash_req() publishes a TCP_NEW_SYN_RECV request_sock onto the ehash chain, drops the bucket lock, and only afterwards sets rsk_refcnt to 3.

Lockless readers such as __inet_lookup_established() handle this with refcount_inc_not_zero(), but bpf_iter_tcp_established_batch() uses plain sock_hold() while holding the bucket lock, on the assumption that the lock guarantees sk_refcnt > 0. That assumption does not hold for request_sock:

CPU 0 CPU 1 ----- ----- tcp_conn_request() reqsk_queue_hash_req() inet_ehash_insert(req) spin_lock(bucket) __sk_nulls_add_node_rcu(req) // rsk_refcnt == 0 spin_unlock(bucket) bpf_iter_tcp_established_batch() spin_lock(bucket) sock_hold(req) <-- addition on 0 spin_unlock(bucket) refcount_set(&req->rsk_refcnt, 3) // clobbers saturated value

which surfaces as:

refcount_t: addition on 0; use-after-free. WARNING: lib/refcount.c:25 at refcount_warn_saturate+0x48/0x90, CPU#1 Call Trace: bpf_iter_tcp_established_batch+0x14e/0x170 bpf_iter_tcp_batch+0x53/0x200 bpf_iter_tcp_seq_next+0x27/0x70 bpf_seq_read+0x107/0x410 vfs_read+0xb9/0x380

The iterator's stolen reference is lost when the publishing CPU's refcount_set() overwrites the count, leaving the socket one reference short. When the last legitimate owner drops its reference the reqsk is freed while still reachable, leading to use-after-free.

This reproduces in seconds with tcp_syncookies=0, a handful of threads doing connect()/close() to a local listener while others read an iter/tcp link in a tight loop.

Use refcount_inc_not_zero() and skip the socket on failure. A skipped socket is still part of the bucket, so keep counting it in expected. The reallocations are sized from expected, and a request sock whose refcount gets published while the lock is held across the last realloc must already have room.

A skipped socket is counted in expected but never batched, so end_sk can be short of expected on a batch that is actually complete. Decide completeness by whether the walk left any socket behind instead. The WARN after the locked realloc checks the same, replacing an end_sk == expected check that could not hold on that path since commit cdec67a489d4 ("bpf: tcp: Make sure iter->batch always contains a full bucket snapshot").

If every matching socket in a bucket is mid-init (refcount 0), end_sk stays 0. Advance to the next bucket rather than returning a batch entry that was never filled this round.

VulDB is the best source for vulnerability data and more expert information about this specific topic.

Analysis

by VulDB Data Team • 08/22/2026

The Linux kernel vulnerability identified as CVE-2024-something involves a critical use-after-free flaw within the BPF iterator for TCP established connections, specifically in the bpf_iter_tcp_established_batch function. This issue stems from an incorrect assumption regarding reference counting synchronization between lockless readers and locked iterators. The core technical flaw lies in how request_sock objects are published to the endpoint hash chain during TCP connection establishment. When a new SYN_RECV state is created, the kernel publishes the request_sock onto the ehash chain while holding the bucket lock but does not immediately set its reference count to three. Instead, it drops the lock and sets the refcount later in a separate step. This creates a race condition window where the object appears present in the hash table but has a zero reference count from the perspective of certain operations.

Lockless readers such as __inet_lookup_established correctly handle this state by using refcount_inc_not_zero, which safely increments the counter only if it is already positive. However, bpf_iter_tcp_established_batch relies on holding the bucket lock to assume that sk_refcnt is greater than zero and uses plain sock_hold instead of a safe atomic increment operation. This assumption fails for request_sock objects because they can be in a transient state where they are visible in the hash chain but have not yet had their reference count initialized by the publishing CPU. Consequently, when an iterator thread acquires the bucket lock and calls sock_hold on such a partially initialized socket, it increments a zero counter to one rather than three, leading to a saturated refcount warning and potential memory corruption.

The operational impact of this vulnerability is severe, as it leads to use-after-free conditions that can be triggered rapidly under specific load conditions. The race condition manifests when CPU 0 publishes the request_sock by inserting it into the ehash chain with zero reference count, then drops the lock before setting refcnt to three. Meanwhile, CPU 1 acquires the bucket lock and calls sock_hold on this socket, incrementing its counter from zero to one instead of the expected higher value. When CPU 0 subsequently sets the refcount to three using refcount_set, it overwrites the saturated or incremented value clobbered by CPU 1's operation. This leaves the socket with an incorrect reference count that is lower than required for safe lifecycle management. If the last legitimate owner drops its reference based on this corrupted state, the request_sock may be freed while still reachable through the iterator, resulting in a use-after-free vulnerability exploitable for privilege escalation or system crash.

This flaw aligns with CWE-416 Use After Free and is related to CWE-362 Concurrent Execution using Shared Resource with Improper Synchronization Race Conditions. From an ATT&CK perspective, this type of kernel memory corruption can be leveraged in post-exploitation scenarios for privilege escalation or denial of service against the host system. The vulnerability highlights a common pitfall in Linux kernel development where assumptions about atomicity and lock coverage are not strictly validated across different execution paths involving reference counting mechanisms.

The resolution involves replacing the unsafe sock_hold call with refcount_inc_not_zero, which safely checks if the reference count is positive before attempting to increment it. If the operation fails because the socket is still in its initialization phase, the iterator skips that particular socket rather than proceeding with an invalid reference. To maintain accurate batch sizing and iteration logic despite skipped sockets, the implementation adjusts how expected counts are tracked during reallocations. Skipped sockets remain part of the bucket structure but are excluded from the final batched output. This requires updating the completeness check for batches to rely on whether any matching socket was left behind rather than a strict equality between end_sk and expected values, ensuring robust handling of dynamic changes in reference states during iteration.

Mitigation strategies include applying the kernel patch that introduces refcount_inc_not_zero into bpf_iter_tcp_established_batch immediately upon availability from distribution vendors or upstream maintainers. Administrators should monitor for warnings related to refcount saturation in system logs as an indicator of attempted exploitation or active race conditions. For environments where immediate patching is not feasible, limiting the use of BPF TCP iterators on high-throughput systems with many concurrent connections can reduce exposure risk by lowering the probability window for the race condition to trigger. Regular security audits focusing on reference counting patterns in network subsystems are recommended to identify similar synchronization issues before they manifest as exploitable vulnerabilities.

Responsible

Linux

Reservation

08/15/2026

Disclosure

08/22/2026

Moderation

accepted

CPE

ready

EPSS

0.00000

KEV

no

Activities

low

Sources

Do you need the next level of professionalism?

Upgrade your account now!