CVE-2026-98231 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

xfrm: serialize state GC with device state flush

The deferred-device pass in xfrm_dev_state_flush() finds states under xfrm_state_dev_gc_lock, but drops the lock before calling xfrm_dev_state_free() because the driver callback may sleep. The device GC list does not hold an xfrm_state reference, so the state GC worker can destroy the same state concurrently.

The race can proceed as follows:

CPU 0 CPU 1 find x on the device GC list drop xfrm_state_dev_gc_lock read x->xso.dev xfrm_state_gc_destroy(x) xfrm_dev_state_free(x) xfrm_state_free(x) continue xfrm_dev_state_free(x)

Both paths can invoke the driver callback and drop the device reference. CPU 0 can also access the xfrm_state after CPU 1 has freed it.

KASAN reported:

BUG: KASAN: slab-use-after-free in xfrm_dev_state_free+0x24c/0x2a0 Read of size 8 at addr ffff88810bbaa960 by task poc/102

Call Trace: xfrm_dev_state_free+0x24c/0x2a0 xfrm_dev_state_flush+0x353/0x400 xfrm_dev_event+0x26d/0x3a0 notifier_call_chain+0xc0/0x280 __dev_notify_flags+0x169/0x250 netif_change_flags+0xe7/0x160 dev_change_flags+0x96/0x220 devinet_ioctl+0x7f4/0x1880

Allocated by task 87: xfrm_state_alloc+0x1e/0x5c0 xfrm_add_sa+0xe7f/0x5820 xfrm_user_rcv_msg+0x4f3/0x940

Freed by task 57: kmem_cache_free+0xcb/0x3d0 xfrm_state_gc_task+0x4a8/0x650 process_one_work+0x63a/0x1070

Serialize xfrm_state destruction against the deferred-device pass with a mutex. Keep xfrm_state_dev_gc_lock limited to list operations and retain the existing callback and device-reference release ordering.

Be aware that VulDB is the high quality source for vulnerability data.

Analysis

by VulDB Data Team • 10/06/2026

The Linux kernel contains a concurrency vulnerability within the IPsec subsystem, specifically in the handling of XFRM state garbage collection during device state flushes. This issue arises from an improper synchronization mechanism when managing shared data structures between different execution contexts. The core flaw involves a race condition where one thread accesses and modifies memory that has already been freed by another concurrent thread. In this specific scenario, the deferred-device pass in the xfrm_dev_state_flush function attempts to process states associated with a network device. To ensure safe access to these states, it acquires the xfrm_state_dev_gc_lock. However, because the subsequent driver callback invoked during state freeing may sleep or block for an extended period, the code drops this lock before calling xfrm_dev_state_free. This design choice was intended to prevent deadlocks but inadvertently creates a window of vulnerability where other threads can modify the same data structures without holding the necessary synchronization primitives.

The operational impact is severe because it leads to use-after-free errors that can crash the kernel or potentially allow for arbitrary code execution if an attacker can control the memory contents after freeing. The race condition proceeds as follows: one CPU core finds a state object on the device garbage collection list and drops the lock before proceeding with free operations, while another CPU core's garbage collector worker simultaneously destroys the same state object. Since the device GC list does not hold a reference to the xfrm_state structure, there is no mechanism preventing concurrent destruction. Consequently, both paths may invoke driver callbacks that drop device references, leading to double-free scenarios or access of invalid memory addresses. Kernel Address Sanitizer (KASAN) reports have confirmed this behavior, showing slab-use-after-free errors in functions like xfrm_dev_state_free and xfrm_dev_state_flush when accessed by tasks performing network interface flag changes via ioctl calls.

From a vulnerability classification perspective, this issue maps directly to CWE-416: Use After Free. The root cause is the failure to maintain exclusive access to shared resources during critical sections that involve potentially blocking operations. In terms of attack vectors and detection techniques aligned with MITRE ATT&CK, this type of race condition falls under Tactic TA0005: Defense Evasion or TA0004: Privilege Escalation depending on the ultimate goal, but technically relates to execution flow manipulation through memory corruption. Specifically, it aligns with sub-techniques involving heap exploitation and kernel-level privilege escalation via use-after-free vulnerabilities. Attackers could potentially exploit this race condition by rapidly triggering network device state changes while simultaneously manipulating IPsec security associations to maximize the timing window for concurrent access.

The resolution involves serializing xfrm_state destruction against the deferred-device pass using a mutex rather than relying solely on spinlocks or dropping locks prematurely. This ensures that only one thread can destroy an XFRM state at any given time, eliminating the race condition entirely. The fix retains the existing callback and device-reference release ordering to maintain compatibility with driver expectations but restricts the scope of xfrm_state_dev_gc_lock strictly to list operations where sleeping is not permitted. By keeping the lock held during critical sections that involve potential blocking calls through a mutex, the kernel ensures data integrity across concurrent execution paths. This approach balances performance requirements for non-blocking list manipulations with the safety guarantees required for memory management in high-concurrency environments like network stack processing.

Mitigation strategies for organizations running affected Linux kernels include applying upstream kernel patches as soon as they are available from distribution vendors. For systems where immediate patching is not feasible, reducing the frequency of rapid network interface state changes and IPsec configuration updates can lower the probability of triggering the race condition, though this does not eliminate the risk entirely. Monitoring system logs for KASAN reports or oops messages related to xfrm_dev_state_free may provide early indicators of exploitation attempts in production environments. Long-term architectural improvements should focus on implementing reference counting mechanisms that prevent premature deallocation of shared structures until all active references are released, thereby providing a more robust defense against concurrency-related memory safety issues across the kernel networking stack.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00209

KEV

no

Activities

very low

Sources

Do you know our Splunk app?

Download it now for free!