CVE-2026-98166 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

drm/ttm: fix swapped-out resources never leaving their bulk_move range

ttm_tt_swapout() returns the number of pages swapped out on success and a negative error code on failure; for a populated ttm it never returns zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm_bo_swapout() infinite LRU walk on swapout failure") moved the bulk_move bookkeeping in ttm_bo_swapout_cb() under "if (!ret)", so the ttm_resource_del_bulk_move_unevictable() / ttm_resource_move_to_lru_tail() pair is now skipped on every successful swapout. The equivalent change for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm_bo_shrink() infinite LRU walk on backup failure") tests "lret > 0", which is what was intended here as well.

Before b2ed01e7ad3d the resource was taken off the bulk_move before the swapout; since then a swapped-out resource stays inside its BO's bulk_move range (and on the manager LRU) although it is unevictable. When it is later freed or the BO leaves the bulk_move (ttm_resource_free(), ttm_bo_set_bulk_move() via amdgpu_vm_bo_del()), ttm_resource_del_bulk_move() skips it because of its !ttm_resource_unevictable() guard, so a range endpoint in pos->first / pos->last is left pointing at freed memory. The next ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() on that cursor is a use-after-free, seen as the resv WARN in ttm_lru_bulk_move_add(), "list_del corruption" in ttm_resource_move_to_lru_tail() or a NULL dereference in ttm_resource_manager_next() -- minutes to hours after a hibernation, or at process exit / reboot following one. Samuel Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the dangling cursor; the missing removal at swapout time is the reason it dangles.

Testing the condition for success restores the removal. On an AMD Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug crashed 5 of 18 hibernation cycles; a function profile of one hibernation showed 336 ttm_tt_swapout() calls and zero ttm_resource_del_bulk_move_unevictable() calls. With this change the removal happens for every swapped-out resource and 12 further cycles were clean.

If you want to get best quality of vulnerability data, you may have to visit VulDB.

Analysis

by VulDB Data Team • 10/07/2026

The vulnerability in question resides within the Linux kernel's Translation Table Maps (TTM) subsystem, specifically affecting how memory resources are managed during swap operations. The core technical flaw is a logic error introduced by commit b2ed01e7ad3d, which altered the control flow of ttm_bo_swapout_cb(). This function handles callbacks for swapping out buffer objects to secondary storage. Previously, the code correctly removed resources from their bulk_move range before initiating the swap operation. However, the patch moved this bookkeeping logic under a conditional check that verifies if the return value indicates failure. Since ttm_tt_swapout() returns the number of pages swapped on success and never zero for populated objects, successful swaps result in a non-zero positive integer. Consequently, the condition evaluating for errors incorrectly treats these successes as failures or simply skips the cleanup block entirely because it expects a specific error code structure that does not align with the actual return value semantics. This results in swapped-out resources remaining erroneously within their buffer object's bulk_move range and on the manager's Least Recently Used (LRU) list, despite being marked as unevictable.

The operational impact of this defect is severe, manifesting primarily as memory safety violations that can lead to system instability or crashes. Because the resource remains in the bulk_move structure while its underlying memory may be freed later via ttm_resource_free() or when the buffer object leaves the bulk move context, a dangling cursor persists within the data structures. When subsequent operations such as ttm_lru_bulk_move_tail() or ttm_resource_add_bulk_move() attempt to interact with this stale cursor, they trigger use-after-free conditions. These manifest in various ways depending on timing and memory state, including reserved object warnings detected by kernel assertions like resv WARNs, list deletion corruption errors within ttm_resource_move_to_lru_tail(), or null pointer dereferences during iteration via ttm_resource_manager_next(). The latency between the initial swap event and the crash can range from minutes to hours, often occurring after hibernation cycles, process exits, or system reboots. This delayed execution makes debugging difficult as the root cause is temporally distant from the symptom manifestation.

From a security perspective, this vulnerability aligns with CWE-416, Use After Free, which occurs when software uses memory after it has been freed, leading to undefined behavior and potential exploitation for arbitrary code execution or denial of service. The attack vector leverages the kernel's internal memory management subsystem rather than external network inputs, classifying it under local privilege escalation vectors if an attacker can trigger specific hibernation patterns or resource allocation sequences that exhaust memory and force swap operations. In terms of MITRE ATT&CK frameworks, this relates to techniques involving exploitation of software flaws for persistence or impact, specifically falling under the category of Exploitation for Impact due to the system crashes observed during testing on AMD Phoenix APUs. The defect highlights critical risks in kernel-level resource management where incorrect state transitions can compromise memory integrity over extended periods.

Mitigation requires applying the upstream fix that corrects the conditional logic within ttm_bo_swapout_cb(). By ensuring that resources are removed from their bulk_move range regardless of whether the swap operation succeeds or fails, provided it is not a fatal error preventing any progress, the dangling cursor issue is resolved. System administrators should update to kernel versions containing this patch, particularly those targeting stable branches like 7.0.y where backports have been applied. Testing indicates that correcting the condition restores proper removal of swapped-out resources, eliminating the use-after-free scenarios and stabilizing hibernation cycles. Organizations relying on Linux kernels for desktop or server environments with significant memory pressure should prioritize this update to prevent intermittent crashes associated with power management states such as suspend-to-hibernate.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00198

KEV

no

Activities

medium

Sources

Do you need the next level of professionalism?

Upgrade your account now!