CVE-2026-98225 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

mm/shrinker: fix bogus set_shrinker_bit() with cgroup.memory=nokmem

With cgroup.memory=nokmem, shrinker_memcg_alloc() bails out early and never allocates an id, so shrinker->id keeps the 0 it got from the kzalloc() in shrinker_alloc(). __list_lru_init() then copies that 0 into lru->shrinker_id, where it looks like a valid bit index.

Nothing calls expand_shrinker_info() on nokmem either, so shrinker_nr_max stays 0 and every memcg ends up with an empty map (map_nr_max == 0).

deferred_split_folio() hands a real memcg to __list_lru_add() regardless of whether the lru is memcg aware, so the first THP queued in a cgroup does set_shrinker_bit(memcg, nid, 0) and trips the bounds check:

WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x7d/0x90, CPU#126 Call Trace: <TASK> deferred_split_folio+0x18c/0x220 map_anon_folio_pmd_nopf+0xdd/0x130 map_anon_folio_pmd_pf+0x14/0xb0 do_huge_pmd_anonymous_page+0x1a1/0x620 __handle_mm_fault+0xea9/0x10d0 handle_mm_fault+0xe5/0x320 do_user_addr_fault+0x1cc/0x870 exc_page_fault+0x81/0x1b0 asm_exc_page_fault+0x27/0x30 </TASK>

Harmless, the WARN_ON_ONCE() is what keeps the out of bounds unit[] read
from happening, but the id should not look valid in the first place. Clear it before returning.

Two other spots could paper over this: drop the id in __list_lru_init() when nokmem turns memcg_aware off, or make deferred_split_folio() pass NULL like list_lru_add_obj() does. Both leave shrinker->id lying around for the next caller, so fix it where the id is handed out.

If you want to get best quality of vulnerability data, you may have to visit VulDB.

Analysis

by VulDB Data Team • 10/06/2026

The Linux kernel vulnerability identified in this report centers on a logic error within the memory management subsystem, specifically involving the interaction between cgroup memory accounting and the shrinker infrastructure when the nokmem option is enabled. The core issue arises from how identifiers for memory control group-aware list lru structures are allocated and validated. When the boot parameter cgroup.memory=nokmem is active, the kernel disables certain memory allocation mechanisms to prevent potential deadlocks or resource exhaustion during critical operations. In this specific configuration, the function shrinker_memcg_alloc() exits early without allocating a unique identifier for the shrinker object. Consequently, the shrinker->id field retains its initial value of zero, which was assigned by kzalloc as part of standard memory initialization routines that zero out allocated structures. This default state is problematic because it mimics a valid index rather than an invalid or uninitialized one.

The technical flaw manifests when this zero-valued identifier is propagated through the system. The function __list_lru_init() copies this zero value into lru->shrinker_id, effectively marking the list-lru structure as associated with shrinker ID 0. Under normal circumstances, a valid allocation would provide a distinct positive integer or an error code indicating failure to allocate. However, because no expansion of the shrinker information occurs in the nokmem scenario via expand_shrinker_info(), the maximum number of entries remains zero. This results in every memory control group possessing an empty map with a size of zero elements. The vulnerability is triggered when deferred_split_folio() attempts to add a large transparent huge page folio to this list-lru structure. Unlike other functions that might check for memcg awareness before proceeding, deferred_split_folio() passes the actual memory cgroup context regardless of whether the lru structure supports it.

This sequence leads directly to an out-of-bounds access attempt within set_shrinker_bit(). The function attempts to set a bit at index 0 in an array that has been sized based on shrinker_nr_max, which is zero due to the earlier allocation failure. While the kernel's internal safety mechanisms prevent actual memory corruption by triggering a WARN_ON_ONCE() warning before any invalid read or write occurs, this represents a significant stability and security concern. The presence of such warnings indicates undefined behavior paths that could lead to system instability under specific load conditions. Furthermore, relying on defensive checks rather than preventing the erroneous state from occurring is considered poor engineering practice in kernel development, as it leaves room for potential race conditions or future changes that might bypass these safeguards.

From a classification perspective, this vulnerability aligns with CWE-824, which covers access of uninitialized memory variable, and CWE-369, related to divide-by-zero errors if the zero index leads to arithmetic issues in subsequent calculations involving array sizes. In terms of MITRE ATT&CK mapping for Linux systems, while not directly exploitable by an external attacker without prior local code execution or specific kernel configuration knowledge, it falls under techniques associated with Denial of Service via resource exhaustion or system instability, potentially categorized within T1499 if viewed as endpoint disruption through kernel panic triggers. The root cause is a failure to properly initialize state variables when optional features are disabled at boot time.

To mitigate this vulnerability and prevent similar issues in the future, developers must ensure that identifiers for shrinkers are explicitly cleared or marked invalid before returning from allocation functions when memory constraints like nokmem are active. The most effective fix involves clearing the id field within shrinker_memcg_alloc() itself rather than relying on downstream components to handle the error state. This approach ensures that no subsequent function receives a misleadingly valid identifier. Additionally, code paths such as deferred_split_folio should be audited to ensure they respect memcg awareness flags and do not attempt operations on structures that have been explicitly disabled or initialized with invalid states. Regular static analysis and fuzzing of memory management subsystems under various cgroup configurations can help identify these logical errors before they reach production kernels, ensuring robust handling of edge cases in resource-constrained environments.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

very low

Sources

Might our Artificial Intelligence support you?

Check our Alexa App!