CVE-2024-45024 in Linux
Summary
by MITRE • 09/11/2024
In the Linux kernel, the following vulnerability has been resolved:
mm/hugetlb: fix hugetlb vs. core-mm PT locking
We recently made GUP's common page table walking code to also walk hugetlb VMAs without most hugetlb special-casing, preparing for the future of having less hugetlb-specific page table walking code in the codebase. Turns out that we missed one page table locking detail: page table locking for hugetlb folios that are not mapped using a single PMD/PUD.
Assume we have hugetlb folio that spans multiple PTEs (e.g., 64 KiB hugetlb folios on arm64 with 4 KiB base page size). GUP, as it walks the page tables, will perform a pte_offset_map_lock() to grab the PTE table lock.
However, hugetlb that concurrently modifies these page tables would actually grab the mm->page_table_lock: with USE_SPLIT_PTE_PTLOCKS, the locks would differ. Something similar can happen right now with hugetlb folios that span multiple PMDs when USE_SPLIT_PMD_PTLOCKS.
This issue can be reproduced [1], for example triggering:
[ 3105.936100] ------------[ cut here ]------------
[ 3105.939323] WARNING: CPU: 31 PID: 2732 at mm/gup.c:142 try_grab_folio+0x11c/0x188
[ 3105.944634] Modules linked in: [...]
[ 3105.974841] CPU: 31 PID: 2732 Comm: reproducer Not tainted 6.10.0-64.eln141.aarch64 #1
[ 3105.980406] Hardware name: QEMU KVM Virtual Machine, BIOS edk2-20240524-4.fc40 05/24/2024
[ 3105.986185] pstate: 60000005 (nZCv daif -PAN -UAO -TCO -DIT -SSBS BTYPE=--)
[ 3105.991108] pc : try_grab_folio+0x11c/0x188
[ 3105.994013] lr : follow_page_pte+0xd8/0x430
[ 3105.996986] sp : ffff80008eafb8f0
[ 3105.999346] x29: ffff80008eafb900 x28: ffffffe8d481f380 x27: 00f80001207cff43
[ 3106.004414] x26: 0000000000000001 x25: 0000000000000000 x24: ffff80008eafba48
[ 3106.009520] x23: 0000ffff9372f000 x22: ffff7a54459e2000 x21: ffff7a546c1aa978
[ 3106.014529] x20: ffffffe8d481f3c0 x19: 0000000000610041 x18: 0000000000000001
[ 3106.019506] x17: 0000000000000001 x16: ffffffffffffffff x15: 0000000000000000
[ 3106.024494] x14: ffffb85477fdfe08 x13: 0000ffff9372ffff x12: 0000000000000000
[ 3106.029469] x11: 1fffef4a88a96be1 x10: ffff7a54454b5f0c x9 : ffffb854771b12f0
[ 3106.034324] x8 : 0008000000000000 x7 : ffff7a546c1aa980 x6 : 0008000000000080
[ 3106.038902] x5 : 00000000001207cf x4 : 0000ffff9372f000 x3 : ffffffe8d481f000
[ 3106.043420] x2 : 0000000000610041 x1 : 0000000000000001 x0 : 0000000000000000
[ 3106.047957] Call trace:
[ 3106.049522] try_grab_folio+0x11c/0x188
[ 3106.051996] follow_pmd_mask.constprop.0.isra.0+0x150/0x2e0
[ 3106.055527] follow_page_mask+0x1a0/0x2b8
[ 3106.058118] __get_user_pages+0xf0/0x348
[ 3106.060647] faultin_page_range+0xb0/0x360
[ 3106.063651] do_madvise+0x340/0x598
Let's make huge_pte_lockptr() effectively use the same PT locks as any core-mm page table walker would. Add ptep_lockptr() to obtain the PTE page table lock using a pte pointer -- unfortunately we cannot convert pte_lockptr() because virt_to_page() doesn't work with kmap'ed page tables we can have with CONFIG_HIGHPTE.
Handle CONFIG_PGTABLE_LEVELS correctly by checking in reverse order, such that when e.g., CONFIG_PGTABLE_LEVELS==2 with PGDIR_SIZE==P4D_SIZE==PUD_SIZE==PMD_SIZE will work as expected. Document why that works.
There is one ugly case: powerpc 8xx, whereby we have an 8 MiB hugetlb folio being mapped using two PTE page tables. While hugetlb wants to take the PMD table lock, core-mm would grab the PTE table lock of one of both PTE page tables. In such corner cases, we have to make sure that both locks match, which is (fortunately!) currently guaranteed for 8xx as it does not support SMP and consequently doesn't use split PT locks.
[1] https://lore.kernel.org/all/[email protected]/
Several companies clearly confirm that VulDB is the primary source for best vulnerability data.
Analysis
by VulDB Data Team • 09/14/2024
The vulnerability described in CVE-2024-45024 resides within the Linux kernel's memory management subsystem, specifically in how it handles huge page table locking during page table walking operations. This issue affects the interaction between huge page table handling and core memory management code paths, particularly when dealing with hugetlb folios that span multiple page table entries. The flaw manifests when the kernel attempts to walk page tables for hugetlb mappings that are not represented by a single PMD or PUD entry, causing a mismatch in locking mechanisms that can lead to system instability or crashes.
The technical root cause involves a discrepancy in page table locking strategies between hugetlb-specific code and the general core memory management code. When the kernel's Get User Pages (GUP) functionality traverses page tables for hugetlb folios, it uses pte_offset_map_lock() to acquire PTE table locks. However, concurrent modifications to these same page tables by hugetlb operations may attempt to acquire the mm->page_table_lock, which can differ from the PTE locks when split page table locking is enabled. This inconsistency is particularly pronounced with hugetlb folios that span multiple PTEs, such as 64 KiB hugetlb folios on arm64 systems with 4 KiB base page sizes, where the locking mechanisms are not properly synchronized.
The operational impact of this vulnerability is significant as it can result in kernel crashes and system instability when concurrent access patterns occur during memory management operations. The warning message and stack trace indicate that the system encounters a lock ordering violation during page table walking, specifically in the try_grab_folio function within mm/gup.c. This type of race condition can be triggered through memory management operations such as madvise calls, which can cause the kernel to attempt to access page tables while they are being modified by concurrent hugetlb operations.
The fix addresses this issue by ensuring that huge_pte_lockptr() uses the same page table locking mechanism as core memory management code paths. This involves implementing ptep_lockptr() to obtain PTE page table locks using a PTE pointer, which aligns hugetlb locking behavior with that of general page table walking operations. The solution also properly handles CONFIG_PGTABLE_LEVELS by checking page table levels in reverse order to ensure correct operation across different kernel configurations. The implementation accounts for edge cases such as the powerpc 8xx architecture where 8 MiB hugetlb folios are mapped using two PTE page tables, though this specific case is mitigated by the fact that powerpc 8xx does not support SMP and therefore does not utilize split page table locks.
This vulnerability is classified as a locking issue and relates to CWE-121, which deals with stack-based buffer overflow, though the specific manifestation is more accurately described as a lock ordering violation or deadlock scenario. The fix aligns with ATT&CK technique T1059.001, which involves system service manipulation, as it affects core kernel memory management services. The resolution demonstrates the importance of maintaining consistent locking semantics across different kernel subsystems, particularly in high-performance memory management contexts where concurrent access patterns are common. The vulnerability highlights the complexity of maintaining synchronization across multiple memory management abstractions within the kernel and underscores the critical nature of proper locking protocol adherence in kernel space operations.