CVE-2026-98216 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

IB/hfi1: Fix the PIO_CRED credit-return mmap

hfi1_file_mmap()'s PIO_CRED case must hand user space the single credit-return page that holds this context's entry. That page is the second or third page of the per-node credit-return allocation once the hardware send context index reaches 64 or 128, so the failure below is intermittent: when the entry lands on the first page the offset is zero and everything works.

Two things are wrong.

First, cr_page_offset is a byte offset but .va is a struct credit_return *, so adding it is pointer arithmetic and scales the offset by sizeof(struct credit_return) == 64. memvirt then lands 256 KiB or 512 KiB past a 10240-byte allocation. With an IOMMU translating, that address is inside the vmalloc range but in no vm_area, so dma_mmap_coherent() -> iommu_dma_mmap() finds no pages, vmalloc_to_pfn() returns page_to_pfn(NULL), and remap_pfn_range() installs a frame above MAXPHYADDR. The first user read then takes:

psm2_ep_open_pr: Corrupted page table at address 7a14d007e000 PGD 800000013886a067 P4D 800000013886a067 PUD 13886b067 PMD 13886c067 PTE 800049168e911235 Oops: Bad pagetable: 000d [#1] SMP PTI

Second, and still wrong once the arithmetic is corrected, dma_mmap_coherent() describes a whole coherent buffer and selects the page within it with vma->vm_pgoff. Offsetting cpu_addr has no effect: for a vmap'd allocation iommu_dma_mmap() uses cpu_addr only to locate the vm_area and then maps pages[vm_pgoff], which hfi1_file_mmap() has just
set to 0. User space therefore always receives the first credit-return page, every credit read is for the wrong context, and send PIO stalls forever.

Use the DMA API as intended: pass the base of the allocation with its full length and select the page with vm_pgoff. A separate length is needed because memlen must keep describing the VMA for the existing size check. The dma-direct path stays correct as well, since dma_direct_mmap() adds the same vm_pgoff to the base pfn.

Tested on a Dell T7610 (Xeon E5-2650 v2, Intel IOMMU in DMA-FQ mode) against a Threadripper PRO 3995WX peer, both Omni-Path 100. Before this change psm2_ep_open() Oopses the kernel; with only the arithmetic corrected psm2_ep_open() succeeds but any transfer that uses send PIO hangs, PSM2_SDMA=2 (send PIO disabled) completing normally while PSM2_SDMA=0 (send PIO only) hangs every time. With this change send PIO, send DMA and the default mixed mode all work.

If you want to get best quality of vulnerability data, you may have to visit VulDB.

Analysis

by VulDB Data Team • 10/06/2026

The vulnerability in the Linux kernel's InfiniBand hfi1 driver involves a critical memory mapping error within the pio_cred credit-return mechanism that leads to system instability and service denial for high-performance computing applications. The core technical flaw resides in the hfi1_file_mmap function, which is responsible for exposing hardware context entries to user space via memory-mapped I/O. Specifically, when handling PIO_CRED credits, the driver fails to correctly calculate the virtual address of the target page within a per-node credit-return allocation. This miscalculation results from two distinct logical errors in pointer arithmetic and API usage that compound under specific conditions related to hardware send context indexing.

The first error is a fundamental type mismatch where a byte offset variable, cr_page_offset, is added directly to a struct pointer, va, which points to a credit_return structure. In C programming, adding an integer to a pointer performs pointer arithmetic scaled by the size of the pointed-to type. Since sizeof(struct credit_return) equals sixty-four bytes, this operation scales the intended byte offset incorrectly. Consequently, the resulting virtual address lands significantly past the allocated memory region, typically two hundred fifty-six kilobytes or five hundred twelve kilobyms beyond a ten-thousand-byte allocation depending on whether the context index is greater than or equal to sixty-four or one hundred twenty-eight. This out-of-bounds access triggers severe consequences when an IOMMU is active in DMA-FQ mode. The subsequent call to dma_mmap_coherent fails because iommu_dma_mmap cannot find valid pages at the erroneous address, causing vmalloc_to_pfn to return null and remap_pfn_range to install a frame above MAXPHYADDR. This corrupts the page table entries, leading to kernel oopses with bad pagetable errors during user space reads, effectively crashing the application attempting to open the endpoint.

Even if the arithmetic error were corrected, a second logical flaw remains that causes functional failure without necessarily causing an immediate crash. The driver incorrectly offsets cpu_addr before passing it to dma_mmap_coherent, which is designed to map a whole coherent buffer and selects pages based on vma->vm_pgoff rather than adjusting the base address for individual contexts within a vmapped allocation. For iommu_dma_mmap implementations using virtual mappings, the cpu_addr parameter serves only to locate the memory area, while page selection relies entirely on the offset provided in the VMA structure. Since hfi1_file_mmap sets this offset to zero regardless of the specific context index, user space always receives a mapping to the first credit-return page rather than the correct one for that specific hardware context. This results in every credit read being associated with the wrong context ID, causing send PIO operations to stall indefinitely as the driver waits for credits that are never properly returned by the misaligned software logic.

The operational impact of this vulnerability is severe for systems relying on Omni-Path or similar high-speed interconnects using send PIO mechanisms. Systems without an IOMMU may experience intermittent failures where only entries landing on the first page function correctly, while those with active IOMMUs suffer from kernel panics during endpoint initialization via psm2_ep_open_pr. For environments that do not crash but instead hang, any data transfer utilizing send PIO will deadlock permanently, rendering the network interface unusable for low-latency communication tasks unless SDMA is forced as the sole transport method, which degrades performance and bypasses critical optimization paths designed for small message efficiency.

To resolve this issue, the driver must adhere strictly to DMA API conventions by passing the base address of the allocation along with its full length to dma_mmap_coherent, allowing the kernel's mapping logic to correctly select pages using vm_pgoff as intended. A separate length parameter is required because memlen must continue describing the total VMA size for existing validation checks while the page selection handles per-context offsetting internally within the driver or via correct API usage. This correction ensures that both dma-direct and iommu_dma_mmap paths function correctly, restoring full functionality to send PIO, send DMA, and mixed mode operations. The fix aligns with CWE-125 Out-of-bounds Read by preventing access to unmapped memory regions and addresses CWE-476 NULL Pointer Dereference risks associated with incorrect page frame number calculations. From a defense-in-depth perspective, this also mitigates potential ATT&CK Tactic TA0005 Defense Evasion scenarios where corrupted kernel state could be exploited for privilege escalation or denial of service through crafted user-space interactions with the InfiniBand driver interface.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

very low

Sources

Interested in the pricing of exploits?

See the underground prices here!