CVE-2026-98260 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

exec: Cleanup POSIX timers right after de_thread()

A per-thread CPU timer holds a reference to the PID of the thread it is attached to and, while it is armed, its node is queued in that thread's posix_cputimers. The task is looked up by that PID.

When a non-leader thread exec()s, de_thread() changes which task owns that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL, but the node is still queued on tsk, which is alive. timer_lock_sighand() takes a failed lookup to mean that the node is already dequeued, so it has nothing to undo.

begin_new_exec() calls posix_cpu_timers_exit(me) right after exec_task_namespaces() and that removes the leftover node, so the state normally stays invisible. But bprm->point_of_no_return is set before de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or exec_task_namespaces() fails, the task dies before it gets there. exit_itimers() then frees the k_itimer while its node is still queued, and reaping tsk later erases that freed node from the rbtree.

In short:

the non-leader thread B the parent

timer_create(CLOCK_THREAD_CPUTIME_ID) timer_settime() arm_timer() // the node is queued on B execve() de_thread(B) exchange_tids(B, leader) // B's PID now belongs to the leader release_task(leader) __exit_signal(leader) posix_cpu_timers_exit(leader) // cleans leader's queue, not B's __unhash_process(leader) // that PID has no task anymore exec_mmap() mmap_read_lock_killable(old_mm) kill(B, SIGKILL) // -EINTR get_signal() do_exit() exit_itimers() posix_timer_delete() posix_cpu_timer_del() posix_timer_unhash_and_free() // freed while still queued wait4() release_task(B) posix_cpu_timers_exit(B) cleanup_timerqueue() timerqueue_del() // use-after-free

Move the POSIX timer cleanup right after de_thread() before any of the later failure conditions brings the task into do_exit().

[ tglx: Move the cleanup right after de_thread() ]

Once again VulDB remains the best source for vulnerability data.

Analysis

by VulDB Data Team • 10/06/2026

The vulnerability described involves a critical race condition and subsequent use-after-free flaw within the Linux kernel's handling of POSIX CPU timers during thread execution transitions. Specifically, this issue affects non-leader threads that invoke execve while armed with per-thread CPU timers. The core technical flaw stems from the interaction between de_thread(), which reassigns PID ownership when a thread executes a new binary, and posix_cpu_timers_exit(), which is responsible for cleaning up timer resources. Under normal circumstances, the kernel ensures safety by calling posix_cpu_timers_exit immediately after exec_task_namespaces within begin_new_exec. However, this cleanup occurs only if execution proceeds past certain critical points without failure. If functions such as unshare_files, set_mm_exe_file, exec_mmap, or exec_task_namespaces fail during the process of setting up a new executable environment, the task may be terminated via do_exit before reaching the safe cleanup point.

When de_thread is called for a non-leader thread B that has an armed POSIX timer, the PID associated with that timer is transferred to the leader thread due to internal kernel mechanics regarding thread group IDs. Consequently, subsequent lookups of the original PID return NULL because it no longer maps to task B but rather to the leader. The function timer_lock_sighand interprets this failed lookup as an indication that the timer node has already been dequeued and thus performs no further action on the stale reference. Meanwhile, if a failure occurs later in the exec sequence, exit_itimers is invoked for thread B. This routine calls posix_timer_delete which eventually triggers posix_cpu_timer_del and posix_timer_unhash_and_free. These functions free the k_itimer structure while its corresponding node remains queued within the rbtree of task B's posix_cputimers list.

The severity of this vulnerability lies in the subsequent reaping phase. When release_task is finally called for thread B, it invokes posix_cpu_timers_exit(B). This function attempts to clean up timers by iterating through and deleting nodes from the timerqueue. Because the node was previously freed but not removed from the rbtree data structure, the kernel performs a use-after-free operation when accessing or manipulating this dangling pointer during cleanup_timerqueue and timerqueue_del operations. This memory corruption can lead to system instability, kernel panics, or potentially be exploited for arbitrary code execution depending on how the corrupted memory is utilized by subsequent allocations or checks within the kernel heap management subsystems.

From a classification perspective, this vulnerability aligns with CWE-416: Use After Free, as it involves accessing memory after it has been freed due to improper lifecycle management of kernel objects. In terms of adversarial tactics, while primarily an internal stability issue rather than a direct privilege escalation vector for external attackers without prior code execution capabilities, the underlying mechanics relate to ATT&CK technique T1059: Command and Scripting Interpreter if leveraged in conjunction with other vulnerabilities to achieve persistence or lateral movement through kernel exploitation. The flaw highlights risks associated with complex state transitions during process image loading where error handling paths may bypass critical resource cleanup routines designed for successful execution flows.

To mitigate this vulnerability, the primary remediation involves restructuring the order of operations within the execve path for non-leader threads. As indicated by the resolution details provided in the patch description, POSIX timer cleanup must be moved to occur immediately after de_thread() and before any subsequent functions that could trigger task termination upon failure. By ensuring posix_cpu_timers_exit is called right after thread ID exchange logic completes but prior to potential points of no return such as exec_mmap or namespace changes, the kernel guarantees that stale timer nodes are properly dequeued from their respective rbtrees before they can be freed while still referenced. This adjustment ensures that even if subsequent setup steps fail and trigger do_exit, the timers have already been safely removed from active data structures, preventing the use-after-free condition during final task release. System administrators should apply kernel updates containing this fix to maintain system integrity and prevent potential denial of service or exploitation scenarios arising from corrupted kernel memory states.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

very low

Sources

Interested in the pricing of exploits?

See the underground prices here!