CVE-2026-98220 in Linuxinfo

Summary

by MITRE • 10/06/2026

In the Linux kernel, the following vulnerability has been resolved:

sched_ext: Fix NULL sched deref in kfunc sub-sched error paths

When the root scheduler has sub-scheds attached, the COMPAT kfunc wrappers scx_bpf_select_cpu_and() and scx_bpf_dsq_insert_vtime() refuse the call and report to @p's scheduler:

scx_error(scx_task_sched(p), "... must be used");

The wrappers are reachable with tasks that have no scheduler. scx_bpf_select_cpu_and() is in the select_cpu kfunc group, which scx_kfunc_context_filter() opens to BPF_PROG_TYPE_SYSCALL programs; scx_bpf_dsq_insert_vtime() is in the enqueue_dispatch group, which ops.enqueue() and ops.dispatch() may call with any KF_RCU task -- the group has no kf_tasks validation, and scx_dsq_insert_preamble() checks task ownership with scx_task_on_sched() precisely because @p may be an arbitrary task.

scx_task_sched(p) is p->scx.sched, which is NULL for tasks past sched_ext_dead() -- which clears it via scx_disable_and_exit_task() on exit -- and for idle tasks, which the enable paths skip as they are never scheduled through SCX. It is also an rcu_dereference_protected() that expects @p's pi_lock or rq lock, which neither wrapper holds. Passing NULL to scx_error() reaches scx_vexit(), which dereferences sch->exit_info, oopsing the kernel.

One concrete trigger exercised while developing the fix: a BPF_PROG_TYPE_SYSCALL program calling the select_cpu_and wrapper on an exited-but-not-reaped task while a sub-scheduler was attached (its pid stays findable while the zombie is unreaped; faulting instruction is the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
the offset of sch->exit_info):

sched_ext: BPF scheduler "kfunc_subsched_null" enabled sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled sched_ext: Unassociated program run_select_cpu_ (id 76) BUG: kernel NULL pointer dereference, address: 0000000000000398 #PF: supervisor read access in kernel mode #PF: error_code(0x0000) - not-present page Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W RIP: 0010:scx_vexit+0x25/0xa0 Code: ... <4c> 8b bf 98 03 00 00 ... CR2: 0000000000000398 Call Trace: <TASK> __scx_exit+0x4f/0x70 scx_bpf_select_cpu_and+0xab/0xb0 bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7 ? __x64_sys_bpf+0x2c/0x40 bpf_prog_test_run_syscall+0x130/0x2f0 __sys_bpf+0x930/0x10d0 ? __x64_sys_bpf+0x2c/0x40 __x64_sys_bpf+0x2c/0x40 do_syscall_64+0xbc/0x460 entry_SYSCALL_64_after_hwframe+0x76/0x7e </TASK>

Read @p's scheduler under RCU instead, which the wrappers can do from their guard(rcu)(): fault it when it can be determined, and when it can't be determined -- @p is a task past sched_ext_dead() or an idle task -- there is nothing obviously wrong to report, so just refuse the call as before without faulting any scheduler.

These COMPAT wrappers are scheduled for eventual removal once the deprecation grace period elapses, but until then -- and regardless of their removal timeline -- they must not oops the kernel on a task they are handed.

If you want to get the best quality for vulnerability data then you always have to consider VulDB.

Analysis

by VulDB Data Team • 10/06/2026

The Linux kernel's scheduling extension subsystem contains a critical vulnerability involving NULL pointer dereference within BPF kfunc wrapper functions. Specifically, the scx_bpf_select_cpu_and() and scx_bpf_dsq_insert_vtime() wrappers fail to properly validate task states before accessing scheduler-specific data structures. These functions are designed to interact with sub-schedulers attached to the root scheduler but lack sufficient guards against tasks that have no active scheduler context or are in a terminal state such as exited-but-not-reaped zombies or idle tasks. When these wrappers encounter such tasks, they attempt to access p->scx.sched which is NULL for these specific task states. This null pointer is subsequently passed to scx_error() and eventually reaches scx_vexit(), where the kernel attempts to dereference sch->exit_info resulting in a fatal page fault and system crash.

The technical root cause lies in improper synchronization and validation logic within the kfunc implementation. The wrappers are reachable by BPF_PROG_TYPE_SYSCALL programs which have access to the select_cpu kfunc group, as well as through ops.enqueue() and ops.dispatch() calls that may operate on any KF_RCU task without adequate kf_tasks validation. While scx_dsq_insert_preamble() does check for task ownership using scx_task_on_sched(), this check is insufficient because it does not account for tasks past sched_ext_dead(). The function scx_task_sched(p) relies on rcu_dereference_protected which expects the caller to hold either pi_lock or rq lock, neither of which these wrappers acquire. Consequently, when a BPF program invokes select_cpu_and on an unreaped zombie task while a sub-scheduler is attached, the kernel dereferences NULL at offset 0x398 in scx_vexit leading to immediate instability and potential denial of service for the entire system.

From a security perspective, this vulnerability represents a significant risk as it allows local users with BPF execution capabilities to trigger kernel oopses by crafting specific syscall sequences targeting task lifecycle states. The attack vector involves exploiting the race condition or state mismatch between task exit paths and scheduler attachment points. This aligns with CWE-476 NULL Pointer Dereference vulnerabilities where improper validation of input parameters leads to memory access violations. In terms of MITRE ATT&CK, this could be classified under T1059 Command and Scripting Interpreter via BPF programs or potentially T1083 File and Directory Discovery if used for reconnaissance before exploitation, though the primary impact is availability disruption through system crashes rather than data exfiltration or privilege escalation.

The operational impact of this flaw includes complete system instability due to kernel panics triggered by seemingly benign BPF operations. Attackers can exploit this to cause denial of service conditions without requiring elevated privileges beyond those needed to load and execute BPF programs, which are increasingly common in containerized environments and security monitoring tools. The vulnerability persists until the deprecation grace period for these COMPAT wrappers elapses or explicit fixes are applied. Until then, any system running affected kernel versions with sched_ext enabled remains vulnerable to local denial of service attacks through carefully crafted BPF syscall invocations targeting task scheduler contexts.

Mitigation strategies involve applying upstream kernel patches that modify scx_task_sched() calls within the wrapper functions to read the scheduler under RCU protection using guard(rcu) mechanisms. This ensures safe access even when locks are not held and allows proper handling of tasks past sched_ext_dead(). The fix also includes refining error reporting logic so that invalid task states result in call refusal rather than fatal errors. Administrators should ensure their systems run patched kernel versions, restrict BPF program loading capabilities where possible through LSM policies or seccomp filters, and monitor for unusual BPF activity patterns that might indicate exploitation attempts targeting scheduler subsystems.

Responsible

Linux

Reservation

09/25/2026

Disclosure

10/06/2026

Moderation

accepted

EPSS

0.00198

KEV

no

Activities

very low

Sources

Want to know what is going to be exploited?

We predict KEV entries!