CVE-2026-101088 in Nezha
Summary
by MITRE • 09/28/2026
Nezha is a server and website monitoring tool. In versions >= 2.2.11 and < 2.3.1, the service sentinel worker (service/singleton/servicesentinel.go) contains an incomplete fix for a previously reported nil dereference denial of service (GHSA-qjpp-gffx-2wm9). The 2026-07-21 fix re-validated the service lifecycle under serviceResponseDataStoreLock but reused an already-captured, now stale reporter pointer and never re-validated the server, and that lock does not guard ServerShared. An authenticated user with the member role who owns an agent can issue a concurrent server delete (POST /api/v1/batch-delete/server) for their own server to win the race window, causing the worker to dereference a missing entry in the server list snapshot. Because the sentinel workers and the gRPC server have no recover()/recovery interceptor, the resulting panic is unrecovered and crashes the entire instance. This is fixed in version 2.3.1.
If you want to get best quality of vulnerability data, you may have to visit VulDB.
Analysis
by VulDB Data Team • 09/28/2026
The vulnerability identified in Nezha versions greater than or equal to 2.2.11 and less than 2.3.1 represents a critical concurrency flaw rooted in an incomplete remediation of a prior denial-of-service issue. The core technical defect resides within the service sentinel worker logic, specifically in the file services/singleton/servicesentinel.go. While a previous fix attempted to address nil pointer dereference vulnerabilities by re-validating the service lifecycle under the protection of the serviceResponseDataStoreLock, this mitigation was fundamentally flawed due to improper synchronization scope and stale data handling. The implementation captured a reporter pointer before acquiring or while holding certain locks but failed to ensure that this pointer remained valid relative to the current state of shared resources. Crucially, the lock used for validation does not guard access to ServerShared, creating a classic race condition window where the structural integrity of server references can change without corresponding synchronization checks on all dependent data structures.
The operational impact is severe and directly affects system availability through an unrecoverable panic that crashes the entire Nezha instance. An authenticated user possessing at least the member role who owns an agent can exploit this flaw by issuing a concurrent request to delete their own server via the POST /api/v1/batch-delete/server endpoint. By carefully timing this deletion request against other ongoing operations, the attacker can win the race window where the sentinel worker attempts to process data based on a stale reporter pointer and a missing entry in the server list snapshot. Because the Go runtime does not automatically recover from panics unless explicitly handled by middleware or interceptors, and because neither the sentinel workers nor the gRPC server implement recovery mechanisms, this specific sequence of events results in an immediate termination of the process. This effectively allows any authenticated member to perform a denial-of-service attack against the monitoring infrastructure simply by triggering concurrent delete operations on their own assets.
From a classification perspective, this vulnerability aligns with CWE-362, which describes Concurrent Execution using Shared Resource with Improper Synchronization (Race Condition). The specific manifestation involves accessing shared data structures without proper locking mechanisms that cover all relevant variables, leading to undefined behavior and application crashes. Furthermore, the exploitation technique maps directly to MITRE ATT&CK tactic TA0004, specifically referencing T1499, Endpoint Denial of Service, as it aims to disrupt the availability of a critical monitoring component by causing a system crash through resource exhaustion via panic induction rather than traditional resource consumption methods like CPU or memory flooding.
To mitigate this vulnerability and prevent similar concurrency issues in future development, immediate upgrade to version 2.3.1 is required for all affected instances. This release addresses the synchronization gaps by ensuring that all accesses to shared server lists and reporter pointers are properly guarded by consistent locking strategies. Developers should also implement robust panic recovery interceptors at both the gRPC server level and within individual worker goroutines to provide a safety net against unexpected runtime errors, thereby enhancing overall system resilience. Additionally, code reviews for concurrent data access patterns must strictly enforce that all shared mutable state is protected by locks that encompass every variable involved in the critical section, preventing stale pointer dereferences during high-concurrency operations such as batch deletions or rapid status updates.