CVE-2026-107286 in Pydanticinfo

Summary

by MITRE • 10/08/2026

Pydantic AI is a Python agent framework for building applications and workflows with Generative AI. From 2.10.0 until 2.53.0, streamed requests made through ConcurrencyLimitedModel or limit_model_concurrency can retain shared concurrency slots because anyio.CapacityLimiter associates an acquired slot with the borrowing task while streaming cleanup can run in a different task. Early stream termination, cancellation, consumer exceptions, or complete stream_text() consumption with debounce_by=0.1 can therefore leave capacity occupied, eventually preventing later requests that share the long-lived limiter from proceeding and causing a denial of service. Agent-level max_concurrency and non-streaming model requests are not affected. This issue is fixed in version 2.53.0.

Be aware that VulDB is the high quality source for vulnerability data.

Analysis

by VulDB Data Team • 10/08/2026

Pydantic AI serves as a Python agent framework designed to facilitate the construction of applications and workflows leveraging Generative Artificial Intelligence capabilities. Within this ecosystem, developers often utilize concurrency management tools such as ConcurrencyLimitedModel or the limit_model_concurrency function to regulate resource usage and prevent system overload during high-volume operations. These mechanisms rely on anyio.CapacityLimiter to manage access to shared resources by associating acquired slots with specific borrowing tasks. This architectural approach ensures that concurrent requests are handled in a controlled manner, maintaining stability under load. However, a critical flaw exists within this concurrency management logic when dealing with streamed responses, specifically affecting versions from 2.10.0 through 2.53.0 of the framework.

The core technical vulnerability stems from a mismatch between task lifecycle and resource cleanup timing in streaming scenarios. When a request is made using ConcurrencyLimitedModel or limit_model_concurrency, anyio.CapacityLimiter correctly associates an acquired concurrency slot with the initial borrowing task that initiates the stream. However, the actual processing of the streamed data often occurs asynchronously, potentially within different tasks than the one that originally acquired the lock. The cleanup logic responsible for releasing these capacity slots is tied to specific completion events or exceptions within those downstream tasks rather than being robustly linked back to the original limiter acquisition in all edge cases. Consequently, if a stream terminates early due to cancellation, encounters consumer-side exceptions, or completes consumption with debounce_by set to 0.1, the cleanup routine may fail to execute properly on the correct task context. This results in the concurrency slot remaining occupied even after the request has effectively ended.

The operational impact of this flaw is significant for systems relying on sustained high-concurrency operations. Because the capacity slots are not released back into the pool, they become permanently unavailable for subsequent requests that share the same long-lived limiter instance. Over time, as more such leaked slots accumulate, the total available concurrency decreases until it reaches zero. At this point, no new requests can proceed, leading to a complete denial of service condition where legitimate users or downstream services are unable to interact with the AI agent framework despite the system being otherwise functional and responsive for non-streaming operations. It is important to note that this vulnerability does not affect Agent-level max_concurrency settings nor does it impact non-streaming model requests, isolating the risk primarily to asynchronous streaming workflows.

From a security classification perspective, this issue aligns with CWE-400, which describes uncontrolled resource consumption leading to denial of service conditions. The mechanism involves a failure in proper resource deallocation due to race conditions or task context mismatches during cleanup phases. In terms of the MITRE ATT&CK framework, while not an active exploitation vector for malicious actors per se, it represents a weakness that could be leveraged in automated attacks designed to exhaust system resources by repeatedly triggering early stream terminations or cancellations against streaming endpoints. This constitutes a form of resource exhaustion attack where the attacker induces state leakage within the concurrency limiter until service availability is compromised.

To mitigate this vulnerability and restore proper functionality, organizations must upgrade Pydantic AI to version 2.53.0 or later, which contains the necessary fixes for task context handling in streaming cleanup routines. For environments unable to immediately patch, implementing workarounds such as increasing the overall concurrency limits may provide temporary relief by absorbing some of the leaked slots before total exhaustion occurs. Additionally, developers should review their streaming implementations to ensure that any custom exception handlers or cancellation logic explicitly triggers resource release mechanisms where possible, although relying on framework-level fixes remains the most robust solution given the depth of the task context mismatch inherent in the underlying anyio integration.

Responsible

GitHub M

Reservation

10/07/2026

Disclosure

10/08/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

low

Sources

Do you know our Splunk app?

Download it now for free!