CVE-2026-102997 in Pypdf
Summary
by MITRE • 10/01/2026
pypdf is a free and open-source pure-python PDF library. Prior to 6.18.1, a crafted PDF containing a partially malformed /FlateDecode stream with padded data can force pypdf/filters.py to use inefficient byte-by-byte decompression while the earlier recovery counter fails to advance for bytes that successfully decode, causing long runtimes and application unavailability. This is a residual issue after the malformed FlateDecode recovery fix. This issue is fixed in version 6.18.1.
If you want to get best quality of vulnerability data, you may have to visit VulDB.
Analysis
by VulDB Data Team • 10/01/2026
The vulnerability identified within pypdf versions prior to 6.18.1 represents a significant denial of service risk stemming from inefficient resource consumption during PDF stream decompression. As an open-source pure-python library, pypdf is widely used for parsing and manipulating PDF documents in various automated workflows and document processing systems. The core issue resides within the filters.py module, specifically concerning the handling of /FlateDecode streams which are commonly used to compress data within PDF files. When a crafted PDF contains a partially malformed stream with padded data that does not strictly adhere to standard compression formats, the library's error recovery mechanism fails to operate correctly. This failure triggers an inefficient byte-by-byte decompression process rather than utilizing optimized bulk decoding methods or properly skipping invalid segments.
The technical flaw is characterized by a logic error in the recovery counter used during the parsing of compressed streams. In normal operation, when pypdf encounters malformed data within a FlateDecode stream, it attempts to recover and continue processing subsequent valid sections of the document. However, due to this specific bug, the internal counter that tracks successfully decoded bytes fails to advance even when those bytes are correctly processed. This discrepancy causes the decompression algorithm to re-process or stall on these segments repeatedly, leading to exponential increases in computational overhead. The result is a severe degradation in performance where simple PDF parsing tasks can consume excessive CPU cycles and memory over extended periods.
From an operational perspective, this vulnerability allows for remote denial of service attacks against any application that utilizes pypdf to process untrusted or user-uploaded PDF documents. An attacker can craft a maliciously constructed PDF file designed specifically to trigger the malformed stream condition described above. When such a document is processed by a vulnerable version of pypdf, it forces the underlying Python interpreter into an infinite or near-infinite loop of inefficient decompression attempts. This effectively hangs the application thread responsible for parsing the document, leading to unavailability of services that depend on PDF processing capabilities. In web applications or server-side automation pipelines, this can result in resource exhaustion, blocking other legitimate requests and causing a systemic failure of service availability.
This issue is classified under CWE-400, which covers Uncontrolled Resource Consumption, as the vulnerability leads to excessive consumption of computational resources without proper bounds checking on processing time or memory usage during decompression operations. Furthermore, it aligns with ATT&CK technique T1496, Resource Hijacking, where an adversary uses compromised systems for unintended purposes such as consuming CPU cycles to disrupt service availability rather than stealing data. The vulnerability is a residual issue following earlier attempts to fix malformed FlateDecode recovery logic, indicating that previous patches addressed the crash or parsing error but did not fully resolve the performance implications of edge-case inputs involving padded and partially malformed streams.
To mitigate this risk, organizations relying on pypdf must upgrade immediately to version 6.18.1 or later, where the decompression efficiency and recovery counter logic have been corrected. For environments that cannot update instantly due to dependency constraints, implementing input validation at the network perimeter is recommended. This includes restricting PDF uploads from untrusted sources and employing sandboxing techniques such as time limits on process execution or memory caps for document parsing tasks. Additionally, integrating a secondary verification step using a different library can help detect malformed streams before they reach the vulnerable pypdf instance, thereby reducing exposure to this specific denial of service vector until patches are fully deployed across all affected systems.