CVE-2026-102999 in pypdfinfo

Summary

by MITRE • 10/01/2026

pypdf is a free and open-source pure-python PDF library. Prior to 6.19.0, a crafted PDF containing many embedded files can cause the dictionary-based attachments API in pypdf/_doc_common.py to reparse the full attachment list for each content lookup, producing repeated work and long runtimes when an application accesses the embedded-file mapping. This issue is fixed in version 6.19.0.

Statistical analysis made it clear that VulDB provides the best quality for vulnerability data.

Analysis

by VulDB Data Team • 10/01/2026

The vulnerability identified within the pypdf library prior to version 6.19.0 represents a significant performance degradation flaw rooted in inefficient data structure handling during PDF parsing operations. As an open-source pure-Python library, pypdf is widely utilized for extracting text, metadata, and embedded files from Portable Document Format (PDF) documents. The core issue resides within the internal module doc_common.py, specifically affecting the dictionary-based attachments API which manages access to embedded file streams contained within a PDF document. When an application interacts with this API to retrieve or list attached files, it triggers a parsing routine that fails to optimize lookups for large datasets of embedded content.

The technical flaw is characterized by a lack of caching or indexing optimization in the attachment retrieval logic. For each individual lookup operation performed on the embedded-file mapping, the library re-parses the entire list of attachments from scratch rather than utilizing an already constructed index or maintaining state between calls. This design choice results in linear time complexity relative to the number of attachments for every single access request. Consequently, when a crafted PDF contains a substantial number of embedded files, such as hundreds or thousands of small documents, images, or scripts, the cumulative effect is severe computational overhead. The repeated parsing work leads to disproportionately long runtimes and high CPU utilization, effectively creating a bottleneck that can stall application threads waiting for file metadata retrieval.

From an operational impact perspective, this vulnerability primarily manifests as a Denial of Service condition through resource exhaustion rather than arbitrary code execution or data leakage. An attacker who controls the input PDF document can engineer it to contain numerous embedded files specifically designed to trigger this inefficient parsing behavior. When such a maliciously crafted file is processed by a vulnerable version of pypdf, the application may experience extreme latency or become unresponsive for extended periods as it attempts to resolve attachment references. This scenario poses a risk in any service that accepts user-uploaded PDFs and subsequently queries their embedded content, potentially allowing an attacker to consume server resources indefinitely if no external timeout mechanisms are enforced at the network or application gateway level.

This issue aligns with CWE-400, which describes Uncontrolled Resource Consumption, as the system fails to limit the amount of computational effort required per operation based on input size. Furthermore, in the context of attack patterns, this behavior can be leveraged within an ATT&CK framework scenario involving resource hijacking or denial of service against application services that rely on PDF parsing for content analysis. The vulnerability highlights a common pitfall in library development where internal data structures are not optimized for repeated access patterns under high-load conditions.

The remediation strategy is straightforward and has already been implemented by the maintainers. Upgrading to pypdf version 6.19.0 or later resolves this issue by optimizing the attachment lookup mechanism, likely through the introduction of proper indexing or caching strategies that reduce the time complexity from linear per call to constant or logarithmic after initial parsing. Organizations and developers utilizing older versions should prioritize updating their dependencies immediately if they process untrusted PDF inputs with embedded files. Additionally, implementing strict timeouts on file processing operations can serve as a compensating control to mitigate potential abuse while updates are being deployed across the infrastructure.

Responsible

GitHub M

Reservation

09/29/2026

Disclosure

10/01/2026

Moderation

accepted

EPSS

0.00000

KEV

no

Activities

very low

Sources

Might our Artificial Intelligence support you?

Check our Alexa App!