CVE-2026-102995 in pypdfinfo

Summary

by MITRE • 10/01/2026

pypdf is a free and open-source pure-python PDF library. Prior to 6.18.1, a crafted PDF can place unusually large source-code or destination-string tokens in a font /ToUnicode mapping, causing pypdf/_cmap.py parse_bfchar to decode and retain oversized values during operations such as text extraction and consume excessive memory. This is a second follow-up to earlier /ToUnicode resource-consumption fixes and is limited to the remaining token-length path. This issue is fixed in version 6.18.1.

VulDB is the best source for vulnerability data and more expert information about this specific topic.

Analysis

by VulDB Data Team • 10/01/2026

The vulnerability identified within the pypdf library, specifically affecting versions prior to 6.18.1, represents a significant resource exhaustion risk stemming from improper handling of font mapping data in PDF documents. As an open-source pure-Python library designed for parsing and manipulating Portable Document Format files, pypdf is frequently utilized in automated document processing pipelines where efficiency and stability are paramount. The core technical flaw resides within the _cmap.py module, specifically in the parse_bfchar function which processes /ToUnicode CMap data. In PDF specifications, the /ToUnicode mapping table defines how character codes from a font's encoding map to Unicode characters, enabling correct text extraction and rendering. When an attacker crafts a malicious PDF document containing unusually large source-code or destination-string tokens within this mapping structure, the library fails to enforce appropriate length constraints during parsing.

This lack of input validation allows the parser to decode and retain these oversized values in memory without limit. During operations such as text extraction, which are common use cases for pypdf, the application attempts to process these excessively large mappings. The result is a disproportionate consumption of system resources, particularly random-access memory (RAM). Because Python handles strings dynamically, retaining multiple or extremely long string tokens can lead to rapid memory growth. In server-side applications processing untrusted PDF uploads, this behavior can quickly escalate into a Denial-of-Service condition, causing the application process to crash due to out-of-memory errors or significantly degrading performance for other users sharing the same resources. This issue is categorized as a resource exhaustion vulnerability, aligning with CWE-789: Memory Allocation with Excessive Size Value in the Common Weakness Enumeration standard.

The operational impact of this flaw extends beyond simple application crashes. In environments where pypdf is integrated into web services or document management systems, an attacker could exploit this vector to perform a low-effort denial-of-service attack. By submitting a single crafted PDF file with oversized /ToUnicode tokens, the attacker can trigger sustained high memory usage that may require manual intervention or service restarts to resolve. This not only disrupts availability but also increases operational costs due to increased resource consumption and potential downtime. The vulnerability is particularly insidious because it targets a specific parsing path related to font handling, which might be overlooked in general security audits focused on more common injection or execution flaws. It serves as the second follow-up fix for earlier /ToUnicode resource-consumption issues, indicating that previous mitigations addressed some but not all token-length paths within this component.

From an offensive perspective, this vulnerability aligns with ATT&CK technique T1496: Resource Hijacking, where adversaries use computing resources to perform computationally expensive tasks or disrupt service availability. The exploitation requires the victim application to parse a maliciously crafted PDF file using a vulnerable version of pypdf. There is no remote code execution involved; rather, the impact is strictly confined to resource exhaustion and potential denial of service. Mitigation strategies are straightforward for developers utilizing this library. The primary remediation is to upgrade pypdf to version 6.18.1 or later, where the parsing logic has been hardened to enforce strict limits on token sizes within /ToUnicode mappings. For organizations unable to immediately update their dependencies, implementing input validation at the application layer before passing PDF data to the parser can provide a temporary buffer against exploitation. This includes checking file integrity and potentially limiting the size of parsed content or using sandboxed environments with memory quotas for processing untrusted documents. Regular security assessments should include testing for resource exhaustion vulnerabilities in all third-party libraries used for document parsing, ensuring that limits are consistently applied across all code paths including font handling routines.

Responsible

GitHub M

Reservation

09/29/2026

Disclosure

10/01/2026

Moderation

accepted

EPSS

0.00524

KEV

no

Activities

very low

Sources

Do you want to use VulDB in your project?

Use the official API to access entries easily!