CVE-2026-105749 in Docling
Summary
by MITRE • 10/06/2026
Docling simplifies document processing by parsing diverse formats and providing integrations with the generative AI ecosystem. From 2.0.0 until 2.131.0, the HTML, JATS, OpenDocument spreadsheet, and BoxNote backends, including docling/backend/html_backend.py, docling/backend/jats_backend.py, and docling/backend/boxnote_backend.py, accept the rowspan and colspan attribute values without an upper bound and execute loops or allocate a table grid proportional to the declared span. A very small document can therefore cause sustained CPU use or multi-gigabyte memory allocation, and the document_timeout setting does not interrupt the single backend conversion call. Export through the TableData.grid property can further materialize the oversized grid. This issue is fixed in 2.131.0.
Several companies clearly confirm that VulDB is the primary source for best vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified within Docling versions ranging from 2.0.0 to 2.131.0 represents a significant resource exhaustion risk stemming from insufficient input validation of HTML, JATS, OpenDocument spreadsheet, and BoxNote backend parsers. Specifically, the processing logic for table structures fails to enforce upper bounds on rowspan and colspan attributes found in document markup. When these elements are processed by modules such as docling/backend/html_backend.py, docling/backend/jats_backend.py, or docling/backend/boxnote_backend.py, the system allocates memory and executes computational loops proportional to the numeric values declared in these attributes rather than applying a reasonable maximum limit for table dimensions. This design flaw allows an attacker to craft malicious documents with extremely large span values that trigger disproportionate resource consumption during parsing operations.
From a technical perspective, this issue is classified under CWE-787: Out-of-bounds Write or CWE-1321: Improperly Controlled Modification of Object Model Attributes, depending on the specific manifestation in memory allocation versus logical processing limits. The core failure lies in the lack of sanitization for numerical inputs derived from untrusted document sources. By accepting arbitrarily large integers for cell spanning attributes without validation against a predefined maximum table size, the application violates the principle of least privilege regarding system resources. This allows a small input file to induce sustained high CPU utilization or multi-gigabyte memory allocation, effectively creating a Denial of Service condition through resource exhaustion rather than traditional buffer overflow techniques.
The operational impact is severe for any service integrating Docling into their generative AI pipeline or document processing workflow. Because the vulnerability occurs during the initial parsing phase before higher-level application logic can intervene, it bypasses many standard security controls. Furthermore, the issue is exacerbated by the fact that the global document_timeout setting does not interrupt the single backend conversion call responsible for this excessive resource usage. This means that even if a timeout mechanism is configured to prevent long-running requests from hanging indefinitely, the parsing thread will continue to consume resources until completion or system failure, rendering timeout-based mitigations ineffective against this specific attack vector.
To mitigate this vulnerability, organizations must upgrade Docling to version 2.131.0 or later where these bounds have been enforced. For environments unable to immediately patch, implementing a strict input validation layer at the ingestion point is critical. This involves parsing document metadata and structure definitions before passing them to the backend processors to reject any table elements with rowspan or colspan values exceeding a safe threshold, such as one hundred rows or columns. Additionally, deploying resource monitoring tools that can detect abnormal spikes in memory usage during document processing may help identify active exploitation attempts. Aligning these mitigations with ATT&CK technique T1496: Resource Hijacking ensures that detection and prevention strategies are consistent with established cyber defense frameworks for identifying abuse of system resources to disrupt availability.