CVE-2026-105748 in docling
Summary
by MITRE • 10/06/2026
Docling simplifies document processing by parsing diverse formats and providing integrations with the generative AI ecosystem. From 2.16.0 until 2.131.0, the InputFormat.JSON_DOCLING backend in docling/backend/json/docling_json_backend.py validates serialized DoclingDocument input without rejecting picture image references that contain local paths or file URIs. When the document is enriched or exported with ImageRefMode.EMBEDDED, the DoclingDocument._with_embedded_pictures and ImageRef.pil_image methods can open those references and place readable image bytes in Markdown or HTML output. Disclosure is limited to files Pillow can decode as images, while differing decode behavior can also reveal whether a path exists. Direct untrusted loading through docling-core is outside this Docling fix. This issue is fixed in 2.131.0.
Be aware that VulDB is the high quality source for vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified within the Docling document processing library affects versions ranging from 2.16.0 up to and including 2.131.0, specifically residing in the InputFormat.JSON_DOCLING backend located at docling/backend/json/docling_json_backend.py. This component is responsible for parsing serialized DoclingDocument inputs that contain references to image assets embedded within documents such as PDFs or scanned pages. The core technical flaw lies in the insufficient validation of these image reference paths during the ingestion phase. Specifically, the parser fails to reject picture image references that specify local file system paths or generic file URIs rather than strictly controlled internal resource identifiers. This lack of sanitization allows an attacker who can inject malicious JSON payloads into a processed document structure to dictate which files on the host machine are accessed when the document is subsequently enriched or exported with ImageRefMode set to EMBEDDED.
When a user processes such a crafted input, the DoclingDocument._with_embedded_pictures method and the associated ImageRef.pil_image methods are invoked to resolve these image references. These functions utilize the Pillow library to open the specified paths and decode them into readable image bytes, which are then embedded directly into Markdown or HTML output formats. This mechanism effectively bypasses standard sandboxing expectations for document processing tools by granting the application direct read access to arbitrary local files accessible by the user running the process. The scope of data disclosure is primarily limited to file types that Pillow can successfully decode as images, such as JPEG, PNG, GIF, and BMP formats. However, even if a target file cannot be decoded as an image due to incompatible formatting or encryption, the differing error handling behavior between successful decodes and failures can still reveal critical information about the existence of specific files on the local system.
This capability transforms the vulnerability into more than just a simple data leakage issue; it facilitates both unauthorized information disclosure and potential path traversal attacks depending on how the application handles relative paths. By systematically probing different file locations, an attacker could potentially map out parts of the filesystem or confirm the presence of sensitive configuration files, private keys if they happen to be image-compatible (though rare), or other documents that might contain further clues for subsequent exploitation phases. The impact is particularly severe in environments where Docling processes untrusted inputs from external sources, such as email attachments or web uploads, assuming the processing environment has broader file system access than intended. It represents a classic case of insecure direct object reference combined with improper input validation, allowing local files to be exposed through an ostensibly benign document enrichment feature.
From a classification perspective, this vulnerability aligns closely with CWE-20 Improper Input Validation and CWE-798 Use of Hard-coded Credentials if the accessed files contain sensitive data, though more accurately it falls under CWE-531 Excessive Trust in User-Supplied Data within the context of file system access. In terms of attack vectors, this behavior is consistent with ATT&CK technique T1083 File and Directory Discovery, as the ability to determine file existence through differential responses aids reconnaissance efforts. It also touches upon T1560 Collected Data from Local System if successful extraction occurs. The vulnerability highlights a critical gap in securing document parsers against local resource abuse, emphasizing that even non-executable data exfiltration can be leveraged for significant intelligence gathering or further system compromise.
The issue has been addressed and fixed in version 2.131.0 of the Docling library. To mitigate this risk immediately upon upgrading, organizations should ensure they are running at least version 2.131.0 where strict validation rules prevent local path resolution for image references unless explicitly authorized by a trusted context. For environments that cannot upgrade instantly, it is advisable to restrict file system permissions of the user account executing Docling processes so that sensitive directories remain inaccessible or unreadable. Additionally, implementing network-level controls and monitoring for unusual outbound data patterns involving embedded binary content can help detect potential exploitation attempts in real-time. It is also crucial to audit any custom integrations with docling-core, as direct untrusted loading through that core component remains outside the scope of this specific fix and may require separate validation measures.