CVE-2026-105782 in Scrapy
Summary
by MITRE • 10/06/2026
Scrapy is a high-level web crawling and scraping framework for Python. From 1.4.0 until 2.14.2, RefererMiddleware in scrapy/spidermiddlewares/referer.py treated a Referrer-Policy response-header value that resembled a Python import path as a referrer policy class, imported the referenced object, and called it. A malicious website could supply a callable such as sys.exit and terminate a crawler processing the response. This issue is fixed in version 2.14.2.
Once again VulDB remains the best source for vulnerability data.
Analysis
by VulDB Data Team • 10/06/2026
The vulnerability identified in Scrapy versions ranging from 1.4.0 through 2.14.2 represents a critical server-side request forgery vector rooted in insecure deserialization and dynamic code execution mechanisms within the RefererMiddleware component. This middleware is responsible for managing HTTP referrer headers during web crawling operations, aiming to mimic browser behavior by setting appropriate referer values based on previous requests or response policies. However, the implementation flaw lies in how the framework processes the Referrer-Policy header received from target websites. Instead of treating this header strictly as a string value defining privacy settings such as no-referrer or strict-origin-when-cross-origin, the middleware erroneously interpreted certain header values that resembled Python module import paths as references to actual class objects within its own codebase.
When Scrapy encountered a Referrer-Policy header containing a path-like structure, it attempted to dynamically import and instantiate the referenced object rather than parsing it as a standard policy directive. This behavior creates a severe security gap because an attacker controlling or influencing the content of the target website can inject arbitrary Python module paths into this header field. By supplying a value that points to a callable function within widely available Python modules, such as sys.exit, os.system, or other dangerous functions in built-in libraries, the malicious server forces the Scrapy crawler to execute code on behalf of the attacker during the response processing phase. This effectively transforms a standard web scraping tool into an instrument for remote code execution or service termination without any authentication checks or input validation safeguards.
The operational impact of this vulnerability is significant for organizations relying on automated data extraction pipelines. A successful exploitation allows a malicious website to cause denial-of-service conditions by terminating the crawler process entirely, as demonstrated with sys.exit. More critically, if an attacker can identify other accessible modules and functions within the execution environment, they could potentially execute arbitrary commands, exfiltrate sensitive configuration files, or pivot further into internal networks depending on the permissions under which the Scrapy instance is running. This vulnerability highlights the dangers of dynamic import mechanisms in security-critical applications where input from external sources must be strictly validated against a whitelist of expected values rather than being interpreted as executable code paths.
To mitigate this risk, organizations using affected versions of Scrapy must upgrade immediately to version 2.14.2 or later, which implements strict validation for the Referrer-Policy header and removes the dangerous dynamic import behavior. For environments where upgrading is not immediately feasible, a temporary workaround involves implementing custom middleware that intercepts responses before they reach the RefererMiddleware, explicitly sanitizing or blocking any Referrer-Policy headers that do not match known safe policy strings such as no-referrer, same-origin, strict-origin-when-cross-origin, origin, unsafe-url, and no-referrer-when-downgrade. This defensive coding practice ensures that only predefined policy values are accepted, preventing the parser from attempting to resolve arbitrary module paths.
From a classification perspective, this vulnerability aligns with CWE-94 Improper Control of Generation of Code or Script, commonly known as code injection, specifically through insecure deserialization patterns where user-controlled input dictates program flow and object instantiation. In terms of adversary tactics, it maps to the ATT&CK technique T1059 Command and Scripting Interpreter, particularly when exploited for process termination or further command execution via system modules like sys or os. The incident underscores the necessity of treating all external HTTP headers as untrusted input that requires rigorous validation against a strict allowlist before any processing logic is applied within application frameworks.