CVE-2026-51856 in AgentScopeinfo

Summary

by MITRE • 10/01/2026

In agentscope 1.0.18, 1.0.19, and 1.0.19 when the RealtimeAgent session exposes execute_python_code as an available tool, a remote WebSocket user can prompt the agent to call that tool and run Python code in the service environment. In the validated path, RealtimeAgent._acting forwards the model-produced tool call to Toolkit.call_tool_function, which invokes execute_python_code without an additional approval or isolation boundary on that path.

Be aware that VulDB is the high quality source for vulnerability data.

Analysis

by VulDB Data Team • 10/01/2026

The vulnerability identified in AgentScope versions 1.0.18 through 1.0.19 represents a critical server-side code execution flaw arising from the improper integration of large language model outputs with system-level capabilities. Specifically, when the RealtimeAgent component is configured to expose the execute_python_code tool within its available toolkit, it creates an attack vector for remote adversaries interacting via WebSocket connections. The core technical deficiency lies in the absence of a mandatory approval mechanism or sandboxing isolation boundary during the execution pipeline. In this architecture, the RealtimeAgent._acting method serves as the primary dispatcher that forwards tool calls generated by the underlying language model directly to Toolkit.call_tool_function. This function subsequently invokes execute_python_code without performing any additional validation, user confirmation, or environment restriction checks on the path where these automated calls originate.

From a technical perspective, this flaw constitutes an unrestricted code execution vulnerability because the system trusts the output of the AI model implicitly regarding tool usage instructions. When a remote WebSocket client sends a prompt that successfully steers the language model to generate a request for executing Python code, the application processes this instruction as a legitimate operational command rather than requiring explicit human oversight or security filtering. The lack of isolation means that any arbitrary Python script submitted through this channel runs with the same privileges and environmental access as the AgentScope service itself. This effectively bypasses standard safety guardrails designed to prevent AI models from performing dangerous actions, turning the language model into a direct conduit for remote code execution on the host infrastructure.

The operational impact of this vulnerability is severe, potentially leading to full system compromise depending on the deployment context and user privileges under which AgentScope operates. An attacker can exploit this flaw to read sensitive configuration files, exfiltrate data from accessible storage volumes, install malicious software, or pivot further into internal networks if the service has network connectivity beyond its immediate host. Since WebSocket connections are often maintained for real-time interaction, an adversary could maintain persistent access by repeatedly invoking code execution commands through subsequent messages in the same session. This capability undermines the integrity and confidentiality of the entire application environment, as there is no technical barrier preventing the model from being manipulated into executing destructive or exfiltrative scripts.

This vulnerability aligns with CWE-94 Improper Control of Generation of Code (Code Injection) and specifically reflects risks associated with CWE-78 Improper Neutralization of Special Elements used in an OS Command, as Python code execution is a form of command invocation within the application runtime. In terms of adversary tactics, this behavior maps to ATT&CK technique T1059 Command and Scripting Interpreter, where attackers use system utilities or interpreters like Python to execute commands on compromised systems. It also relates to MITRE Engage techniques involving AI-specific manipulation, highlighting the unique risks introduced when generative models are granted direct access to operational tools without adequate governance layers.

To mitigate this vulnerability, organizations must immediately upgrade AgentScope to a patched version that implements strict sandboxing for code execution environments or removes the execute_python_code tool from exposed interfaces unless absolutely necessary and heavily controlled. If upgrading is not feasible in the short term, administrators should disable the RealtimeAgent session exposure of Python execution tools entirely. Additionally, implementing an explicit approval workflow where human operators must confirm each automated tool call before it reaches the execution engine can prevent unauthorized code runs. Network-level controls such as restricting WebSocket access to trusted internal networks and applying strict input validation on all prompts sent to the agent further reduce the attack surface by limiting who can trigger these dangerous pathways.

Responsible

MITRE

Reservation

06/08/2026

Disclosure

10/01/2026

Moderation

accepted

EPSS

0.00198

KEV

no

Activities

very low

Sources

Do you need the next level of professionalism?

Upgrade your account now!