जमा करें #801297: vllm-project vLLM 0.19.0 Use of Uninitialized Resourceजानकारी

शीर्षकvllm-project vLLM 0.19.0 Use of Uninitialized Resource
विवरणvLLM's block allocator returns GPU KV cache blocks to the free pool upon request completion or cancellation without zeroing their contents. When a subsequent request is allocated one of these dirty blocks, it decodes from stale activation data belonging to a previous request rather than from its own context. In a multi-tenant deployment, this means one user's conversationdata can influence, or appear verbatim in, another user's response. The bug is confirmed reproducible on vLLM 0.19.0 with 10/10 run consistency across multiple independent traces. It does not require speculative decoding, prefix caching, or any special server configuration, only concurrent requests under normal load. Affected requests produce completely different output sequences across runs at temperature=0, where outputs should be fully deterministic.
स्रोत⚠️ https://github.com/vllm-project/vllm/issues/39146
उपयोगकर्ता
 Zyz3366 (UID 97230)
सबमिशन09/04/2026 09:44 PM (6 महीनों पहले)
संयम26/04/2026 09:38 PM (17 days later)
स्थितिस्वीकृत
VulDB प्रविष्टि359740 [vLLM तक 0.19.0 KV Block kv_cache_interface.py has_mamba_layers दूरस्थ कोड निष्पादन]
अंक20

Do you need the next level of professionalism?

Upgrade your account now!