Submit #997933: https://github.com/Tiiny-AI/PowerInfer/issues/284 PowerInfer commit 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 (main, 2026-05-11) Out-of-bounds Readinfo

Titlehttps://github.com/Tiiny-AI/PowerInfer/issues/284 PowerInfer commit 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 (main, 2026-05-11) Out-of-bounds Read
DescriptionA vulnerability was found in PowerInfer commit 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 (main, 2026-05-11) and classified as high 7.1. Affected is the FFN GPU-split loader component, specifically create_striped_mat_to_gpu() in the llama.cpp file. The manipulation of the argument int32 row indices stored in the gpu_bucket tensors of an auto-loaded <model>.generated.gpuidx sidecar file leads to an out-of-bounds read. The attack is local: the victim must run the inference binary on a crafted model package (model file plus sidecar, distributed as a directory). User interaction is required because the victim starts the load. The sidecar is loaded merely because the file exists next to the model; only split.vram_capacity and the tensor count are validated, while each int32 index is used unchecked as host_mat_row->data = (char*)src->data + host_i * row_data_size and then passed as the cudaMemcpy source address via ggml_cuda_cpy_1d. An index such as 0x7FFFFFFF (or any negative value) places the read up to INT32_MAX * row_data_size bytes outside the tensor buffer, crashing the process (ASan heap-buffer-overflow READ / SEGV); copied out-of-bounds host memory becomes part of the device FFN weights used by subsequent inference. Technical Details - Affected file/function: llama.cpp / create_striped_mat_to_gpu() (llama.cpp:2898, vulnerable loop at llama.cpp:2926-2934); auto-load at llm_load_gpu_split_with_budget() (llama.cpp:3088-3097); only checks at llama_gpu_split_loader ctor / check_vram_allocable() (llama.cpp:2816) and load_gpu_idx_for_model() (llama.cpp:2820) - Vulnerable parameter: int32 entries of gpu_bucket tensors in <model>.generated.gpuidx - Attack vector: Local - Privileges required: None - Trigger condition: run ./main -m model.gguf (--vram-budget 8 or the default -1 budget) with a GPU present while a crafted .generated.gpuidx sits next to the model file; --disable-gpu-index is the only opt-out Impact - Confidentiality: High (out-of-bounds host memory is copied into device FFN weights and can surface through inference outputs; conservative estimate) - Integrity: None - Availability: High (process crash during model load) CVSS v3.1 Score: 7.1 (High) Vector: CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:N/A:H Timeline - Discovered: 2026-09-24 - Vendor notified: [pending] - Patch released: [pending] - Public disclosure: [pending] Countermeasure Validate every gpu_bucket entry before pointer arithmetic (reject host_i < 0 or host_i >= src->ne[1]); treat the auto-loaded .generated.gpuidx sidecar as untrusted input and bind it to the model file (e.g. hash check) instead of trusting it on existence alone.
Source⚠️ https://github.com/Tiiny-AI/PowerInfer/issues/284
User
 Dem0 (UID 82596)
Submission09/24/2026 17:57 (17 days ago)
Moderation10/11/2026 21:28 (17 days later)
StatusAccepted
VulDB entry416779 [Tiiny-AI PowerInfer up to 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 FFN GPU-split Loader ggml-cuda.cu create_striped_mat_to_gpu out-of-bounds]
Points20

Might our Artificial Intelligence support you?

Check our Alexa App!