| Title | https://github.com/Tiiny-AI/PowerInfer/issues/284 PowerInfer commit 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 (main, 2026-05-11) Out-of-bounds Read |
|---|
| Description | A vulnerability was found in PowerInfer commit 8bd56d69906c9d2dba4d3bf6899763401e01a9a4
(main, 2026-05-11) and classified as high 7.1. Affected is the FFN GPU-split loader
component, specifically create_striped_mat_to_gpu() in the llama.cpp file. The
manipulation of the argument int32 row indices stored in the gpu_bucket tensors of an
auto-loaded <model>.generated.gpuidx sidecar file leads to an out-of-bounds read.
The attack is local: the victim must run the inference binary on a crafted model
package (model file plus sidecar, distributed as a directory). User interaction is
required because the victim starts the load. The sidecar is loaded merely because the
file exists next to the model; only split.vram_capacity and the tensor count are
validated, while each int32 index is used unchecked as host_mat_row->data =
(char*)src->data + host_i * row_data_size and then passed as the cudaMemcpy source
address via ggml_cuda_cpy_1d. An index such as 0x7FFFFFFF (or any negative value)
places the read up to INT32_MAX * row_data_size bytes outside the tensor buffer,
crashing the process (ASan heap-buffer-overflow READ / SEGV); copied out-of-bounds
host memory becomes part of the device FFN weights used by subsequent inference.
Technical Details
- Affected file/function: llama.cpp / create_striped_mat_to_gpu() (llama.cpp:2898,
vulnerable loop at llama.cpp:2926-2934); auto-load at llm_load_gpu_split_with_budget()
(llama.cpp:3088-3097); only checks at llama_gpu_split_loader ctor / check_vram_allocable()
(llama.cpp:2816) and load_gpu_idx_for_model() (llama.cpp:2820)
- Vulnerable parameter: int32 entries of gpu_bucket tensors in <model>.generated.gpuidx
- Attack vector: Local
- Privileges required: None
- Trigger condition: run ./main -m model.gguf (--vram-budget 8 or the default -1 budget)
with a GPU present while a crafted .generated.gpuidx sits next to the model file;
--disable-gpu-index is the only opt-out
Impact
- Confidentiality: High (out-of-bounds host memory is copied into device FFN weights and
can surface through inference outputs; conservative estimate)
- Integrity: None
- Availability: High (process crash during model load)
CVSS v3.1
Score: 7.1 (High)
Vector: CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:N/A:H
Timeline
- Discovered: 2026-09-24
- Vendor notified: [pending]
- Patch released: [pending]
- Public disclosure: [pending]
Countermeasure
Validate every gpu_bucket entry before pointer arithmetic (reject host_i < 0 or
host_i >= src->ne[1]); treat the auto-loaded .generated.gpuidx sidecar as untrusted
input and bind it to the model file (e.g. hash check) instead of trusting it on
existence alone. |
|---|
| Source | ⚠️ https://github.com/Tiiny-AI/PowerInfer/issues/284 |
|---|
| User | Dem0 (UID 82596) |
|---|
| Submission | 09/24/2026 17:57 (17 days ago) |
|---|
| Moderation | 10/11/2026 21:28 (17 days later) |
|---|
| Status | Accepted |
|---|
| VulDB entry | 416779 [Tiiny-AI PowerInfer up to 8bd56d69906c9d2dba4d3bf6899763401e01a9a4 FFN GPU-split Loader ggml-cuda.cu create_striped_mat_to_gpu out-of-bounds] |
|---|
| Points | 20 |
|---|