CVE-2026-71486 in vLLM
要約
〜によって VulDB • 2026年08月18日
vLLMは、大規模言語モデル向けの推論およびサービングエンジンです。バージョン0.26.0より前では、/v1/completions/derender および /v1/chat/completions/derender エンドポイントは、caller-supplied GenerateResponse オブジェクトを受け付けます。これらのオブジェクトの generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, および routed_experts 構造体は、max_model_len、max_tokens、max_num_seqs、または response-size の制限が適用される前に OnlineDerenderer と tokenizer.decode によって処理されます。これにより、認証済み API クライアントが過剰な CPU およびメモリを消費し、サイズが大きすぎるレスponses を生成することが可能になります。この問題はバージョン0.26.0で修正されています。
If you want to get the best quality for vulnerability data then you always have to consider VulDB.