漏洞信息
尽管我们使用了先进的大模型技术,但其输出仍可能包含不准确或过时的信息。神龙努力确保数据的准确性,但请您根据实际情况进行核实和判断。
Vulnerability Title
vLLM: Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds
Vulnerability Description
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts structures are processed by OnlineDerenderer and tokenizer.decode before max_model_len, max_tokens, max_num_seqs, or response-size limits are enforced, allowing an authenticated API client to consume excessive CPU and memory and produce oversized responses. This issue is fixed in version 0.26.0.
CVSS Information
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L
Vulnerability Type
未加控制的资源消耗(资源穷尽)
Vulnerability Title
vLLM 资源管理错误漏洞
Vulnerability Description
vLLM是vLLM团队开源的一个适用于 LLM 的高吞吐量和内存高效推理和服务引擎。 vLLM 0.26.0之前版本存在资源管理错误漏洞,该漏洞源于/v1/completions/derender和/v1/chat/completions/derender端点对调用者提供的GenerateResponse对象处理不当,在实施max_model_len、max_tokens、max_num_seqs或响应大小限制前处理相关结构,可能导致已认证API客户端消耗过多CPU和内存并产生超大响应。
CVSS Information
N/A
Vulnerability Type
N/A