vLLM是vLLM团队开源的一个适用于 LLM 的高吞吐量和内存高效推理和服务引擎。 vLLM 0.26.0之前版本存在资源管理错误漏洞,该漏洞源于/v1/completions/derender和/v1/chat/completions/derender端点对调用者提供的GenerateResponse对象处理不当,在实施max_model_len、max_tokens、max_num_seqs或响应大小限制前处理相关结构,可能导致已认证API客户端消耗过多CPU和内存并产生超大响应。
| Vendor | Product | Version Range | Status |
|---|---|---|---|
| vllm-project | vllm | < 0.26.0 |
affected |
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | < 0.26.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POCNo comments yet