vLLM 是用于大型语言模型推理和服务的引擎。在 0.30.0 版本之前,/score 和 /rerank 端点的 flash 晚期交互评分机制从由调用者控制的 X-Request-Id 头部字段中派生每个工作进程(worker)的 query_key 值。一个并发请求如果复用了受害者的标识符,可以覆盖缓存中的查询嵌入(query embedding),导致受害者的文档被针对攻击者的查询进行评分;此外,共享使用的计数器也可能引发晚期交互缓存未命中错误。此问题已在 0.30.0 版本中得到修复。
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | < 030.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POC| CVE-2026-105753 | 6.5 MEDIUM | vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reu |
| CVE-2026-105754 | 6.5 MEDIUM | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features |
| CVE-2026-105756 | 6.5 MEDIUM | vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de |
| CVE-2026-105757 | 6.5 MEDIUM | vLLM: Structured-output request errors escape the request boundary and terminate the share |
| CVE-2026-105759 | 5.9 MEDIUM | vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i |
| CVE-2026-105760 | 5.3 MEDIUM | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion |
| CVE-2026-105758 | 5.3 MEDIUM | vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the |
| CVE-2026-105752 | 3.1 LOW | vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache |
No comments yet