vLLM 是一个用于大语言模型的推理和服务引擎。在 0.30.0 版本之前,Harmony 工具续传请求通过 “POST /v1/responses” 提交时,会重新构建下一轮的引擎输入,但未保留 值,从而将续传前缀写入全局无盐缓存命名空间,即使调用方启用了加盐机制。在启用前缀缓存(默认开启)的部署环境中,已认证的租户若能重构受害者的低熵工具后处理历史,则可提交相同的续传请求,并利用 计数判断该前缀是否曾被处理过,从而破坏加盐前缀缓存所期望的租户隔离机制。此问题已在 0.30.0 版本中得到修复。
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | < 0.30.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POC| CVE-2026-105753 | 6.5 MEDIUM | vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reu |
| CVE-2026-105754 | 6.5 MEDIUM | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features |
| CVE-2026-105756 | 6.5 MEDIUM | vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de |
| CVE-2026-105757 | 6.5 MEDIUM | vLLM: Structured-output request errors escape the request boundary and terminate the share |
| CVE-2026-105759 | 5.9 MEDIUM | vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i |
| CVE-2026-105760 | 5.3 MEDIUM | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion |
| CVE-2026-105758 | 5.3 MEDIUM | vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the |
| CVE-2026-105755 | 4.2 MEDIUM | vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled req |
No comments yet