vLLM 是一款为大语言模型提供推理和服务的引擎。在版本 0.28.0 之前,默认的镜像多模态 LRU 缓存存在一个问题:在多模态渲染过程中、引擎准入检查之前,前端发送者缓存可能会提交一个媒体哈希值;而如果该请求随后被拒绝,引擎接收者缓存将永远不会收到对应的有效载荷。当后续请求复用相同的媒体哈希时,MultiModalProcessorSenderCache 不会发送任何有效载荷,而 MultiModalReceiverCache 则会触发一个断言失败,报错信息为“Expected a cached item”(期
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | < 0.28.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POC| CVE-2026-105754 | 6.5 MEDIUM | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features |
| CVE-2026-105756 | 6.5 MEDIUM | vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de |
| CVE-2026-105757 | 6.5 MEDIUM | vLLM: Structured-output request errors escape the request boundary and terminate the share |
| CVE-2026-105759 | 5.9 MEDIUM | vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i |
| CVE-2026-105760 | 5.3 MEDIUM | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion |
| CVE-2026-105758 | 5.3 MEDIUM | vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the |
| CVE-2026-105755 | 4.2 MEDIUM | vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled req |
| CVE-2026-105752 | 3.1 LOW | vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache |
No comments yet