vLLM 是一个用于大型语言模型的推理和服务引擎。在 0.30.0 版本之前,结构化输出请求的失败可能绕过请求级别的验证,进而触发 EngineCore 中的致命错误路径。具体的漏洞场景包括: 每个请求的后端不匹配可能导致重新抛出语法编译异常; 由 ngram_gpu 推测解码模式生成的填充可能向引导验证传入负数 token; Rust 前端可能接受 Python 前端所拒绝的空结构化输出值。 这些缺陷使得普通的约束生成请求能够终止共享引擎。该问题已在 0.30.0 版本中得到修复。
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | < 0.30.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POC| CVE-2026-105753 | 6.5 MEDIUM | vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reu |
| CVE-2026-105754 | 6.5 MEDIUM | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features |
| CVE-2026-105756 | 6.5 MEDIUM | vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de |
| CVE-2026-105759 | 5.9 MEDIUM | vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i |
| CVE-2026-105760 | 5.3 MEDIUM | vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion |
| CVE-2026-105758 | 5.3 MEDIUM | vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the |
| CVE-2026-105755 | 4.2 MEDIUM | vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled req |
| CVE-2026-105752 | 3.1 LOW | vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache |
No comments yet