vLLM 是专为大语言模型构建的推理与 Serving 引擎。在版本 0.30.0 之前,攻击者可通过请求级别的 media_io_kwargs 字段选择 GLMGA 视频后端,并为 fps 和 max_frames 选项传入极大的值,且系统未设置严格的工作量上限。GLMGA 会构造并去重一个由攻击者控制大小的预解码帧索引列表,从而导致一个紧凑的请求和极小的有效视频即可在共享的媒体加载执行器中消耗不成比例的 CPU 时间和内存资源。该问题已在 0.30.0 版本中修复。
Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.
| Vendor | Product | Affected Versions | CPE | Subscribe |
|---|---|---|---|---|
| vllm-project | vllm | >= 0.23.0rc2, < 0.30.0 | - |
|
| # | POC Description | Source Link | Shenlong Link |
|---|
No public POC found.
Login to generate AI POC| CVE-2026-105753 | 6.5 MEDIUM | vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reu |
| CVE-2026-105754 | 6.5 MEDIUM | vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features |
| CVE-2026-105756 | 6.5 MEDIUM | vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de |
| CVE-2026-105757 | 6.5 MEDIUM | vLLM: Structured-output request errors escape the request boundary and terminate the share |
| CVE-2026-105759 | 5.9 MEDIUM | vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i |
| CVE-2026-105758 | 5.3 MEDIUM | vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the |
| CVE-2026-105755 | 4.2 MEDIUM | vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled req |
| CVE-2026-105752 | 3.1 LOW | vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache |
No comments yet