Goal Reached Thanks to every supporter — we hit 100%!

Goal: 1000 CNY · Raised: 1359 CNY

100%

CVE-2026-105757— vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)

Quick assessment

Affected
vllm-project vllm
Exploitation
No confirmed in-the-wild exploitation; assess based on exposure
Recommended action
Check the vendor advisory and references for a fixed version. If immediate upgrade is impossible, restrict exposure and increase monitoring.

vLLM 是一个用于大型语言模型的推理和服务引擎。在 0.30.0 版本之前,结构化输出请求的失败可能绕过请求级别的验证,进而触发 EngineCore 中的致命错误路径。具体的漏洞场景包括: 每个请求的后端不匹配可能导致重新抛出语法编译异常; 由 ngram_gpu 推测解码模式生成的填充可能向引导验证传入负数 token; Rust 前端可能接受 Python 前端所拒绝的空结构化输出值。 这些缺陷使得普通的约束生成请求能够终止共享引擎。该问题已在 0.30.0 版本中得到修复。

CVSS 6.5 · Medium
Get alerts for future matching vulnerabilities Log in to subscribe

I. Basic Information for CVE-2026-105757

Vulnerability Information

Have questions about the vulnerability? See if Shenlong's analysis helps!
View Shenlong Deep Dive ↗

Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.

Vulnerability Title
vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Source: CVE Program / CVE List V5
Vulnerability Description
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
Source: CVE Program / CVE List V5
CVSS Information
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Source: CVE Program / CVE List V5
Vulnerability Type
输入验证不恰当
Source: CVE Program / CVE List V5

Affected Products

Vendor Product Affected Versions CPE Subscribe
vllm-project vllm < 0.30.0 -

II. Public POCs for CVE-2026-105757

# POC Description Source Link Shenlong Link
AI-Generated POC Premium

No public POC found.

Login to generate AI POC

III. Intelligence Information for CVE-2026-105757

请登录查看更多情报信息。

Other References for CVE-2026-105757 (4)

Same Patch Batch · vllm-project · 2026-10-05 · 9 CVEs total

CVE-2026-105753 6.5 MEDIUM vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reu
CVE-2026-105754 6.5 MEDIUM vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
CVE-2026-105756 6.5 MEDIUM vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP de
CVE-2026-105759 5.9 MEDIUM vLLM: Unbounded Prometheus label cardinality from attacker-controlled HTTP method tokens i
CVE-2026-105760 5.3 MEDIUM vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
CVE-2026-105758 5.3 MEDIUM vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the
CVE-2026-105755 4.2 MEDIUM vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled req
CVE-2026-105752 3.1 LOW vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache

IV. Related Vulnerabilities

V. Comments for CVE-2026-105757

No comments yet


Leave a comment