Goal Reached Thanks to every supporter — we hit 100%!

Goal: 1000 CNY · Raised: 1359 CNY

100%

CVE-2026-100651— vllm before 0.29.0 Denial of Service via Decoder Prompt Length Bypass

Quick assessment

Affected
vllm-project vllm
Exploitation
No confirmed in-the-wild exploitation; assess based on exposure
Recommended action
Check the vendor advisory and references for a fixed version. If immediate upgrade is impossible, restrict exposure and increase monitoring.

vLLM 在 0.29.0 版本之前,在解耦式服务(disaggregated serving)端点 /inference/v1/generate 上未能强制执行解码器提示词长度验证。当请求中包含“features”(多模态)负载时,vllm/entrypoints/serve/disagg/serving.py 会直接根据调用方提供的 token_ids 构建多模态 EngineInput,且未对 GenerateRequest.token_ids(定义于 vllm/entrypoints/serve/disag

CVSS 6.5 · Medium EPSS 0.31% · P21
Get alerts for future matching vulnerabilities Log in to subscribe

I. Basic Information for CVE-2026-100651

Vulnerability Information

Have questions about the vulnerability? See if Shenlong's analysis helps!
View Shenlong Deep Dive ↗

Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.

Vulnerability Title
vllm before 0.29.0 Denial of Service via Decoder Prompt Length Bypass
Source: CVE Program / CVE List V5
Vulnerability Description
vLLM before 0.29.0 fails to enforce decoder prompt-length validation on the disaggregated serving endpoint /inference/v1/generate. When the request contains a 'features' (multimodal) payload, vllm/entrypoints/serve/disagg/serving.py builds a multimodal EngineInput directly from the caller-supplied token_ids, and GenerateRequest.token_ids (vllm/entrypoints/serve/disagg/protocol.py) is not checked against model_config.max_model_len. For multimodal processors that report skip_prompt_length_check=True (for example Nemotron Parse, Whisper, and FireRedLID), InputProcessor._validate_prompt_len() returns immediately for both encoder and decoder prompts, so an overlong prompt becomes an EngineCoreRequest and reaches the worker input-batch copy into a fixed max_model_len-wide NumPy row. A client able to reach the endpoint on an affected model configuration can therefore submit an overlong token_ids list to trigger a worker failure and denial of service. Fixed in 0.29.0.
Source: CVE Program / CVE List V5
CVSS Information
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
Source: CVE Program / CVE List V5
Vulnerability Type
未加控制的资源消耗(资源穷尽)
Source: CVE Program / CVE List V5

Affected Products

Vendor Product Affected Versions CPE Subscribe
vllm-project vllm 0 ~ 0.29.0 -

II. Public POCs for CVE-2026-100651

# POC Description Source Link Shenlong Link
AI-Generated POC Premium

No public POC found.

Login to generate AI POC

III. Intelligence Information for CVE-2026-100651

请登录查看更多情报信息。

Other References for CVE-2026-100651 (1)

Other References for CVE-2026-100651 (1)

Same Patch Batch · vllm-project · 2026-09-26 · 8 CVEs total

CVE-2026-100654 6.5 MEDIUM vLLM before 0.29.0 Denial of Service via out-of-range stop_token_ids
CVE-2026-100650 6.5 MEDIUM vLLM before 0.29.0 Resource Exhaustion via Unbounded Media Materialization
CVE-2026-100653 6.5 MEDIUM vLLM 0.22.1 before 0.28.0 Incomplete Artifact Pin Propagation
CVE-2026-100652 5.9 MEDIUM vLLM 0.22.0 through 0.23.0 Denial of Service via stop_token_ids
CVE-2026-100647 5.3 MEDIUM vLLM before 0.29.0 CPU Exhaustion via unbounded cache_salt
CVE-2026-100648 5.3 MEDIUM vllm before 0.29.0 Uncontrolled Resource Consumption via Audio Decoding
CVE-2026-100649 3.7 LOW vLLM before 0.29.0 Resource Limit Bypass via Sampler Subclass

IV. Related Vulnerabilities

V. Comments for CVE-2026-100651

No comments yet


Leave a comment