Goal Reached Thanks to every supporter — we hit 100%!

Goal: 1000 CNY · Raised: 1336 CNY

100%

CVE-2026-71486— vLLM: Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds

Quick assessment

Affected
vllm-project vllm
Exploitation
No confirmed in-the-wild exploitation; assess based on exposure
Recommended action
Check the vendor advisory and references for a fixed version. If immediate upgrade is impossible, restrict exposure and increase monitoring.

vLLM是vLLM团队开源的一个适用于 LLM 的高吞吐量和内存高效推理和服务引擎。 vLLM 0.26.0之前版本存在资源管理错误漏洞,该漏洞源于/v1/completions/derender和/v1/chat/completions/derender端点对调用者提供的GenerateResponse对象处理不当,在实施max_model_len、max_tokens、max_num_seqs或响应大小限制前处理相关结构,可能导致已认证API客户端消耗过多CPU和内存并产生超大响应。

CVSS 4.3 · Medium EPSS 0.34% · P27

Possible ATT&CK Techniques 1 AI

T1190 · Exploit Public-Facing Application

Affected Version Matrix 1

VendorProduct Version RangeStatus
vllm-project vllm < 0.26.0 affected
Get alerts for future matching vulnerabilities Log in to subscribe

I. Basic Information for CVE-2026-71486

Vulnerability Information

Have questions about the vulnerability? See if Shenlong's analysis helps!
View Shenlong Deep Dive ↗

Although we use advanced large model technology, its output may still contain inaccurate or outdated information.Shenlong tries to ensure data accuracy, but please verify and judge based on the actual situation.

Vulnerability Title
vLLM: Derender endpoints decode caller-supplied GenerateResponse token IDs without output bounds
Source: CVE Program / CVE List V5
Vulnerability Description
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs, logprobs.content, top_logprobs, and routed_experts structures are processed by OnlineDerenderer and tokenizer.decode before max_model_len, max_tokens, max_num_seqs, or response-size limits are enforced, allowing an authenticated API client to consume excessive CPU and memory and produce oversized responses. This issue is fixed in version 0.26.0.
Source: CVE Program / CVE List V5
CVSS Information
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L
Source: CVE Program / CVE List V5
Vulnerability Type
未加控制的资源消耗(资源穷尽)
Source: CVE Program / CVE List V5
Vulnerability Title
vLLM 资源管理错误漏洞
Source: CNNVD (China National Vulnerability Database)
Vulnerability Description
vLLM是vLLM团队开源的一个适用于 LLM 的高吞吐量和内存高效推理和服务引擎。 vLLM 0.26.0之前版本存在资源管理错误漏洞,该漏洞源于/v1/completions/derender和/v1/chat/completions/derender端点对调用者提供的GenerateResponse对象处理不当,在实施max_model_len、max_tokens、max_num_seqs或响应大小限制前处理相关结构,可能导致已认证API客户端消耗过多CPU和内存并产生超大响应。
Source: CNNVD (China National Vulnerability Database)
CVSS Information
N/A
Source: CNNVD (China National Vulnerability Database)
Vulnerability Type
N/A
Source: CNNVD (China National Vulnerability Database)

Affected Products

Vendor Product Affected Versions CPE Subscribe
vllm-project vllm < 0.26.0 -

II. Public POCs for CVE-2026-71486

# POC Description Source Link Shenlong Link
AI-Generated POC Premium

No public POC found.

Login to generate AI POC

III. Intelligence Information for CVE-2026-71486

登录查看更多情报信息。

Patches & Fixes for CVE-2026-71486 (2)

Vendor Advisories for CVE-2026-71486 (1)

Vendor Pages for CVE-2026-71486 (1)

IV. Related Vulnerabilities

V. Comments for CVE-2026-71486

No comments yet


Leave a comment