Description
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
Published: 2026-10-05
Score: 6.5 Medium
EPSS: < 1% Very Low
KEV: No
Impact: n/a
Action: n/a
AI Analysis

Analysis and contextual insights are available on OpenCVE Cloud.

Remediation

No vendor fix or workaround currently provided.

Additional remediation guidance may be available on OpenCVE Cloud.

Tracking

Sign in to view the affected projects.

Advisories
Source ID Title
Github GHSA Github GHSA GHSA-85xf-c7hm-whqw vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
History

Tue, 06 Oct 2026 14:30:00 +0000

Type Values Removed Values Added
Metrics ssvc

{'options': {'Automatable': 'no', 'Exploitation': 'none', 'Technical Impact': 'partial'}, 'version': '2.0.3'}


Tue, 06 Oct 2026 12:15:00 +0000

Type Values Removed Values Added
References
Metrics threat_severity

None

threat_severity

Moderate


Tue, 06 Oct 2026 01:15:00 +0000

Type Values Removed Values Added
First Time appeared Vllm-project
Vllm-project vllm
Vendors & Products Vllm-project
Vllm-project vllm

Mon, 05 Oct 2026 23:00:00 +0000

Type Values Removed Values Added
Description vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
Title vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Weaknesses CWE-20
CWE-248
CWE-755
References
Metrics cvssV3_1

{'score': 6.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H'}


Subscriptions

Vllm-project Vllm
cve-icon MITRE

Status: PUBLISHED

Assigner: GitHub_M

Published:

Updated: 2026-10-06T13:20:25.829Z

Reserved: 2026-10-05T19:11:07.947Z

Link: CVE-2026-105757

cve-icon Vulnrichment

Updated: 2026-10-06T13:20:18.839Z

cve-icon NVD

Status : Awaiting Analysis

Published: 2026-10-05T23:17:02.467

Modified: 2026-10-06T15:19:48.933

Link: CVE-2026-105757

cve-icon Redhat

Severity : Moderate

Publid Date: 2026-10-05T22:52:05Z

Links: CVE-2026-105757 - Bugzilla

cve-icon OpenCVE Enrichment

Updated: 2026-10-06T01:00:09Z

Weaknesses
  • CWE-20

    Improper Input Validation

  • CWE-248

    Uncaught Exception

  • CWE-755

    Improper Handling of Exceptional Conditions