GHSA-2823-qmq8-rwvjMediumCVSS 6.5

vLLM: Loose `cache_salt` validation lets a single request kill EngineCore on LMCache-MP deployments — uncaught downstream `ValueError` denial of service

Published
October 5, 2026
Last Modified
October 5, 2026

🔗 CVE IDs covered (1)

📋 Description

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the loose cache_salt validator and unguarded scheduling-path lookup reach.

Summary

vLLM's OpenAI-compatible request models (Completions, Chat Completions, Responses) accept a client-supplied cache_salt field and validate it only as "must be a non-empty string" — no character or length restrictions. On a deployment with the built-in LMCache-MP KV connector enabled, that value is stored verbatim on the request tracker and forwarded unguarded as a keyword argument into the scheduler's per-step cache lookup. The downstream LMCache library applies a stricter check in IPCCacheServerKey.__post_init__, which raises ValueError for any cache_salt containing @, /, \, or NUL (or longer than 128 characters).

Neither the LMCache-MP connector lookup call site nor Scheduler.schedule() wraps that call in a request-scoped try/except, so the ValueError propagates uncaught into EngineCore's top-level handler, which treats any uncaught exception as fatal and kills the whole engine process. A single publicly reachable request with, for example, cache_salt="/" therefore takes down the engine for all concurrent users. vLLM's boundary validator is looser than the downstream consumer's, and the gap is never converted into a request-scoped failure on the scheduling path.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The three OpenAI check_cache_salt_support validators are identical; the completion one is representative — a non-empty-string test with no character or length bound:

# vllm/entrypoints/openai/completion/protocol.py Lines 503-509
        if data.get("cache_salt") is not None and (
            not isinstance(data["cache_salt"], str) or not data["cache_salt"]
        ):
            raise VLLMValidationError(
                "Parameter 'cache_salt' must be a non-empty string if provided.",
                parameter="cache_salt",
            )

On the scheduling path the connector is invoked with no surrounding try — a ValueError from the downstream salt check propagates straight out of schedule():

# vllm/v1/core/sched/scheduler.py Lines 736-742
                    # Get externally-cached tokens if using a KVConnector.
                    if self.connector is not None:
                        ext_tokens, load_kv_async = (
                            self.connector.get_num_new_matched_tokens(
                                request, num_new_local_computed_tokens
                            )
                        )

run_engine_core's generic handler — the next except up the stack — treats that as fatal, marks the engine dead, and re-raises:

# vllm/v1/engine/core.py Lines 1229-1235
        except Exception as e:
            if engine_core is None:
                logger.exception("EngineCore failed to start.")
            else:
                logger.exception("EngineCore encountered a fatal error.")
                engine_core._send_engine_dead()
            raise e

Impact

Availability only. cache_salt is an attacker-controlled, publicly reachable request field that vLLM validates too loosely. A value such as "/" passes vLLM's check, reaches the stricter downstream validator, and its ValueError is never converted into a request-scoped failure — instead it kills the EngineCore process, a denial of service for every concurrent request on that server (HTTP failures, then /health failing).

Applicability: the built-in LMCache-MP KV connector must be enabled (lmcache >= 0.4.4), which is itself an opt-in KV-connector boundary. On such deployments no other special configuration is required, and the crash is a resource-availability failure rather than expected behavior of the opt-in feature.

Suggested Fix

Two independent fixes:

  1. Tighten admission — add one shared validate_cache_salt() helper (for example in vllm/entrypoints/openai/engine/protocol.py) that matches or exceeds the downstream IPCCacheServerKey rules — reject @, /, \, NUL, and >128-character salts at the HTTP boundary with a 4xx — and route every request model that exposes cache_salt through it: the three check_cache_salt_support validators above, the pooling base request, and the token-in-token-out GenerateRequest (which has no validator today).
  2. Defense in depth — wrap the LMCache-MP lookup call reached from Scheduler.schedule() so a downstream validator ValueError becomes a request-scoped failure instead of an EngineCore-fatal exception. This is the fix that also covers future divergence between vLLM's and LMCache's salt rules; the right failure semantics (fail the request vs. fall back to a cold lookup) is a maintainer design call.

The core of fix 1 is a single shared helper; each check_cache_salt_support body then becomes validate_cache_salt(data.get("cache_salt")), and GenerateRequest gains an equivalent mode="before" validator:

# vllm/entrypoints/openai/engine/protocol.py — new shared helper
_CACHE_SALT_FORBIDDEN_CHARS = frozenset("@/\\\x00")
_MAX_CACHE_SALT_LENGTH = 128


def validate_cache_salt(cache_salt: object) -> None:
    """Validate cache salts before they reach downstream cache backends."""
    if cache_salt is None:
        return
    if not isinstance(cache_salt, str) or not cache_salt:
        raise VLLMValidationError(
            "Parameter 'cache_salt' must be a non-empty string if provided.",
            parameter="cache_salt",
        )
    if len(cache_salt) > _MAX_CACHE_SALT_LENGTH or any(
        char in _CACHE_SALT_FORBIDDEN_CHARS for char in cache_salt
    ):
        raise VLLMValidationError(
            "Parameter 'cache_salt' must be at most 128 characters and must "
            "not contain '@', '/', '\\\\', or NUL.",
            parameter="cache_salt",
        )

This distinguishes the finding from GHSA-6qc9-v4r8-22xg: that advisory fixed only the guided_json/xgrammar trigger of the same schedule()-into-run_engine_core fatal-handler family, so the cache_salt admission gap and the schedule()-level defense-in-depth (fix 2) survive its published fix. A patch implementing fix 1 across all five request models, with a regression test covering the rejected ("/", 129-char) and accepted salt shapes, applies to v0.25.1 with line offsets and no fuzz.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51444

🎯 Affected products1

  • pip/vllm:< 0.30.0

🔗 References (5)