GHSA-935w-9g4m-p28pLowCVSS 3.1

vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle

Published
October 6, 2026
Last Modified
October 6, 2026

🔗 CVE IDs covered (1)

📋 Description

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.

Summary

On the GPT-OSS "Harmony" path (POST /v1/responses), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries request.cache_salt, but the tool-continuation re-submission rebuilds the engine input via tokens_input(token_ids) with no cache_salt. The continuation prefix is therefore cached in the global unsalted namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that cache_salt is documented to prevent.

Silently dropping a preserved salt after the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.

This is distinct from GHSA-4qjh-9fv9-r85r (CVE-2025-46570): that advisory is the prefix-cache membership oracle for which cache_salt is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact cached_tokens_per_turn counts from a different sink (the Responses serving continuation, not general TTFT timing).

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The tool-continuation re-submission rebuilds the engine input with no cache_salt:

# vllm/entrypoints/openai/responses/serving.py Lines 711-715
            if isinstance(context, HarmonyContext):
                token_ids = context.render_for_completion()
                engine_input = tokens_input(token_ids)

                sampling_params.max_tokens = max_model_len - len(token_ids)

Contrast with the correct turn-1 call, which does preserve the caller's salt:

# vllm/entrypoints/openai/responses/serving.py Lines 754-755
        prompt_token_ids = render_for_completion(messages)
        engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)

tokens_input stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:

# vllm/inputs/engine.py Lines 51-66
def tokens_input(
    prompt_token_ids: list[int],
    *,
    prompt: str | None = None,
    cache_salt: str | None = None,
) -> TokensInput:
    """
    Construct [`TokensInput`][vllm.inputs.engine.TokensInput]
    from optional values.
    """
    inputs = TokensInput(type="token", prompt_token_ids=prompt_token_ids)

    if prompt is not None:
        inputs["prompt"] = prompt
    if cache_salt is not None:
        inputs["cache_salt"] = cache_salt

Impact

An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle cache_salt is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.

Preconditions: a GPT-OSS Harmony model on /v1/responses; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets cache_salt and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The AC:H metric reflects that guessable-history precondition.

Suggested Fix

Propagate request.cache_salt into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call tokens_input(token_ids, cache_salt=request.cache_salt), mirroring the correct turn-1 call. Carry the salt on the HarmonyContext (thread the originating request into the context) so no continuation path can omit it:

# vllm/entrypoints/openai/responses/serving.py
             if isinstance(context, HarmonyContext):
                 token_ids = context.render_for_completion()
-                engine_input = tokens_input(token_ids)
+                engine_input = tokens_input(
+                    token_ids,
+                    cache_salt=(
+                        context.request.cache_salt
+                        if context.request is not None
+                        else None
+                    ),
+                )

with HarmonyContext.__init__ gaining a request: ResponsesRequest | None = None parameter (stored as self.request) that _create_responses passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.

Suggested regression test: assert cached_tokens_per_turn == 0 for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818

🎯 Affected products1

  • pip/vllm:< 0.30.0

🔗 References (6)