vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle
🔗 CVE IDs covered (1)
📋 Description
Affected
- Ecosystem / package: pip /
vllm - Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit
752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the tool-continuation re-submission has omitted the salt.
Summary
On the GPT-OSS "Harmony" path (POST /v1/responses), a request that uses a built-in or MCP tool runs as a multi-turn loop: after each tool call vLLM re-renders the full next-turn Harmony prompt and re-submits it to the engine. Turn 1 correctly carries request.cache_salt, but the tool-continuation re-submission rebuilds the engine input via tokens_input(token_ids) with no cache_salt. The continuation prefix is therefore cached in the global unsalted namespace even though the caller opted into salting. A second tenant who can guess the low-entropy post-tool history submits the reconstructed continuation (unsalted) and reads exact per-turn cached-token counts from the Responses usage — restoring the prompt-membership oracle that cache_salt is documented to prevent.
Silently dropping a preserved salt after the supported tool workflow is enabled is a broken isolation control: the caller enabled salting and every turn should stay isolated, but continuation turns leak into the shared cache.
This is distinct from GHSA-4qjh-9fv9-r85r (CVE-2025-46570): that advisory is the prefix-cache membership oracle for which cache_salt is the documented mitigation, and its PR-17045 fix does not close this site — the Harmony tool continuation silently drops the preserved salt, caching in the unsalted namespace and leaking exact cached_tokens_per_turn counts from a different sink (the Responses serving continuation, not general TTFT timing).
Affected code
Links pinned to the confirmed commit 752a3a504485 (v0.25.1):
- The drop (sink):
vllm/entrypoints/openai/responses/serving.py#L712-L713—token_ids = context.render_for_completion()thenengine_input = tokens_input(token_ids), with nocache_salt. - Correct turn-1 call for contrast:
vllm/entrypoints/openai/responses/serving.py#L755—tokens_input(prompt_token_ids, cache_salt=request.cache_salt). tokens_inputstores the salt only if passed:vllm/inputs/engine.py#L51-L66(if cache_salt is not None: inputs["cache_salt"] = cache_salt).- The engine request copies only the current input's salt:
vllm/v1/engine/input_processor.py#L380(cache_salt=decoder_inputs.get("cache_salt")→Nonefor the continuation). - Prefix-cache hashing keys on the salt only when present:
vllm/v1/core/kv_cache_utils.py#L560-L561([request.cache_salt] if (start_token_idx == 0 and request.cache_salt) else []). - The oracle the attacker reads:
vllm/entrypoints/openai/responses/serving.py#L909(cached_tokens_per_turn). - The documented control being defeated:
vllm/entrypoints/openai/responses/protocol.py#L235(cache_saltfield).
The tool-continuation re-submission rebuilds the engine input with no cache_salt:
# vllm/entrypoints/openai/responses/serving.py Lines 711-715
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
engine_input = tokens_input(token_ids)
sampling_params.max_tokens = max_model_len - len(token_ids)
Contrast with the correct turn-1 call, which does preserve the caller's salt:
# vllm/entrypoints/openai/responses/serving.py Lines 754-755
prompt_token_ids = render_for_completion(messages)
engine_input = tokens_input(prompt_token_ids, cache_salt=request.cache_salt)
tokens_input stores the salt on the engine input only when it is passed, so the continuation input carries none and lands in the unsalted namespace:
# vllm/inputs/engine.py Lines 51-66
def tokens_input(
prompt_token_ids: list[int],
*,
prompt: str | None = None,
cache_salt: str | None = None,
) -> TokensInput:
"""
Construct [`TokensInput`][vllm.inputs.engine.TokensInput]
from optional values.
"""
inputs = TokensInput(type="token", prompt_token_ids=prompt_token_ids)
if prompt is not None:
inputs["prompt"] = prompt
if cache_salt is not None:
inputs["cache_salt"] = cache_salt
Impact
An authenticated tenant of a shared deployment can recover whether a guessed post-tool prompt or history was processed by another tenant, with exact cached-token counts rather than noisy latency — the exact prompt-membership oracle cache_salt is documented to prevent. It defeats the multi-user prefix-cache isolation guarantee for salted Harmony tool sessions.
Preconditions: a GPT-OSS Harmony model on /v1/responses; prefix caching enabled (default); an operator-enabled built-in or MCP tool server; the victim sets cache_salt and triggers at least one tool continuation; and the attacker can reconstruct the post-tool history closely enough to match the token prefix. The AC:H metric reflects that guessable-history precondition.
Suggested Fix
Propagate request.cache_salt into every Harmony (and Parsable) tool-continuation re-submission — at the continuation call site call tokens_input(token_ids, cache_salt=request.cache_salt), mirroring the correct turn-1 call. Carry the salt on the HarmonyContext (thread the originating request into the context) so no continuation path can omit it:
# vllm/entrypoints/openai/responses/serving.py
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
- engine_input = tokens_input(token_ids)
+ engine_input = tokens_input(
+ token_ids,
+ cache_salt=(
+ context.request.cache_salt
+ if context.request is not None
+ else None
+ ),
+ )
with HarmonyContext.__init__ gaining a request: ResponsesRequest | None = None parameter (stored as self.request) that _create_responses passes when constructing the context. The continuation prefix is then cached in the victim's salted namespace, mirroring turn 1.
Suggested regression test: assert cached_tokens_per_turn == 0 for a different-salt probe against a salted victim continuation (the four-way control from the proof of concept).
Credit
Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)
This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.
Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51818
🎯 Affected products1
- pip/vllm:< 0.30.0
🔗 References (6)
- https://github.com/vllm-project/vllm/security/advisories/GHSA-935w-9g4m-p28p
- https://github.com/vllm-project/vllm/pull/50195
- https://github.com/vllm-project/vllm/pull/51818
- https://github.com/vllm-project/vllm/commit/6a2a2bb02b563b83f946012959fd3927984d072a
- https://github.com/vllm-project/vllm/releases/tag/v0.30.0
- https://github.com/advisories/GHSA-935w-9g4m-p28p