GHSA-2phq-3phc-84pxMediumCVSS 4.2

vLLM: Flash late-interaction scoring caches query embeddings under a caller-controlled request id — cross-request integrity break and induced errors on `/score` and `/rerank`

Published
October 5, 2026
Last Modified
October 5, 2026

🔗 CVE IDs covered (1)

📋 Description

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the flash late-interaction query cache reaches.

Summary

On late-interaction /score and /rerank deployments with flash late interaction enabled (the default for supported models), the worker caches per-request query embeddings under a key derived from the caller-controlled X-Request-Id header. A second concurrent request that reuses the victim's header value replaces the victim's cached query embedding before document scoring — so the victim's documents are scored against the attacker's query. Because the data-parallel router pins all requests sharing a cache key to the same engine, the collision is deterministic for an attacker who reuses the victim's X-Request-Id. Depending on timing, one request can also consume the shared use counter and force the other request into a late-interaction cache-miss error.

This is a remotely reachable, request-controlled cross-request integrity break on the standard scoring and reranking endpoints. It requires only that flash late interaction be enabled, which is the default for supported models.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The caller-controlled header enters as the request id, and the flash late-interaction path derives the worker cache key directly from it:

# vllm/entrypoints/serve/engine/serving.py Lines 116-126
    @staticmethod
    def _base_request_id(
        raw_request: Request | None, default: str | None = None
    ) -> str | None:
        """Pulls the request id to use from a header, if provided"""
        if raw_request is not None and (
            (req_id := raw_request.headers.get("X-Request-Id")) is not None
        ):
            return req_id

        return random_uuid() if default is None else default
# vllm/entrypoints/pooling/scoring/serving.py Lines 207-212
        n_queries = ctx.n_queries
        n_docs = len(ctx.engine_inputs) - n_queries
        query_engine_inputs = ctx.engine_inputs[:n_queries]

        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
        query_uses = [n_docs if n_queries == 1 else 1] * n_queries

The worker then stores and reads the query embedding under that string with no check that the reader owns the entry — a colliding key returns another request's cached query, or (once the use counter is exhausted) raises a cache-miss error:

# vllm/v1/worker/gpu/pool/late_interaction_runner.py Lines 91-107
            if mode == LATE_INTERACTION_MODE_CACHE_QUERY:
                assert query_uses is not None
                # `output` can be a view into the current step's hidden-states
                # buffer, so clone it before storing across scheduling steps.
                self._query_cache[query_key] = output.clone()
                self._query_uses[query_key] = query_uses
                outputs[i] = torch.zeros((), device=output.device, dtype=torch.float32)
                continue

            if mode == LATE_INTERACTION_MODE_SCORE_DOC:
                query_output = self._query_cache.get(query_key)
                if query_output is None:
                    raise ValueError(
                        "late-interaction query cache miss for key "
                        f"{query_key!r}. Ensure query requests are executed "
                        "before their paired document requests."
                    )

The bug is specific to the flash late-interaction path. Non-flash late-interaction scoring computes MaxSim directly from one request's in-memory outputs and does not create a cross-request worker cache key.

Impact

A network client of the standard scoring API can, on a flash late-interaction /score or /rerank deployment:

  1. Corrupt another user's results — by reusing the victim's X-Request-Id, the attacker's query embedding overwrites the victim's cached entry, so the victim's documents are scored against the attacker's query (a cross-request integrity break).
  2. Induce errors — depending on timing, one request consumes the shared use counter and forces the other request into a late-interaction cache-miss error.

Both consequences follow deterministically from reusing the victim's header value, because same-key work is pinned to one engine. This is reachable through normal request handling and does not depend on any trusted inter-node network.

Suggested Fix

Derive the flash late-interaction query-cache key from a server-generated, unforgeable per-request identifier (a random_uuid() namespace) rather than the caller-supplied X-Request-Id, and thread that key through the PoolingServeContext to the doc-scoring pass so both passes reuse the same key and a caller cannot address another request's cache entry.

In vllm/entrypoints/pooling/scoring/serving.py, the encode-queries pass mints a fresh namespace and stashes the keys on the context:

-        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+        query_namespace = random_uuid()
+        query_keys = [
+            f"late-interaction-{query_namespace}-query-{i}" for i in range(n_queries)
+        ]
+        ctx.late_interaction_query_keys = query_keys

and the encode-docs pass reads those stored keys instead of re-deriving them from ctx.request_id:

-        query_keys = [f"{ctx.request_id}-query-{i}" for i in range(n_queries)]
+        query_keys = ctx.late_interaction_query_keys
+        if query_keys is None:
+            raise RuntimeError("Late-interaction query keys were not initialized.")

This requires adding the late_interaction_query_keys: list[str] | None = None field to PoolingServeContext (vllm/entrypoints/pooling/typing.py). Because the namespace is a server-generated UUID, colliding X-Request-Id values no longer produce a shared cache key; a regression test asserting exactly that (colliding request ids yield distinct query-cache keys) accompanies the change.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51445

🎯 Affected products1

  • pip/vllm:< 0.30.0

🔗 References (5)