GHSA-ph3r-5jfg-f84fMediumCVSS 6.5

vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core

Published
October 6, 2026
Last Modified
October 6, 2026

🔗 CVE IDs covered (1)

📋 Description

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.

Summary

vLLM's default multimodal cache (mm_processor_cache_type="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that get_and_update() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1.

That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on max_model_len after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mm_item is not None, f"Expected a cached item for {mm_hash=}".

This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The P1 receiver sink — the assertion that fires when P1 receives None for a hash it never cached:

# vllm/multimodal/cache.py Lines 651-663
    @override
    def get_and_update_item(
        self,
        mm_item: MultiModalKwargsItem | None,
        mm_hash: str,
    ) -> MultiModalKwargsItem:
        if (cached_item := self._cache.get(mm_hash)) is not None:
            return cached_item

        assert mm_item is not None, f"Expected a cached item for {mm_hash=}"

        self._cache[mm_hash] = mm_item
        return mm_item

The P0 sender — on a hit it drops the payload (returns None) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:

# vllm/multimodal/cache.py Lines 409-422
    @override
    def get_and_update_item(
        self,
        mm_item: MultiModalProcessorCacheInItem,
        mm_hash: str,
    ) -> MultiModalProcessorCacheOutItem:
        if (cached_item := self._cache.get(mm_hash)) is not None:
            return None, cached_item.prompt_updates

        assert mm_item is not None, f"Expected a cached item for {mm_hash=}"

        self._cache[mm_hash] = MultiModalProcessorCacheItemMetadata(*mm_item)

        return mm_item

Impact

A remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a later request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.

On this revision the failure is scoped as a request-level preprocessing error (the engine core catches around preprocess_add_request); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored lru cache.

Suggested Fix

Two complementary changes:

  1. Make the mirrored commit atomic with admission — insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection.
  2. Defense in depth — convert the P1 receiver assert mm_item is not None into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.

The core of the rollback half: wrap the post-render length check so a ValueError rejection discards the P0 entries the render just committed, before re-raising. Add a discard_sender_cache_item() on the processor cache (no-op default, pop on the sender) and a Renderer.discard_mm_cache_entries() that walks a rendered request's mm_hashes:

# vllm/entrypoints/openai/chat_completion/serving.py (_create_chat_completion)
-            max_tokens = get_max_tokens(
-                max_model_len,
-                ...,
-                truncate_prompt_tokens=request.truncate_prompt_tokens,
-            )
+            try:
+                max_tokens = get_max_tokens(
+                    max_model_len,
+                    ...,
+                    truncate_prompt_tokens=request.truncate_prompt_tokens,
+                )
+            except ValueError:
+                for rendered_input in engine_inputs:
+                    if mm_hashes := rendered_input.get("mm_hashes"):
+                        self.renderer.discard_mm_cache_entries(mm_hashes)
+                raise
# vllm/multimodal/cache.py (MultiModalProcessorSenderCache)
+    @override
+    def discard_sender_cache_item(self, mm_hash: str) -> None:
+        self._cache.pop(mm_hash, None)

This closes the max_model_len rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 assert a checked, request-scoped error) is recommended.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897

🎯 Affected products1

  • pip/vllm:< 0.28.0

🔗 References (6)