GHSA-x6mc-67gf-chw4MediumCVSS 5.3

vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach

Published
October 5, 2026
Last Modified
October 5, 2026

🔗 CVE IDs covered (1)

📋 Description

Summary

An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level media_io_kwargs.video.max_frames and fps knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated POST /tokenize.

The num_frames ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read num_frames at all. The same knobs were already capped upstream for GLMGAVideoBackend as an accepted security fix in 8b6de0eb9 (PR #54935, merged 2026-09-04); that cap never reached Qwen.

Details

Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first)

GHSA-vxqj-p4gw-9h4c reported that request-level media_io_kwargs.video.num_frames overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping num_frames inside VideoMediaIO.merge_kwargs.

That clamp does not reach the Qwen samplers. Qwen2VLVideoBackend and Qwen3VLVideoBackend do not read num_frames at all — Qwen2VLVideoBackend's own docstring says so ("num_frames is ignored (fps-driven, like the Qwen3-VL loader)"). They bound on max_frames, read from the same merged dict with no ceiling:

# vllm/multimodal/video.py — Qwen3VLVideoBackend.compute_frames_index_to_sample
min_frames = kwargs.get("min_frames", 4)
max_frames = kwargs.get("max_frames", 768)
num_frames = int(total_frames_num / original_fps * fps)
num_frames = min(max(num_frames, min_frames), max_frames, total_frames_num)

With max_frames raised from the request, the only remaining bound is total_frames_num — every frame in the container.

I applied PR #51969's patch locally and re-ran both paths through the real merge layer (merge_media_io_kwargs → VideoMediaIO.merge_kwargs → MediaConnector), against main @ b23433088:

| request media_io_kwargs.video | merged kwargs after #51969 | frames decoded | peak RSS | |---|---|---|---| | (absent) | None | 32 | 706 MiB | | {"video_backend":"opencv","num_frames":-1} | {…,"num_frames":32} | 32 — fixed | 706 MiB | | {"video_backend":"qwen3_vl"} | {…,"num_frames":32} | 60 | 779 MiB | | {"video_backend":"qwen3_vl","max_frames":1e9,"fps":1e6} | {…,"max_frames":1000000000,"fps":1000000,"num_frames":32} | 900 — survives | 2 994 MiB |

#51969 does exactly what it claims for num_frames; the clamp writes num_frames: 32 into the merged dict and the Qwen sampler ignores it, while max_frames and fps pass through untouched.

The codebase already has the fix pattern, on other backends

This is not a new control being proposed. Commit 8b6de0eb9 — "[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion" (PR #54935, merged 2026-09-04, same author as #51969) — caps precisely these two knobs for GLMGAVideoBackend:

_MAX_FRAMES: ClassVar[int] = 640
_MAX_FPS: ClassVar[int] = 30
...
target_fps = min(target.fps, cls._MAX_FPS)
max_frames = min(kwargs.get("max_frames", cls._MAX_FRAMES), cls._MAX_FRAMES)

Its description states the root cause as "both target_fps and max_frames are controllable via request-level media_io_kwargs", and says it follows "the pattern established by GLM46VVideoBackend" (which caps via _MAX_FRAME_COUNT_DYNAMIC = 640 and _MAX_DURATION = 2400).

So two backends cap request-controlled fps/max_frames as an accepted security measure. Qwen2VLVideoBackend and Qwen3VLVideoBackend — the most widely deployed video models on vLLM — cap neither. In GLMGA the uncapped knobs sized an intermediate index list; in the Qwen samplers they size the decoded frame buffer, which is larger by the per-frame pixel count.

Reachability — default configuration, no authentication

  • media_io_kwargs is a request body field on ChatCompletionRequest and is carried by /v1/chat/completions, /v1/embeddings, /v1/responses, /tokenize and /invocations. No flag gates it.
  • api_key defaults to None (vllm/entrypoints/launchers/cli_args.py), and AuthenticationMiddleware is installed only when a key is configured — a default vllm serve is entirely unauthenticated.
  • Even with --api-key set, GUARDED_PREFIX = ("/v1", "/v2", "/inference", "/cohere") (vllm/entrypoints/serve/middleware/authenticate.py:11), so /invocations — which validates the same ChatCompletionRequest body — and /tokenize — which performs full media ingestion — remain unauthenticated.
  • MediaConnector.fetch_video applies the model's registered sampler only when the request did not name one (if "video_backend" not in video_io_kwargs), so video_backend: "qwen3_vl" is selectable on any deployment; request-level selection of stock sampler subclasses is already established as reachable by GHSA-j682-9xp5-rrf3. On a Qwen deployment no video_backend key is needed at all.
  • Default-on for any video-capable Qwen model.

The Rust frontend is not affected. rust/src/server/src/routes/openai/chat_completions/validate.rs:90 rejects media_io_kwargs with "media_io_kwargs is not supported." No second front is needed.

Impact

Unauthenticated remote denial of service by memory exhaustion of the API-server process. The decode runs in the frontend during chat parsing, before scheduling or admission control, so every tenant on the instance is affected. --limit-mm-per-prompt does not apply — it bounds media items, not frames within an item. CWE-770 / CWE-400.

Measured over unauthenticated HTTP: 1.43 MiB request body, 74 extra bytes of JSON, server peak RSS 2 271 → 13 629 MiB. In-process, the same request takes the decode from 60 to 900 frames. The frame count is bounded only by total_frames_num, which is the attacker's choice of video, and decoded bytes are frames × H × W × 3.

Context, measured on the num_frames path (GHSA-vxqj-p4gw-9h4c's path), not this one — these figures show what an unbounded frame count costs once the source video is chosen for it, and they transfer to this path because both converge on the same _read_frames_no_recovery allocation:

  • 5.74 MiB request body → 9.27 GiB decoded, 9.73 GiB of new resident memory (1 736×).
  • 117.6 MiB video/jpeg payload → process OOM-killed: Out of memory: Killed process 182667 (python) total-vm:38665704kB, anon-rss:21121616kB

The two were not re-run at the larger sizes through the Qwen sampler; the 900-frame figure above is what I measured on this path.

Suggested fix

Extend the ceiling to the sampler-side knobs rather than clamping the single num_frames key. Two options, either acceptable:

  1. Per-backend caps, matching 8b6de0eb9. Give Qwen2VLVideoBackend and Qwen3VLVideoBackend the _MAX_FRAMES / _MAX_FPS treatment already applied to GLMGAVideoBackend, so kwargs.get("max_frames", …) and target.fps are clamped to class constants. Smallest change; consistent with the accepted precedent. It leaves Glm5NextVideoBackend, Molmo2VideoBackend, NemotronVLVideoBackend, DynamicVideoBackend and OpenCVDynamicOpenPanguVideoBackend to be audited one by one, which is the current trajectory (#55727, #56207, #56390).
  2. Strip the frame-count knobs at the merge boundary. In VideoMediaIO.merge_kwargs, drop max_frames / min_frames / fps from runtime_kwargs the way hw_decoders, pool_size, device and unconfigured GPU backends are already dropped there. That treats the whole frame-count family as startup-only in one place and is robust to future sampler subclasses, at the cost of removing a request-level knob some users may rely on. A clamp-don't-strip variant (request may lower, never raise) preserves the feature.

Option 2 composes with #51969 and needs no per-backend audit; I would favour it, but option 1 is the more conservative change and matches what has already been merged.

Affected versions

>= 0.24.0, through v0.29.1rc0 and main @ b23433088.

Lower bound established by probing release tags through the GitHub contents API; no version below is inferred, each was read out of the file at that tag.

Confirmed present — Qwen samplers reading unclamped max_frames (max_frames = kwargs.get("max_frames", 768) inside Qwen2VLVideoBackend / Qwen3VLVideoBackend in vllm/multimodal/video.py):

| ref | Qwen3VLVideoBackend present | unclamped max_frames | |---|---|---| | v0.23.0 | no (class does not exist) | n/a | | v0.24.0 | yes | yes | | v0.25.0, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.29.1rc0 | yes | yes | | main @ b23433088 | yes | yes |

Confirmed present — request-level media_io_kwargs (field on the chat request model, and VideoMediaIO.merge_kwargs present): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.

Confirmed absent — any num_frames ceiling clamp (PR #51969 unmerged): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.

Not resolved, and why:

  1. The introducing commit/PR for Qwen2VLVideoBackend / Qwen3VLVideoBackend. My clone is shallow (--depth=300), so git log -S'max_frames' cannot reach it; the v0.23.0 → v0.24.0 boundary above is the tightest bound I established by probing release tags.
  2. The rc tags between v0.23.0 and v0.24.0, to tighten the lower bound to a specific release candidate.
  3. Whether a max_frames knob on a differently named pre-v0.24.0 backend is separately affected — at v0.23.0 GLMGAVideoBackend already carried max_frames = kwargs.get("max_frames", 640) with no clamp, and that clamp was only added on 2026-09-04 by 8b6de0eb9. Versions between are plausibly affected through GLMGA rather than Qwen; I did not test that path.
  4. Whether GHSA-vxqj-p4gw-9h4c's own affected range differs, which I cannot see — the advisory returns 404 to me.

No dependency versions are asserted anywhere in this report.

Weaknesses of this report, stated upfront

  • No GPU was used. This host has no CUDA device, so the HTTP results come from the in-tree GPU-less render server rather than a full vllm serve. That server runs the real frontend — the same request model, the same merge_media_io_kwargs → VideoMediaIO.merge_kwargs → MediaConnector.fetch_video chain, the same /tokenize route — and media decoding happens entirely in the frontend, so I do not believe the engine's presence changes the result. I have not confirmed that on a GPU deployment, and a reviewer may reasonably want that repeated under vllm serve.
  • The in-process measurements (the #51969 comparison table) were taken by importing the tree directly at b23433088, verified by __file__, with no vllm wheel installed.
  • Amplification figures are a floor. The OpenCV build available here offers only mp4v/XVID, giving ~2 200× compression on static content. An attacker using x264/x265 would do materially better for the same frame count.
  • The two large figures in the Impact section (9.73 GiB RSS; the OOM kill) were measured on the num_frames path, not this one, and are labelled as such.
  • Per-frame pixels remain bounded by VLLM_MAX_IMAGE_PIXELS; this concerns the unbounded frame count, which multiplies it.
  • VLLM_MAX_MEDIA_DOWNLOAD_SIZE_MB (default 256) caps the compressed size of an HTTP-fetched video but not a data: URI, which MediaConnector.load_from_url dispatches before any size logic is reached.

Prepared with AI assistance (Claude), per the repository's contributing guidance on disclosing AI-assisted contributions. All findings were verified by executing vLLM's own code at the commits cited; the analysis and the claims are my own.

Reported by Eva Crystal / 0xiviel (XSource Security).

🎯 Affected products1

  • pip/vllm:>= 0.24.0, < 0.30.0

🔗 References (5)