vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled max_frames, which the num_frames ceiling does not reach
🔗 CVE IDs covered (1)
📋 Description
Summary
An unauthenticated remote attacker can exhaust the memory of the vLLM API-server process by raising the request-level media_io_kwargs.video.max_frames and fps knobs on any deployment serving a Qwen2-VL or Qwen3-VL model. 74 extra bytes of JSON took the server's peak RSS from 2 271 MiB to 13 629 MiB over unauthenticated POST /tokenize.
The num_frames ceiling reported in GHSA-vxqj-p4gw-9h4c and fixed by open PR #51969 does not reach this path: the Qwen samplers do not read num_frames at all. The same knobs were already capped upstream for GLMGAVideoBackend as an accepted security fix in 8b6de0eb9 (PR #54935, merged 2026-09-04); that cap never reached Qwen.
Details
Relationship to GHSA-vxqj-p4gw-9h4c and PR #51969 (read this first)
GHSA-vxqj-p4gw-9h4c reported that request-level media_io_kwargs.video.num_frames overrides the engine frame-count ceiling, and open PR #51969 fixes it by clamping num_frames inside VideoMediaIO.merge_kwargs.
That clamp does not reach the Qwen samplers. Qwen2VLVideoBackend and Qwen3VLVideoBackend do not read num_frames at all — Qwen2VLVideoBackend's own docstring says so ("num_frames is ignored (fps-driven, like the Qwen3-VL loader)"). They bound on max_frames, read from the same merged dict with no ceiling:
# vllm/multimodal/video.py — Qwen3VLVideoBackend.compute_frames_index_to_sample
min_frames = kwargs.get("min_frames", 4)
max_frames = kwargs.get("max_frames", 768)
num_frames = int(total_frames_num / original_fps * fps)
num_frames = min(max(num_frames, min_frames), max_frames, total_frames_num)
With max_frames raised from the request, the only remaining bound is total_frames_num — every frame in the container.
I applied PR #51969's patch locally and re-ran both paths through the real merge layer (merge_media_io_kwargs → VideoMediaIO.merge_kwargs → MediaConnector), against main @ b23433088:
| request media_io_kwargs.video | merged kwargs after #51969 | frames decoded | peak RSS |
|---|---|---|---|
| (absent) | None | 32 | 706 MiB |
| {"video_backend":"opencv","num_frames":-1} | {…,"num_frames":32} | 32 — fixed | 706 MiB |
| {"video_backend":"qwen3_vl"} | {…,"num_frames":32} | 60 | 779 MiB |
| {"video_backend":"qwen3_vl","max_frames":1e9,"fps":1e6} | {…,"max_frames":1000000000,"fps":1000000,"num_frames":32} | 900 — survives | 2 994 MiB |
#51969 does exactly what it claims for num_frames; the clamp writes num_frames: 32 into the merged dict and the Qwen sampler ignores it, while max_frames and fps pass through untouched.
The codebase already has the fix pattern, on other backends
This is not a new control being proposed. Commit 8b6de0eb9 — "[Security] Cap GLMGA video sampling to prevent request-driven resource exhaustion" (PR #54935, merged 2026-09-04, same author as #51969) — caps precisely these two knobs for GLMGAVideoBackend:
_MAX_FRAMES: ClassVar[int] = 640
_MAX_FPS: ClassVar[int] = 30
...
target_fps = min(target.fps, cls._MAX_FPS)
max_frames = min(kwargs.get("max_frames", cls._MAX_FRAMES), cls._MAX_FRAMES)
Its description states the root cause as "both target_fps and max_frames are controllable via request-level media_io_kwargs", and says it follows "the pattern established by GLM46VVideoBackend" (which caps via _MAX_FRAME_COUNT_DYNAMIC = 640 and _MAX_DURATION = 2400).
So two backends cap request-controlled fps/max_frames as an accepted security measure. Qwen2VLVideoBackend and Qwen3VLVideoBackend — the most widely deployed video models on vLLM — cap neither. In GLMGA the uncapped knobs sized an intermediate index list; in the Qwen samplers they size the decoded frame buffer, which is larger by the per-frame pixel count.
Reachability — default configuration, no authentication
media_io_kwargsis a request body field onChatCompletionRequestand is carried by/v1/chat/completions,/v1/embeddings,/v1/responses,/tokenizeand/invocations. No flag gates it.api_keydefaults toNone(vllm/entrypoints/launchers/cli_args.py), andAuthenticationMiddlewareis installed only when a key is configured — a defaultvllm serveis entirely unauthenticated.- Even with
--api-keyset,GUARDED_PREFIX = ("/v1", "/v2", "/inference", "/cohere")(vllm/entrypoints/serve/middleware/authenticate.py:11), so/invocations— which validates the sameChatCompletionRequestbody — and/tokenize— which performs full media ingestion — remain unauthenticated. MediaConnector.fetch_videoapplies the model's registered sampler only when the request did not name one (if "video_backend" not in video_io_kwargs), sovideo_backend: "qwen3_vl"is selectable on any deployment; request-level selection of stock sampler subclasses is already established as reachable by GHSA-j682-9xp5-rrf3. On a Qwen deployment novideo_backendkey is needed at all.- Default-on for any video-capable Qwen model.
The Rust frontend is not affected. rust/src/server/src/routes/openai/chat_completions/validate.rs:90 rejects media_io_kwargs with "media_io_kwargs is not supported." No second front is needed.
Impact
Unauthenticated remote denial of service by memory exhaustion of the API-server process. The decode runs in the frontend during chat parsing, before scheduling or admission control, so every tenant on the instance is affected. --limit-mm-per-prompt does not apply — it bounds media items, not frames within an item. CWE-770 / CWE-400.
Measured over unauthenticated HTTP: 1.43 MiB request body, 74 extra bytes of JSON, server peak RSS 2 271 → 13 629 MiB. In-process, the same request takes the decode from 60 to 900 frames. The frame count is bounded only by total_frames_num, which is the attacker's choice of video, and decoded bytes are frames × H × W × 3.
Context, measured on the num_frames path (GHSA-vxqj-p4gw-9h4c's path), not this one — these figures show what an unbounded frame count costs once the source video is chosen for it, and they transfer to this path because both converge on the same _read_frames_no_recovery allocation:
- 5.74 MiB request body → 9.27 GiB decoded, 9.73 GiB of new resident memory (1 736×).
- 117.6 MiB
video/jpegpayload → process OOM-killed:Out of memory: Killed process 182667 (python) total-vm:38665704kB, anon-rss:21121616kB
The two were not re-run at the larger sizes through the Qwen sampler; the 900-frame figure above is what I measured on this path.
Suggested fix
Extend the ceiling to the sampler-side knobs rather than clamping the single num_frames key. Two options, either acceptable:
- Per-backend caps, matching
8b6de0eb9. GiveQwen2VLVideoBackendandQwen3VLVideoBackendthe_MAX_FRAMES/_MAX_FPStreatment already applied toGLMGAVideoBackend, sokwargs.get("max_frames", …)andtarget.fpsare clamped to class constants. Smallest change; consistent with the accepted precedent. It leavesGlm5NextVideoBackend,Molmo2VideoBackend,NemotronVLVideoBackend,DynamicVideoBackendandOpenCVDynamicOpenPanguVideoBackendto be audited one by one, which is the current trajectory (#55727, #56207, #56390). - Strip the frame-count knobs at the merge boundary. In
VideoMediaIO.merge_kwargs, dropmax_frames/min_frames/fpsfromruntime_kwargsthe wayhw_decoders,pool_size,deviceand unconfigured GPU backends are already dropped there. That treats the whole frame-count family as startup-only in one place and is robust to future sampler subclasses, at the cost of removing a request-level knob some users may rely on. A clamp-don't-strip variant (request may lower, never raise) preserves the feature.
Option 2 composes with #51969 and needs no per-backend audit; I would favour it, but option 1 is the more conservative change and matches what has already been merged.
Affected versions
>= 0.24.0, through v0.29.1rc0 and main @ b23433088.
Lower bound established by probing release tags through the GitHub contents API; no version below is inferred, each was read out of the file at that tag.
Confirmed present — Qwen samplers reading unclamped max_frames (max_frames = kwargs.get("max_frames", 768) inside Qwen2VLVideoBackend / Qwen3VLVideoBackend in vllm/multimodal/video.py):
| ref | Qwen3VLVideoBackend present | unclamped max_frames |
|---|---|---|
| v0.23.0 | no (class does not exist) | n/a |
| v0.24.0 | yes | yes |
| v0.25.0, v0.26.0, v0.27.0, v0.28.0, v0.29.0, v0.29.1rc0 | yes | yes |
| main @ b23433088 | yes | yes |
Confirmed present — request-level media_io_kwargs (field on the chat request model, and VideoMediaIO.merge_kwargs present): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.
Confirmed absent — any num_frames ceiling clamp (PR #51969 unmerged): every ref probed, v0.19.0 through v0.29.1rc0 and main @ b23433088.
Not resolved, and why:
- The introducing commit/PR for
Qwen2VLVideoBackend/Qwen3VLVideoBackend. My clone is shallow (--depth=300), sogit log -S'max_frames'cannot reach it; the v0.23.0 → v0.24.0 boundary above is the tightest bound I established by probing release tags. - The rc tags between v0.23.0 and v0.24.0, to tighten the lower bound to a specific release candidate.
- Whether a
max_framesknob on a differently named pre-v0.24.0 backend is separately affected — at v0.23.0GLMGAVideoBackendalready carriedmax_frames = kwargs.get("max_frames", 640)with no clamp, and that clamp was only added on 2026-09-04 by8b6de0eb9. Versions between are plausibly affected through GLMGA rather than Qwen; I did not test that path. - Whether GHSA-vxqj-p4gw-9h4c's own affected range differs, which I cannot see — the advisory returns 404 to me.
No dependency versions are asserted anywhere in this report.
Weaknesses of this report, stated upfront
- No GPU was used. This host has no CUDA device, so the HTTP results come from the in-tree GPU-less render server rather than a full
vllm serve. That server runs the real frontend — the same request model, the samemerge_media_io_kwargs→VideoMediaIO.merge_kwargs→MediaConnector.fetch_videochain, the same/tokenizeroute — and media decoding happens entirely in the frontend, so I do not believe the engine's presence changes the result. I have not confirmed that on a GPU deployment, and a reviewer may reasonably want that repeated undervllm serve. - The in-process measurements (the #51969 comparison table) were taken by importing the tree directly at
b23433088, verified by__file__, with novllmwheel installed. - Amplification figures are a floor. The OpenCV build available here offers only
mp4v/XVID, giving ~2 200× compression on static content. An attacker using x264/x265 would do materially better for the same frame count. - The two large figures in the Impact section (9.73 GiB RSS; the OOM kill) were measured on the
num_framespath, not this one, and are labelled as such. - Per-frame pixels remain bounded by
VLLM_MAX_IMAGE_PIXELS; this concerns the unbounded frame count, which multiplies it. VLLM_MAX_MEDIA_DOWNLOAD_SIZE_MB(default 256) caps the compressed size of an HTTP-fetched video but not adata:URI, whichMediaConnector.load_from_urldispatches before any size logic is reached.
Prepared with AI assistance (Claude), per the repository's contributing guidance on disclosing AI-assisted contributions. All findings were verified by executing vLLM's own code at the commits cited; the analysis and the claims are my own.
Reported by Eva Crystal / 0xiviel (XSource Security).
🎯 Affected products1
- pip/vllm:>= 0.24.0, < 0.30.0
🔗 References (5)
- https://github.com/vllm-project/vllm/security/advisories/GHSA-x6mc-67gf-chw4
- https://github.com/vllm-project/vllm/pull/56729
- https://github.com/vllm-project/vllm/commit/ea723c81c3ea26425cb69503a5d5e90822a04a45
- https://github.com/vllm-project/vllm/releases/tag/v0.30.0
- https://github.com/advisories/GHSA-x6mc-67gf-chw4