vllm
PyPI88 known CVEs affecting this package
Aggregated from OSV, GitHub Security Advisories, NVD, and vendor advisories. Each CVE links to its full detail page with vendor advisories, patches, fixed versions, and remediation guidance.
CVEs affecting vllmpage 2 of 2
- CVE-2026-44222MEDIUMCVSS 6.5EG 6.5fixed in 0.20.02026-05-12
vulnerable: 0.10.0 ... 0.9.2 (46 versions)
vLLM is an inference and serving engine for large language models (LLMs). From 0.6.1 to before 0.20.0, there is a a Token Injection vulnerability in vLLM’s multimodal processing. Unauthenticated, text-only prompts that spell special toke…
- CVE-2026-44223MEDIUMCVSS 6.5EG 6.5fixed in 0.20.02026-05-12
vulnerable: 0.18.0, 0.18.1, 0.19.0, 0.19.1
vLLM is an inference and serving engine for large language models (LLMs). From 0.18.0 to before 0.20.0, the extract_hidden_states speculative decoding proposer in vLLM returns a tensor with an incorrect shape after the first decode step, c…
- CVE-2026-47155MEDIUMCVSS 6.5EG 6.5fixed in 0.22.02026-06-10
vulnerable: 0.0.1 ... 0.9.2 (86 versions)
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.0, vLLM's revision pinning controls do not consistently apply to all artifacts loaded for a model. A deployment that supplies --revision or --code-revi…
- CVE-2026-48746CRITICALCVSS 9.1EG 9.1fixed in 0.22.02026-06-16
vulnerable: 0.10.0 ... 0.9.2 (68 versions)
vLLM is an inference and serving engine for large language models (LLMs). From 0.3.0 until 0.22.0, a vulnerability in ASGI web servers and starlette's trust on those web servers enables an authentication bypass of the OpenAI API Authentica…
- CVE-2026-53923HIGHCVSS 7.5EG 7.5fixed in 0.24.02026-06-17
vulnerable: 0.10.0 ... 0.9.2 (55 versions)
vLLM is an inference and serving engine for large language models (LLMs). From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels (csrc/quantization/gguf/gguf_kernel.cu) causes partial tensor p…
- CVE-2026-54232HIGHCVSS 8.8EG 8.8fixed in 0.22.12026-06-22
vulnerable: 0.0.1 ... 0.9.2 (87 versions)
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.22.1, the vLLM Dockerfile is vulnerable to a dependency confusion attack through the flashinfer-jit-cache package. The package is installed from a custom …
- CVE-2026-54233MEDIUMCVSS 6.5EG 6.5fixed in 0.24.02026-06-17
vulnerable: 0.0.1 ... 0.9.2 (89 versions)
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.23.1rc0, vLLM's /v1/audio/transcriptions endpoint limits compressed upload size but not decoded PCM output. A 25MB OPUS file expands to ~14.9GB of float32…
- CVE-2026-54234HIGHCVSS 7.5EG 7.5fixed in 0.24.02026-07-06
vulnerable: 0.17.1 ... 0.23.0 (12 versions)
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, a frontend-legal multi-request speculative decoding workload can cause the rejection sampler to produce a recovered token equal to the m…
- CVE-2026-54235MEDIUMCVSS 6.5EG 6.5fixed in 0.24.02026-06-17
vulnerable: 0.0.1 ... 0.9.2 (89 versions)
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.23.1rc0, ll temperature validation gates use comparison operators (<, >), which silently evaluate to False for NaN and for positive Infinity in Python's I…
- CVE-2026-54236MEDIUMCVSS 5.3EG 5.3fixed in 0.24.02026-06-17
vulnerable: 0.0.1 ... 0.9.2 (89 versions)
vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.23.1rc0, the fix for CVE-2026-22778, which introduced a sanitize_message helper that strips object-repr memory addresses from error messages before they r…
- CVE-2026-5497HIGHCVSS 7.5EG 7.5fixed in 0.19.02026-06-11
vulnerable: 0.10.0 ... 0.9.2 (29 versions)
vLLM versions 0.8.0 and later are vulnerable to an Out-of-Memory (OOM) Denial of Service (DoS) attack due to unbounded frame count processing in the `VideoMediaIO.load_base64()` method. When processing `video/jpeg` data URLs, the method sp…
- CVE-2026-55514MEDIUMCVSS 6.5EG 6.5fixed in 0.24.02026-07-06
vulnerable: 0.12.0 ... 0.23.0 (20 versions)
vLLM is a library for LLM inference and serving. From 0.12.0 to before 0.24.0, sending a pure prompt embeds payload in a /v1/completions request with a model using M-RoPE causes EngineCore to fail an assertion and fatally crash, shutting d…
- CVE-2026-55574HIGHCVSS 7.5EG 7.5fixed in 0.24.02026-07-06
vulnerable: 0.0.1 ... 0.9.2 (89 versions)
vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Prior to 0.24.0, the structured_outputs.regex API parameter passes a user-supplied regular expression string directly to the grammar compiler backends wi…
- CVE-2026-55646MEDIUMCVSS 6.5EG 6.5fixed in 0.24.02026-07-06
vulnerable: 0.22.0, 0.22.1, 0.23.0
vLLM is an inference and serving engine for large language models. From 0.22.0 to 0.23.0, the /v1/audio/transcriptions and /v1/audio/translations routes call request.file.read() to fully materialize an uploaded audio file into memory befor…
- CVE-2026-56340HIGHCVSS 7.5EG 8.8fixed in 0.13.02026-06-20
vulnerable: 0.10.2, 0.11.0, 0.11.1, 0.11.2, 0.12.0
vLLM versions >= 0.10.2 and < 0.13.0 are missing sparse tensor validation in multimodal embeddings processing. Because PyTorch disables sparse tensor invariant checks by default, an attacker can submit crafted embedding requests with malfo…
- CVE-2026-57173MEDIUMCVSS 6.5EG 6.5fixed in 0.24.02026-09-16
vulnerable: 0.0.1 ... 0.9.2 (89 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATI…
- CVE-2026-69147MEDIUMCVSS 6.5EG 6.5fixed in 0.28.02026-09-16
vulnerable: 0.0.1 ... 0.9.2 (95 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards th…
- CVE-2026-7141MEDIUMCVSS 5.6EG 5.6fixed in 0.19.12026-04-27
vulnerable: 0.0.1 ... 0.9.2 (81 versions)
A vulnerability was found in vLLM up to 0.19.0. The affected element is the function has_mamba_layers of the file vllm/v1/kv_cache_interface.py of the component KV Block Handler. Performing a manipulation results in uninitialized resource.…
- CVE-2026-71486MEDIUMCVSS 4.3EG 4.3fixed in 0.26.02026-08-17
vulnerable: 0.0.1 ... 0.9.2 (92 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices,…
- CVE-2026-73555MEDIUMCVSS 5.3EG 5.3fixed in 0.26.02026-08-13
vulnerable: 0.0.1 ... 0.9.2 (92 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the validation_exception_handler in vllm/entrypoints/openai/server_utils.py converts FastAPI RequestValidationError objects with str(exc), and sanitize_mes…
- CVE-2026-73556MEDIUMCVSS 5.3EG 5.3fixed in 0.26.02026-08-13
vulnerable: 0.0.1 ... 0.9.2 (92 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py is passed to lmformatenforcer.RegexParser without compile…
- CVE-2026-73557MEDIUMCVSS 6.3EG 6.3fixed in 0.26.02026-08-13
vulnerable: 0.21.0 ... 0.25.1 (7 versions)
vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose process-global save, enable, a…
- CVE-2026-73558MEDIUMCVSS 5.3EG 5.3fixed in 0.27.02026-08-13
vulnerable: 0.0.1 ... 0.9.2 (93 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input, allowing a request…
- CVE-2026-73559MEDIUMCVSS 6.5EG 6.5fixed in 0.26.02026-08-13
vulnerable: 0.19.0 ... 0.25.1 (12 versions)
vLLM is an inference and serving engine for large language models. From 0.19.0 until 0.26.0, the /v1/completions CompletionRequest.prompt field in vllm/entrypoints/openai/completion/protocol.py accepts an unbounded list[str] or list[list[i…
- CVE-2026-73560MEDIUMCVSS 6.5EG 6.5fixed in 0.26.02026-08-17
vulnerable: 0.0.1 ... 0.9.2 (92 versions)
vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the MiMoV2OmniMultiModalProcessor in vllm/transformers_utils/processors/mimo_v2_omni.py passes attacker-controlled image and audio strings through _fetch_i…
- CVE-2026-78684MEDIUMCVSS 5.3EG 5.3fixed in 0.27.02026-08-25
vulnerable: 0.0.1 ... 0.9.2 (93 versions)
vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the process-wide GPU decode poo…
- CVE-2026-90553HIGHCVSS 7.8EG 7.8fixed in 0.28.02026-09-12
vulnerable: 0.0.1 ... 0.9.2 (95 versions)
vLLM before 0.28.0 contains a remote code execution vulnerability in the LlavaOnevision2 processor loader that ignores the trust_remote_code parameter when loading remote processor classes. Attackers can craft a malicious model with arbitr…
- CVE-2026-93436HIGHCVSS 7.5EG 7.5fixed in 0.30.02026-09-17
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without…
- CVE-2026-93592HIGHCVSS 7.5EG 7.5fixed in 0.28.02026-09-18
vulnerable: 0.0.1 ... 0.9.2 (95 versions)
vLLM versions before 0.28.0 fail to validate the lower bound of token IDs in the /v1/embeddings and /pooling endpoints, allowing unauthenticated attackers to crash the engine by submitting negative token IDs. A single request with a negati…
- CVE-2026-93840MEDIUMCVSS 5.3EG 5.3fixed in 0.29.02026-09-18
vulnerable: 0.0.1 ... 0.9.2 (96 versions)
vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, …
- CVE-2026-93841MEDIUMCVSS 5.3EG 5.3fixed in 0.30.02026-09-18
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal …
- CVE-2026-93989MEDIUMCVSS 4.3EG 4.3fixed in 0.30.02026-09-19
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of co…
- CVE-2026-94622HIGHCVSS 7.5EG 7.5fixed in 0.30.02026-09-21
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM versions through 0.29.0 contain a denial of service vulnerability in the NIXL connector's metadata handling for prefill/decode disaggregated deployments. Attackers can send requests with incomplete kv_transfer_params dictionary entrie…
- CVE-2026-94623HIGHCVSS 7.5EG 7.5fixed in 0.30.02026-09-21
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 contains a denial of service vulnerability in the NIXL connector's prefix caching implementation that fails to properly validate block counts across multi-prompt completion requests in prefill/decode disaggregated deplo…
- CVE-2026-94624HIGHCVSS 7.5EG 7.5fixed in 0.30.02026-09-21
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 contains a denial of service vulnerability in P2P KV offloading when OffloadingConnector is configured with TieringOffloadingSpec and a peer-to-peer secondary tier. Attackers can supply arbitrary remote host and port va…
- CVE-2026-94625MEDIUMCVSS 5.3EG 5.3fixed in 0.30.02026-09-21
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 contains a resource exhaustion vulnerability in MooncakeConnector where rejected prefill requests create ownerless transfer placeholders that are never reclaimed. Attackers can send rejected requests to exhaust sender t…
- CVE-2026-94626HIGHCVSS 7.5EG 7.5fixed in 0.30.02026-09-21
vulnerable: 0.0.1 ... 0.9.2 (97 versions)
vLLM through 0.29.0 fails to validate the tp_size parameter in kv_transfer_params on OpenAI-compatible completion endpoints, allowing attackers to allocate unbounded memory. Attackers can supply arbitrary tp_size values in prefill/decode d…
- CVE-2026-9540MEDIUMCVSS 5.3EG 5.32026-05-26
vulnerable: 0.0.1 ... 0.9.2 (81 versions)
A vulnerability was identified in vllm-project vllm 0.19.0. This issue affects some unknown processing of the component OpenAI-compatible Serving Path. Such manipulation leads to denial of service. It is possible to launch the attack remot…
Check whether vllm is used in your infrastructure
EchelonGraph scans your cloud and SBOMs to map every package to your actual deployments. See blast radius for vllm CVEs against the assets you own.
Book a Demo →