10 matches found
EUVD-2026-92848
vLLM: Qwen2-VL / Qwen3-VL video samplers bound on request-controlled maxframes, which the numframes ceiling does not reach...
EUVD-2026-92851
vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion...
CVE-2026-105760
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level mediaiokwargs field to select the GLMGA video backend and supply large values for the fps and maxframes options without a strict work ceiling. GLMGA constructs and deduplicates a...
CVE-2026-105758
vLLM is an inference and serving engine for large language models. From 0.24.0 until 0.30.0, the Qwen2VLVideoBackend and Qwen3VLVideoBackend classes accept request-level values for the mediaiokwargs.video.maxframes and mediaiokwargs.video.fps fields without enforcing server-side ceilings. An...
CVE-2026-105760 vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level mediaiokwargs field to select the GLMGA video backend and supply large values for the fps and maxframes options without a strict work ceiling. GLMGA constructs and deduplicates a...
CVE-2026-105760 vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level mediaiokwargs field to select the GLMGA video backend and supply large values for the fps and maxframes options without a strict work ceiling. GLMGA constructs and deduplicates a...
CVE-2026-105760 vLLM: GLMGA video sampling permits request-driven CPU and memory exhaustion
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, a caller can use the request-level mediaiokwargs field to select the GLMGA video backend and supply large values for the fps and maxframes options without a strict work ceiling. GLMGA constructs and deduplicates a...
CVE-2026-105760
The vLLM inference and serving engine for large language models is vulnerable to CPU and memory exhaustion in versions prior to 0.30.0 . An attacker can use the media_io_kwargs field to select the GLMGA video backend and provide large values for the fps and max_frames options. Because there is no...
GHSA-8PW2-6JV3-MJ5J vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Summary Current vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
Allocation of Resources Without Limits or Throttling
Overview vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs Affected versions of this package are vulnerable to Allocation of Resources Without Limits or Throttling via VideoMediaIO.mergekwargs in vllm/multimodal/media/video.py, where a request-level...