19 matches found
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
SummaryCurrent vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
CVE-2026-100649
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits an...
EUVD-2026-87743
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits an...
CVE-2026-100649 vLLM before 0.29.0 Resource Limit Bypass via Sampler Subclass
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits an...
CVE-2026-100649
Versions of vLLM prior to 0.29.0 are vulnerable to a resource-limit bypass within the PyNvVideoCodec decoder allocation. The root cause is sampler subclass shadowing , which allows for independent counter increments rather than global tracking. An unauthenticated remote attacker can exploit this ...
CVE-2026-100649 vLLM before 0.29.0 Resource Limit Bypass via Sampler Subclass
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits an...
CVE-2026-100649: Allocation of Resources Without Limits or Throttling
vLLM before 0.29.0 contains a resource-limit bypass vulnerability in PyNvVideoCodec decoder allocation where sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select different sampler subclasses in video requests to exceed configured decoder limits an...
PT-2026-99320
Name of the Vulnerable Software and Affected Versions vLLM versions prior to 0.29.0 Description A resource-limit bypass exists in the PyNvVideoCodec decoder allocation. This issue occurs because sampler subclass shadowing allows independent counter increments. Unauthenticated attackers can select...
EUVD-2026-80935
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation...
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Summary Current vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
GHSA-8PW2-6JV3-MJ5J vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Summary Current vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
CVE-2026-69147
A flaw was found in vLLM, an inference and serving engine for large language models. An attacker can exploit this by submitting specially crafted video requests that force the use of the PyNvVideoCodec GPU decoder. This bypasses the engine's static GPU memory reservation, allowing the attacker to...
CVE-2026-69147
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...
CVE-2026-69147 vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...
CVE-2026-69147 vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...
CVE-2026-69147
vLLM (inference/serving engine for LLMs) prior to 0.28.0 allows a request-selected pynvvideocodec backend to bypass the engine's static VRAM reservation. Specifically, MediaConnector.fetch_video forwards a request-level video_backend override to VideoMediaIO even when startup configuration select...
CVE-2026-69147 vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...
PT-2026-93891
Name of the Vulnerable Software and Affected Versions vLLM versions prior to 0.28.0 Description An issue exists where the engine fails to properly budget GPU memory when a user specifies a GPU decoder at request time. Specifically, request bodies for Chat Completions and Responses can set media i...
CVE-2026-69147: Uncontrolled Resource Consumption
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...