4 matches found
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
SummaryCurrent vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Summary Current vLLM main lets an inference request choose the PyNvVideoCodec GPU video decoder through mediaiokwargs.video.videobackend, but engine GPU memory reservation is computed only from static startup configuration and VLLMVIDEOLOADERBACKEND. If the server starts with the default...
CVE-2026-69147
vLLM (inference/serving engine for LLMs) prior to 0.28.0 allows a request-selected pynvvideocodec backend to bypass the engine's static VRAM reservation. Specifically, MediaConnector.fetch_video forwards a request-level video_backend override to VideoMediaIO even when startup configuration select...
CVE-2026-69147 vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...