Lucene search
+L

5 matches found

NVD
NVD
added 2026/08/13 3:20 p.m.7 views

CVE-2026-73558

vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x 2 d in activationkernels.cu can cause actandmulkernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or...

5.3CVSS0.00261EPSS
SaveExploits0References5
ATTACKERKB
ATTACKERKB
added 2026/08/13 3:06 p.m.5 views

CVE-2026-73558

vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x 2 d in activationkernels.cu can cause actandmulkernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or...

5.3CVSS5.3AI score0.00261EPSS
SaveExploits0References6Affected Software1
Cvelist
Cvelist
added 2026/08/13 3:06 p.m.28 views

CVE-2026-73558 vLLM: Cross-User Data Leak Vulnerability

vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x 2 d in activationkernels.cu can cause actandmulkernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or...

5.3CVSS0.00261EPSS
SaveExploits0References5
OSV
OSV
added 2026/08/13 3:06 p.m.11 views

CVE-2026-73558 vLLM: Cross-User Data Leak Vulnerability

vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x 2 d in activationkernels.cu can cause actandmulkernel to consume another batched user's input, allowing a request processed in the same inference batch to receive a partial or...

5.3CVSS5.5AI score
SaveExploits0References7
Github Security Blog
Github Security Blog
added 2026/06/17 2:03 p.m.35 views

vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving

Summary Integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels csrc/quantization/gguf/ggufkernel.cu causes partial tensor processing. The output tensor is allocated at full size via torch::empty uninitialized memory, but the dequantize CUDA kernel processes only a truncated...

7.5CVSS5.6AI score0.00281EPSS
SaveExploits0References8Affected Software1
Rows per page
Query Builder