Lucene search
+L

4 matches found

Packet Storm News
Packet Storm News
•added 2026/08/10 12:00 a.m.•15 views

Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference

The key-value KV cache is the primary throughput optimization in modern large language model LLM inference, enabling prefix reuse across requests. In multi-tenant deployments this cache is shared across tenants, creating a timing side channel: an adversarial tenant can reconstruct another tenant'...

5.2AI score
SaveExploits0
RedhatCVE
RedhatCVE
•added 2026/06/29 4:39 a.m.•32 views

CVE-2026-53923

A flaw was found in vLLM. Integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels leads to partial tensor processing. This results in the output tensor retaining previously used GPU memory, which, in multi-tenant inference deployments, can expose sensitive tensor data from other...

7.5CVSS5.7AI score0.00484EPSS
SaveExploits0References6
NVD
NVD
•added 2026/06/22 11:16 p.m.•25 views

CVE-2026-53923

vLLM is an inference and serving engine for large language models LLMs. From 0.5.5 until 0.23.1rc0, integer truncation of tensor dimensions in vLLM's GGUF dequantize kernels csrc/quantization/gguf/ggufkernel.cu causes partial tensor processing. The output tensor is allocated at full size via...

7.5CVSS0.00484EPSS
SaveExploits0References3
Positive Technologies
Positive Technologies
•added 2026/06/17 12:00 a.m.•43 views

PT-2026-50472

Name of the Vulnerable Software and Affected Versions vLLM versions 0.5.5 through 0.23.1rc0 Description Integer truncation of tensor dimensions in GGUF dequantize kernels within csrc/quantization/gguf/gguf kernel.cu leads to partial tensor processing. The output tensor is allocated at full size...

7.5CVSS5.8AI score0.00484EPSS
SaveExploits0References15
Rows per page
Query Builder