Lucene search
+L

2 matches found

PyPA
PyPA
added 2025/05/29 5:15 p.m.18 views

PYSEC-2025-53

vLLM is an inference and serving engine for large language models LLMs. Prior to version 0.9.0, when a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT Time to First Token. These timing differences...

2.6CVSS6.8AI score0.00257EPSS
SaveExploits0References4Affected Software1
Snyk
Snyk
added 2025/05/29 4:43 p.m.3 views

Timing Attack

Overview vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs Affected versions of this package are vulnerable to Timing Attack due to the PageAttention mechanism. An attacker can observe timing differences to infer details about the processed data by analyzing...

6.3CVSS6.9AI score0.00257EPSS
SaveExploits0References2
Rows per page
Query Builder