48 matches found
CacheTrap
CacheTrap Repository for "CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs" , IEEE/ACM International Conference on Computer-Aided Design ICCAD, 2026. arXiv CacheTrap searches for a single bit flip in the KV cache of a fine-tuned LLM classifier and measures the resulting attack succe...
CVE-2026-103765: Missing Authentication for Critical Function
Mooncake through 0.3.13.post1 contains a missing authentication vulnerability in the HTTP metadata server /metadata handler that allows unauthenticated attackers to read, overwrite, and delete transfer engine metadata keys. Attackers can poison segment descriptors such as tcpdataport or re-create...
CVE-2026-103765 Mooncake through 0.3.13.post1 Missing Authentication in HTTP Metadata Server
Mooncake through 0.3.13.post1 contains a missing authentication vulnerability in the HTTP metadata server /metadata handler that allows unauthenticated attackers to read, overwrite, and delete transfer engine metadata keys. Attackers can poison segment descriptors such as tcpdataport or re-create...
CVE-2026-103764 Mooncake transfer engine before 0.3.13 Unauthenticated Arbitrary Memory Read/Write via TCP Transport
Mooncake transfer engine before 0.3.13 contains an untrusted pointer dereference in ServerSession::readHeader that allows unauthenticated attackers to read and write arbitrary process memory via the TCP transport data port. Attackers can send a crafted SessionHeader with arbitrary addr and size...
PT-2026-104057
Mooncake transfer engine before 0.3.13 contains an untrusted pointer dereference in ServerSession::readHeader that allows unauthenticated attackers to read and write arbitrary process memory via the TCP transport data port. Attackers can send a crafted SessionHeader with arbitrary addr and size...
SUSE CVE-2026-94627
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
CVE-2026-94627
A flaw was found in vLLM. A remote attacker can exploit a vulnerability in the Mooncake connector's management of GPU Key-Value KV cache block ownership. By submitting completion requests with multiple prompts that share a single transfer ID, attackers can cause orphaned KV cache blocks to...
EUVD-2026-84221
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
Improper Handling of Exceptional Conditions
Overview vllm is an A high-throughput and memory-efficient inference and serving engine for LLMs Affected versions of this package are vulnerable to Improper Handling of Exceptional Conditions in mooncakeconnector.py via the Mooncake KV transfer connector, when a remote KV cache load fails during...
CVE-2026-94627 vLLM through 0.29.0 GPU KV Cache Leak via Mooncake Transfer ID Collision
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
CVE-2026-94627 vLLM through 0.29.0 GPU KV Cache Leak via Mooncake Transfer ID Collision
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
CVE-2026-94627 vLLM through 0.29.0 GPU KV Cache Leak via Mooncake Transfer ID Collision
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
PT-2026-96317
Name of the Vulnerable Software and Affected Versions vLLM versions prior to 0.29.1 Description The Mooncake connector fails to properly manage GPU KV cache block ownership during prefill/decode disaggregated deployments when concurrent child requests share a single transfer ID. An attacker can...
CVE-2026-94627: Missing Release of Memory after Effective Lifetime
vLLM Mooncake connector through 0.29.0 fails to properly manage GPU KV cache block ownership when concurrent child requests share a single transfer ID in prefill/decode disaggregated deployments. Attackers can trigger GPU memory exhaustion by submitting completion requests with multiple prompts,...
CVE-2026-69147 vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set mediaiokwargs.video.videobackend to pynvvideocodec, and MediaConnector.fetchvideo forwards that choice to VideoMediaIO even when startup configuration...
Characterizing Contention-Induced Reliability Collapse in KV-Cache Timing Side Channels for Multi-Tenant LLM Serving
Shared key--value KV cache reuse improves large language model LLM serving, but it can also create a timing side channel that reveals whether a prefix is already cached. Previous work shows that such attacks are possible, but their reliability under realistic multi-tenant contention is less...
llama-cpp-security-patches — Updated!
llama.cpp Security Patches Security patches for unpatched vulnerabilities in llama.cpp, discovered by Cyera Research. Background Between July 2025 and June 2026, we reported 10 vulnerabilities to the llama.cpp project through GitHub Security Advisories and MITRE. All advisories were closed by the...
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference
The key-value KV cache is the primary throughput optimization in modern large language model LLM inference, enabling prefix reuse across requests. In multi-tenant deployments this cache is shared across tenants, creating a timing side channel: an adversarial tenant can reconstruct another tenant'...
SUSE CVE-2026-43629
llama.cpp builds b4882 through b9058 contain a heap buffer overflow vulnerability in the KV cache state restore path where the statereaddata function computes write size without overflow checking, allowing attackers with write access to the slotsavepath directory to corrupt heap memory. Attackers...
EUVD-2026-54284
llama.cpp builds b4882 through b9058 contain a heap buffer overflow vulnerability in the KV cache state restore path where the statereaddata function computes write size without overflow checking, allowing attackers with write access to the slotsavepath directory to corrupt heap memory. Attackers...