1 matches found
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference
The key-value KV cache is the primary throughput optimization in modern large language model LLM inference, enabling prefix reuse across requests. In multi-tenant deployments this cache is shared across tenants, creating a timing side channel: an adversarial tenant can reconstruct another tenant'...
5.2AI score
SaveExploits0
20