1 matches found
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems Via Deterministic Side Channels
Modern large language models LLMs exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens. Researchers have leveraged this property to optimize LLM serving systems by omitting weight accesses and computations pertaining to inactive neurons...
5.5AI score
SaveExploits0
20