1 matches found
TrojanStego: Your Language Model Can Secretly Be a Steganographic Privacy Leaking Agent
As large language models LLMs become integrated into sensitive workflows, concerns grow over their potential to leak confidential information. We propose TrojanStego, a novel threat model in which an adversary fine-tunes an LLM to embed sensitive context information into natural-looking outputs v...
6.5AI score
SaveExploits0
20