2 matches found
Trust Me, I'M Your Developer: Self-Issued Authentication in Large Language Models
Large language model LLM security has largely focused on role-playing jailbreaks, with less attention to what happens when a user asks an LLM to verify an identity claim through a test designed by the model itself. We study this behavior through a staged developer-identity experiment with ChatGPT...
5.5AI score
SaveExploits0
Don't blab
If you’re worried that your conversations are being monitored, old fashioned “coded talk” also works to disguise the meaning of your conversations. Asking “how are the kids?” rather than “how’s the progress on our new chip design?” may be enough to throw attackers off the scent. Images via...
4.5AI score
SaveExploits0References1
20