Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/11/08 12:0 a.m.12 views

Injecting Falsehoods: Adversarial Man-In-The-Middle Attacks Undermining Factual Recall in LLMs

LLMs are now an integral part of information retrieval. As such, their role as question answering chatbots raises significant concerns due to their shown vulnerability to adversarial man-in-the-middle MitM attacks. Here, we propose the first principled attack evaluation on LLM factual memory unde...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/07/06 12:0 a.m.4 views

Emergent Misalignment As Prompt Sensitivity: a Research Note

Betley et al. 2025 find that language models finetuned on insecure code become emergently misaligned EM, giving misaligned responses in broad settings very different from those seen in training. However, it remains unclear as to why emergent misalignment occurs. We evaluate insecure models across...

7.1AI score
SaveExploits0
Rows per page
Query Builder