Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/11/21 12:0 a.m.8 views

Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models

Modern large language models LLMs are typically secured by auditing data, prompts, and refusal policies, while treating the forward pass as an implementation detail. We show that intermediate activations in decoder-only LLMs form a vulnerable attack surface for behavioral control. Building on...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/19 12:0 a.m.10 views

Probing the Robustness of Large Language Models Safety to Latent Perturbations

Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that this stems from the shallow nature of existing...

6.9AI score
SaveExploits0
Rows per page
Query Builder