Lucene search
+L

2 matches found

Packet Storm News
Packet Storm News
added 2025/06/19 12:0 a.m.8 views

Probing the Robustness of Large Language Models Safety to Latent Perturbations

Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models. We argue that this stems from the shallow nature of existing...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/18 12:0 a.m.8 views

Private Statistical Estimation Via Truncation

We introduce a novel framework for differentially private DP statistical estimation via data truncation, addressing a key challenge in DP estimation when the data support is unbounded. Traditional approaches rely on problem-specific sensitivity analysis, limiting their applicability. By leveragin...

6.9AI score
SaveExploits0
Rows per page
Query Builder