Lucene search
+L

6 matches found

Kitploit
Kitploit
added 2026/08/27 3:33 a.m.4 views

IF-Guide

IF-Guide: 영향 함수 기반 LLM 탈독화 NeurIPS 2025 이 저장소는 우리 논문에서 소개한 LLM 탈독화 기법인 IF-Guide의 코드를 포함합니다: IF-Guide: 영향 함수 기반 LLM 탈독화 Zachary Coalson , Juhan Bae, Nicholas Carlini, Sanghyun Hong TL;DR 우리의 방법을 사용하면 유해한 학습 예제를 식별하고 사전 학습 또는 미세 조정 중에 이를 억제하여 LLM을 탈독화할 수 있습니다! 초록 우리는 학습 데이터가 대규모 언어 모델에서 유해한 행동의 출현에...

5.4AI score
SaveExploits0References4
Packet Storm News
Packet Storm News
added 2025/12/10 12:00 a.m.20 views

Chasing Shadows: Pitfalls in LLM Security Research

Large language models LLMs are increasingly prevalent in security research. Their unique characteristics, however, introduce challenges that undermine established paradigms of reproducibility, rigor, and evaluation. Prior work has identified common pitfalls in traditional machine learning researc...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/09/25 12:00 a.m.11 views

Vision Transformers: the Threat of Realistic Adversarial Patches

The increasing reliance on machine learning systems has made their security a critical concern. Evasion attacks enable adversaries to manipulate the decision-making processes of AI systems, potentially causing security breaches or misclassification of targets. Vision Transformers ViTs have gained...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/09 12:00 a.m.9 views

IF-GUIDE: Influence Function-Guided Detoxification of LLMs

We study how training data contributes to the emergence of toxic behaviors in large-language models. Most prior work on reducing model toxicity adopts $reactive$ approaches, such as fine-tuning pre-trained and potentially toxic models to align them with human values. In contrast, we propose a...

7.1AI score
SaveExploits0
FireEye
FireEye
added 2021/01/21 12:00 a.m.66 views

Training Transformers for Cyber Security Tasks: A Case Study on Malicious URL Prediction

Highlights Perform a case study on using Transformer models to solve cyber security problems Train a Transformer model to detect malicious URLs under multiple training regimes Compare our model against other deep learning methods, and show it performs on-par with other top-scoring models Identify...

0.1AI score
SaveExploits0References13
FireEye
FireEye
added 2020/08/05 12:00 a.m.21 views

Repurposing Neural Networks to Generate Synthetic Media for Information Operations

FireEye’s Data Science and Information Operations Analysis teams released this blog post to coincide with our Black Hat USA 2020 Briefing, which details how open source, pre-trained neural networks can be leveraged to generate synthetic media for malicious purposes. To summarize our presentation,...

0.6AI score
SaveExploits0References21
Rows per page
Query Builder