Lucene search
+L

28 matches found

Packet Storm News
Packet Storm News
added 2025/08/16 12:0 a.m.5 views

Mitigating Jailbreaks with Intent-Aware LLMs

Despite extensive safety-tuning, large language models LLMs remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off between safety and task performance. In this work, we propose Intent-FT, a simple and lightweight fine-tuning approach that...

7.2AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/08/04 12:0 a.m.8 views

PentestJudge: Judging Agent Behavior against Operational Requirements

We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model LLM-as-judge with access to tools that allow it to consume arbitrary trajectories of agent states and tool call history to determine whether a security agent's...

6.7AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/07/26 12:0 a.m.7 views

SDD: Self-Degraded Defense against Malicious Fine-Tuning

Open-source Large Language Models LLMs often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuni...

7.1AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/06/06 12:0 a.m.7 views

Stealix: Model Stealing Via Prompt Evolution

Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing intellectual property and exposing sensitive information. Recent methods that use pre-trained diffusion models for data...

6.8AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/05/26 12:0 a.m.7 views

Capability-Based Scaling Laws for LLM Red-Teaming

As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into a weak-to-strong problem, where target models surpass...

6.9AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/05/22 12:0 a.m.4 views

CAIN: Hijacking LLM-Humans Conversations Via a Two-Stage Malicious System Prompt Generation and Refining Framework

Large language models LLMs have advanced many applications, but are also known to be vulnerable to adversarial attacks. In this work, we introduce a novel security threat: hijacking AI-human conversations by manipulating LLMs' system prompts to produce malicious answers only to specific targeted...

7.1AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/05/15 12:0 a.m.5 views

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

The integration of large language models LLMs into cyber security applications presents significant opportunities, such as enhancing threat analysis and malware detection, but can also introduce critical risks and safety concerns, including personal data leakage and automated generation of new...

7.5AI score
Exploits0
Packet Storm News
Packet Storm News
added 2025/04/23 12:0 a.m.6 views

AiXamine: Simplified LLM Safety and Security

Evaluating Large Language Models LLMs for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address this challenge, we present aiXamine, a comprehensive black-box evaluation...

7.5AI score
Exploits0
Rows per page
Query Builder