Lucene search
+L

33 matches found

Packet Storm News
Packet Storm News
•added 2025/10/12 12:00 a.m.•80 views

SASER: Stego Attacks on Open-Source LLMs

Open-source large language models LLMs have demonstrated considerable dominance over proprietary LLMs in resolving neural processing tasks, thanks to the collaborative and sharing nature. Although full access to source codes, model parameters, and training data lays the groundwork for transparenc...

6.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/10/11 12:00 a.m.•17 views

ArtPerception: ASCII Art-Based Jailbreak on LLMs with Recognition Pre-Test

The integration of Large Language Models LLMs into computer applications has introduced transformative capabilities but also significant security challenges. Existing safety alignments, which primarily focus on semantic interpretation, leave LLMs vulnerable to attacks that use non-standard data...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/09/18 12:00 a.m.•20 views

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism Via Probabilistically Ablating Refusal Direction

Jailbreak attacks pose persistent threats to large language models LLMs. Current safety alignment methods have attempted to address these issues, but they experience two significant limitations: insufficient safety alignment depth and unrobust internal defense mechanisms. These limitations make...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/09/04 12:00 a.m.•18 views

Between a Rock and a Hard Place: Exploiting Ethical Reasoning to Jailbreak LLMs

Large language models LLMs have undergone safety alignment efforts to mitigate harmful outputs. However, as LLMs become more sophisticated in reasoning, their intelligence may introduce new security risks. While traditional jailbreak attacks relied on singlestep attacks, multi-turn jailbreak...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/18 12:00 a.m.•24 views

Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes

The advancement of AI technologies, particularly Large Language Models LLMs, has transformed computing while introducing new security and privacy risks. Prior research shows that cybercriminals are increasingly leveraging uncensored LLMs ULLMs as backends for malicious services. Understanding the...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/16 12:00 a.m.•15 views

Mitigating Jailbreaks with Intent-Aware LLMs

Despite extensive safety-tuning, large language models LLMs remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off between safety and task performance. In this work, we propose Intent-FT, a simple and lightweight fine-tuning approach that...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/04 12:00 a.m.•21 views

PentestJudge: Judging Agent Behavior against Operational Requirements

We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model LLM-as-judge with access to tools that allow it to consume arbitrary trajectories of agent states and tool call history to determine whether a security agent's...

6.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/07/26 12:00 a.m.•17 views

SDD: Self-Degraded Defense against Malicious Fine-Tuning

Open-source Large Language Models LLMs often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuni...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/06 12:00 a.m.•16 views

Stealix: Model Stealing Via Prompt Evolution

Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing intellectual property and exposing sensitive information. Recent methods that use pre-trained diffusion models for data...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/26 12:00 a.m.•13 views

Capability-Based Scaling Laws for LLM Red-Teaming

As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into a weak-to-strong problem, where target models surpass...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/22 12:00 a.m.•15 views

CAIN: Hijacking LLM-Humans Conversations Via a Two-Stage Malicious System Prompt Generation and Refining Framework

Large language models LLMs have advanced many applications, but are also known to be vulnerable to adversarial attacks. In this work, we introduce a novel security threat: hijacking AI-human conversations by manipulating LLMs' system prompts to produce malicious answers only to specific targeted...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/05/15 12:00 a.m.•17 views

Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data

The integration of large language models LLMs into cyber security applications presents significant opportunities, such as enhancing threat analysis and malware detection, but can also introduce critical risks and safety concerns, including personal data leakage and automated generation of new...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/04/23 12:00 a.m.•18 views

AiXamine: Simplified LLM Safety and Security

Evaluating Large Language Models LLMs for safety and security remains a complex task, often requiring users to navigate a fragmented landscape of ad hoc benchmarks, datasets, metrics, and reporting formats. To address this challenge, we present aiXamine, a comprehensive black-box evaluation...

7.5AI score
SaveExploits0
Rows per page
Query Builder