Lucene search
+L

170 matches found

Packet Storm News
Packet Storm News
added 2025/06/14 12:00 a.m.13 views

Step-By-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models

Whitepaper called Step-By-Step Reasoning Attack: Revealing 'Erased' Knowledge In Large Language Models...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/13 12:00 a.m.8 views

Towards Understanding the Cognitive Habits of Large Reasoning Models

Large Reasoning Models LRMs, which autonomously produce a reasoning Chain of Thought CoT before producing final responses, offer a promising approach to interpreting and monitoring model behaviors. Inspired by the observation that certain CoT patterns -- e.g., "Wait, did I miss anything?'' --...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/12 12:00 a.m.6 views

Chain-Of-Code Collapse: Reasoning Failures in LLMs Via Adversarial Prompting in Code Generation

Large Language Models LLMs have achieved remarkable success in tasks requiring complex reasoning, such as code generation, mathematical problem solving, and algorithmic synthesis -- especially when aided by reasoning tokens and Chain-of-Thought prompting. Yet, a core question remains: do these...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/08 12:00 a.m.11 views

SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows

Large language models LLMs have seen widespread success in code generation tasks for different scenarios, both everyday and professional. However current LLMs, despite producing functional code, do not prioritize security and may generate code with exploitable vulnerabilities. In this work, we...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/06/08 12:00 a.m.10 views

HauntAttack: When Attack Follows Reasoning As a Shadow

Emerging Large Reasoning Models LRMs consistently excel in mathematical and reasoning tasks, showcasing exceptional capabilities. However, the enhancement of reasoning abilities and the exposure of their internal reasoning processes introduce new safety vulnerabilities. One intriguing concern is:...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/31 12:00 a.m.8 views

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences

LLM generated code often contains security issues. We address two key challenges in improving secure code generation. First, obtaining high quality training data covering a broad set of security issues is critical. To address this, we introduce a method for distilling a preference dataset of...

7.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/25 12:00 a.m.9 views

ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast \& Slow Reasoning for Robust Agent Defense

LLM Agents are becoming central to intelligent systems. However, their deployment raises serious safety concerns. Existing defenses largely rely on "Safety Checks", which struggle to capture the complex semantic risks posed by harmful user inputs or unsafe agent behaviors - creating a significant...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/25 12:00 a.m.14 views

CoTGuard: Using Chain-Of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems

As large language models LLMs evolve into autonomous agents capable of collaborative reasoning and task execution, multi-agent LLM systems have emerged as a powerful paradigm for solving complex problems. However, these systems pose new challenges for copyright protection, particularly when...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/21 12:00 a.m.18 views

SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning

Large Reasoning Models LRMs introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs,...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/20 12:00 a.m.11 views

Robust and Efficient AI-Based Attack Recovery in Autonomous Drones

We introduce an autonomous attack recovery architecture to add common sense reasoning to plan a recovery action after an attack is detected. We outline use-cases of our architecture using drones, and then discuss how to implement this architecture efficiently and securely in edge devices...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/16 12:00 a.m.15 views

GuardReasoner-VL: Safeguarding VLMs Via Reinforced Reasoning

To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL. First, we construct GuardReasoner-VLTrain, a reasoning...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/16 12:00 a.m.11 views

AutoRAN: Weak-To-Strong Jailbreaking of Large Reasoning Models

This paper presents AutoRAN, the first automated, weak-to-strong jailbreak attack framework targeting large reasoning models LRMs. At its core, AutoRAN leverages a weak, less-aligned reasoning model to simulate the target model's high-level reasoning structures, generates narrative prompts, and...

7.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/16 12:00 a.m.12 views

DMind Benchmark: toward a Holistic Assessment of LLM Capabilities across the Web3 Domain

Large Language Models LLMs have achieved impressive performance in diverse natural language processing tasks, but specialized domains such as Web3 present new challenges and require more tailored evaluation. Despite the significant user base and capital flows in Web3, encompassing smart contracts...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/10 12:00 a.m.16 views

Practical Reasoning Interruption Attacks on Reasoning Large Language Models

Reasoning large language models RLLMs have demonstrated outstanding performance across a variety of tasks, yet they also expose numerous security vulnerabilities. Most of these vulnerabilities have centered on the generation of unsafe content. However, recent work has identified a distinct...

7.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/09 12:00 a.m.12 views

System Prompt Poisoning: Persistent Attacks on Large Language Models beyond User Injection

Large language models LLMs have gained widespread adoption across diverse applications due to their impressive generative capabilities. Their plug-and-play nature enables both developers and end users to interact with these models through simple prompts. However, as LLMs become more integrated in...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/07 12:00 a.m.10 views

DMRL: Data- and Model-Aware Reward Learning for Data Extraction

Large language models LLMs are inherently vulnerable to unintended privacy breaches. Consequently, systematic red-teaming research is essential for developing robust defense mechanisms. However, current data extraction methods suffer from several limitations: 1 rely on dataset duplicates...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/05/06 12:00 a.m.11 views

The Steganographic Potentials of Language Models

The potential for large language models LLMs to hide messages within plain text steganography poses a challenge to detection and thwarting of unaligned AI agents, and undermines faithfulness of LLMs reasoning. We explore the steganographic capabilities of LLMs fine-tuned via reinforcement learnin...

6.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/04/28 12:00 a.m.12 views

The Automation Advantage in AI Red Teaming

This paper analyzes Large Language Model LLM security vulnerabilities based on data from Crucible, encompassing 214,271 attack attempts by 1,674 users across 30 LLM challenges. Our findings reveal automated approaches significantly outperform manual techniques 69.5% vs 47.6% success rate, despite...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/04/26 12:00 a.m.12 views

CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges

Large language models LLMs have demonstrated remarkable capabilities, especially the recent advancements in reasoning, such as o1 and o3, pushing the boundaries of AI. Despite these impressive achievements in mathematics and coding, the reasoning abilities of LLMs in domains requiring cryptograph...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/04/26 12:00 a.m.9 views

Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control

Large language models LLMs have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those who control the models. To understand how this "censorship...

7.2AI score
SaveExploits0
Rows per page
Query Builder