Lucene search
+L

5396 matches found

Packet Storm News
Packet Storm News
•added 2025/06/21 12:00 a.m.•19 views

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

We introduce AIRTBench, an AI red teaming benchmark for evaluating language models' ability to autonomously discover and exploit Artificial Intelligence and Machine Learning AI/ML security vulnerabilities. The benchmark consists of 70 realistic black-box capture-the-flag CTF challenges from the...

7.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/21 12:00 a.m.•16 views

Doppelgänger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack

Since the advent of large language models, prompt engineering now enables the rapid, low-effort creation of diverse autonomous agents that are already in widespread use. Yet this convenience raises urgent concerns about the safety, robustness, and behavioral consistency of the underlying prompts,...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/21 12:00 a.m.•14 views

Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers

We study privacy leakage in the reasoning traces of large reasoning models used as personal agents. Unlike final outputs, reasoning traces are often assumed to be internal and safe. We challenge this assumption by showing that reasoning traces frequently contain sensitive user data, which can be...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/19 12:00 a.m.•14 views

Probe Before You Talk: Towards Black-Box Defense against Backdoor Unalignment for Large Language Models

Backdoor unalignment attacks against Large Language Models LLMs enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service...

7.4AI score
SaveExploits0
HackRead
HackRead
•added 2025/06/18 4:19 p.m.•25 views

AgentSmith Flaw in LangSmith’s Prompt Hub Exposed User API Keys, Data

A CVSS 8.8 AgentSmith flaw in LangSmith's Prompt Hub exposed AI agents to data theft and LLM manipulation. Learn how malicious AI agents could steal API keys and hijack LLM responses. Fix deployed...

7.2AI score
SaveExploits0
AstraLinux
AstraLinux
•added 2025/06/16 11:28 a.m.•13 views

Astra Linux – Vulnerability in Firefox

By first using the AI chatbot in one tab and then activating it in another tab, the document title from the previous tab would be leaked into the chat prompt. This vulnerability was fixed in Firefox 137...

5.3CVSS7.7AI score0.00291EPSS
SaveExploits0References3
Packet Storm News
Packet Storm News
•added 2025/06/15 12:00 a.m.•12 views

The Safety Reminder: a Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

As Vision-Language Models VLMs demonstrate increasing capabilities across real-world applications such as code generation and chatbot assistance, ensuring their safety has become paramount. Unlike traditional Large Language Models LLMs, VLMs face unique vulnerabilities due to their multimodal...

7.5AI score
SaveExploits0
RedhatCVE
RedhatCVE
•added 2025/06/13 6:15 p.m.•21 views

CVE-2025-49150

Cursor is a code editor built for programming with AI. Prior to 0.51.0, by default, the setting json.schemaDownload.enable was set to True. This means that by writing a JSON file, an attacker can trigger an arbitrary HTTP GET request that does not require user confirmation. Since the Cursor Agent...

5.9CVSS5.8AI score0.00383EPSS
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2025/06/13 12:00 a.m.•17 views

AgentVigil: Generic Black-Box Red-Teaming for Indirect Prompt Injection against LLM Agents

The strong planning and reasoning capabilities of Large Language Models LLMs have fostered the development of agent-based systems capable of leveraging external tools and interacting with increasingly complex environments. However, these powerful features also introduce a critical security risk:...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•13 views

LLMs Cannot Reliably Judge (Yet?): a Comprehensive Assessment on the Robustness of LLM-As-A-Judge

Large Language Models LLMs have demonstrated remarkable intelligence across various tasks, which has inspired the development and widespread adoption of LLM-as-a-Judge systems for automated model testing, such as red teaming and benchmarking. However, these systems are susceptible to adversarial...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•14 views

LLMail-Inject: a Dataset from a Realistic Adaptive Prompt Injection Challenge

Indirect Prompt Injection attacks exploit the inherent limitation of Large Language Models LLMs to distinguish between instructions and data in their inputs. Despite numerous defense proposals, the systematic evaluation against adaptive adversaries remains limited, even when successful attacks ca...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•16 views

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

In this work, we show that some machine unlearning methods may fail when subjected to straightforward prompt attacks. We systematically evaluate eight unlearning techniques across three model families, and employ output-based, logit-based, and probe analysis to determine to what extent supposedly...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/11 12:00 a.m.•16 views

Design Patterns for Securing LLM Agents against Prompt Injections

As AI agents powered by Large Language Models LLMs become increasingly versatile and capable of addressing a broad spectrum of tasks, ensuring their security has become a critical challenge. Among the most pressing threats are prompt injection attacks, which exploit the agent's resilience on...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/10 12:00 a.m.•14 views

DAVSP: Safety Alignment for Large Vision-Language Models Via Deep Aligned Visual Safety Prompt

Large Vision-Language Models LVLMs have achieved impressive progress across various applications but remain vulnerable to malicious queries that exploit the visual modality. Existing alignment approaches typically fail to resist malicious queries while preserving utility on benign ones effectivel...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/10 12:00 a.m.•14 views

Evaluation Empirique De La Sécurisation Et De L'Alignement De ChatGPT Et Gemini: Analyse Comparative Des Vulnérabilités Par Expérimentations De Jailbreaks

Large Language models LLMs are transforming digital usage, particularly in text generation, image creation, information retrieval and code development. ChatGPT, launched by OpenAI in November 2022, quickly became a reference, prompting the emergence of competitors such as Google's Gemini. However...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/10 12:00 a.m.•15 views

Securing Generative AI Agentic Workflows: Risks, Mitigation, and a Proposed Firewall Architecture

Generative Artificial Intelligence GenAI presents significant advancements but also introduces novel security challenges, particularly within agentic workflows where AI agents operate autonomously. These risks escalate in multi-agent systems due to increased interaction complexity. This paper...

7.3AI score
SaveExploits0
Rapid7 Vulnerability Database (full)
Rapid7 Vulnerability Database (full)
•added 2025/06/07 12:00 a.m.•5 views

CVE-2025-49619: Improper Neutralization of Special Elements Used in a Template Engine

Skyvern through 0.1.85 is vulnerable to server-side template injection SSTI in the Prompt field of workflow blocks such as the Navigation v2 Block. Improper sanitization of Jinja2 template input allows authenticated users to inject crafted expressions that are evaluated on the server, leading to...

8.5CVSS8.2AI score0.19972EPSS
SaveExploits6References1
Mend
Mend
•added 2025/06/07 12:00 a.m.•4 views

CVE-2025-49619

Skyvern through 0.1.85 is vulnerable to server-side template injection SSTI in the Prompt field of workflow blocks such as the Navigation v2 Block. Improper sanitization of Jinja2 template input allows authenticated users to inject crafted expressions that are evaluated on the server, leading to...

8.5CVSS0.19972EPSS
SaveExploits6References11
Packet Storm News
Packet Storm News
•added 2025/06/06 12:00 a.m.•16 views

Stealix: Model Stealing Via Prompt Evolution

Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing intellectual property and exposing sensitive information. Recent methods that use pre-trained diffusion models for data...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/05 12:00 a.m.•38 views

Sentinel: SOTA Model to Protect against Prompt Injections

Large Language Models LLMs are increasingly powerful but remain vulnerable to prompt injection attacks, where malicious inputs cause the model to deviate from its intended instructions. This paper introduces Sentinel, a novel detection model, qualifire/prompt-injection-sentinel, based on the...

7.2AI score
SaveExploits0
Rows per page
Query Builder