Lucene search
+L

33 matches found

Kitploit
Kitploit
added 2026/09/12 3:34 a.m.5 views

redeval

RedEval - LLM 안전성 평가 프레임워크 체계적인 공격attack 및 거부refusal 테스트를 통해 대규모 언어 모델LLM의 안전성을 평가하기 위한 포괄적인 프레임워크입니다. RedEval은 적대적 프롬프트와 유해 콘텐츠에 대한 LLM의 견고성을 평가하기 위한 통합되고 안전하며 확장 가능한 플랫폼을 제공합니다. 이는 LLM의 종합적인 레드팀링red teaming을 위한 범용 데이터셋인 논문 "RedBench: A Universal Dataset for Comprehensive Red Teaming of Large...

6AI score
SaveExploits0References1
Kitploit
Kitploit
added 2026/09/10 8:11 p.m.8 views

BrokenHill

!\ Broken Hill ロゴ \https://assets.kitploit.com/production/public/readmes/44608/c14154675c6f33f5c3887887afcdb39351c82f269ee2772cf31847106a20a871.jpg Broken Hill Broken Hill は、「Universal and Transferable Adversarial Attacks on Aligned Language Models」という論文(Andy Zou、Zifan Wang、Nicholas Carlini、Milad...

5.9AI score
SaveExploits0References5
Kitploit
Kitploit
added 2026/09/08 2:45 p.m.9 views

T-MAP

T-MAP: 궤적 인식 진화 탐색을 통한 LLM 에이전트 레드팀 테스트 T-MAP 은 MCP 서버에서 LLM 에이전트를 레드팀 테스트하기 위한 궤적 인식 진화 탐색 프레임워크입니다. 실행 궤적을 기반으로 적대적 프롬프트를 반복적으로 생성하고 변형하여 다양한 위험 범주와 공격 스타일에 걸쳐 에이전트의 취약점 지형을 매핑합니다. 🔧 설정 root@kitploit: pip install -r requirements.txt 요구 사항: Python 3.11+, 공격자 및 대상 모델용 API 키, 하나 이상의 MCP 서버에 대한 액세스...

6AI score
SaveExploits0
The Hacker News
The Hacker News
added 2026/07/16 8:42 a.m.27 views

OpenAI’s GPT-Red Automates Prompt Injection Testing to Harden GPT-5.6 Sol

OpenAI has disclosed details of GPT-Red , an internal automated red-teaming model that scales prompt injection vulnerability discovery with an aim to fix issues before the tools are deployed widely. "GPT‑Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt...

6.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/06/10 12:00 a.m.32 views

Smarter Saboteurs, Better Fixers: Scaling and Security in Linear Multi-Agent Workflows

As LLM-based multi-agent systems MAS are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/26 12:00 a.m.34 views

Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security

Large Language Models LLMs are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/11 12:00 a.m.15 views

When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications

Natural language interfaces to structured databases are becoming increasingly common, largely due to advances in large language models LLMs that enable users to query data using conversational input rather than formal query languages such as SQL. While this paradigm significantly improves usabili...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/09 12:00 a.m.17 views

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security beyond Binary Scoring

Jailbreak attacks -- adversarial prompts that bypass LLM alignment through purely linguistic manipulation -- pose a growing operational security threat, yet the field lacks large-scale, reproducible infrastructure for generating, categorizing, and evaluating them systematically. This paper...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/05/06 12:00 a.m.23 views

Information Theoretic Adversarial Training of Large Language Models

Large language models LLMs remain vulnerable to adversarial prompting despite advances in alignment and safety, often exhibiting harmful behaviors under novel attack strategies. While adversarial training can improve robustness, existing approaches are computationally expensive and difficult to...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/04/20 12:00 a.m.37 views

ARES: Adaptive Red-Teaming and End-To-End Repair of Policy-Reward System

Reinforcement Learning from Human Feedback RLHF is central to aligning Large Language Models LLMs, yet it introduces a critical vulnerability: an imperfect Reward Model RM can become a single point of failure when it fails to penalize unsafe behaviors. While existing red-teaming approaches...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/04/19 12:00 a.m.81 views

GuardPhish: Securing Open-Source LLMs from Phishing Abuse

The rapid adoption of open-source Large Language Models LLMs in offline and enterprise environments has introduced a largely unexamined security risk like susceptibility to adversarial phishing prompts under static safety configurations. In this work, we systematically investigate this...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/03/21 12:00 a.m.32 views

T-MAP: Red-Teaming LLM Agents with Trajectory-Aware Evolutionary Search

While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models LLMs, such approaches fail to capture agent-specific vulnerabilities that emerge through multi-step tool execution, particularly in rapidly growing ecosystems such as the Model Context Protoc...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/03/17 12:00 a.m.10 views

Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework

Large Language Models increasingly power critical infrastructure from healthcare to finance, yet their vulnerability to adversarial manipulation threatens system integrity and user safety. Despite growing deployment, no comprehensive comparative security assessment exists across major LLM...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2026/02/24 12:00 a.m.21 views

AdapTools: Adaptive Tool-Based Indirect Prompt Injection Attacks on Agentic LLMs

The integration of external data services e.g., Model Context Protocol, MCP has made large language model-based agents increasingly powerful for complex task execution. However, this advancement introduces critical security vulnerabilities, particularly indirect prompt injection IPI attacks...

6AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/12/30 12:00 a.m.9 views

Jailbreaking Attacks Vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?

As large language models LLMs are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful outputs, present significant risks, with existing studies reporting high success rates in evading common LLMs. However,...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/12/18 12:00 a.m.15 views

Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models

This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model LLM safety evaluation from a constrained example-based approach to a more expansive and effective policy-based framework. By leveraging an attack LLM to generate a high volume of...

7.1AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/12/07 12:00 a.m.58 views

ThinkTrap: Denial-Of-Service Attacks against Black-Box LLM Services Via Infinite Thinking

Large Language Models LLMs have become foundational components in a wide range of applications, including natural language understanding and generation, embodied intelligence, and scientific discovery. As their computational requirements continue to grow, these models are increasingly deployed as...

6.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/11/14 12:00 a.m.13 views

NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks

Jailbreak attacks designed to bypass safety mechanisms pose a serious threat by prompting LLMs to generate harmful or inappropriate content, despite alignment with ethical guidelines. Crafting universal filtering rules remains difficult due to their inherent dependence on specific contexts. To...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/11/04 12:00 a.m.18 views

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

Large Language Models LLMs remain vulnerable to jailbreaking attacks where adversarial prompts elicit harmful outputs, yet most evaluations focus on single-turn interactions while real-world attacks unfold through adaptive multi-turn conversations. We present AutoAdv, a training-free framework fo...

7.2AI score
SaveExploits0
Packet Storm News
Packet Storm News
added 2025/08/09 12:00 a.m.9 views

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection

Ensuring LLM alignment is critical to information security as AI models become increasingly widespread and integrated in society. Unfortunately, many defenses against adversarial attacks and jailbreaking on LLMs cannot adapt quickly to new attacks, degrade model responses to benign prompts, or...

6.7AI score
SaveExploits0
Rows per page
Query Builder