Lucene search
+L

29 matches found

Kitploit
Kitploit
•added 2026/10/10 12:13 p.m.•24 views

BoxPwnr-Traces

BoxPwnr-Traces BoxPwnr ट्रेसेस और बेंचमार्क परिणाम कई सुरक्षा प्लेटफार्मों पर। प्रत्येक ट्रेस में पूर्ण LLM इंटरेक्शन, निष्पादित कमांड्स, एक मार्कडाउन रिपोर्ट + अटैक ग्राफ, उपयोग किए गए आँकड़े और कॉन्फ़िगरेशन शामिल हैं। लीडरबोर्ड ब्राउज़ करें, एक इंटरैक्टिव वेब व्यूअर में रन को रीप्ले करें, और...

6AI score
SaveExploits0References1
Kitploit
Kitploit
•added 2026/10/10 7:38 a.m.•16 views

DeepTrap

DeepTrap الإنجليزية | 中文 تقييم أمني للعالم المفتوح لوكلاء OpenClaw في سياقات تنفيذ عدائية. DeepTrap هو معيار أمني benchmark لتقييم ما إذا كان بإمكان وكلاء OpenClaw إكمال مهام مستخدم سليمة مع مقاومة ضغوط سياق التنفيذ الخبيثة: ملفات مساحة عمل مسمومة، ومهارات محقونة، وبيانات وصفية مضللة للأدوات،...

6.2AI score
SaveExploits0References3
Kitploit
Kitploit
•added 2026/10/07 10:35 p.m.•12 views

Ajar

Ajar : mesurer le privilège ouvert dans les défenses d'agents Un benchmark de sécurité d'agents rapporte deux nombres, le succès des attaques et l'utilité bénigne, et tous deux sont lus sur des exécutions qui ont eu lieu. Ni l'un ni l'autre ne dit ce que la défense se tenait prête à autoriser sur...

6.3AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/09/10 12:00 a.m.•29 views

HOL Guard AI Agent Runtime Security Benchmark

HOL Guard Benchmark is an open source fixture-based benchmark for evaluating security controls applied to AI coding agents. It defines deterministic scenarios covering secret access, risky shell execution, tool poisoning, approval handling, safe operations, and receipt generation across Codex CLI...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/08/06 12:00 a.m.•17 views

Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security Baselining

Research and Education Networks RENs serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their normal traffic patterns which are characterized by massive, bursty "elephant flows" are statistically indistinguishable from volumetric attacks such as DD...

5.4AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/07/30 12:00 a.m.•21 views

Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents

Large Language Model LLM-driven multimodal agents are increasingly deployed to execute autonomous tasks via continuous audio interaction. While this paradigm enhances interaction naturalness, it introduces a critical yet under-explored attack surface, as audio inputs inevitably contain...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/07/29 12:00 a.m.•25 views

Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

Enterprises are moving toward autonomous cyber defense: agentic AI that builds situational awareness of an organization's security state and reasons from it to assessments, decisions, and actions. This rests on a holistic view of the enterprise's security state, the continuous, cross-vendor pictu...

5.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/05/17 12:00 a.m.•56 views

ContraFix: Agentic Vulnerability Repair Via Differential Runtime Evidence and Skill Reuse

Large language model LLM agents are increasingly used for automated vulnerability repair AVR, where repository-level reasoning enables them to inspect context and produce source-code patches. However, recent empirical results show that these agents still struggle with real-world vulnerabilities...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/05/05 12:00 a.m.•27 views

ARGUS: Defending LLM Agents against Context-Aware Prompt Injection

The rise of Large Language Model LLM agents, augmented with tool use, skills, and external knowledge, has introduced new security risks. Among them, prompt injection attacks, where adversaries embed malicious instructions into the agent workflow, have emerged as the primary threat. However,...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/02/23 12:00 a.m.•19 views

Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks

LLM agents are evolving rapidly, powered by code execution, tools, and the recently introduced agent skills feature. Skills allow users to extend LLM applications with specialized third-party code, knowledge, and instructions. Although this can extend agent capabilities to new domains, it creates...

5.9AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/02/02 12:00 a.m.•11 views

Benchmarking Large Language Models for Zero-Shot and Few-Shot Phishing URL Detection

The Uniform Resource Locator URL, introduced in a connectivity-first era to define access and locate resources, remains historically limited, lacking future-proof mechanisms for security, trust, or resilience against fraud and abuse, despite the introduction of reactive protections like HTTPS...

5.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/11/19 12:00 a.m.•20 views

Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments

The rapid proliferation of Multimodal Large Language Models MLLMs has introduced unprecedented security challenges, particularly in phishing detection within academic environments. Academic institutions and researchers are high-value targets, facing dynamic, multilingual, and context-dependent...

6.6AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/08/17 12:00 a.m.•19 views

MCPSecBench: a Systematic Security Benchmark and Playground for Testing Model Context Protocols

Large Language Models LLMs are increasingly integrated into real-world applications via the Model Context Protocol MCP, a universal, open standard for connecting AI agents with data sources and external tools. While MCP enhances the capabilities of LLM-based agents, it also introduces new securit...

7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/21 12:00 a.m.•14 views

RAS-Eval: a Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments

The rapid deployment of Large language model LLM agents in critical domains like healthcare and finance necessitates robust security frameworks. To address the absence of standardized evaluation benchmarks for these agents in dynamic environments, we introduce RAS-Eval, a comprehensive security...

7.3AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/09 12:00 a.m.•17 views

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges

The widespread adoption of Large Language Models LLMs has heightened concerns about their security, particularly their vulnerability to jailbreak attacks that leverage crafted prompts to generate malicious outputs. While prior research has been conducted on general security capabilities of LLMs,...

7.1AI score
SaveExploits0
OpenVAS
OpenVAS
•added 2025/05/07 12:00 a.m.•10 views

Configure the dmesg Access Permission Properly

The permission to access dmesg information is restricted. Unprivileged users cannot view system information. This prevents any one from obtaining sensitive information and attacking the system. Only processes with the CAPSYSLOG capability are allowed to access kernel logs. In this way, the least...

6.5AI score
SaveExploits0References4
OpenVAS
OpenVAS
•added 2025/05/07 12:00 a.m.•12 views

Configure Proper Policies for OUTPUT of nftables

There are two occasions in which a server sends outgoing packets: 1. The local host process proactively connects to an external server, for example, performing an HTTP access, or sending data to a log server. 2. The local host responds to the external access to the local services. If no policy is...

6.7AI score
SaveExploits0References2
OpenVAS
OpenVAS
•added 2025/05/07 12:00 a.m.•19 views

Disable SysRq

SysRq enables users with physical access to access dangerous system-level commands in a computer. Therefore, it is advised to restrict the usage of the SysRq function. If SysRq is not disabled, you can use the keyboard to trigger SysRq. As a result, commands may be directly sent to the kernel,...

6.8AI score
SaveExploits0References4
OpenVAS
OpenVAS
•added 2025/05/07 12:00 a.m.•6 views

Configure a Proper Number of Concurrent Unauthenticated SSH Connections

Without knowing the password, an attacker can set up a large number of concurrent connections that have not been authenticated to consume system resources. The number of concurrent unauthenticated SSH connections is not configured in openEuler by default. You are advised to configure the upper...

6.9AI score
SaveExploits0References3
OpenVAS
OpenVAS
•added 2025/05/07 12:00 a.m.•7 views

Do Not Use X11 Forwarding

The X11 forwarding function of SSH allows the GUI program of the remote host to be executed on the local host. If the X11 forwarding function is enabled, the attack surface is expanded and other users on the X11 server may attack the local host. If the function is not required in the service...

6.7AI score
SaveExploits0References3
Rows per page
Query Builder