639 matches found
attack-attention
Attacking Attention of Foundation Models Effectively Disrupts Downstream Tasks Official PyTorch implementation of the paper "Attacking Attention of Foundation Models Effectively Disrupts Downstream Tasks" , accepted in the Adversarial Machine Learning on Computer Vision: Foundation Models + X ADV...
deep-pwning
Deep-pwning is a lightweight framework for experimenting with machine learning models with the goal of evaluating their robustness against a motivated adversary. Note that deep-pwning in its current state is no where close to maturity or completion. It is meant to be experimented with, expanded...
SecML
SecML: A library for Secure and Explainable Machine Learning SecML is an open-source Python library for the security evaluation of Machine Learning ML algorithms. It comes with a set of powerful features: Wide range of supported ML algorithms. All supervised learning algorithms supported by...
ai-security-battle
🤖 AI Security Battle System Red Team vs Blue Team: Adversarial ML Combat Simulation A comprehensive adversarial machine learning security simulation where Red Team attackers battle against Blue Team defenders in a 30-minute AI/ML security warfare scenario. 🎯 Project Overview This system simulates...
augustus
Augustus - LLM vulnerability scanner for prompt injection, jailbreak, and adversarial attack testing Augustus - LLM Vulnerability Scanner Test large language models against 210+ adversarial attacks covering prompt injection, jailbreaks, encoding exploits, and data extraction. Augustus is a Go-bas...
Sec5GLoc
Sec5GLoc: Securing 5G Indoor Localization via Adversary-Resilient Deep Learning Architecture Our work addresses security and privacy challenges in 5G indoor localization by proposing an adversary-resilient deep learning architecture. Sec5GLoc combines Channel Impulse Response CIR fingerprinting...
Ruse
Ruse Aplicación móvil basada en la cámara que intenta modificar fotos para preservar su utilidad para los humanos, al mismo tiempo que las vuelve inutilizables para sistemas de reconocimiento facial. Instalación 1 Método sencillo: esperar y descargar la aplicación desde la tienda de aplicaciones...
metta
Metta Metta is an information security preparedness tool. This project uses Redis/Celery, python, and vagrant with virtualbox to do adversarial simulation. This allows you to test mostly your host based instrumentation but may also allow you to test any network based detection and controls...
metta
Metta Metta es una herramienta de preparación para la seguridad de la información. Este proyecto utiliza Redis/Celery, Python y Vagrant con VirtualBox para realizar simulación adversarial. Esto le permite probar principalmente su instrumentación basada en el host, pero también puede permitirle...
OpenAttack
Documentación • Características y Usos • Ejemplos de Uso • Modelos de Ataque • Diseño del Toolkit OpenAttack es un kit de herramientas de ataque adversarial textual de código abierto basado en Python, que maneja todo el proceso de ataque adversarial textual, incluyendo preprocesamiento de texto,...
alibi-detect
Alibi Detect is a source-available Python library focused on outlier , adversarial and drift detection. The package aims to cover both online and offline detectors for tabular data, text, images and time series. Both TensorFlow and PyTorch backends are supported for drift detection. Documentation...
offensive-ai-compilation
Offensive AI Compilation A curated list of useful resources that cover Offensive AI. 📁 Contents 📁 🚫 Abuse 🚫 🧠 Adversarial Machine Learning 🧠 ⚡ Attacks ⚡ 🔒 Extraction 🔒 ⚠️ Limitations ⚠️ 🛡️ Defensive actions 🛡️ 🔗 Useful links 🔗 ⬅️ Inversion or inference ⬅️ 🛡️ Defensive actions 🛡️ 🔗 Useful links 🔗 💉...
PacketPatch
PacketPatch: Practical Generation and Deployment of Adversarial Packets for Byte-Feature-Based Encrypted Traffic Classification Paper : PacketPatch: Practical Generation and Deployment of Adversarial Packets for Byte-Feature-Based Encrypted Traffic Classification Authors : Yuwei Xu, Yuanyuan Xu,...
ActBench
ActBench ActBench is a self-evolving benchmark of behavioral safety in cowork agents. It defines behavioral safety as whether an agent's execution remains within the permissions and state changes required by a benign task, and evaluates realized behavioral risk from execution trajectories rather...
Awesome-MoAI-Security
Awesome Mobile On-Device AI Security Seguridad de IA en Dispositivos Móviles SoK: Landscape de Ataques y Defensas de los Sistemas de IA en Dispositivos Móviles Los sistemas de IA en dispositivos móviles ejecutan modelos de IA localmente a través de frameworks de ML como LiteRT/TFLite , Core ML ,...
security-audit-skill
security-audit A coding-agent skill that turns your agent into a security auditor. It orchestrates multiple parallel agents through a six-phase pipeline -- recon, hunting, validation, reporting, structured output, and independent verification -- to find exploitable vulnerabilities with real impac...
QuasarNix
QuasarNix: Adversarially Robust Living-off-the-Land Reverse-Shell Detection Dmitrijs Trizna · Luca Demetrio · Battista Biggio · Fabio Roli Overview Linux living-off-the-land LOTL reverse shells abuse legitimate binaries bash, python, nc, … to establish covert outbound connections, making...
DeepTrap
DeepTrap English | 中文 Open-world security evaluation for OpenClaw agents under adversarial execution contexts. DeepTrap is a security benchmark for evaluating whether OpenClaw agents can complete benign user tasks while resisting malicious execution-context pressure: poisoned workspace files,...
cerebro-red-v2
CEREBRO-RED v2 Research Edition Autonomous Local LLM Red Teaming Suite A research-grade framework for automated vulnerability discovery in local LLMs using Agentic Fuzzing and Adaptive Adversarial Mutation AAM. Research Goals Implement PAIR Algorithm Prompt Automatic Iterative Refinement from...
StealthRL
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors Paper arXiv Demo Model Hugging Face Benchmark Dataset Hugging Face Abstract AI-text detectors are increasingly used in high-stakes settings, yet their robustness to meaning-preserving adversarial...