24160 matches found
safety
!NOTE Come and join us at SafetyCLI. We are hiring for various roles. Table of Contents Table of Contents Introduction Key Features Getting Started GitHub Action Command Line Interface 1. Installation 2. Log In or Register 3. Running Your First Scan Basic Commands Service-Level Agreement SLA...
xmloxide
xmloxide A pure Rust reimplementation of libxml2 — the de facto standard XML/HTML parsing library in the open-source world. libxml2 became officially unmaintained in December 2025 with known security issues. xmloxide aims to be a memory-safe, high-performance replacement that passes the same...
nano-analyzer
Nano-analyzer A minimal LLM-powered zero-day vulnerability scanner byAISLE. Research prototype for demonstration purposes. This is a simple, single-file harness that is able to detect real zero-day vulnerabilities. Note that it is a prototype, biased towards C/C++ memory safety bugs, and will...
autoguardrails
autoguardrails Open source by Santander AI Lab. An LLM / AI-safety guardrail research library / evaluation harness autoresearch-style: it searches over a single mutable policy.md surface to minimize attack success rate ASR against a fixed evaluation suite, with a benign-pass floor. Part of...
clr
CLR C hecker of L ifetimes and other R efinement types for Zig Video: https://www.youtube.com/watch?v=mf0WzTOe-40 Sponsorship: https://buymeacoffee.com/dnautics discuss on hn: https://news.ycombinator.com/item?id=42923829 discuss on lobste.rs: https://lobste.rs/s/9sitsj/clrcheckerforlifetimesothe...
meta-ai-support-prompt
Meta AI Support Assistant System Prompt Extracted system prompt from Meta's AI Support Assistant on June 1, 2026. Files system-prompt.md — Extracted system prompt ⚠️ Disclaimer & Legal Notice Purpose This repository is published strictly for educational and authorized security research purposes...
mcpsafetywarden
MCP safety warden is a proxy server that wraps any MCP server and adds behavioral profiling, security scanning, risk gating, and safe execution to its tools. Contents Overview Prerequisites Installation Configuration MCP Integration CLI Reference Auxiliary Integrations Development Testing Fu...
heretic
Heretic: Fully automatic censorship removal for language models Heretic is a tool that removes censorship aka "safety alignment" from transformer-based language models without expensive post-training. It combines an advanced implementation of directional ablation, also known as "abliteration"...
kalitorify
Transparent Proxy through Tor for Kali Linux About kalitorify kalitorify is a shell script for Kali Linux which use iptables settings to create a Transparent Proxy through the Tor Network , the program also allows you to perform various checks like checking the Tor Exit Node i.e. your public IP...
GuardReasoner-VL
GuardReasoner-VL: Sécuriser les VLM grâce au raisonnement renforcé Yue Liu, Shengfang Zhai, Mingzhe Du Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi 1National University of Singapore, 2Nanyang Technological University Pour renforcer l...
AutoRAN-public
🧠 AutoRAN : Détournement automatisé du raisonnement de sécurité dans les grands modèles de raisonnement AutoRAN est un détournement automatisé du raisonnement de sécurité qui exploite des modèles auxiliaires secondaires moins alignés pour simuler des traces de raisonnement, générer des prompts...
saffron
Saffron-1: Масштабирование инференса для обеспечения безопасности LLM 📖 Статья 🛠️ Зависимости Код был протестирован со следующими зависимостями: Python 3.12.3 CUDA 12.2 typingextensions==4.14.0 numpy==2.2.6 torch==2.5.1 huggingfacehub==0.30.2 accelerate==1.1.1 datasets==3.1.0 evaluate==0.4.3...
safety
!NOTE Rejoignez-nous chez SafetyCLI. Nous recrutons pour divers postes. Table des matières Table des matières Introduction Fonctionnalités clés Pour commencer Action GitHub Interface en ligne de commande 1. Installation 2. Connexion ou inscription 3. Exécution de votre première analyse Commandes...
redeval
RedEval - фреймворк для оценки безопасности LLM Комплексный фреймворк для оценки безопасности больших языковых моделей LLM с помощью систематического тестирования атак и отказов. RedEval предоставляет унифицированную, безопасную и расширяемую платформу для оценки устойчивости LLM к состязательным...
CUAHarm
Оценка вредоносности агентов, использующих компьютер 🤗 Hugging Face 📄 Paper 📋 Введение CUAHarm — это бенчмарк, предназначенный для оценки рисков безопасности агентов, использующих компьютер CUA — ИИ-агентов, которые могут автономно управлять компьютером для выполнения многошаговых действий...
CVE-2026-39113
CVE-2026-39113: Переполнение буфера в куче в опциональном расширении SQLAR для SQLite Executive Summary CVE-2026-39113 — это переполнение буфера в куче в опциональном расширении SQLAR для SQLite. В приложении, загрузившем это расширение, атакующий, способный вызвать sqlaruncompress с управляемым...
aegis-audio-defense
AEGIS: Audio Endogenous Guarding via Internal Signals Against Large Audio-Language Model Jailbreaks Submitted to ICASSP 2027. AEGIS keeps the target audio-language model frozen and adds a mid-layer risk gate that conditionally drives late-layer safety adapters, with closed-loop scaling at...
santa
Santa !NOTE As of 2025, Santa is no longer maintained by Google. We encourage existing users to migrate to an actively maintained fork of Santa, such as https://github.com/northpolesec/santa. Santa is a binary and file access authorization system for macOS. It consists of a system extension that...
safety-awareness
Transfer Safety Awareness for Cross-Modal Safety Drift Официальная реализация Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models EMNLP 2026 Findings. Этот репозиторий содержит конвейер направления осведомлённости о безопасности Qwen3-VL. Он извлекает...
ShadowMem
ShadowMem: Защита LLM-агентов от долгосрочных угроз с помощью теневой памяти Этот репозиторий содержит официальный код к статье Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory. ShadowMem — это защитный фреймворк, который поддерживает выделенную агентную память,...