23488 matches found
heretic
Heretic: Fully automatic censorship removal for language models...
nano-analyzer
Nano-analyzer A minimal LLM-powered zero-day vulnerability scanner byAISLE. Research prototype for demonstration purposes. This is a simple, single-file harness that is able to detect real zero-day vulnerabilities. Note that it is a prototype, biased towards C/C++ memory safety bugs, and will...
safety
!NOTE Come and join us at SafetyCLI. We are hiring for various roles. Table of Contents Table of Contents Introduction Key Features Getting Started GitHub Action Command Line Interface 1. Installation 2. Log In or Register...
meta-ai-support-prompt
Meta AI Support Assistant System Prompt Extracted system prompt from Meta's AI Support Assistant on June 1, 2026. Files system-prompt.md — Extracted system prompt ⚠️ Disclaimer & Legal Notice Purpose This repository is published strictly for educational and authorized security research purposes...
saffron
Saffron-1: Inference Scaling for LLM Safety Assurance 📖 Paper 🛠️ Dependencies The code was tested under the following dependencies: Python 3.12.3 CUDA 12.2 typingextensions==4.14.0 numpy==2.2.6 torch==2.5.1 huggingfacehub==0.30.2 accelerate==1.1.1 datasets==3.1.0 evaluate==0.4.3...
AutoRAN-public
🧠 AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models AutoRAN is an automated Hijacking of Safety Reasoning that leverages less-aligned secondary auxiliary models to simulate reasoning traces, generate narrative prompts, and iteratively refine those prompts to bypass safety...
autoguardrails
autoguardrails Open source by Santander AI Lab. An LLM / AI-safety guardrail research library / evaluation harness autoresearch-style: it searches over a single mutable policy.md surface to minimize attack success rate ASR against a fixed evaluation suite, with a benign-pass floor. Part of...
redeval
RedEval - LLM Safety Evaluation Framework A comprehensive framework for evaluating the safety of Large Language Models LLMs through systematic attack and refusal testing. RedEval provides a unified, secure, and extensible platform for assessing LLM robustness against adversarial prompts and harmf...
GuardReasoner-VL
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning Yue Liu, Shengfang Zhai, Mingzhe Du Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi 1National University of Singapore, 2Nanyang Technological University To enhance the safety ...
xmloxide
xmloxide A pure Rust reimplementation of libxml2 — the de facto standard XML/HTML parsing library in the open-source world. libxml2 became officially unmaintained in December 2025 with known security issues. xmloxide aims to be a memory-safe, high-performance replacement that passes the same...
mcpsafetywarden
MCP safety warden is a proxy server that wraps any MCP server and adds behavioral profiling, security scanning, risk gating, and safe execution to its tools...
clr
CLR C hecker of L ifetimes and other R efinement types for Zig Video: https://www.youtube.com/watch?v=mf0WzTOe-40 Sponsorship: https://buymeacoffee.com/dnautics discuss on hn: https://news.ycombinator.com/item?id=42923829 discuss on lobste.rs: https://lobste.rs/s/9sitsj/clrcheckerforlifetimesothe...
ActBench
ActBench ActBench is a self-evolving benchmark of behavioral safety in cowork agents. It defines behavioral safety as whether an agent's execution remains within the permissions and state changes required by a benign task, and evaluates realized behavioral risk from execution trajectories rather...
mlkem-native
mlkem-native mlkem-native is a secure, fast, and portable C901 implementation of ML-KEM2. It is a fork of the ML-KEM reference implementation3. All C code in mlkem/src/ and mlkem/src/fips202/ is proved memory-safe no memory overflow and type-safe no integer overflow using CBMC4. All AArch64 and...
safety
!NOTE Únete a nosotros en SafetyCLI. Estamos contratando para varios puestos. Índice Índice Introducción Características principales Primeros pasos GitHub Action Interfaz de línea de comandos 1. Instalación 2. Iniciar sesión o registrarse...
kalitorify
Transparent Proxy through Tor for Kali Linux About kalitorify kalitorify is a shell script for Kali Linux which use iptables settings to create a Transparent Proxy through the Tor Network , the program also allows you to perform various checks like checking the Tor Exit Node i.e. your public IP...
EUVD-2026-76467
In the Linux kernel, the following vulnerability has been resolved: module: validate string table section types In elfvaliditycachesechdrs, section sizes and offsets are validated, unless the section type is SHTNULL or SHTNOBITS. Later, elfvaliditycachesecstrings and elfvaliditycacheindexstr acce...
CUAHarm
Measuring Harmfulness of Computer-Using Agents 🤗 Hugging Face 📄 Paper 📋 Introduction CUAHarm is a benchmark designed to evaluate the safety risks of Computer-Using Agents CUAs - AI agents that can autonomously control computers to perform multi-step actions. Key Features 🔍 CUAHarm Dataset : A...
CVE-2026-78547
CVE-2026-78547 is an Out-of-Bounds Write (CWE-787) in Citrix Workspace app for Windows , disclosed by Citrix on 2026-09-08. Exploitation requires physical access to the target system and low privileges, with high attack complexity. The vulnerability can lead to high impact on integrity and availa...
CVE-2026-39113
CVE-2026-39113: Heap Buffer Overflow in SQLite's Optional SQLAR Extension Executive Summary CVE-2026-39113 is a heap buffer overflow in SQLite's optional SQLAR extension. In an application that has loaded the extension, an attacker who can invoke sqlaruncompress with a controlled compressed blob...