23505 matches found
heretic
Heretic: Eliminación totalmente automática de la censura para modelos de lenguaje...
AutoRAN-public
🧠 AutoRAN: Secuestro automatizado del razonamiento de seguridad en grandes modelos de razonamiento AutoRAN es un secuestro automatizado del razonamiento de seguridad que aprovecha modelos auxiliares secundarios menos alineados para simular trazas de razonamiento, generar prompts narrativos y...
CVE-2026-92238
CVE-2026-92238 affects Thunderbird , the Mozilla email client. The vulnerability involves ambiguous parsing of mail headers , where a maliciously constructed header could cause multiple fields to be parsed as a single field, or lead to potential memory safety violations . The issue was fixed in T...
EUVD-2026-79242
A maliciously constructed mail header could lead to multiple fields being parsed as one, or potential memory safety violations. This vulnerability was fixed in Thunderbird 156 and Thunderbird 140.16...
CVE-2026-92238 Ambiguous parsing of mail headers
A maliciously constructed mail header could lead to multiple fields being parsed as one, or potential memory safety violations. This vulnerability was fixed in Thunderbird 156 and Thunderbird 140.16...
saffron
Saffron-1: Inference Scaling for LLM Safety Assurance 📖 Paper 🛠️ Dependencies The code was tested under the following dependencies: Python 3.12.3 CUDA 12.2 typingextensions==4.14.0 numpy==2.2.6 torch==2.5.1 huggingfacehub==0.30.2 accelerate==1.1.1 datasets==3.1.0 evaluate==0.4.3...
redeval
RedEval - LLM Safety Evaluation Framework A comprehensive framework for evaluating the safety of Large Language Models LLMs through systematic attack and refusal testing. RedEval provides a unified, secure, and extensible platform for assessing LLM robustness against adversarial prompts and harmf...
autoguardrails
autoguardrails Código abierto por Santander AI Lab. Una biblioteca / arnés de evaluación de investigación en seguridad de LLM / IA estilo autoresearch: busca sobre una única superficie mutable policy.md para minimizar la tasa de éxito de ataques ASR frente a un conjunto de evaluación fijo, con un...
clr
CLR C hecker of L ifetimes and other R efinement types for Zig Video: https://www.youtube.com/watch?v=mf0WzTOe-40 Sponsorship: https://buymeacoffee.com/dnautics discuss on hn: https://news.ycombinator.com/item?id=42923829 discuss on lobste.rs: https://lobste.rs/s/9sitsj/clrcheckerforlifetimesothe...
GuardReasoner-VL
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning Yue Liu, Shengfang Zhai, Mingzhe Du Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi 1National University of Singapore, 2Nanyang Technological University To enhance the safety ...
mcpsafetywarden
MCP safety warden 是一个代理服务器,它包装任何 MCP 服务器,并为其工具添加行为分析、安全扫描、风险门控和安全执行。 目录 概述 前提条件 安装 配置 MCP 集成 CLI 参考 辅助安全工具集成 开发 测试 进一步阅读 !IMPORTANT MCP 安全是一个活跃的研究领域。最近的调查列举了许多特定于协议的威胁类别,涵盖工具投毒、提示注入、rug-pull 攻击、供应链破坏、凭据泄露以及整个服务器生命周期中的组合攻击。参见 保护 MCP OpenReview,现状与威胁 arXiv,当 MCP 服务器攻击 arXiv,以及 MCP...
xmloxide
xmloxide 一个纯 Rust 实现的 libxml2 —— 开源世界中事实标准的 XML/HTML 解析库。 libxml2 于 2025 年 12 月正式停止维护,并存在已知的安全问题。xmloxide 旨在成为一个内存安全、高性能的替代方案,并通过相同的符合性测试套件。 特性 内存安全 —— 基于竞技场的树结构,公共 API 中零 unsafe 符合标准 —— W3C XML 符合性测试套件 100% 通过率(1727/1727 个适用测试) 错误恢复 —— 解析格式错误的 XML 仍能生成可用的树,与 libxml2 类似 多解析 API —— DOM 树、SAX2...
CVE-2026-92024
Use-after-free in the SVG component. This vulnerability was fixed in F...
EUVD-2026-78487
Use-after-free in the SVG component. This vulnerability was fixed in Firefox 156, Firefox ESR 115.41, Firefox ESR 140.16, Firefox ESR 153.3, Thunderbird 156, and Thunderbird 140.16...
CVE-2026-92024
Use-after-free in the SVG component. This vulnerability was fixed in Firefox 156, Firefox ESR 115.41, Firefox ESR 140.16, Firefox ESR 153.3, Thunderbird 156, and Thunderbird 140.16...
CVE-2026-92016
Use-after-free in the Disability Access APIs component. This vulnerability was fixed in Firefox 156, Firefox ESR 140.16, Firefox ESR 153.3, Thunderbird 156, and Thunderbird 140.16...
nano-analyzer
Nano-analyzer 由AISLE 提供的极简 LLM 驱动的零日漏洞扫描器。 仅为演示目的的研究原型。 这是一个简单的单文件工具,能够检测真实的零日漏洞。请注意,这是一个原型,偏向于 C/C++ 内存安全漏洞,并且会产生误报。我们本着开放研究的精神原样分享它——请留意其中可能存在的尖锐问题。 功能说明 Nano-analyzer 是一个简单的单文件 Python 扫描器,它将源代码通过一个三阶段 LLM 流水线处理: 1. 上下文生成 — 模型为文件撰写安全简报:它的功能、不受信任数据的流向、存在的缓冲区及其大小。 2. 漏洞扫描 —...
ActBench
ActBench ActBench is a self-evolving benchmark of behavioral safety in cowork agents. It defines behavioral safety as whether an agent's execution remains within the permissions and state changes required by a benign task, and evaluates realized behavioral risk from execution trajectories rather...
meta-ai-support-prompt
Prompt del sistema del Asistente de Soporte de IA de Meta Prompt del sistema extraído del Asistente de Soporte de IA de Meta el 1 de junio de 2026. Archivos system-prompt.md — Prompt del sistema extraído ⚠️ Descargo de responsabilidad y aviso legal Propósito Este repositorio se publica...
gemini-2.5-pro-nf-tables-red-teaming
Caso de Estudio de Red Teaming de Gemini 2.5 Pro sobre nftables CVE-2023-32233 Investigación de Seguridad de LLM | Divulgación Responsable | Evaluación de Alineación de IA Investigador : Niranj R Mahaswar Destawell Programa de Recompensas por Vulnerabilidades de Google AI : 889286 — Fuera de...