3 matches found
autoguardrails
autoguardrails Santander AI Lab의 오픈 소스입니다. LLM/AI 안전 가드레일 연구 라이브러리/평가 하네스autoresearch 스타일입니다: 단일 변경 가능한 policy.md 표면을 검색하여 고정된 평가 스위트에 대한 공격 성공률ASR을 최소화하고, 무해 통과 기준benign-pass floor이 있습니다. Santander AI Open Source의 일부 — Banco Santander의 오픈 소스 AI 프로젝트 santander.com입니다. autoguardrails는 Karpathy의...
DREAM: Dynamic Red-Teaming across Environments for AI Models
Large Language Models LLMs are increasingly used in agentic systems, where their interactions with diverse tools and environments create complex, multi-stage safety challenges. However, existing benchmarks mostly rely on static, single-turn assessments that miss vulnerabilities from adaptive,...
The vulnerability of the ACL-policy search mechanism based on application prefixing by the Nomad orchestrator allows attackers to bypass existing security mechanisms.
The vulnerability of the ACL-policy-based search mechanism of the Nomad application lies in the improper assignment of access control rules. Exploiting this vulnerability allows a malicious actor to bypass existing security mechanisms by creating tasks with special prefix names...