32 matches found
delirium-ai-safety-benchmark
डेलिरियम AI सुरक्षा बेंचमार्क एफेक्टिव कॉन्टेक्स्चुअल इरोज़न ACE और संबंधित लिमिनल अटैक वेक्टरों के प्रति LLM भेद्यता को मापने के लिए एक नैदानिक ढांचा। डेलिरियम एक शोषण उपकरण नहीं है। यह एक मानकीकृत बेंचमार्क है जिसे उस सटीक क्षण का पता लगाने के लिए डिज़ाइन किया गया है जब भाषा मॉडल का ध्यान भार...
SURE: Framework for Safety to Construct Trustworthy AI
Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains. As the use of these large language models increases, people are increasingly concerned about AI safety and demand that large language...
Will AI kill us all within the next decade?
The Wall Street Journal reports that concerns are rising inside AI labs that competition is pushing tech companies to race toward self-improving models that could spiral out of human control. Jacob Coxon, an AI researcher who has worked at Anthropic and OpenAI, said: “The people building AI...
OpenAI Pauses Frontier RL Training as It Tightens Defenses Against Unsafe AI Behavior
OpenAI on Tuesday revealed that it paused reinforcement learning RL training for its latest artificial intelligence AI models for two weeks while it shored up additional defenses and increased the scope of its monitoring to avert another Hugging Face-like incident. "As models become more capable,...
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards...
LLMs and Contextual Integrity
I have been thinking a lot about AI and integrity. Part of that is contextual integrity. I recently found two papers on the topic. "CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs": Abstract: Large Language Models LLMs increasingly use persistent memory...
If the Markets Reject OpenAI and Anthropic, the US Should Nationalize Them
This essay was written with Nathan E. Sanders, and originally appeared inThe Guardian. OpenAI, and then Anthropic, were each formed by AI developers who feared unrestrained corporate AI development--specifically, that companies like Google and Meta would steer the technology towards deleterious,...
Enhancing AI security through global AI red teaming
In this article 1. Building a global alliance 2. What the research focuses on 3. Why this matters Most AI safety testing still happens inside the walls of individual organizations. That has resulted in a fundamental disconnect: many of the highest-risk failure modes in modern AI systems require...
ChatGPT produced graphic violent images that shocked researchers
AI assistants like ChatGPT are supposed to be safe to use, with appropriate guardrails to stop people creating harmful content. However, a British AI security firm just figured out how to make ChatGPT produce explicit material. Mindgard, a company that tests AI engines for weaknesses, found that ...
FLARE-AI: Flaw Reporting for AI
Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify flaws often do not know what or where to report, and groups who receive reports rarely share them with other relevan...
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses...
Malicious code in intel-ai-safety-explainer (PyPI)
--- -= Per source details. Do not edit below this line.=- Source: kam193 7561bb0b816a4521b6de43bce01afa55516a7201b6daa7696de4924623557f90 Installing the package or importing the module exfiltrates basic information about the host, and the package has no other purpose. --- Category: PROBABLYPENTES...
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
Affordances and permissions are promising and timely safety levers for mitigating Loss of Control LoC threats in high-stakes deployment contexts, such as national security. Deployers in defense and intelligence could rely on several approaches to identify which affordances and permissions should ...
SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations
Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes critical vulnerabilities to prompt injection attacks where...
Whispering poetry at AI can make it break its own rules
Most of the big AI makers don't like people using their models for unsavory activity. Ask one of the mainstream AI models how to make a bomb or create nerve gas and you'll get the standard "I don't help people do harmful things" response. That has spawned a cat-and-mouse game of people who try to...
Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging
Vision Transformers ViTs have emerged as powerful architectures in medical image analysis, excelling in tasks such as disease detection, segmentation, and classification. However, their reliance on large, attention-driven models makes them vulnerable to hardware-level attacks. In this paper, we...
A week in security (October 20 – October 26)
Last week on Malwarebytes Labs: Is AI moving faster than its safety net? Thousands of online stores at risk as SessionReaper attacks spread Apple may have to open its walled garden to outside app stores Meta boosts scam protection on WhatsApp and Messenger Home Depot Halloween phish gives users a...
Replit AI Agent Deletes Sensitive Data Despite Explicit Instructions
Replit AI agent deleted data from 1,200+ executives and companies without permission, raising concerns about AI safety and control in live environments...
Meta AI chatbot bug could have allowed anyone to see private conversations
A researcher has disclosed to TechCrunch that he received a $10,000 bounty for reporting a bug that let anyone access private prompts and responses with the Meta AI chatbot. On June 13, we reported that the Meta AI app publicly exposes user conversations, often without users realizing it. In thes...
Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Cybersecurity, Privacy & AI Safety: a Comprehensive Survey, Roadmap & Implementation Blueprint
Large Language Models LLMs and generative AI GenAI systems such as ChatGPT, Claude, Gemini, LLaMA, and Copilot, developed by OpenAI, Anthropic, Google, Meta, and Microsoft are reshaping digital platforms and app ecosystems while introducing key challenges in cybersecurity, privacy, and platform...