Lucene search
+L

32 matches found

Kitploit
Kitploit
•added 2026/10/08 12:11 p.m.•12 views

delirium-ai-safety-benchmark

डेलिरियम AI सुरक्षा बेंचमार्क एफेक्टिव कॉन्टेक्स्चुअल इरोज़न ACE और संबंधित लिमिनल अटैक वेक्टरों के प्रति LLM भेद्यता को मापने के लिए एक नैदानिक ढांचा। डेलिरियम एक शोषण उपकरण नहीं है। यह एक मानकीकृत बेंचमार्क है जिसे उस सटीक क्षण का पता लगाने के लिए डिज़ाइन किया गया है जब भाषा मॉडल का ध्यान भार...

6.2AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/09/29 12:00 a.m.•9 views

SURE: Framework for Safety to Construct Trustworthy AI

Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains. As the use of these large language models increases, people are increasingly concerned about AI safety and demand that large language...

5.8AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2026/09/10 12:18 p.m.•27 views

Will AI kill us all within the next decade?

The Wall Street Journal reports that concerns are rising inside AI labs that competition is pushing tech companies to race toward self-improving models that could spiral out of human control. Jacob Coxon, an AI researcher who has worked at Anthropic and OpenAI, said: “The people building AI...

5.8AI score
SaveExploits0
The Hacker News
The Hacker News
•added 2026/08/19 6:06 p.m.•27 views

OpenAI Pauses Frontier RL Training as It Tightens Defenses Against Unsafe AI Behavior

OpenAI on Tuesday revealed that it paused reinforcement learning RL training for its latest artificial intelligence AI models for two weeks while it shored up additional defenses and increased the scope of its monitoring to avert another Hugging Face-like incident. "As models become more capable,...

6AI score
SaveExploits0
Wired Threat Level
Wired Threat Level
•added 2026/08/18 6:33 p.m.•65 views

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards...

5.5AI score
SaveExploits0
Schneier on Security
Schneier on Security
•added 2026/08/18 10:40 a.m.•9 views

LLMs and Contextual Integrity

I have been thinking a lot about AI and integrity. Part of that is contextual integrity. I recently found two papers on the topic. "CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs": Abstract: Large Language Models LLMs increasingly use persistent memory...

5.7AI score
SaveExploits0
Schneier on Security
Schneier on Security
•added 2026/08/14 11:03 a.m.•18 views

If the Markets Reject OpenAI and Anthropic, the US Should Nationalize Them

This essay was written with Nathan E. Sanders, and originally appeared inThe Guardian. OpenAI, and then Anthropic, were each formed by AI developers who feared unrestrained corporate AI development--specifically, that companies like Google and Meta would steer the technology towards deleterious,...

5.8AI score
SaveExploits0
Microsoft Secure
Microsoft Secure
•added 2026/07/27 4:25 p.m.•24 views

Enhancing AI security through global AI red teaming

In this article 1. Building a global alliance 2. What the research focuses on 3. Why this matters Most AI safety testing still happens inside the walls of individual organizations. That has resulted in a fundamental disconnect: many of the highest-risk failure modes in modern AI systems require...

5.9AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2026/07/01 9:10 a.m.•39 views

ChatGPT produced graphic violent images that shocked researchers

AI assistants like ChatGPT are supposed to be safe to use, with appropriate guardrails to stop people creating harmful content. However, a British AI security firm just figured out how to make ChatGPT produce explicit material. Mindgard, a company that tests AI engines for weaknesses, found that ...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/30 12:00 a.m.•16 views

FLARE-AI: Flaw Reporting for AI

Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify flaws often do not know what or where to report, and groups who receive reports rarely share them with other relevan...

5.7AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/06/23 12:00 a.m.•27 views

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses...

5.8AI score
SaveExploits0
OSSF Malicious Packages
OSSF Malicious Packages
•added 2026/06/15 5:40 p.m.•27 views

Malicious code in intel-ai-safety-explainer (PyPI)

--- -= Per source details. Do not edit below this line.=- Source: kam193 7561bb0b816a4521b6de43bce01afa55516a7201b6daa7696de4924623557f90 Installing the package or importing the module exfiltrates basic information about the host, and the package has no other purpose. --- Category: PROBABLYPENTES...

5.6AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/05/20 12:00 a.m.•20 views

Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security

Affordances and permissions are promising and timely safety levers for mitigating Loss of Control LoC threats in high-stakes deployment contexts, such as national security. Deployers in defense and intelligence could rely on several approaches to identify which affordances and permissions should ...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2026/01/12 12:00 a.m.•16 views

SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations

Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes critical vulnerabilities to prompt injection attacks where...

7.7AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2025/12/02 2:18 p.m.•20 views

Whispering poetry at AI can make it break its own rules

Most of the big AI makers don't like people using their models for unsavory activity. Ask one of the mainstream AI models how to make a bomb or create nerve gas and you'll get the standard "I don't help people do harmful things" response. That has spawned a cat-and-mouse game of people who try to...

7.5AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/10/28 12:00 a.m.•19 views

Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging

Vision Transformers ViTs have emerged as powerful architectures in medical image analysis, excelling in tasks such as disease detection, segmentation, and classification. However, their reliance on large, attention-driven models makes them vulnerable to hardware-level attacks. In this paper, we...

7.2AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2025/10/27 7:15 a.m.•27 views

A week in security (October 20 – October 26)

Last week on Malwarebytes Labs: Is AI moving faster than its safety net? Thousands of online stores at risk as SessionReaper attacks spread Apple may have to open its walled garden to outside app stores Meta boosts scam protection on WhatsApp and Messenger Home Depot Halloween phish gives users a...

7.2AI score
SaveExploits0
HackRead
HackRead
•added 2025/07/24 4:37 p.m.•13 views

Replit AI Agent Deletes Sensitive Data Despite Explicit Instructions

Replit AI agent deleted data from 1,200+ executives and companies without permission, raising concerns about AI safety and control in live environments...

7.5AI score
SaveExploits0
Malwarebytes
Malwarebytes
•added 2025/07/17 2:03 p.m.•20 views

Meta AI chatbot bug could have allowed anyone to see private conversations

A researcher has disclosed to TechCrunch that he received a $10,000 bounty for reporting a bug that let anyone access private prompts and responses with the Meta AI chatbot. On June 13, we reported that the Meta AI app publicly exposes user conversations, often without users realizing it. In thes...

6.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/06/22 12:00 a.m.•18 views

Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Cybersecurity, Privacy & AI Safety: a Comprehensive Survey, Roadmap & Implementation Blueprint

Large Language Models LLMs and generative AI GenAI systems such as ChatGPT, Claude, Gemini, LLaMA, and Copilot, developed by OpenAI, Anthropic, Google, Meta, and Microsoft are reshaping digital platforms and app ecosystems while introducing key challenges in cybersecurity, privacy, and platform...

6.7AI score
SaveExploits0
Rows per page
Query Builder