Lucene search
+L

7 matches found

Kitploit
Kitploit
•added 2026/10/01 4:24 p.m.•16 views

inspect_petri

Inspect Petri Bienvenido a Inspect Petri, un agente de auditoría que permite la supervisión e interacción automatizada con modelos de lenguaje para detectar posibles problemas de alineación, manipulación de recompensas y otros comportamientos preocupantes. Petri te ayuda a probar rápidamente...

6.2AI score
SaveExploits0References2
The Hacker News
The Hacker News
•added 2026/08/27 6:36 p.m.•17 views

OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face

OpenAI on Wednesday revealed that reward hacking was a key driver behind the artificial intelligence AI-powered hack of Hugging Face last month, adding that it found evidence of misaligned behavior as early as late May. The incident, the company said, took place during cybersecurity evaluations o...

7.8CVSS7.3AI score0.00709EPSS
SaveExploits0
Positive Technologies
Positive Technologies
•added 2026/05/15 12:00 a.m.•22 views

PT-2026-41417

Claude Mythos Preview case studies also, read your transcripts! https://t.co/drNlAH5mLE "Mythos demonstrates its bug reproduction and exploitation capabilities on CVE-2024-051912, an in-the-wild exploited bug that has no public report nor a working PoC whatsoever in the public domain. This bug ha...

5.8AI score
SaveExploits0References1
Packet Storm News
Packet Storm News
•added 2026/05/12 12:00 a.m.•44 views

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a score without performing the intended task, emerges spontaneously in frontier models without overfitting. We argue tha...

5.8AI score
SaveExploits0
Packet Storm News
Packet Storm News
•added 2025/11/20 12:00 a.m.•65 views

Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models

The growing misuse of Vision-Language Models VLMs has led providers to deploy multiple safeguards, including alignment tuning, system prompts, and content moderation. However, the real-world robustness of these defenses against adversarial attacks remains underexplored. We introduce Multi-Faceted...

7.3AI score
SaveExploits0
Schneier on Security
Schneier on Security
•added 2025/08/20 11:02 a.m.•14 views

Subverting AIOps Systems Through Poisoned Input Data

In this input integrity attack against an AI system, researchers were able to fool AIOps tools: AIOps refers to the use of LLM-based agents to gather and analyze application telemetry, including system logs, performance metrics, traces, and alerts, to detect problems and then suggest or carry out...

7.2AI score
SaveExploits0
Schneier on Security
Schneier on Security
•added 2021/04/26 11:06 a.m.•58 views

When AIs Start Hacking

If you dont have enough to worry about already, consider a world where AIs are hackers. Hacking is as old as humanity. We are creative problem solvers. We exploit loopholes, manipulate systems, and strive for more influence, power, and wealth. To date, hacking has exclusively been a human activit...

6.8AI score
SaveExploits0
Rows per page
Query Builder