120 matches found
Ajar
Ajar: Measuring Open Privilege in Agent Defenses An agent-security benchmark reports two numbers, attack success and benign utility, and both are read off runs that happened. Neither says what the defense stood ready to allow on the paths no run took. Ajar asks it directly: for each benign task i...
Enhancing Privacy, Neglecting Harms: An Analysis of Real-World Digital Privacy Incidents
Privacy-enhancing technologies PETs have emerged as a technical means for providing individuals with greater control over their information. Yet despite the growing deployment of PETs, people continue to experience privacy harms. In this work, we revisit our understanding of privacy incidents and...
MAL-2026-14552 Malicious code in mathkitlite (PyPI)
--- -= Per source details. Do not edit below this line.=- Source: kam193 c91a8962fd66464487bcdb5e2c217990cfdea55df22e7166c11703f91dfb7845 Installing the package or importing the module exfiltrates basic information about the host, and the package has no other purpose. --- Category: PROBABLYPENTES...
Meta ordered to pay $942 million over harm to children
A New Mexico court has ordered Meta to pay a total of $942 million after finding that Facebook and Instagram harmed young users and that the company misled consumers about the safety of its platforms. Reportedly, the decision combines a $375 million civil-penalty verdict from March with a newly...
Invisible Ink Threats: Adversarial Goals behind Legitimate Tasks in Computer-Use Agents
Computer-use agents CUAs, which empower large language models to autonomously operate operating systems and the web, are increasingly vulnerable to indirect prompt injection attacks. A widely adopted defense is the human-in-the-loop paradigm, in which the agent pauses for explicit user confirmati...
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Aligned language models refuse harmful requests, but a one-line prefill "Sure, here is" strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones 0.91-0.98, while...
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Safety evaluation of large language models LLMs relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete...
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the information a defender most needs, namely how harmful the resulting action was. We introduce an...
MAL-2026-6314 Malicious code in @zynkit/probe (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector 5f7da2d08466718d66bb9a80478987eab5041e19ee53f80d6968e6dc3069edf2 No concrete indicators of installer-side harm were observed in this package version. No lifecycle scripts performing remote fetch-and-execute, no...
Malicious code in @zynkit/probe (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector 5f7da2d08466718d66bb9a80478987eab5041e19ee53f80d6968e6dc3069edf2 No concrete indicators of installer-side harm were observed in this package version. No lifecycle scripts performing remote fetch-and-execute, no...
Malicious code in @frostnode/waitfor (npm)
@frostnode/waitfor malicious versions 0.9.0, 0.10.3, 0.10.4, and 0.10.5, published by [email protected] is a trojanized npm package belonging to the wshu.net credential-stealer campaign. The campaign published trojanized look-alike utility packages across 12+ scopes whose publisher accoun...
A Red-Team Study of Anthropic Fable 5 and Opus 4.8 Models
We evaluate the adversarial robustness of two frontier large language models LLMs developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm taxonomy. Using the HackAgent red-teaming framework, hundre...
Fourth Frontier Frontier X Mobile Application, Frontier X2
ADVISORY SUMMARY Successful exploitation of this vulnerability could allow an attacker to read and write arbitrary handle values and change clinical readings, which could result in taking control of the device and lead to patient harm. 2. RECOMMENDED PRACTICES CISA recommends users take...
CVE-2026-44410
This vulnerability stems from a business logic flaw.Attackers can exploit legitimate application functions in unintended and abnormal ways, deviating from the designer's expectations, to carry out malicious attacks...
Malicious code in @onerjs/smart-filters-blocks (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector e772d7a844409df378591a5a587c7cc8045e0ec0e8cb493912f0da8fa594c169 This package is published as @onerjs/smart-filters-blocks but its README, repository URL git+https://github.com/BabylonJS/Babylon.js.git, description...
MAL-2026-4629 Malicious code in openmct-couch-plugin (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector ce8eff366d17efa64bf8605941d009d01cf7a24aaf011af30faec449fc4a2e28 On npm install, the package's preinstall script runs node index.js and then curls the output of hostname && whoami to...
MAL-2026-4634 Malicious code in osep-react-antd (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector 9373e8880ad89854cc168b48a36c59bd72abfaf220e08fb751b948f0c4d8ddfb package.json declares preinstall: node index.js, which runs automatically on npm install. index.js collects host identifiers os.hostname,...
Malicious code in moneykit-cardano-demo (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector e6186e5ec8b6cea4f1cec3b4284cf09f2e317dd7d745fb5f88e15b355497d08e package.json declares preinstall: node index.js, which fires automatically on npm install. index.js collects host identifiers and OS files —...
MAL-2026-4445 Malicious code in @signetai/signet-memory-openclaw (npm)
--- -= Per source details. Do not edit below this line.=- Source: amazon-inspector c0ad0b94bf8d4fa9787ee3ae220aa20bc7dde35d2f7a90d2a8e19a8781b8a552 On plugin registration in 'full' mode the package reads the installer's Claude Code OAuth credential from /.claude/.credentials.json and...