8598 matches found
AuraForge: Scaling Security Supervision for Training Coding Agents
Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a...
No One Architecture Fits All: A Cross-Environment Evaluation of Hierarchical Red Team Agents
Autonomous red team agents increasingly stress-test AI-enabled cyber defenses by planning strategy and executing multistage attacks. Reinforcement learning RL and large language models LLMs offer complementary mechanisms for the planning and execution such agents require, and prior work has...
angr 10.0.1
angr is an open-source binary analysis platform for Python. It combines both static and dynamic symbolic "concolic" analysis, providing tools to solve a variety of tasks...
Karma Pro 0.12
Karma Pro is an open source code review tool written in Swift that can assist code reviewers with a multitude of useful tools. Karma Pro is a macOS source-code security scanner AST base and Heuristics that statically analyses projects in multiple languages. It's backed by an ML classifier trained...
OpenSSL Toolkit 3.5.9
OpenSSL is a robust, fully featured Open Source toolkit implementing the Secure Sockets Layer and Transport Layer Security protocols with full-strength cryptography world-wide. This is the 3.6 release...
Security Properties of Neural Networks As Decision Problems
Certifying a deployed neural network raises decision problems that the verification literature has not classified: whether the model carries a backdoor planted in its training data, whether a fault in its stored parameters can drive it into an unsafe state, whether its output leaks a private part...
MediaWiki Wikibase Extension 1.46.0 Cross Site Scripting
MediaWiki Wikibase Extension versions before 1.46.1, 1.45.5, and 1.43.10 contain a reflected cross site scripting vulnerability in Special:SetLabel language validation. The supplied proof of concept submits a harmless marker and detects unescaped reflection...
Avesis 202608201331 Open Redirect
Avesis version 202608201331 and earlier contains an open redirect vulnerability in the /home/changeculture language-switch endpoint via the returnUrl parameter...
Verifiable Quantum Advantage Based on Polynomials with Planted Structures
A central question in the theory of quantum advantage is whether there are quantum advantage protocols with similar resource requirements as random circuit sampling that are also verifiable just from the classical outputs of the quantum computation. Here, we develop the idea of simulation secrets...
Cybersecurity in Edge Computing: A Trust-Aware Federated Hybrid Intrusion Detection Framework
Edge computing has emerged as a critical computing paradigm in modern distributed systems by migrating data processing closer to end users and Internet of Things IoT devices. While this paradigm decentralizes processes, minimizes latency, and reduces backhaul bandwidth congestion, it exponentiall...
Compression Footprints As Security Signals for Model-Poisoning Defense in Federated Learning
Lossy compression is widely used in Federated Learning FL but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's response as a security signal: the input-dependent distortion and payload behavior...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
Robust detection of AI-generated text under deployment conditions is challenging: distribution shifts across domains and generators, adversarial perturbations of the input surface, and the absence of target-domain labels for threshold calibration all degrade detectors that perform well in-domain...
On the Simplest Quantum-Secure Block Cipher
Pseudorandom permutations are ubiquitous in theoretical and applied cryptography. PRPs that offer security even against adversaries making quantum queries are of increasing interest, and used in applications ranging from constructing pseudorandom unitaries to separating SZK from BQP. A successful...
MADBench: Benchmarking the Security of Multi-Agent Debate
Multi-agent debate MAD can improve large language model LLM reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect...
Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents
LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authori...
APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry
Large language model LLM agents could help security operations centers SOCs investigate advanced persistent threats APTs by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection,...
ActionGuard: Tool Call Authorization under Poisoned Skills
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data...
Aletheia: Permission-Minimality Testing for Coding-Agent Rules
Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia translates requested...
Link Inference Attack on Privacy-Preserving Knowledge Graphs
Knowledge Graphs KGs are widely used to store and share structured information across sensitive domains such as healthcare, fi- nance, and social networks. A common privacy practice is to delete sen- sitive relations before publishing the graph, under the assumption that removing edges is...
Xalgorix Autonomous AI Pentesting Agent 4.6.124
Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported - so you get proof, not a pile of maybes to triage. Self-hosted, private, and bring-your-own-LLM. Built in Go + TypeScript...
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models Via Structured Code Completion
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained predominantly on natural...
Path-Finding, Orbit State Preparation, and the Security of Invariant Quantum Money
The security of quantum money from knots, and of its generalization to invariant money, is based on the assumption that path-finding, exhibiting a sequence of moves between two equivalent objects, is hard. No proof of security from that assumption alone is known. The existing proofs add...
OpenSSL Toolkit 4.0.3
OpenSSL is a robust, fully featured Open Source toolkit implementing the Secure Sockets Layer and Transport Layer Security protocols with full-strength cryptography world-wide. This is the 3.6 release...
Do Defenses against LLM Extraction Work across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Large language models LLMs deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access...
Learning Normal Diffusion Dynamics for Backdoor Defense in Text-To-Image Models
Backdoor attacks pose a serious threat to the secure deployment of text-to-image T2I diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack...
Joern 4.0.643
Joern is the bug hunter's workbench. With this tool, you can uncover attack surface, sloppy coding practices, and variants of known vulnerabilities using an interactive code analysis shell. Joern supports C, C++, LLVM bitcode, x86 binaries via Ghidra, JVM bytecode via Soot, and Javascript...
Natural Barriers to Quantum Extraction: On the Post-Quantum (In)Security of (O)EKE and Masny-Rindal OT
Encrypted key exchange EKE, introduced by Bellovin and Merritt IEEE S&P 1992, and Masny-Rindal OT, introduced by Masny and Rindal ACM CCS 2019, are highly-efficient methods for compiling essentially any KEM into advanced cryptographic protocols, namely password-authenticated key exchange PAKE and...
From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model
Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a natural channel for indirect prompt injection. Existing agent...
Helol Tunnel: Covert Channel Exploitation of TLS Extensibility and Privacy Features
Covert channels exploiting network protocols for data exfiltration and command-and-control C2 are integral parts of modern cyberattacks. In search of a significant covert channel within the fabric of the Internet, we targeted the combinatorial properties of the Client Hello CHLO packets in the...
Breaking the Bounded Entanglement Barrier for Quantum Position Verification
Position verification, introduced by Chandran et al. SIAM J. Computing 2014, allows verifiers to test a prover's claimed position by an interactive protocol. Classical position verification is impossible. Even for Quantum Position Verification QPV, there always exists an LOCC Local Operations and...
On Evaluating Quantum Kernel Robustness for Low-Resource Cross-Corpus Audio Deepfake Detection
Synthetic speech detection is critical for audio security, but performance can degrade when labeled data are scarce and evaluation conditions differ from training. This study examines quantum kernel methods and lightweight neural models for cross-corpus audio deepfake detection under limited...
Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
Recently, Vision-Language-Action VLA models have revolutionized robotic manipulation by seamlessly integrating visual perception, language understanding, and action generation in an end-to-end learning framework. However, since these models are designed to interact directly with the physical worl...
Sapien: A Stateful Policy Engine for Autonomous AI Agents
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a poli...
Iron Mountain enVision 260654 Authenticated SQL Injection
Iron Mountain enVision versions before 260655 contain an authenticated SQL injection vulnerability in the DocumentModule document-listing and detailed-search sorting parameter...
OpenSSL Toolkit 3.6.5
OpenSSL is a robust, fully featured Open Source toolkit implementing the Secure Sockets Layer and Transport Layer Security protocols with full-strength cryptography world-wide. This is the 3.6 release...
Faithful Dual-Constrained Erasure for Robust LLM Safety Alignment
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models LLMs. However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicio...
Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents
Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The...
Camaleon CMS 2.9.2 Authorization Bypass
Camaleon CMS versions through 2.9.2 contain an authorization bypass in the Media Crop Handler that can allow a media-management user to set another same-site user's avatar via the savedavatar parameter...
ZoneClaw: Mitigating Persistent Memory Attacks by Establishing Memory-Zoning in OpenClaw-Style Computer-Use Agents
Computer-use agents increasingly operate as long-running assistants through persistent workspace memory, which OpenClaw-style CUAs realize as automatically reloaded files that hold user instructions, system summaries, and external claims at the same privilege level. Here, remembering a claim...
Semi-Quantum Cryptography with Certified Deletion
Certified deletion allows a client to upload encrypted data to a server as a quantum state, then later request that the server delete their data and detect whether the server complies. If verification passes, then the data on the server will remain hidden even if the decryption key is later leake...
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show th...
Certified Randomness with Optimal Rate
The generation of certified random bits is an emerging near-term application of quantum computers. Potential applications, such as randomness beacons and CRS generation, require nearly uniform randomness, whose rate the ratio of min-entropy to bitlength is 1. However, existing protocols for...
Backdoor Containment Via Expert Quarantine and Shutdown in LLMs
Backdoored large language models LLMs can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either...
ABSENTIA: Detecting Broken Access Control Vulnerabilities in Web Applications
Broken access control, the failure of authorization, is one of the most prevalent web security risks. Unlike injection, a flow of untrusted input into a dangerous operation, authorization is a relation: who may act on what, not how data moves. Each application decides that relation for itself, so...
Selective Channel Restoration for Backdoored Vision-Language Models
Vision-language models VLMs exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these...
ModalFidelity: Routing Modalities for Deepfake Detection on a Budget
Deepfakes no longer need to fake a whole video. Generators that read the transcript now alter only the few seconds in which a video's meaning turns, so a forgery hides in a small, unknown fraction of the video. Yet detectors still read every one-second window of both the audio and image streams,...
A Competing-Hazards Systematization of Loss of Control in Autonomous Agents
Leading AI developers have reported agents acting beyond their approved limits, which a United Nations panel described as an early warning of loss of human control. Yet incident reports and agent-safety evaluations describe these events differently, making it difficult to compare failures, trace...
Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
When a prompt injection attack succeeds, a Large Language Model LLM abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually...
The Geometry of Harmfulness in Multi-Turn Attacks
Large language models LLMs remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-tur...