143 matches found
CyberBattleSim
CyberBattleSim April 8th, 2021: See the announcement on the Microsoft Security Blog. CyberBattleSim is an experimentation research platform to investigate the interaction of automated agents operating in a simulated abstract enterprise network environment. The simulation provides a high-level...
Pesidious
Malware Mutation using Deep Reinforcement Learning and GANs The purpose of the tool is to use artificial intelligence to mutate a malware PE32 only sample to bypass AI powered classifiers while keeping its functionality intact. In the past, notable work has been done in this domain with researche...
AutoPentest-DRL
AutoPentest-DRL: Automated Penetration Testing Using Deep Reinforcement Learning AutoPentest-DRL is an automated penetration testing framework based on Deep Reinforcement Learning DRL techniques. AutoPentest-DRL can determine the most appropriate attack path for a given logical network, and can...
GuardReasoner-VL
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning Yue Liu, Shengfang Zhai, Mingzhe Du Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi 1National University of Singapore, 2Nanyang Technological University To enhance the safety ...
PISmith
PIForge An Open Framework for RL-based Prompt Injection Red Teaming 💻 Code · 🤗 Models · 📜 Papers PIForge is the shared codebase for PISmith and Climbing the Hill. PISmith addresses the sparse reward problem in prompt injection red teaming, helping RL attackers explore and learn from rare successf...
StealthRL
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors Paper arXiv Demo Model Hugging Face Benchmark Dataset Hugging Face Abstract AI-text detectors are increasingly used in high-stakes settings, yet their robustness to meaning-preserving adversarial...
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on...
Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift
Quantum key distribution QKD links are provisioned from security analyses of stationary channels, whereas the devices that determine the channel drift between recalibrations. Whether an eavesdropper who cannot alter the channel's own noise gains by following that drift has not been quantified...
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning under Fully Homomorphic Encryption Constraints
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption FHE offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applyin...
CVE-2026-103270
LightLLM through 1.2.0 mounts reinforcement learning control routes on the public HTTP API without authentication checks. Unauthenticated attackers can call endpoints like /pausegeneration, /abortrequest, /flushcache, and /initweightsupdategroup to disrupt inference operations and wedge workers o...
CVE-2026-103270: Missing Authentication for Critical Function
LightLLM through 1.2.0 mounts reinforcement learning control routes on the public HTTP API without authentication checks. Unauthenticated attackers can call endpoints like /pausegeneration, /abortrequest, /flushcache, and /initweightsupdategroup to disrupt inference operations and wedge workers o...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
SecureVibe: Making Vibe Coding More Secure
As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective...
Distillation Defenses Easily Break after Reinforcement Learning
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train i.e., "distill" their own model...
AI-Based Vulnerability Assessment Capability and Cyber Attack Graph Analysis
Cyber threats targeting mission-critical infrastructure are becoming more sophisticated while the barrier to launching attacks continues to fall. Traditional point solutions like antivirus and firewalls are reactive and fail to address the combinatorial complexity of modern attack surfaces. This...
RAISE: Reinforcing Access Control Policy Synthesis in LLMs Via Symbolic Evaluation
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first...
Climbing the Hill: Prompt Injection Red-Teaming against Frontier Models with Curriculum Reinforcement Learning
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning RL to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such ...
JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
Models trained with reinforcement learning for calibrated decisions RLCD, such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial...
A Cyber Range Evaluation of Autonomous Network Incident Response Agents
We test the performance of agents for automated network intrusion response in a cyber range intended for human operator training. The range implements an emulated networking environment with a variable network topology, red-team emulation and simulated user agents. The goal of the defensive agent...
The MAL Simulator: Cyber Operations Simulation Based on Attack and Defense Graphs
We have developed the MAL Simulator, a cyber operation simulator based on the Meta Attack Language MAL. The MAL Simulator is intended for decision-driven cyber attack and defense simulations, for system analysis and the development of automated agents. By building the simulator around an attack...