8683 matches found
Multi-View Fusion for Encrypted C2 Detection: A Leakage-Controlled Measurement Study of Evaluation Pitfalls
Command-and-control C2 traffic increasingly hides within TLS, so defenders now apply machine learning to traffic metadata. Many studies assume that combining two metadata views, namely flow statistics and TLS handshake fingerprints, improves both accuracy and robustness. We tested this assumption...
The like Trap: Multi-Stage Poisoning against Agents in Similarity-Based Recommendation Systems
With recent advancements in large language models LLMs and LLM-based agents, these agents are becoming increasingly autonomous and gaining broader access to act on users' behalf on the internet. However, the vulnerability of automated agents deployed on social media platforms e.g., for managing a...
Agentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking
IC3 is a state-of-the-art algorithm for hardware model checking that proves safety properties by incrementally constructing an inductive invariant consisting of a set of lemmas. Its effectiveness depends on generalization heuristics that identify useful lemmas and guide proof search. However, man...
Stage-Supervised Latent Reasoning for Single-Shot JavaScript Deobfuscation
JavaScript obfuscation is widely used to protect code, but it also makes program analysis and security review substantially harder. Existing LLM-based deobfuscation methods usually treat the task as one-step translation, ignoring the staged structure of practical deobfuscation pipelines. This WIP...
Topological Signatures of Cyber-Attack Classes in Natural Visibility Graph Representations of Network Traffic
Natural Visibility Graph NVG-based representations provide a promising approach for capturing structural patterns in sequential network traffic. However, whether different cyber-attack classes exhibit distinctive topological signatures in such representations remains insufficiently understood. Th...
ACTS: A Multi-Tier Benchmark Evaluating LLM Cipher Identification under Controlled Blind Conditions
We introduce ACTS Artifacts in Cipher Testing Suite, a reproducible benchmark that isolates cryptanalytic ability through tiered metadata deprivation Tier-1: full metadata; Tier-2: filename only; Tier-3: completely blind and tests forced reasoning Tier-5: chain-of-thought, code-as-reasoning,...
Backdoors in Learning-Based Industrial Robotic Arm Manipulation: An Empirical Security Study
Learning-based models e.g., visuomotor and Vision-Language-Action VLA are increasingly explored for industrial robotic manipulation, where model predictions are directly translated into physical actions. This tight coupling between model behavior and physical execution makes hidden security...
Reliable Federated TinyML Deployment for IoT Security
The growing deployment of Internet of Things IoT devices has increased the need for privacy-preserving intrusion detection systems that operate directly on resource-constrained hardware. Federated Learning enables collaborative model training without sharing raw data, but conventional federated...
Privacy Leakage through AI-Mediated Analysis of Smartphone Data
Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text da...
Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
Large language models LLMs are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function...
Decoding the Legalese: A Scalable and Quantitative Framework for Analyzing Corporate Privacy Policies
Even though privacy policies are the primary mechanism organizations use to disclose how they collect, process, and share personal data, they are difficult for average users to interpret, perhaps by design, due to their verbosity and dense legal language. Importantly, there is a lack of...
Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-Of-Thought in Frontier Models
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to...
Rouxii: Exploiting Honeypots with Deception-Aware AI Pentesters
Honeypots are designed to deceive attackers, and recent work shows they can also derail autonomous LLM-based pentesters. These evaluations, however, largely consider attackers unaware of the deception they face. We study the opposite setting: an autonomous attacker explicitly equipped to recogniz...
HYDRA: Proactive Android Malware Drift Adaptation Via Hierarchical Graph Contrastive Learning
Concept drift, driven by the rapid evolution of Android malware, severely degrades the performance of machine learning detectors. Current adaptation strategies are often reactive, responding only after performance has dropped and imposing a significant manual annotation burden, or they are...
Design and Evaluation of a Controlled Post-Alert Incident Orchestration and Response Subsystem Using a Rule Engine and a Local Large Language Model
This paper presents a controlled post-alert incident orchestration and response subsystem for educational information systems. The architecture separates deterministic classification, contextual analysis, human approval, and technical execution. A Rule Engine determines severity and selects the...
On the Security and Privacy of LLMs in Mobility
The mobility sector is undergoing a paradigm shift driven by advances in Generative Artificial Intelligence. With a global market valued at approximately 2.9 trillion dollars annually, considering only cars, the integration of these technologies has the potential to impact more than 1.5 billion...
Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-In-Depth, and the Case for Human-AI Collaboration
Cybersecurity literature has extensively documented the operational benefits of artificial intelligence AI for threat detection, incident response, and prevention, while raising qualitative concerns about over-automation, algorithmic bias, and analyst-skill erosion. What remains largely absent is...
Check Point Security Management R82.20 Compromise Checks
These shell scripts check Check Point Security Management logs and temporary directories for indicators associated with CVE-2026-93616...
WordPress 7.1.0 Comment Cross Site Scripting
WordPress versions through 7.1.0 contain an unauthenticated stored cross site scripting vulnerability in comment processing. The supplied proof of concept safely identifies potentially affected installations by detecting publicly exposed version information...
Confidence-Guided Cross-Modal Knowledge Transfer for Multimodal Anomaly Detection in Microservice Systems
Accurate anomaly detection is essential for reliable and secure operations of microservice systems. While an increasing number of studies have shifted from unimodal modeling to multimodal interaction and fusion, effectively leveraging reliable cross-modal information remains challenging. The...
On the Construction of Trapdoor Claw-Free Functions with Certifiable Key
Trapdoor claw-free functions TCFs underpin much of classical-quantum cryptographic interaction, yet every TCF-based protocol states its guarantees relative to an honestly generated key. We give a family-agnostic abstraction of key certification for noisy TCF constructions, built on two notions: a...
Secure ISAC with Sensing Privacy under Eavesdropper Uncertainty
This paper investigates the joint protection of confidential data and legitimate-user directional information in integrated sensing and communication ISAC networks. We consider a multiuser downlink in which passive multi-antenna eavesdroppers Eves attempt to decode confidential signals while...
Adaptive Traffic Camouflage: Causal and Resource-Aware Defense against IoT Fingerprinting
Encryption hides IoT payloads, but traffic shape can still reveal device identity through packet sizes, timing, direction, and packetization. We present Adaptive Traffic Camouflage, a causal, leakage-aware controller that characterizes traffic-shape leakage without runtime device labels and selec...
How It's Made: Uncovering Detection Engineering Processes for Network Intrusion Detection Rules
Many Security Operations Centers rely on signature-based Network Intrusion Detection Systems like Suricata, yet detection rule engineering remains understudied. We investigate this process by introducing SuriCap, a platform for rule engineering exercises, and hosting CTF-style workshops where 60...
IronCurtain 0.14.0
IronCurtain is an early-stage research project exploring how to make AI agents safe enough to be genuinely useful. It is a runtime for autonomous AI agents, where security policy is derived from a human-readable constitution. APIs, configuration formats, and architecture may change...
tcpdump 4.99.7
tcpdump allows you to dump the traffic on a network. It can be used to print out the headers and/or contents of packets on a network interface that matches a given expression. You can use this tool to track down network problems, to detect many attacks, or to monitor the network activities...
Falco 0.45.0
Sysdig Falco is a behavioral activity monitoring agent that is open source and comes with native support for containers. Falco lets you define highly granular rules to check for activities involving file and network activity, process execution, IPC, and much more, using a flexible syntax. Falco...
Joern 4.0.633
Joern is the bug hunter's workbench. With this tool, you can uncover attack surface, sloppy coding practices, and variants of known vulnerabilities using an interactive code analysis shell. Joern supports C, C++, LLVM bitcode, x86 binaries via Ghidra, JVM bytecode via Soot, and Javascript...
Specter NFC Reader / Skimmer Bug-Sweep 3.0.1
Specter turns your Flipper Zero into a pocket counter-surveillance bug-sweep for active 13.56 MHz NFC readers - a hidden card skimmer slipped into a payment terminal, a covert reader behind a door panel, a rogue logger taped under a desk. It passively senses the RF carrier that any powered-on...
Formally Modeling the Terrapin Attack on SSH
The Terrapin attack against SSH channel integrity USENIX Security 2024 used a novel attack vector: attacks on the channel state. Surprisingly, not all AEAD modes of SSH were equally affected by this attack, and it remained an open question if "unaffected" meant "secure". Existing formal models fo...
COBRA: A Content-Agnostic Framework for Zero-Day Detection of Suspicious Domains
The use of malicious domains is central to cyberattacks such as phishing, malware distribution, impersonation, and fraudulent transactions. Because domains are inexpensive to register and easy to deploy at scale, they remain one of the most common and damaging tools used in cybercrime across...
Unlocking Cross-Scenario Physical Layer Security: A Mixture-Of-Experts Framework with Generative Diffusion Models
The future 6G networks are expected to incorporate a proliferation of wireless services in diverse environments, which presents a significant challenge for information security. Conventionally optimization always requires recalculation and learning strategy often suffers poor generalization, whic...
Solidity Meets LLMs: A Transformer-Based Approach to Smart Contract Vulnerability Detection
The growing adoption of blockchain technologies, particularly the Ethereum platform, has amplified the critical role of smart contracts in decentralized applications. However, the increasing complexity and financial value of these contracts make them prime targets for cyber attacks. In this work,...
Divide and Doubt: Diverse Distributed Poisoning for Retrieval-Augmented Generation
Multi-passage corpus poisoning often repeats one target claim across similar documents, creating correlated lexical and semantic patterns that similarity- and conflict-aware defenses can suppress jointly. We introduce DnD Divide and Doubt, a targeted attack based on two principles: distributing...
Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection
Web applications are increasingly targeted by cyberattacks that exploit HTTP requests to evade security mechanisms. Traditional web application firewalls WAFs rely on rule-based approaches that often exhibit high false positive rates and limited adaptability. Recent studies have explored machine...
Kubernetes Misconfigurations in the Wild: Taxonomy, Evolution, and Automated Repair with Large Language Models
Kubernetes is widely used to orchestrate cloud-native applications, yet its declarative configuration model often introduces security misconfigurations that threaten system reliability. Despite available detection tools, misconfiguration patterns and scalable remediation remain insufficiently...
Ajar: Measuring Open Privilege in Agent Defenses
A language model agent acts through the tools it is given. The data it reads while working on a task can redirect what it does with those tools. A growing set of techniques for safe and secure agent execution therefore sits between the agent and its tools, aiming to enforce access control,...
CISA: CVE Program: Establishing a Quality Era Framework
The Common Vulnerabilities and Exposures CVE Program is the global standard for identifying and cataloging publicly disclosed cybersecurity vulnerabilities. Through a collaborative, federated model and strong community engagement the program has grown to support the timely identification and...
C-To-Rust Fallacy: Automatic Refactoring != Memory Security
Rust has emerged as the leading system programming language, offering strong memory and type safety guarantees without compromising performance. This positions it as a compelling alternative to traditional languages like C and C++, which are susceptible to memory security bugs. However, manually...
SR-Fraud: An Outcome-Supervised Reflective LLM Agent Framework for Non-Stationary Payment Fraud Detection
Real-time payment fraud detection is a non-stationary streaming prediction problem: adversaries adapt before supervised labels mature, and localized burst attacks can cause losses before retraining. Production systems typically rely on tabular classifiers and rules, which can struggle to capture...
Mitigating Front-Running Attacks through Fair and Resilient Transaction Dissemination
In modern blockchains, efficient, fair, and faulttolerant information dissemination is critical for performance and security. Several stages of the transaction lifecycle are affected, from the creation and dissemination of transactions to the dissemination of blocks in the consensus layer. Mempoo...
IronCurtain memory-mcp-server/v0.2.1
IronCurtain is an early-stage research project exploring how to make AI agents safe enough to be genuinely useful. It is a runtime for autonomous AI agents, where security policy is derived from a human-readable constitution. APIs, configuration formats, and architecture may change...
Evaluating Coding Agents on Kernel Exploit Generation
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels...
Partition-Matched Evaluation of Community Features under Distribution Shift in Android Malware Function-Call Graphs
Graph-based Android malware classifiers can lose accuracy under malware-type or family shifts. We test whether mesoscopic organization in function-call graphs provides shift-stable information beyond local degree profiles LDP, global statistics, lightweight metadata, and size-matched random...
Key Reconciliation with RC-LDPC/Error Estimation for Satellite-Based FSO/QKD Systems
Satellite-based free-space optics FSO quantum key distribution QKD systems have recently attracted significant research interest due to their potential to enable globally secured applications. However, the inherent uncertainty of FSO channels, caused by weather conditions and satellite mobility,...
Rethinking Backdoor Repair Evaluation: Distinguishing Aggregate Clean Utility from Benign Performance Preservation
Backdoor repair aims to suppress malicious behavior in compromised models while preserving benign task performance. Existing studies typically evaluate these objectives using Attack Success Rate ASR and Overall Clean Accuracy, but aggregate clean accuracy can obscure substantial degradation...
Dynamic Conformance Testing of WebGPU through Specification-Driven Mutation
WebGPU is a low-level graphics and compute API that exposes modern GPU functionality to web applications. While the official WebGPU Conformance Test Suite CTS focuses on well-formed usage under the WebGPU specification, it is not designed to stress implementations with semantic edge cases or...
RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-Specific Refutation
Retrieval augmented generation RAG systems have emerged as the dominant architecture for grounding large language model LLM outputs in verifiable external knowledge, yet their structural reliance on a dynamic retrieval pipeline introduces a largely unexplored class of adversarial vulnerability...
Quantum ROP: Using Quantum Algorithms for ROP Chain Selection in Exploit Construction
The quantum computing threat to cybersecurity is nowadays predominantly framed around Shor's algorithm and its eventual capacity to break asymmetric cryptography. Beyond cryptanalysis, however, quantum computing may also enable other capabilities in offensive security. This work explores one such...
SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation
Evaluation of large language models LLMs for safety, security, and privacy SSP relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on...