8740 matches found
GraftyVul: Synthesising Insecure Programs through Real-World Vulnerability Grafting
Vulnerability datasets underpin a wide range of security research, including vulnerability detection, automated remediation, and secure code generation. However, existing datasets sacrifice at least one of three desirable properties: diversity of language or vulnerability type,...
Exploiting Per-Core Leakage: Electromagnetic Side-Channel Monitoring of Multicore Architectures
Multicore processors are increasingly adopted in embedded systems to meet growing performance demands. However, physical side-channel analysis of multicore architectures remains underexplored, as obtaining usable leakage is inherently challenging. Consequently, side-channel security research on...
Joern 4.0.614
Joern is the bug hunter's workbench. With this tool, you can uncover attack surface, sloppy coding practices, and variants of known vulnerabilities using an interactive code analysis shell. Joern supports C, C++, LLVM bitcode, x86 binaries via Ghidra, JVM bytecode via Soot, and Javascript...
Quantum-Based Solutions for Security Enhancement in Open Radio Access Networks
Open Radio Access Networks O-RAN introduce unprecedented flexibility, interoperability, and intelligence into next-generation wireless systems, but their disaggregated and software-defined architecture also expands the attack surface and creates new security vulnerabilities. Conventional...
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-As-Code
Large language models are increasingly used to author Infrastructure-as-Code IaC, where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but without a human baseline they cannot determine whether models...
Managing Inherent Risk: On the Conceptualization of Risk in Defense Systems
Certain defense systems are, by nature, deployed in a civilian environment in order to serve a defensive function for that environment. However, the risk involved with their deployment and operation poses a challenge for the public acceptance of these systems. Compared to safety engineering for...
Enhancing Web Application Firewalls with Machine Learning for SQL Injection Detection
Detecting SQL Injection SQLi attacks ranks among the most critical challenges in web application security. This research conducted a systematic literature review to identify the research gaps in this domain and responsively designed and optimised a DistilBERT-Stacked Ensemble pipeline to improve...
Breaking Darknet CAPTCHAs with General Purpose LLM
Our work evaluates the effectiveness of automated methods for solving CAPTCHA challenges commonly encountered in darknet environments. These CAPTCHAs are typically designed to operate without JavaScript, resulting in distinct characteristics compared to mainstream CAPTCHA systems. Our study...
Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection
The Android ecosystem faces persistent and rapidly evolving malware threats. Existing machine learning detectors are vulnerable to concept drift because they rely on implementation-specific features whose distributions change over time. Large language models LLMs offer strong semantic understandi...
The Impact of Magma: A Ground-Truth Fuzzing Benchmark
Magma is an open-source and ground-truth fuzzing benchmark that enables uniform fuzzer evaluation and comparison. Magma was originally released with a research paper published at ACM SIGMETRICS 2021. This short paper explains the motivation, the design, and the impact of Magma, with a description...
CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?
Prompt injection attacks on Large Language Model LLM agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign scope. While current defenses effectively counter known...
Enhancing Web Application Firewalls with BERT-GNN for SQL Injection Detection
Detecting sophisticated SQL Injection SQLi attacks remains among the most critical challenges in web applications security. This research study has resulted in an optimised hybrid BERT-GNN pipeline with improved detection accuracy and robustness while reducing false-positive and false-negative...
TagZilla: Automated Owner and Abuse Type Tagging for Indicators of Compromise in Threat Reports
Cyber Threat Intelligence CTI reports often describe Indicators of Compromise IoCs such as IP addresses, URLs, file hashes, and cryptocurrency wallets involved in cyberattacks. Those IoCs are typically described in the unstructured report's text, or listed at the end of the report with little...
LongPIBench: A Long-Context Benchmark for Prompt Injection
Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and defenses in long-context settings largely unexplored. This gap leads to a...
REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
To determine the real-world effectiveness of machine learning based malware detection, it is vital to evaluate its robustness against highly capable adversaries. However, state-of-the-art attacks do not effectively model realistic adversaries, as they often assume access to privileged information...
When Verified Source Becomes Attack Input: Defending Smart Contracts against LLM-Based Vulnerability Scanning
Smart contracts are financial programs deployed on blockchains to manage digital assets. To build trust with users and investors, smart contract projects typically publish their source code on blockchain explorers and verify it against the deployed bytecode, making the on-chain program accessible...
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
Software and systems security workflows are typically procedural: analysts inspect heterogeneous artifacts, form hypotheses, invoke tools, interpret outputs, and revise plans. Large language model LLM-based agents, which can plan, use tools, retain state, and revise actions across multi-step...
Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study
Despite extensive safety alignment, large language models LLMs remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts...
LAAF: A Layered Accountability Architecture Framework for LLM Applications
Large Language Models LLMs operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility ...
EBPF-Based Cybersecurity Mechanisms: A Systematic Literature Review
Extended Berkeley Packet Filter eBPF has emerged as a kernel-level framework enabling dynamic security enforcement in modern operating systems. While eBPF's cybersecurity potential has attracted significant attention, existing work remains fragmented across domains, evaluation methodologies, and...
From Security Events to Conflict States: A Three-Layer Cyber Defense Scenario Model for Enhanced Cyber Situational Awareness
Cyber defense in mission-critical environments requires integrated approaches capable of representing adversarial progression, defender-side uncertainty, mission impact, and defensive decision support within a unified framework. In operational domains, defenders must continuously estimate the...
NethServer WebTop 1.5.6 Cross Site Scripting
NethServer WebTop module versions 1.5.6 and below suffer from two persistent cross site scripting vulnerabilities...
Cyber-Electromagnetic Anomaly Detection through Time-Series Analysis
Military operations benefit from the coordination between kinetic and non-kinetic domains. In particular, the coordination of cyber operations and electromagnetic warfare has become increasingly relevant for gaining operational advantage. This coordination is also relevant for Cyber Situational...
Thresholding Post-Quantum Signatures
Threshold signature schemes distribute the signing process among $T$ parties out of $N$. They enable a variety of applications and their research is also motivated by a recent NIST call. However, applications are dominated by pre-quantum signatures, which are more efficient but not secure in the...
Decoupling Is a Necessity: Transformation-Agnostic Decompiled Code Recovery under Optimization and Obfuscation
Reverse engineering is essential for software security analysis and vulnerability detection. Decompilation, the process of lifting binaries to high-level pseudocode, is central to this task. However, production binaries are hostile environments: aggressive compiler optimizations and adversarial...
Metamorphism: A Mathematical Challenge for Antivirus Technology
Metamorphic viruses, currently the most advanced computer viruses in the wild, have the unique ability of mutating their own code virtually into infinitely many highly dissimilar copies of themselves that nevertheless have the same functionality. This ability -- metamorphism -- together with othe...
When Context Gets Root: Privilege Escalation in LLM Harnesses
Instruction hierarchy is a model-side defense that assigns instructions different levels of privilege according to their sources. These levels constrain which content may direct model behavior. During agent execution, however, agent harnesses construct context for each model invocation. This...
When Relationships Break: Interpreting Network Traffic Anomalies Via Dependency Violations
Current research on security monitoring is increasingly focusing on machine-learning-based approaches, but caveats remain. In addition to huge computational overhead, one concern is the lack of insights into "why" alerts are raised. Existing interpretability approaches rely on feature attribution...
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent...
ContextLeak: Exfiltrating LLM Agent Context Via Malicious Tools
Exfiltrating an LLM agent's runtime context -- such as the user prompt, execution trajectory, and tool list -- poses severe security and privacy risks to users. Such attacks can be carried out via malicious tools and typically require three conditions: 1 the agent selects the malicious tool for...
Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering
Integrating Large Language Models LLMs into the Software Development Life Cycle SDLC can improve developer productivity, but it also introduces security, privacy, and compliance risks during model selection. Regulations and frameworks such as the EU AI Act, the NIST AI Risk Management Framework...
A Catalog of User Authentication Patterns
Security patterns are intended to support the design and development of secure software systems. However, although established catalogs of security patterns exist, their practical application remains limited. In particular, despite these catalogs, concrete patterns for common security controls su...
OWASP Top 10 for LLM Applications 2026
OWASP Top 10 for LLM Applications 2026 is the latest community-driven guide to the most critical security risks facing applications powered by large language models. Developed by hundreds of AI security experts, this edition introduces updated rankings, expanded threat coverage, and new research...
Joern 4.0.612
Joern is the bug hunter's workbench. With this tool, you can uncover attack surface, sloppy coding practices, and variants of known vulnerabilities using an interactive code analysis shell. Joern supports C, C++, LLVM bitcode, x86 binaries via Ghidra, JVM bytecode via Soot, and Javascript...
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlle...
SysComb: Fine-Grained Transparent System Call Filtering for Attack Surface Reduction
Restricting the system calls available to applications shrinks the kernel's attack surface and greatly mitigates the impact of compromised programs. Recent approaches showcase techniques to generate system call filters, however, all existing solutions require either kernel or application...
SPA: Securing Persistent LLM Agents across Queries with Plan-First Information-Flow Control
Large language model LLM agents increasingly operate over untrusted webpages, documents, tools, and persistent states while exercising authority over security-sensitive resources. Existing defenses typically protect either planning or individual tool interactions, but persistent agents face a...
Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted providers can keep these files secret while selling access to task results, making the skill itself a valuable target. Existing disclosure defenses can block request...
Low-ASR Backdoors: Exploiting Attack Success Rate Reduction and Attacker-Defender Asymmetry
Backdoor attacks are among the most effective and stealthy attacks in deep learning. Existing attacks and defenses are largely designed and evaluated under the assumption that successful backdoors exhibit high Attack Success Rates ASRs. In this paper, we show that this assumption creates a...
FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Evaluating the ability of large language models LLMs to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlo...
Are We Shooting Flies with Cannons? Trade-Off Analysis for AI-Based 5G Intrusion Detection
The increasing adoption of Artificial Intelligence AI in network intrusion detection raises the question of whether complex and computationally expensive models are justified for this task. In this work, we investigate the trade-off between detection performance and computational cost for intrusi...
The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action...
X-WAD: EXplainable Web Anomaly Detection
The rapid growth of web-based services, particularly API-driven architectures, reflects an increasing reliance on distributed systems, exposing sensitive data to security risks and making the adoption of automated defensive mechanisms essential. In this context, where benign traffic predominates ...
Nmap 7.99 Memory Exhaustion
Nmap version 7.99 and earlier contain a loop with an unreachable exit condition in the Packet:parseoptions method of nselib/packet.lua. A remote host that is the target of a scan can exhaust the memory of the scanning Nmap process and terminate it by replying with a packet that carries a...
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
Capture-the-Flag CTF benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with...
Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs
This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment TEE on NVIDIA B200 GPUs, using Intel Trust Domain Extensions TDX confidential VMs together with NVIDIA Confidential Computing CC on Blackwell GPUs. The...
A Hybrid Security Framework for Mini-Programs: Visual UI Compliance and Network Risk Assessment
With the continuous development of the WeChat ecosystem, WeChat Mini Programs, due to their advantages of not requiring installation, using little memory, and being ready to use instantly, have seen a surge in user numbers and have now become an indispensable service carrier in mobile internet...
Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation
Large Language Models LLMs generate register-transfer-level RTL code with rapidly improving functional correctness. Security of LLM-generated code, however, has been studied mainly for software, where flaws can still be patched after deployment. Insecure RTL offers no such remedy once taped out...
CodeQL 2.26.4
Discover vulnerabilities across a codebase with CodeQL, an industry-leading semantic code analysis engine. CodeQL lets you query code as though it were data. Write a query to find all variants of a vulnerability, eradicating it forever. Then share your query to help others do the same...
CISA: Vulnerability Review for Fiscal Years 2024 and 2025
Most compromises do not rely on advanced techniques or cutting-edge tools. Cyber threat actors scan the internet looking for exposed, well-known software vulnerabilities to exploit. Basic security failures enable most compromises and organizations can reduce their risk by addressing these...