135 matches found
better_opts_attacks
¿Puedo tener su atención? Rompiendo defensas de inyección de prompts basadas en fine-tuning mediante ataques conscientes de la arquitectura Este repositorio contiene el código para ejecutar los ataques ASTRA y ASTRA++ que rompen SecAlign++, SecAlign, StruQ. Este repositorio también contiene algun...
DeeCLIP
DeeCLIP: Un marco robusto y generalizable basado en Transformer para detectar imágenes generadas por IA ☀️ Si este trabajo te resulta útil para tu investigación, ¡por favor, dale una estrella a nuestro repositorio y cita nuestro artículo! ☀️ TODO Estamos trabajando duro en los siguientes elementos:...
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on...
refusal-relocates
Quick start · Usage · CLI · Results · Reproduce · FAQ · Citation News 2026-09-28 Code released, with the run records and scripts that rebuild the paper's tables. Overview A few dozen harmful examples can remove an aligned model's refusal of harmful requests. Prior work localizes safety to specifi...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
Backdoor Purification for LoRA-Tuned LLMs Via Null-Space Projection
With the rapid adoption of large language models LLMs and parameter-efficient fine-tuning PEFT methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access...
SecureVibe: Making Vibe Coding More Secure
As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective...
Backdoor Mitigation in Decentralized LLM Fine-Tuning
Decentralized large language model LLM fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This...
Selective Channel Restoration for Backdoored Vision-Language Models
Vision-language models VLMs exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these...
RAISE: Reinforcing Access Control Policy Synthesis in LLMs Via Symbolic Evaluation
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first...
Exploring Automated Vulnerability Identification in JavaScript Code Using Large Language Models
JavaScript powers approximately 98.8% of all websites, making vulnerabilities in its code a significant security risk, yet existing detection approaches such as Static Application Security Testing SAST tools often fail to identify many real-world vulnerabilities when applied to isolated code...
Benchmarking LLMs for Threat Level Determination
The fast progress of large language models LLMs opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset...
Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text
Synthetic data is increasingly used to train large language models LLMs, yet its security implications remain poorly understood. Prior work on subliminal learning suggests that models can inherit behavioral traits from seemingly unrelated training data. In this work, we investigate whether such...
Improving LLM-Based SSH Honeypots through Prompting and Fine-Tuning
LLM-based SSH honeypots often use closed cloud LLMs because they give strong shell realism, but cloud models create deployment problems. These include no stable versioning, provider-side changes, attacker-driven cost, and model decommissioning. Local open-weight models avoid these problems, but...
Backdoor Decontamination Dynamics in LLM Agents
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a kno...
Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
Large Language Models LLMs are increasingly fine-tuned for critical-domain Question-Answering QA, yet choosing which small model to adapt, before paying the cost of adaptation, remains difficult. Fine-tuning can improve domain alignment, but it may also erode prior knowledge, weaken...
Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions
Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code. However, these models remain vulnerable to backdoor attacks, where malicious fine-tuning data covertly implants unsafe behaviors. Despite advances in defensive...
Decomposing Memorization Reduction in Privacy-Preserving Fine-Tuning of SLMs for CSIRTs
CSIRTs increasingly fine tune language models on vulnerability scan records, but these records expose internal network topology and create privacy risks under regulations such as GDPR and LGPD. We present the first empirical study of how DP SGD and HMAC pseudonymization interact when fine tuning...
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy...
Empirical Evaluation of Large Language Models for Migration of Code Fragments to Post-Quantum Cryptography
The transition to post-quantum cryptography PQC requires not only replacing vulnerable cryptographic primitives, but also refactoring the surrounding software logic. While existing PQC migration frameworks provide organizational guidance, practical code-level remediation remains largely manual an...