5061 matches found
llm-guard
!WARNING このプロジェクトはアーカイブされました。 このプロジェクトおよび Hugging Face 上の関連モデルは、もはや積極的な開発やメンテナンスは行われていません。 LLM Guard - LLMインタラクションのためのセキュリティツールキット Protect AI による LLM Guard は、大規模言語モデル(LLM)のセキュリティを強化するために設計された包括的なツールです。 ドキュメント | プレイグラウンド | チェンジログ LLM Guardとは?...
Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-To-Text Jailbreaks
Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The...
The Poisoned Conversation: Privacy-Leaking Watermarks in Unified Multimodal Models
Multimodal models are increasingly shifting toward unified architectures that understand and generate text, images, and other modalities within a shared conversational context. This design enables fluid interaction across modalities, but it also changes the privacy threat model: Information...
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with...
AutoDP-LLM: Automating Data Pre-Processing for Intrusion Detection Systems Using Large Language Models
The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems IDS. However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative...
Static Bootstrap Placement for Encrypted Language Model Decoding
Language models increasingly serve prompts that carry private data, and secure inference under homomorphic encryption lets a client outsource the computation without revealing the prompt. Existing secure inference systems run a forward pass without consuming a token under encryption, and generati...
Forecasting Cybersecurity Incidents Using Geopolitical Data and Large Language Models
Predicting security incidents is a profound task critical for informing proactive defensive measures and cyber-insurance policies. Prior work tackling this problem mainly utilized structured, manually defined features based on network measurements e.g., protocol misconfigurations. Still, despite...
The vulnerability of the parseKeyString() function in the models/asymkey/ssh_key_parse.go file of the Git repository management system Gitea allows a attacker to cause a service failure.
The vulnerability of the parseKeyString function in the models/asymkey/sshkeyparse.go file of the Git repository management system Gitea is related to access control deficiencies. Exploiting this vulnerability could allow an attacker to cause service failures...
SecJev: Bringing Security Expertise to System One Decision Models
Security workflows need models that turn complex observations and explicit policies into decisions. System One models introduced by Jev return typed predictions and probabilities; security specialization supplies the domain expertise behind those predictions. We introduce SecJev, to our knowledge...
Picklescan is vulnerable to RCE through missing detection when calling numpy.f2py.crackfortran.myeval
SummaryPicklescan uses numpy.f2py.crackfortran.myeval, which is a function in numpy to execute remote pickle files. DetailsThe attack payload executes in the following steps:- First, the attacker crafts the payload by calling the numpy.f2py.crackfortran.myeval function in its reduce method- Then,...
PYSEC-2026-4132 Picklescan has a missing detection when calling built-in python profile.Profile.runctx
Summary Using profile.Profile.runctx, which is a built-in python library function to execute remote pickle file. Details The attack payload executes in the following steps: First, the attacker craft the payload by calling to profile.Profile.runctx function in reduce method Then when the victim...
PYSEC-2026-4131 Picklescan is vulnerable to RCE through missing detection when calling numpy.f2py.crackfortran.myeval
Summary Picklescan uses numpy.f2py.crackfortran.myeval, which is a function in numpy to execute remote pickle files. Details The attack payload executes in the following steps: - First, the attacker crafts the payload by calling the numpy.f2py.crackfortran.myeval function in its reduce method -...
PT-2026-105881
Summary Using profile.Profile.runctx, which is a built-in python library function to execute remote pickle file. Details The attack payload executes in the following steps: First, the attacker craft the payload by calling to profile.Profile.runctx function in reduce method Then when the victim...
The Innocent Courier: Covert Exfiltration through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models LLMs and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt...
Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks
Recent work has shown that large language models LLMs can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We...
Sleeping Secrets: How Fine-Tuning Reawakens Privacy Risks in Language Models
Beyond adapting Large Language Models LLMs to specialized applications, fine-tuning has recently been shown to recover private information that is no longer accessible through direct queries. Previous fine-tuning recovery attacks, however, require genuine private supervision drawn from the same...
Learning Normal Diffusion Dynamics for Backdoor Defense in Text-To-Image Models
Backdoor attacks pose a serious threat to the secure deployment of text-to-image T2I diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack...
Faithful Dual-Constrained Erasure for Robust LLM Safety Alignment
Machine unlearning has emerged as a crucial mechanism for removing hazardous knowledge and enforcing safety alignment in Large Language Models LLMs. However, recent studies reveal a persistent security risk: unlearned models remain highly vulnerable to retraining attacks, where suppressed malicio...
Towards Hierarchical Cyber Defense with Large Language Models: From Planning to Execution
An autonomous cyber defender trained with reinforcement learning RL is typically tied to the network on which it was trained, limiting its ability to generalize as network scale changes. Hierarchical RL reduces decision complexity by separating strategic targeting from tactical execution, but it...
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show th...