39669 matches found
Security-Enhanced Seed-Based Weight Quantization for Large Language Models
Large language models LLMs incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity ...
The Geometry of Harmfulness in Multi-Turn Attacks
Large language models LLMs remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-tur...
[SECURITY] [DLA 4799-1] lxml security update
Debian LTS Advisory DLA-4799-1 [email protected] https://www.debian.org/lts/security/ Guilhem Moulin September 28, 2026 https://wiki.debian.org/LTS Package : lxml Version : 4.9.2-1+deb12u1 CVE ID : CVE-2026-28348 CVE-2026-28350 CVE-2026-41066 CVE-2026-49825 Multiple vulnerabilities were...
MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual...
Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks
Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this...
Distillation Defenses Easily Break after Reinforcement Learning
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train i.e., "distill" their own model...
Share-Borne AI Virus: Memory-Hopping Attacks across LLM Agents
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent...
TEE Anchor: Cross-TEE Organizational Endorsement for Mitigating TEE Physical Attacks
In 2025, practical physical-access attacks against TEEs, such as TEE.fail and Battering RAM, were disclosed, posing a serious threat to current TEEs. The more strictly the target machine is guarded, the harder such attacks are to mount. Residing under trusted management has therefore emerged as a...
Inference-Layer Security: Defending against Adversarial Inference and Infrastructure Abuse
A Technical Report: Operating a large language model LLM as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and...
Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness
Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservatio...
Dual Lattice Attacks for Bounded Distance Decoding, Revisited
Analyses of dual lattice attacks have often assumed that the individual scores associated with short dual vectors are mutually independent. Laarhoven-Walter used this heuristic to derive explicit trade-offs between the target radius and query time for bounded distance decoding BDD with...
No Free Efficiency: Revisiting the Trade-Off between Training Efficiency and Model Vulnerability
Training efficiency has become the central driver of recent progress in foundation models. To overcome the massive computational and data requirements of large-scale training, researchers increasingly adopt strategies such as selective data sampling, efficient pre-training, and simplified...
Frame the Adversary: A Structure-Aware Attack Methodology
Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical an...
AGATE: Provenance-Based Runtime Defense against Compositional Attacks on LLM Agents
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness...
FeatMark: Feature-Level Watermark Protection against Mimicry Attacks with Diffusion Models
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies sho...
How Much Must a Private Mempool Hide? Exact Leakage Thresholds for Sandwich Attacks
Private and encrypted mempools hide pending transactions to stop sandwich attacks and other forms of maximal extractable value MEV, but what they hide is rarely everything: a transaction's pair, direction, and a coarse range for its size can still leak. How much leakage makes sandwiching pay? We...
No Place to Hide: An Analysis on Protected Order Flow Sandwich Attacks
Front-running has long plagued Ethereum's public mempool, earning it the nickname of a "dark forest", where predators lurk for profitable transactions. In response, Ethereum and other blockchain ecosystems increasingly rely on private RPCs and native protections to shield transactions from...
Fast Frame Rate Estimation in Electromagnetic Side-Channel Attacks on Public Systems
Frame refresh rate estimation is a fundamental step in identifying compromising harmonic frequencies in electromagnetic side-channel attacks. Methods based on discrete linear autocorrelation DLA are robust across different scenarios and require a computational complexity of $ON\log N$ for a signa...
Backdoors in Learning-Based Industrial Robotic Arm Manipulation: An Empirical Security Study
Learning-based models e.g., visuomotor and Vision-Language-Action VLA are increasingly explored for industrial robotic manipulation, where model predictions are directly translated into physical actions. This tight coupling between model behavior and physical execution makes hidden security...
CVE-2026-77582: Observable Timing Discrepancy
Tinyauth is an authentication and authorization server. Prior to 5.1.0, Tinyauth exposes a remotely observable timing difference between authentication attempts for existing and nonexistent local usernames. internal/controller/usercontroller.go loginHandler and...