344 matches found
gradient-untangler
gradient-untangler A local research harness that searches for the exact tokens that make an open-weight language model start its answer the way you specify. Open-weight models ship as ordinary files: config.json, a tokenizer, and one or more .safetensors or .bin shards. A Hugging Face id is only ...
Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Models
Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale. Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with...
Does the Readout Bypass Leak the Input? A Feature-Visibility Audit of Hybrid Quantum-Classical Models
Readout-side residual hybrids concatenate raw inputs with measured quantum features. Under single-example gradient sharing, a biased first linear layer admits standard analytic recovery of its input, so the bypass exposes raw coordinates without requiring inversion of the quantum circuit. We audi...
Does the Unsafe Gradient Survive a Conversation? on the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the...
Forensic-Aware Continual Adaptation for Image Forgery Localization
The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization IFL methods can accurately localize manipulated regions but are often unable to adapt to newly emerging forgeries. In real-world forensic scenarios, data typically...
Inference-Layer Security: Defending against Adversarial Inference and Infrastructure Abuse
A Technical Report: Operating a large language model LLM as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of service, and...
Decoding Guardrails: XAI-Guided Perturbation Analysis of Prompt Injection Detection
Large language models LLMs are increasingly deployed in production systems, raising concerns about their exposure to adversarial manipulation through prompt injection and jailbreak attacks. Classifier-based guardrails, such as Prompt Guard 2, are widely used as a first line of defense against suc...
Towards TEE-Certified DP: Verifiable Differentially Private Training on Legacy GPUs
Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy DP emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier shou...
Astra Linux – Vulnerability in libvncserver
LibVNCClient is a library for easy implementation of a VNC client. In versions 0.9.15 and earlier, LibVNCClient’s Tight encoding decoder used fixed-size 2048-pixel scratch buffers for the Gradient filter. However, it did not reject Tight rectangles whose width exceeded 2048 pixels. A malicious VN...
Cascading Gradient Inversion Via LT-Code Inspired Peeling in Federated Learning
Federated learning shares model updates rather than raw data, yet these updates can be inverted to reconstruct the clients' training data. Analytic reconstruction attacks, which invert a gradient in closed form, degrade as the batch grows: prior single-round attacks recover only about half of a...
Federated Attack Campaign Detection Via Contrastive Encoding of Threat Indicators in Gradient Updates
Detecting orchestrated cyberattack campaigns that span multiple organizations traditionally requires sharing sensitive telemetry and threat intelligence across institutional boundaries and country borders, a barrier that Federated Learning removes by training shared threat detectors directly on...
MGASA-2026-0315 Updated libvncserver packages fix security vulnerabilities
The updated packages fix security vulnerabilities: Heap Out-of-Bounds Read in HandleUltraZipBPP due to unchecked subrectangle count. CVE-2026-32853 NULL pointer dereferences in httpd proxy handlers via malformed CONNECT/GET requests. CVE-2026-32854 LibVNCClient Tight Gradient decoding allows...
A Multi-Model Hybrid Defense Approach against White-Box Adversarial Attacks in Computer Network Traffic
It is crucial to safeguard computer networks from evolving network security threats and unknown cyberattacks. An essential tool for protecting computer networks against unknown cyber threats is Network Intrusion Detection System NIDS. However, NIDS faces a major security concern due to its...
MGASA-2026-0259 Updated libreoffice packages fix security vulnerabilities
The updated packages fix security vulnerabilities: Heap buffer overflow in DXF polyline import. CVE-2026-6039 Heap use-after-free in ODF number-format blank-width parsing. CVE-2026-6040 Heap buffer overflow in EMF+ gradient brush import. CVE-2026-6045 Stack buffer overflow in PPT presentation...
Updated libreoffice packages fix security vulnerabilities
The updated packages fix security vulnerabilities: Heap buffer overflow in DXF polyline import. CVE-2026-6039 Heap use-after-free in ODF number-format blank-width parsing. CVE-2026-6040 Heap buffer overflow in EMF+ gradient brush import. CVE-2026-6045 Stack buffer overflow in PPT presentation...
Evaluating Frontier AI Agents As Autonomous Clinical Security Auditors
Clinical AI models can expose patients to harm when adversarial vulnerabilities go undetected, yet formal security auditing requires statistical expertise, specialized tools, and significant time. We present an open evaluation task, built on METR Task Standard v0.3.0, that tests whether frontier ...
TensorFlow has Heap-buffer-overflow in AvgPoolGrad
Impactpythonimport osos.environ'TFENABLEONEDNNOPTS' = '0'import tensorflow as tfprinttf.versionwith tf.device"CPU": ksize = 1, 40, 128, 1 strides = 1, 128, 128, 30 padding = "SAME" dataformat = "NHWC" originputshape = 11, 9, 78, 9 grad = tf.saturatecasttf.random.uniform16, 16, 16, 16, minval=-128...
PYSEC-2026-3366 TensorFlow has Floating Point Exception in AvgPoolGrad with XLA
Impact If the stride and window size are not positive for tf.rawops.AvgPoolGrad, it can give an FPE. python import tensorflow as tf import numpy as np @tf.functionjitcompile=True def test: y = tf.rawops.AvgPoolGradoriginputshape=1,0,0,0, grad=0.39117979, ksize=1,0,0,0, strides=1,0,0,0,...
PYSEC-2026-3268 Overflow in `ResizeNearestNeighborGrad`
Impact When tf.rawops.ResizeNearestNeighborGrad is given a large size input, it overflows. import tensorflow as tf aligncorners = True halfpixelcenters = False grads = tf.constant1, shape=1,8,16,3, dtype=tf.float16 size = tf.constant1879048192,1879048192, shape=2, dtype=tf.int32...
`CHECK` fail via inputs in `SparseFillEmptyRowsGrad`
ImpactIf SparseFillEmptyRowsGrad is given empty inputs, TensorFlow will crash.pythonimport tensorflow as tftf.rawops.SparseFillEmptyRowsGrad reverseindexmap=, gradvalues=, name=None PatchesWe have patched the issue in GitHub commit af4a6a3c8b95022c351edae94560acc61253a1b8.The fix will be included...