1654 matches found
DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents
When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator needs three facts: where the attack entered, which steps it...
AutoTrans: AI-Assisted Automatic Translation of Security Assertions for RISC-V Processors
Reusing a set of verified security assertions across RISC-V processor targets remains one of the most expensive bottlenecks in hardware security verification. Manual translation takes hours per assertion. Raw LLM translation is fast but unreliable, introducing signal hallucination, where the mode...
Understanding the Security Boundary of Obfuscation-Based On-Device LLM Protection
Trusted Execution Environments TEEs offer a promising mechanism for safeguarding the intellectual property of on-device Large Language Models LLMs. To overcome the inherent computational bottlenecks of TEEs, existing TEE-Shielded LLM Partition TSLP methods apply efficient obfuscation schemes to...
Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents
Large language model LLM agents are increasingly applied to penetration testing, but we still know little about what they can do or how they fail. We compare two PentestGPT-based systems: a legacy human-in-the-loop system running the open-weight Kimi K2.5, and a newer autonomous system running...
Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation
Static analysis remains a cornerstone of software security, yet the effectiveness of tools such as CodeQL is often limited by the substantial manual effort required to develop high-coverage query suites. While large language models LLMs have emerged as a potential solution for automated code...
Ensembling LLMs for AI-Augmented Cybersecurity Software Requirements Generation
Translating high-level controls from security standards into concrete, system-specific requirements is central to cybersecurity requirements engineering. Large language models LLMs can accelerate this labor-intensive, recall-sensitive task, but any single run is unreliable: it misses valid...
PrivAudit: A Dual-Lens Auditing Framework for Website Privacy Practices under the CCPA
Five years after the enforcement of the California Consumer Privacy Act CCPA, understanding how website privacy practices evolve at scale in response to regulation remains a key challenge for both researchers and regulators. Prior work and regulatory efforts have focused on manual and case-specif...
CVE-2026-78230
AshAi exposes Ash read actions to language-model tool calls. The read tool accepts an aggregate result type min, max, sum, avg that builds an ad-hoc Ash.Query.Aggregate over a named field and returns its raw value. Ash field policies redact forbidden fields on returned records replacing them with...
CVE-2026-78230 AshAi aggregate tool can read field-policy-protected fields
AshAi exposes Ash read actions to language-model tool calls. The read tool accepts an aggregate result type min, max, sum, avg that builds an ad-hoc Ash.Query.Aggregate over a named field and returns its raw value. Ash field policies redact forbidden fields on returned records replacing them with...
Stealing AI Reasoning Traces
Interesting research: "Stealing Reasoning Traces from Proprietary LLM APIs": Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces...
Scrapegraph-ai v2.2.4
🚀 Looking for an even faster and simpler way to scrape at scale only 5 lines of code? Check out our enhanced version at ScrapeGraphAI.com! 🚀 🕷️ ScrapeGraphAI: You Only Scrape Once English | 中文 | 日本語 | 한국어 | Русский | Türkçe | Deutsch | Español | français | Português | Italiano ScrapeGraphAI is a...
llm-supply-chain-backdoor-lab
Large Model Supply Chain Backdoor · Replicating the Attack...
LLMSec-AV: A Vulnerability Taxonomy and LLM-Driven Software Weakness Discovery Framework for Autonomous Vehicles
Automated vehicles rely on millions of lines of safety-critical software, yet general-purpose analyzers do not understand which code can affect vehicle motion. This study asks whether large language models LLMs with explicit automated-vehicle AV security knowledge improve weakness detection beyon...
Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems
Long-running language-model agents depend on persistent memory. Many agent-memory systems preserve history through soft revocation: a contradicted fact is marked invalid and retained rather than deleted. However, whether that mark is enforced at retrieval time is unexamined. In this paper, we...
Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
Large language model LLM benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as...
HoneyRoute: Honeypot-Model Routing for Adversarial LLM Serving
We introduce HoneyRoute, an inference-serving layer that detects whether an incoming request is malicious and, if so, routes it to a dedicated honeypot model, shielding production while the adversary's interaction is continuously harvested for intelligence. Existing defenses embed traps inside...
ACEA: An Adversarial Co-Evolution Arena for Head-To-Head Red-Team and Blue-Team LLM Testing
Automated red-team attacks and blue-team defenses for large language models LLMs are advancing quickly. However, attackers and defenders are built and tested in isolation, and the resulting scores are hard to trust. To tackle this, we present ACEA Adversarial Co-Evolution Arena, a platform that...
PT-2026-87175
AshAi exposes Ash read actions to language-model tool calls. The read tool accepts an aggregate result type min, max, sum, avg that builds an ad-hoc Ash.Query.Aggregate over a named field and returns its raw value. Ash field policies redact forbidden fields on returned records replacing them with...
xalgorix v4.6.72
Xalgorix — Open-source AI pentester that proves vulnerabilities Most scanners detect. Xalgorix proves. An autonomous LLM agent works a full pentest methodology, then an independent verifier re-exploits every finding before it's reported — so you get proof, not a pile of maybes to triage...
VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities
The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependen...