44 matches found
SAFE
SAFE SAFE performs controlled, repository-aware security assessment of Semgrep and Trivy findings in research artifacts. It supports two independent classification tasks: direct binary prediction SECURITYRELEVANT or NONSECURITY and the detailed multiclass contextual taxonomy three labels — see...
RLCDAlignBench
Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures This repository holds RLCDAlignBench and the code behind the paper. The benchmark measures whether a detector can tell when a language model's output is an alignment failure. It has 44...
Agentic-CLIP-Benchmark
Agentic-CLIP-Benchmark: Zero-Shot Evaluation on CIFAR-10 -green.svg This project implements a robust, automated pipeline for evaluating OpenAI's CLIP ViT-B/32 model on the full CIFAR-10 test set 10,000 images. Developed using an AI Native workflow Trae IDE , it achieves a high-precision zero-shot...
Benchmarking LLMs for Threat Level Determination
The fast progress of large language models LLMs opens new opportunities in the management of cyber threat intelligence, but their reliability for operational tasks remains unclear. In this work, we benchmark LLMs on the task of threat level determination. First, we construct a curated dataset...
Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection
The Android ecosystem faces persistent and rapidly evolving malware threats. Existing machine learning detectors are vulnerable to concept drift because they rely on implementation-specific features whose distributions change over time. Large language models LLMs offer strong semantic understandi...
LLMs for Zero-Shot Threat Detection Via Structured Risk Indicators
We propose a two-stage large language model LLM framework for zero-shot detection of insider threats and advanced persistent threats APTs from heterogeneous security logs. The framework models user activity as chronological timelines and incorporates retrieval-augmented generation RAG to provide...
From Evaluation to Optimisation: Hierarchy-Aware Training Signals for CWE Prediction in Python
The original ALPHA benchmark introduced a taxonomy-aware penalty for evaluating CWE-level vulnerability prediction in Python and proposed that the penalty could theoretically also serve as a training signal. This paper provides that validation. We compare three delivery mechanisms: supervised...
TGCM: Topic-Guided Generative Disentanglement of Interleaved APT Technique Sequences
In enterprise environments, multiple Advanced Persistent Threat APT campaigns often unfold concurrently, producing audit logs in which attack techniques across actors sources are interleaved over time. This setting naturally gives rise to an Unknown-K Interleaved Sequence Demixing UKISD problem:...
Data-Centric Benchmarking of Exploit Generation in LLMs: Understanding the Impact of Fine-Tuning
We study the task of CVE-conditioned exploit generation, where a model drafts proof-of-concept PoC exploits given software vulnerability context. We adopt a data-centric approach, constructing a high-quality dataset via multi-stage preprocessing and introducing a scalable evaluation framework wit...
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
The deployment of large language models LLMs in Swiss financial and regulatory contexts demands empirical evidence of both production reliability and adversarial security, dimensions not jointly operationalized in existing Swiss-focused evaluation frameworks. This paper introduces Swiss-Bench 003...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
NASimJax: GPU-Accelerated Policy Learning Framework for Penetration Testing
Penetration testing, the practice of simulating cyberattacks to identify vulnerabilities, is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Training reinforcement learning RL policies for this domain faces a fundamental...
The vulnerability of the Zero Shot scFoundation software lies in the presence of a vulnerability in the borrowed component, allowing a perpetrator to execute arbitrary code.
The vulnerability of the Zero Shot scFoundation software is related to the presence of a vulnerability in the borrowed component. Exploiting this vulnerability allows a remote attacker to execute arbitrary code...
EUVD-2026-10577
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
EUVD-2026-10578
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654 GitHub: Zero Shot SCFoundation Remote Code Execution Vulnerability
...
CVE-2026-23654 GitHub: Zero Shot SCFoundation Remote Code Execution Vulnerability
...
Microsoft GitHub Repo: Zero Shot scFoundation 安全漏洞
Microsoft GitHub Repo: Zero Shot scFoundation is a biological information research code base owned by Microsoft Corporation. There are security vulnerabilities present in Microsoft GitHub Repo: Zero Shot scFoundation. Attackers can exploit these vulnerabilities to execute code remotely...