38 matches found
RLCDAlignBench
Just Ask Jev: 較正された意思決定のための強化学習によるAIアラインメント失敗のゼロショット検出器 このリポジトリには RLCDAlignBench と論文の基盤となるコードが含まれています。 このベンチマークは、言語モデルの出力がアラインメント失敗であるかどうかを検出器が判別できるかを測定します。 10種類の失敗タイプ (追従性、ジェイルブレイク、欺瞞、プロンプトインジェクション、幻覚、プライバシー侵害、社会的バイアス、報酬ハッキング、不確実性の隠蔽、権力追求)にわたる44のベンチマーク と5つのターゲットモデル を含み、7,193のラベル付き検出インスタンス を提供します...
SAFE
SAFE SAFEは、研究アーティファクトにおけるSemgrepおよびTrivyの検出結果に対して、リポジトリを認識した制御されたセキュリティ評価を実行します。 これは、binary直接予測(SECURITYRELEVANTまたはNONSECURITY)と、詳細なmulticlassコンテキスト分類法(3つのラベル — 3つのラベルを参照)という2つの独立した分類タスクをサポートします。各タスクは、ゼロショットモードまたはエージェントモードで実行できます。 必要なものは以下のみです: 1. セミコロン区切りの検出結果CSV。 2...
Agentic-CLIP-Benchmark
Agentic-CLIP-Benchmark: CIFAR-10におけるゼロショット評価 -green.svg このプロジェクトは、OpenAIのCLIP ViT-B/32モデルを CIFAR-10テストセット全体(10,000枚の画像)で評価するための堅牢で自動化されたパイプラインを実装しています。AIネイティブワークフロー(Trae IDE)を用いて開発され、高精度なゼロショット精度 88.80%を達成しています。 📊 パフォーマンス概要 Top-1精度 : 88.80%(ゼロショット) モデル : openai/clip-vit-base-patch32(Safetensors使用...
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
The deployment of large language models LLMs in Swiss financial and regulatory contexts demands empirical evidence of both production reliability and adversarial security, dimensions not jointly operationalized in existing Swiss-focused evaluation frameworks. This paper introduces Swiss-Bench 003...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
NASimJax: GPU-Accelerated Policy Learning Framework for Penetration Testing
Penetration testing, the practice of simulating cyberattacks to identify vulnerabilities, is a complex sequential decision-making task that is inherently partially observable and features large action spaces. Training reinforcement learning RL policies for this domain faces a fundamental...
The vulnerability of the Zero Shot scFoundation software lies in the presence of a vulnerability in the borrowed component, allowing a perpetrator to execute arbitrary code.
The vulnerability of the Zero Shot scFoundation software is related to the presence of a vulnerability in the borrowed component. Exploiting this vulnerability allows a remote attacker to execute arbitrary code...
EUVD-2026-10577
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
EUVD-2026-10578
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
CVE-2026-23654 GitHub: Zero Shot SCFoundation Remote Code Execution Vulnerability
...
CVE-2026-23654 GitHub: Zero Shot SCFoundation Remote Code Execution Vulnerability
...
Microsoft GitHub Repo: Zero Shot scFoundation 安全漏洞
Microsoft GitHub Repo: Zero Shot scFoundation is a biological information research code base owned by Microsoft Corporation. There are security vulnerabilities present in Microsoft GitHub Repo: Zero Shot scFoundation. Attackers can exploit these vulnerabilities to execute code remotely...
CVE-2026-23654: Dependency on Vulnerable Third-Party Component
Dependency on vulnerable third-party component in GitHub Repo: zero-shot-scfoundation allows an unauthorized attacker to execute code over a network...
SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection
Large language models LLMs have shown remarkable capabilities in natural language processing tasks, yet their application in hardware security verification remains limited due to scarcity of publicly available hardware description language HDL datasets. This knowledge gap constrains LLM performan...
MultiVer: Zero-Shot Multi-Agent Vulnerability Detection
We present MultiVer, a zero-shot multi-agent system for vulnerability detection that achieves state-of-the-art recall without fine-tuning. A four-agent ensemble security, correctness, performance, style with union voting achieves 82.7% recall on PyVul, exceeding fine-tuned GPT-3.5 81.3% by 1.4...
LLM-FS: Zero-Shot Feature Selection for Effective and Interpretable Malware Detection
Feature selection FS remains essential for building accurate and interpretable detection models, particularly in high-dimensional malware datasets. Conventional FS methods such as Extra Trees, Variance Threshold, Tree-based models, Chi-Squared tests, ANOVA, Random Selection, and Sequential...
Benchmarking Large Language Models for Zero-Shot and Few-Shot Phishing URL Detection
The Uniform Resource Locator URL, introduced in a connectivity-first era to define access and locate resources, remains historically limited, lacking future-proof mechanisms for security, trust, or resilience against fraud and abuse, despite the introduction of reactive protections like HTTPS...
Lightweight LLMs for Network Attack Detection in IoT Networks
The rapid growth of Internet of Things IoT devices has increased the scale and diversity of cyberattacks, exposing limitations in traditional intrusion detection systems. Classical machine learning ML models such as Random Forest and Support Vector Machine perform well on known attacks but requir...