6 matches found
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training
Despite recent advances, frontier large language model LLM agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available agents can already aid attackers, who only need to find one exploitable weakness, while defenders must continuously...
Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection
Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the function to be classified correctly. Agentic reinforcement learni...
EUVD-2026-46150
A flaw was found in Red Hat Quay's repository-level mirror configuration feature. The POST and PUT handlers in endpoints/api/mirror.py accept an externalreference parameter without SSRF validation, unlike the organization-level mirror handlers which apply validateexternalregistryurl. A repository...
PT-2026-61830
Name of the Vulnerable Software and Affected Versions Red Hat Quay affected versions not specified Description A flaw exists in the repository-level mirror configuration feature. The POST and PUT handlers in the endpoints/api/mirror.py file accept an external reference parameter without Server-Si...
An Empirical Study of Security Calibration in Large Language Models for Code
Large Language Models LLMs are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do models know when their generated code is insecure? This property, known as calibration, measures whether a model's confidence aligns with the true...
A.S.E: a Repository-Level Benchmark for Evaluating Security in AI-Generated Code
The increasing adoption of large language models LLMs in software engineering necessitates rigorous security evaluation of their generated code. However, existing benchmarks are inadequate, as they focus on isolated code snippets, employ unstable evaluation methods that lack reproducibility, and...