1 matches found
WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks specify only functional objectives without explicit security requirements. WeSCE consists of 400 executable programs derived from real-world code, covering...
5.5AI score
SaveExploits0
20