3 matches found
Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios
Large Language Models LLMs are widely used for code generation, yet their security behavior in realistic development workflows remains underexplored. Existing benchmarks often rely on explicitly specified security requirements, failing to capture real-world scenarios where prompts are frequently...
AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems
Agentic AI systems retrieve private context, invoke tools, write files, call external services, coordinate with other agents, and may act without human approval. Existing bill of materials artifacts improve transparency for dependencies, model metadata, and training provenance, but leave an agent...
SOSBENCH: Benchmarking Safety Alignment on Scientific Knowledge
Large language models LLMs exhibit advancing capabilities in complex tasks, such as reasoning and graduate-level question answering, yet their resilience against misuse, particularly involving scientifically sophisticated risks, remains underexplored. Existing safety benchmarks typically focus...