1 matches found
AURA-Eval: Evaluation Framework for Acting under Risk Awareness in LLM Agent Trajectories
LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combini...
6.1AI score
SaveExploits0
20